On-call is how software teams make sure that when production breaks at 3am, someone who can fix it is reachable, and getting it right is as much about humans as about tooling.
On-call is the practice of designating engineers to be available outside normal working hours to respond to production incidents, so that whenever something breaks, someone with the access and authority to fix it can be reached quickly.
On-call runs on a rotation: a schedule assigns which engineer is responsible for responding during a given period, and alerts route to that person through a paging tool. If the primary on-call does not acknowledge within a set time, an escalation policy pages a secondary, then a manager, so an alert is never lost. Rotations hand off between shifts, and larger organisations run follow-the-sun coverage across time zones so no one is paged in the middle of their night. The mechanics are simple; making them humane is the hard part.
Systems run around the clock and incidents do not wait for business hours, so someone has to be reachable. But on-call imposes a real human cost: interrupted sleep, constant low-level stress, and burnout, especially when alert quality is poor and people are woken for things that were not actionable. Badly run on-call drives away exactly the experienced engineers a team can least afford to lose, which is why on-call health is an operational concern, not just a scheduling detail.
Sustainable on-call comes from a few disciplines: every alert that pages a human should be actionable, with a clear owner and a runbook; rotations should be fair and shifts a reasonable length; escalation paths should be clear; and, most powerfully, the volume of incidents that reach a human at all should be reduced. The best way to improve on-call is to page people less, by cutting noise through correlation and by resolving the common, well-understood incidents automatically before they ever become a page.
Ops Singularity improves on-call by attacking its root cause: Sentinel AI correlates alerts into real incidents and resolves the common ones autonomously through governed Action Tickets, so on-call engineers are paged for genuinely novel problems instead of a storm of noise.
It means an engineer is designated to be available, often outside normal hours, to respond to production incidents, reached through a paging tool with escalation to backups if they do not acknowledge.
Reduce the number of non-actionable pages, keep rotations fair and shifts reasonable, provide clear runbooks, and, most effectively, resolve common incidents automatically so fewer reach a human at all.
A rule that pages a secondary on-call, and then a manager, if the primary does not acknowledge an alert within a set time, so no incident is missed.
Ops Singularity turns open telemetry into autonomous, governed resolution. See it on your own stack.