Who does it: In Indian companies with 5–50 engineers, there is no dedicated SRE or incident manager. Incident response falls to whoever is on-call—which is whoever the CTO pointed at on a list. The on-call engineer is usually a full-stack developer who also handles feature work.
What they use, in order of prevalence:
- WhatsApp group — one per service or product, 10–40 people. An alert fires, someone posts "is anyone seeing the DB latency spike?" Three people reply with screenshots from different dashboards. Someone makes a call to roll back. No one writes down what happened.
- Shared Excel or Google Sheet — columns for Incident ID, Time, Service, Description, Assignee, Status. Updated manually by whoever is free. After the incident, the sheet is rarely touched again.
- Phone/SMS for P0 pages — for database down or payment failures, the CTO or senior dev gets a phone call at 2 AM. The caller's number is in someone's personal contacts.
- Monitoring tools in isolation — AWS CloudWatch, Datadog, Grafana, or cheaper tools like Healthchecks.io send alerts into the WhatsApp group. The tool watches; no structured response happens.
- Ticketing systems retrofitted — Jira or Freshdesk used for incidents because the company already pays for it. The workflow is: create ticket, assign, resolve. No runbook, no post-mortem, no cascade analysis.
- 3–5 hours per major incident in coordination overhead: who is on-call, what broke, who is investigating, when did we resolve. The actual fix might take 20 minutes. The coordination takes the rest.
- Midnight false alarms — WhatsApp alerts fire for non-critical items. Engineers get desensitized. Real P0 alerts get missed because everything looks the same.
- Post-mortem never happens — 90% of incidents have no documented root cause. The same incident recurs 6 weeks later. The cost of the second incident is fully preventable.
- Knowledge lives in people's heads — the senior dev who fixed the 3 AM outage in June has since left. The next one starts from scratch.
- On-call burnout — engineers managing incidents via WhatsApp for 12–18 months report higher attrition. No structured escalation means senior engineers handle incidents junior engineers could route.