Run an end-to-end test
Run this before pointing production alerts at Shankh, and again after changing sources, channels, teams, schedules or escalation.
Before you start
- A workspace with members and at least one team.
- An alert source connected, or the webhook ingest URL and key.
- A connected, tested notification channel.
- An on-call assignment covering right now.
- Escalation levels configured for the severity you will test.
Trigger a test alert
- Use a low severity so you do not disturb a real rotation.
- For AWS, breach a test threshold on a non-production resource.
- For the webhook, send a test POST with Info or Warning severity.
- Note the time you triggered it.
Verify ingestion
- Open Incidents.
- Confirm a new incident appears in the expected workspace.
- Confirm the service, severity and source are correct.
- Confirm the trigger time matches when you sent it.
Verify notification and escalation
- Confirm the on-call engineer received the notification on the channel configured in step 1.
- Wait past the step-1 delay without acknowledging.
- Confirm escalation advances to step 2.
- Confirm the second team, person or channel was reached.
Verify acknowledgement and resolution
- Acknowledge as the on-call engineer and confirm the state changes to Acknowledged.
- Confirm the acknowledging user, time and note are recorded, and escalation stops.
- Resolve with a resolution note and confirm the state changes to Resolved.
- Open Reports and confirm the test incident is counted with an MTTA and MTTR.
Record the result somewhere your team can see it, then repeat after any change to routing or channels.
If a step fails
| Failed step | Where to look |
|---|---|
| No incident created | Alert source setup, webhook key, alert rule enabled |
| Incident but no notification | Channel setup and testing, escalation step channels |
| Notification to the wrong person | Workspace assignment on the alert, teams on the level, on-call calendar |
| Escalation did not advance | Escalation levels, step delays, incident state |
| Nobody was reached at all | An uncovered calendar period, or a level pointing at an empty team |
| Metrics missing | Confirm the incident was resolved rather than left active |
When a level has nobody: if nobody is on call, Shankh skips Level 0 for Level 1 — and if no Level 1 exists, the incident notifies nobody. If a level exists but resolves to zero members, Shankh falls back to the workspace's primary contact and all its Admins by push only, plus a single Slack post. If neither of those exists either, the incident is silent. The chain keeps looping back to Level 0 until somebody acknowledges.