Source checkpoints
Inventory
Document the current operating model before changing alert routes:- Users, teams, service owners, and backup responders
- Rotations, overrides, holidays, calendars, and time zones
- Escalation chains, route rules, grouping behavior, and severity mapping
- Integrations, ChatOps hooks, runbooks, notification preferences, and audit needs
- Alert sources that must be parallel-run during the cutover window
Migration pattern
1
Pick one service
Start with one service that has clear ownership and enough alert traffic to validate the
workflow.
2
Rebuild responder coverage
Recreate the Grafana OnCall rotation in Squasher, then verify the schedule, current engineer,
calendar feed, runbook link, and severity routing.
3
Mirror alerts
Send copied alerts to the new path while Grafana OnCall OSS remains visible. Compare
acknowledgements, escalations, and duplicate noise.
4
Connect Squasher context
Add Squasher monitors, log or error ingestion, status workflows, and incident links so
responders have evidence after alerts arrive.
5
Cut over deliberately
Move the service route, name the rollback owner for the first week, then retire duplicate
notifications after one deploy cycle.
Onboarding checklist
- Confirm every responder can sign in to the new incident workflow.
- Verify time zones, contact methods, notification preferences, and quiet hours.
- Test SMS, push, email, Slack, and phone delivery in the tool responsible for paging.
- Run one shadow week with both old and new routes visible.
- Review missed, duplicate, and delayed notifications before switching the next service.
- Train responders on Squasher incident pages, monitor history, AI triage, and status update workflows.
Squasher setup
- Use status monitoring for monitor-driven incidents and status-page updates.
- Use on-call to recreate and verify Grafana OnCall schedules before cutover.
- Use the Query Guide API for script and agent discovery of logs, traces, monitors, and incidents.
- Use OpenTelemetry, log drains, or SDKs to preserve investigation context.
- Use status pages when customers need public updates during incidents.
Cutover validation
- A test alert reaches the right responder and escalation path.
- The responder can find the linked Squasher incident, evidence, logs, traces, monitor checks, and runbook.
- Status updates are published only by the owners you expect.
- Rollback ownership is documented for the first week after cutover.
- Duplicate Grafana OnCall routes are removed only after the overlap window is complete.