Uptime — The On-Call Triage Game

You are on call. Twelve alerts, four possible responses, one error budget. Roll back, restart, page or snooze — free, no sign-up.

Uptime🎯 Score 0
  • ⏱ Error budget43/43 min
  • 🫂 Team▮▮▮▮▮5/5
  • 🟢 Uptime100.000%

Twelve alerts. Read the facts, pick one response. Not everything is an emergency.

How to play
  • No customer impact? Snooze it. Nothing internal is worth waking someone for.
  • Customer impact and a deploy in the last 30 minutes? Roll back — undoing a change beats diagnosing it.
  • Customer impact, nothing deployed recently, but a current runbook? Restart, and follow the runbook.
  • Customer impact and neither? Page the owner. Watch for old deploys and stale runbooks — they do not count.

How it works

You are on call and the only thing you control is what you do about each alert. A shift deals twelve of them, one at a time, and each card states its facts: whether a customer can see it, when the service last shipped, and whether a runbook covers the symptom. From those three facts exactly one of four responses is right. The rule is on screen from the first card and never changes — what changes is how carefully you have to read, because a deploy six hours ago is not the suspect and a runbook written for the old architecture is not a runbook. Two meters can end your shift early: the error budget, which a snoozed customer-visible incident burns through fast, and your team's patience, which every needless 3am page spends. Runs on your device, nothing is sent anywhere, and there is nothing to sign up for.

A game about incident judgement, not an operations manual. Real on-call has more than four options and far worse hours.

Frequently asked questions

What is the actual rule?

If a customer cannot see it, snooze it until business hours. If they can: roll back when a deploy went out in the last thirty minutes, restart when a current runbook covers the symptom, and page the service owner when neither applies. That order matters — undoing a recent change is faster than diagnosing it, which is why a fresh deploy outranks a runbook.

Why is snoozing punished so much harder than paging?

Because the mistakes are not equally bad. Sleeping through something customers can see burns error budget the entire time you are asleep. Waking a colleague for something a runbook already covers costs goodwill and nothing else — annoying, recoverable, and not an outage. The scoring is the lesson.

Are the alerts random?

The deck is generated from a seed, so a shift is reproducible, but the mix is deliberate: no single response is ever right more than about a third of the time, so no habit beats reading. The subtle cards — an old deploy, a stale runbook — arrive later in the shift, which is the difficulty curve.

What does the uptime figure mean?

A month of 99.9% availability allows about 43 minutes of downtime, and that is your starting error budget. Every wrong call spends some of it, and the percentage shown is what your SLO dashboard would read at the end of the month. Three decimal places, because that is where a 99.9% target lives.

Related tools

Embed this game

Add this free game to your own site: