Previous incidents
3 total · all resolved
Uptime Runner — scripted major outage
Detected — 5.00s into gameplay, automated monitoring (it was in the code) ·
Acknowledged — instantly, the on-call engineer was already looking at it (the on-call engineer is me) ·
Resolved — after exactly 2.5s. MTTR: 2.5s. Players resumed exactly where they left off.
Post-mortem: root cause was a candidate demonstrating he understands incident communication. No action items. Will happen again for every new player.
2020 — Career stability outage
Detected — engineer decides to found a coliving company. During a pandemic. In Latin America. ·
Impact — total loss of salary, office chair, and certainty. ·
Resolved — 3 cities, 40 staff and $800K raised later. Root cause accepted as a feature. Duration: 5 years and counting.
Candidate motivation — brief degradation
Detected — Belgium loses a football match. ·
Impact — motivation dropped to 97% for approximately one evening. ·
Resolved — automatically, by morning. Recurrence expected; impact contained.