how it works
Instead of a whiteboard,
they get a live incident.
One candidate, one AI teammate, one system that's breaking. Nobody's drawing boxes from memory; they have to actually fix it, with the clock running. It's a new kind of interview, so here's the whole run, start to finish, and what lands on your desk after.
It starts vague, on purpose.
The candidate gets one alert: search is slow in us-east. That's the whole brief. The architecture diagram, the logs, the metrics, last week's deploys, all of it exists, but none of it is on screen yet. Real incidents don't arrive with a dashboard open to the answer. So the first thing you learn about a candidate is whether they can turn a vague page into good questions.
SearchApiLatencyHigh: p99 above 20 ms SLO for 12 m (now 9,214 ms)
The questions are the interview.
There's an AI teammate in the room, and it opens with one question: “Where do you want to start?” It has everything: the architecture diagram, the repo, the deploy history, logs and metrics for every service. But it only shows what the candidate asks for. So what you end up with is a record of their thinking: what they asked for, in what order, and how many dead ends they hit on the way to the answer.
And they can ask about anything, not just the scripted stuff. If they check database latency and it's fine, the teammate says so. That's a real answer too: knowing what isn't broken is half of debugging.
They say the fix. The system responds.
“Roll back that deploy.” “Add backoff to the retry loop.” “Put a cache in front of search.” The teammate does it, and the system reacts the way a real one would: latency comes back down or it doesn't, the new cache shows up in the metrics, and a fix that misses the point changes nothing.
Every scenario has traps. A fix that only treats the symptom buys a few quiet minutes, then it breaks again. Nobody fails for forgetting syntax, and nobody passes by sounding confident. You find out who actually fixes the cause before they're on your on-call rotation.
No AI opinion in the pass/fail.
Every scenario has one real root cause, decided when we build it. Pass or fail is measured against that. AI grades some of the writing, but it's clearly labeled and it never decides the verdict.
The pass bar
Passing means two things happened: they named the root cause, and their fix actually brought the system back, with latency under the limit and holding when time runs out. You can't argue with a recovered system.
Partial credit, defined up front
Isolated: they found where it's breaking, not yet why. Mitigated: they stopped the bleeding and said it wasn't the real fix. Root-caused: they named the exact defect. Getting close counts as close, not as zero.
The path they took
Everything they asked for, everything they did, and every dead end, timestamped from page to fix. This is where you actually read the investigation. It's a record of what happened, not a score of how they seemed.
The postmortem
They finish with a short writeup: what broke, how they fixed it, and how they'd stop it happening again. AI grades the writing against a rubric and links every point back to the transcript. It adds color to the report. It never touches the verdict.
Everyone applying for the role runs the same scenario against the same bar, so you can compare candidates fairly, and if a decision ever gets challenged, the evidence is right there. The report shows your team what happened; you decide what matters.
How it fits into your hiring loop.
It sits in front of your phone screen, and nothing else in your process changes. Your engineers only meet candidates who have already fixed something real.
- 01
Import the role
Paste the job description or pick your stack: Kafka, Redis, Postgres. The incident runs on systems that look like yours, generalist by default and stack-specific when you want to test for fluency.
- 02
Send one link
It runs in the browser. No setup, no scheduling. The candidate can take it tonight instead of playing calendar tetris with your interviewers.
- 03
The incident runs
Under an hour, timeboxed like real on-call. Everything gets captured as it happens: the conversation, what they asked for, what they did, and the postmortem.
- 04
Read the report in three minutes
Verdict on top, the path they took under it, full transcript attached. You decide who meets your engineers.
The screen costs your engineers nothing. A whiteboard round burns a senior hour to produce one opinion; here, that hour goes to candidates who already cleared a real bar.
The only prep that works is having done the job.
There's a whole industry for rehearsing interviews: question banks, mock interviews, the diagrams everyone memorizes. It exists because most interviews ask the same questions, so the answers can be studied. And all a rehearsed pass tells you is that someone is good at interviews.
None of that works here. The incident starts vague, the system only shows what they ask for, and no two candidates take the same path. There's no answer key to study, because the path is the answer: what they asked, what they did, and how they explained it after.
So the people who pass are the people who have actually been paged at 2 AM and worked out what broke. Grinding doesn't get you through it. You end up hiring people who know what they're doing.
A new interview for an
era of rehearsed answers.
Screen for real engineering judgment without spending an engineer-hour.
Want to try it yourself?