how it works

Instead of a whiteboard,
they get a live incident.

One candidate, one AI teammate, one system that's breaking. Nobody's drawing boxes from memory; they have to actually fix it, with the clock running. It's a new kind of interview, so here's the whole run, start to finish, and what lands on your desk after.

01 · THE PAGE

It starts vague, on purpose.

The candidate gets one alert: search is slow in us-east. That's the whole brief. The architecture diagram, the logs, the metrics, last week's deploys, all of it exists, but none of it is on screen yet. Real incidents don't arrive with a dashboard open to the answer. So the first thing you learn about a candidate is whether they can turn a vague page into good questions.

02 · THE INVESTIGATION

The questions are the interview.

There's an AI teammate in the room, and it opens with one question: “Where do you want to start?” It has everything: the architecture diagram, the repo, the deploy history, logs and metrics for every service. But it only shows what the candidate asks for. So what you end up with is a record of their thinking: what they asked for, in what order, and how many dead ends they hit on the way to the answer.

And they can ask about anything, not just the scripted stuff. If they check database latency and it's fine, the teammate says so. That's a real answer too: knowing what isn't broken is half of debugging.

Assessment #4127search-apiIncidentPass
Three asks (regional log comparison, deploy history, rollback) and the root cause is found in 11 minutes 42 seconds. The path itself is the score.
03 · THE FIX

They say the fix. The system responds.

“Roll back that deploy.” “Add backoff to the retry loop.” “Put a cache in front of search.” The teammate does it, and the system reacts the way a real one would: latency comes back down or it doesn't, the new cache shows up in the metrics, and a fix that misses the point changes nothing.

Every scenario has traps. A fix that only treats the symptom buys a few quiet minutes, then it breaks again. Nobody fails for forgetting syntax, and nobody passes by sounding confident. You find out who actually fixes the cause before they're on your on-call rotation.

04 · THE SCORE

No AI opinion in the pass/fail.

Every scenario has one real root cause, decided when we build it. Pass or fail is measured against that. AI grades some of the writing, but it's clearly labeled and it never decides the verdict.

PASS / FAIL

The pass bar

Passing means two things happened: they named the root cause, and their fix actually brought the system back, with latency under the limit and holding when time runs out. You can't argue with a recovered system.

CHECKPOINTS

Partial credit, defined up front

Isolated: they found where it's breaking, not yet why. Mitigated: they stopped the bleeding and said it wasn't the real fix. Root-caused: they named the exact defect. Getting close counts as close, not as zero.

THE PATH

The path they took

Everything they asked for, everything they did, and every dead end, timestamped from page to fix. This is where you actually read the investigation. It's a record of what happened, not a score of how they seemed.

POSTMORTEM

The postmortem

They finish with a short writeup: what broke, how they fixed it, and how they'd stop it happening again. AI grades the writing against a rubric and links every point back to the transcript. It adds color to the report. It never touches the verdict.

Everyone applying for the role runs the same scenario against the same bar, so you can compare candidates fairly, and if a decision ever gets challenged, the evidence is right there. The report shows your team what happened; you decide what matters.

05 · RUNNING IT

How it fits into your hiring loop.

It sits in front of your phone screen, and nothing else in your process changes. Your engineers only meet candidates who have already fixed something real.

  1. 01

    Import the role

    Paste the job description or pick your stack: Kafka, Redis, Postgres. The incident runs on systems that look like yours, generalist by default and stack-specific when you want to test for fluency.

  2. 02

    Send one link

    It runs in the browser. No setup, no scheduling. The candidate can take it tonight instead of playing calendar tetris with your interviewers.

  3. 03

    The incident runs

    Under an hour, timeboxed like real on-call. Everything gets captured as it happens: the conversation, what they asked for, what they did, and the postmortem.

  4. 04

    Read the report in three minutes

    Verdict on top, the path they took under it, full transcript attached. You decide who meets your engineers.

The screen costs your engineers nothing. A whiteboard round burns a senior hour to produce one opinion; here, that hour goes to candidates who already cleared a real bar.

The only prep that works is having done the job.

There's a whole industry for rehearsing interviews: question banks, mock interviews, the diagrams everyone memorizes. It exists because most interviews ask the same questions, so the answers can be studied. And all a rehearsed pass tells you is that someone is good at interviews.

None of that works here. The incident starts vague, the system only shows what they ask for, and no two candidates take the same path. There's no answer key to study, because the path is the answer: what they asked, what they did, and how they explained it after.

So the people who pass are the people who have actually been paged at 2 AM and worked out what broke. Grinding doesn't get you through it. You end up hiring people who know what they're doing.

A new interview for an
era of rehearsed answers.

Screen for real engineering judgment without spending an engineer-hour.

Want to try it yourself?

Free practice incidents for candidates at launch. No spam, one email when it opens.