4 August 2026 · 7 min
Your interview loop is testing a skill that no longer exists
Take-homes are dead and live coding is contaminated. Most teams have quietly stopped trusting their own process. This is how I’d rebuild it.
I was the Bar Raiser at Marshmallow (the person the founders made responsible for the hiring standard) while engineering went from under ten people to over fifty. I ran more than fifty interviews in six months at the peak of that, and then rebuilt a hiring process from scratch again at Vita Mojo. I mention it so you know this isn’t a think-piece. It’s what I’d do on Monday.
Most interview loops were designed to answer the question “can this person produce working code?” That question has become close to worthless, because the answer for nearly every candidate is now yes.
Every signal you had has been degraded
- The take-home is dead. It measures whether someone can prompt an assistant and tidy the output. You’re not learning nothing, but you’re not learning what you think.
- Live coding with AI banned tests a skill the person is unlikely to use again. It also tells the candidate you haven’t thought about the job you’re hiring for, and good engineers notice.
- Live coding with AI allowed is better. But if you haven’t decided what you’re watching for, you’ll sit there for an hour and come away with an impression rather than evidence.
- Algorithm puzzles were always a weak proxy. Now they measure very little.
Most teams know this. A common response is to keep the loop and stop believing its output. That’s the worst of both worlds: you still spend the hours, then discount the result and hire on gut feel.
Stop testing production. Test judgement.
The job is now less about typing the code and more about deciding whether code is right, and being accountable when it isn’t. So the loop should put candidates in the position of deciding, and make them defend the decision.
The question is no longer “can you write this?” It’s “here’s something that was written. Would you merge it, and what happens if you’re wrong?”
The loop I’d build
There are four stages. The second is the change that matters most. If you only take one thing from this piece, take that.
1. Screen, largely unchanged. Thirty minutes on what they’ve built and what they own. The one addition I’d make is a direct question about how AI has changed their working week. I’d expect the answers to separate people who have thought about it from people performing enthusiasm.
2. A code review exercise. I think this is now the most useful hour in the loop. Hand them a substantial diff that was genuinely AI-generated, in your stack, solving a plausible ticket. Plant three or four problems of different kinds: a subtle correctness bug, a security issue, an over-engineered abstraction that will cost you in a year, and something that’s wrong for your context in a way only a careful reader would catch. Ask what they’d merge, what they’d change, what they’d reject outright, and why.
This is the stage I’d expect to separate candidates most clearly. A weak reviewer comments on formatting and variable names. A strong one finds the correctness bug. A very strong one tells you what they’d need to know about the system before they’d approve it at all. It fits in an hour and it’s fair to prepare for. It’s also close to what the job now is.
3. Bring-your-own-tools pairing. A real ticket, in something close to a real codebase, with whatever assistant they normally use. The solution matters less than how they get there. Watch how they direct the tool, whether they check what comes back, whether they notice when it’s confidently wrong, and whether they can explain code they accepted ninety seconds ago. Incidentally, provide the tooling and a licence. Making candidates use their own paid subscriptions selects for people who can afford them.
4. Debugging, not greenfield. Give them something broken in a system they didn’t write. Greenfield problems flatter the tools. Debugging shows whether someone can build a mental model of code they’re meeting for the first time, which is now a large part of senior engineering.
You need someone who owns the bar
None of this survives a hiring push unless a specific person owns the standard, and it can’t be the hiring manager. A manager who is three months behind on headcount isn’t a neutral judge of whether a candidate clears the bar. That’s about incentives rather than character.
The Amazon Bar Raiser model is a cheap fix. A trained person from outside the hiring team sits in the loop with the power to block a hire, and their only job there is to protect the standard. At Marshmallow it let us scale fast without scaling badly. Engineering grew more than fivefold in under two years and the bar held, because someone was accountable for it.
It works at ten people too. You don’t need a programme, just one person with a veto and the standing to use it.
If you don’t have a senior engineer yet, the problem is harder. Nobody in the room can judge the code review exercise, so for the first few hires the bar has to be borrowed from outside. Write it down as you go, so the first senior engineer you hire inherits a standard rather than inventing one under pressure.
Write the bar down
Most teams can’t tell you whether a candidate met their bar, because nobody has written it down. What does a mid-level engineer do that a junior doesn’t? What does someone have to show to be hired as senior here? Not at Google, but here, in this codebase, with this team.
Until that’s on a page, every offer is a judgement call made by tired people under time pressure, and your bar is only as high as the last person you were desperate to hire. The fix is scorecards, a written levelling framework, and debriefs where evidence counts for more than impressions. None of it is glamorous, but it’s most of the work.
What to stop doing on Monday
- Take-homes used as a filter. Either drop them or make them AI-allowed and interrogate the result in person.
- “No AI” rules. They can’t be enforced, and they tell your best candidates you haven’t understood the job.
- Algorithm puzzles.
- Any stage where nobody can say in advance what a pass looks like.
I don’t think the teams that get this right over the next two years will be the ones paying the most. I think they’ll be the ones who worked out what they’re actually assessing, while everyone else was still arguing about whether candidates should be allowed to use the tools they’ll use every day in the job.
If you’re investing in a company and want to know what’s really been built, or you’re a founder hiring engineers without a senior one to judge them, this is the work I do.