Should you allow Claude Code in technical interviews?
Este artículo aún no está traducido, así que se muestra en inglés.
Allow AI agents when the job uses them, set the task and token budget to match, tell candidates upfront, and review the transcript alongside the code.
Allow Claude Code or Codex in a technical interview when the job involves working with coding agents, and leave them off when you want to see someone reason through code alone. If you allow them, design the task so an agent can't finish it in one prompt, cap the tokens, tell candidates exactly what is allowed and recorded, and review the agent transcript alongside the code.
Why the question matters now
Many engineering teams now use agents every day. Banning them in interviews tests a way of working the candidate won't use on the job. Allowing them without thought gives you a test the agent passes, not the person. Neither default is right for every role, so make the call per interview.
How do you decide whether to allow AI agents?
Ask what the job looks like six months in.
| Allow agents when | Leave them off when |
|---|---|
| The team uses Claude Code, Codex or similar daily | The role is early-career and you need to see fundamentals |
| The task is a realistic change in a real codebase | The task is small enough that an agent solves it outright |
| You care about review, testing and judgment | You're specifically assessing debugging by hand |
| You want to see how they delegate and verify | You can't explain to candidates how it will be judged |
Whichever you choose, apply it to every candidate for that role. Mixing rules within one hiring round makes the results impossible to compare.
Hiding the rule doesn't work
"No AI" in a take-home is hard to enforce: a candidate can run an agent on another machine. Pretending otherwise rewards the people who ignore the rule. If you need a no-AI signal, get it in a live session where you can see the work happen.
What signal do you get with agents on?
The code alone tells you less, because an agent may have written most of it. The useful signal moves to how the candidate worked:
- Framing. Did they explain the problem to the agent clearly, with the right context and constraints?
- Reading. Did they read what the agent produced, or accept it whole?
- Verifying. Did they run the tests, add new ones, and check edge cases the agent missed?
- Correcting. When the agent went the wrong way, did they notice and steer it back?
- Ownership. Can they explain every line in the final diff?
A candidate who prompts once and submits shows you very little. One who breaks the task down, pushes back on a weak suggestion and adds a test the agent forgot is showing you how they'll work on the team.
Design the task for agents
An agent-on task needs room for judgment. Good shapes:
- A bug in a medium codebase where the cause is two files away from the symptom.
- A feature with a trade-off, such as consistency against speed, where the candidate has to choose and explain.
- A change with existing tests that must keep passing, so careless edits show up.
Avoid textbook algorithms. Agents have seen them all.
Set a token budget
A budget protects your bill and keeps the playing field level: one candidate shouldn't get ten times the agent time of another. It also nudges candidates to think before they prompt.
Size it to the task. Run the task yourself with the agent and note what you spent, then allow two to three times that. Tell candidates the budget exists and show them how much is left, so running out is never a surprise.
Be upfront with candidates
Before they start, candidates should know:
- Which agents are available and how to start them.
- That their prompts, the agent's replies and the commands it runs are recorded and shown to reviewers.
- That their code is sent to the AI provider under your organization's account.
- Whether there's a token budget.
- How the work will be judged, for example "we care about how you verify the agent's output".
Saying this out loud removes the guessing game and gets you a more honest picture of how they work.
Review the transcript, not just the diff
With agents on, the transcript is part of the work. Read it like a pairing session:
- Skim the candidate's prompts first. They tell you how the person thought about the problem.
- Look for the moments they rejected or corrected the agent.
- Check whether the tests in the diff were written by the candidate, the agent, or both, and whether they were run.
- Don't mark someone down for using the agent a lot. Mark them down for not checking its work.
More on this in How to review take-home assignments fairly.
Doing this in Kendor
In Kendor, candidates run claude (Claude Code) or codex (Codex) in the sandbox terminal, against the real project. They run on your organization's own Anthropic or OpenAI account. See Turn on Claude Code and Codex.
Connect your keys
An organization owner opens Integrations and connects Claude Code and Codex. Claude Code takes an Anthropic API key, or Identity federation with your Claude Console so no key is stored. Codex takes an OpenAI API key. Keys are encrypted at rest, and Kendor only shows their last four characters. Only owners can change them. Details in Organization keys.
Keys never enter the sandbox
The sandbox holds only a session token, not a vendor key. Every call from the agent goes through Kendor's gateway, which checks the session and the budget, adds your organization's key, and forwards the call to Anthropic or OpenAI. A candidate can't read or copy your key from the terminal.
Turn agents on for a screen
In the screen's settings, turn on AI assistants in the terminal. It's off by default. Then set a Token budget per candidate in million tokens: what the assistant reads and writes, added up across all challenges, with cached context not counted. Leave it empty for No limit.
For the whole organization, the Monthly token cap in the integration's settings limits what all candidates together may spend in a calendar month. See Token budgets.
What candidates see
Before they choose Begin screen, candidates are told that AI assistants are available, which command to run, that their workspace files and prompts go to your AI provider under your account, and that the conversation is saved and shown to reviewers. While they work, they see their token use against the budget, and a warning when they're close. Once the budget is used up, the agent stops working for that attempt.
Live coding interviews
For a live coding interview, the setting AI assistants in the terminal sets the default, and Claude Code and Codex in the terminal in the invite dialog changes it for one session. In the room, an interviewer can flip AI on / AI off for everyone. Live sessions have no per-candidate budget; the organization's monthly cap still applies.
Review the transcript
On the review page, the AI agents tab shows each conversation turn by turn: Prompt, Agent, Thinking, each tool call and its Tool output. The Terminal tab replays what the candidate and the agent ran, and the header shows the tokens the attempt used against its budget. While an attempt is still in progress, reviewers can watch the terminal live.
Note. Kendor doesn't record the candidate's screen or webcam. What reviewers see is the terminal, the agent conversation, the code and the activity in the editor.