AI agents in the terminal
Claude Code and Codex in the sandbox. On your keys, within your budget.
If the job involves working with an agent, screen for it. Candidates run Claude Code or Codex in their sandbox terminal, billed to your organization's own Anthropic or OpenAI key. The key never enters the sandbox, and you can cap each attempt with a token budget.
01
Your key never enters the sandbox.
Claude Code and Codex in the sandbox talk to Kendor's gateway with a session token that opens nothing on its own. The gateway checks the session and its budget, adds your organization's key and forwards the call to Anthropic or OpenAI. There's nothing in the environment for a candidate to print or copy.
- Keys are encrypted at rest, and only organization owners can change them
- Anthropic can connect through identity federation, with no long-lived key at all
- The gateway only reaches Anthropic and OpenAI. It isn't a way out to the internet
$ env | grep -E 'ANTHROPIC|OPENAI'ANTHROPIC_BASE_URL=http://kendor-gateway:4102/anthropicANTHROPIC_AUTH_TOKEN=eyJhbGciOiJIUzI1NiIs…OPENAI_BASE_URL=http://kendor-gateway:4102/openai/v1OPENAI_API_KEY=eyJhbGciOiJIUzI1NiIs…a session token, not a vendor key02
A budget per candidate, a cap per month.
Give each screen a token budget per candidate, counted across all its challenges, and set a monthly cap for the whole organization. Candidates see a ring in their header fill toward the budget, and the gateway stops the agent as soon as either limit is reached.
- Counts what the agent reads and writes, not cached context
- Usage is billed to your own Anthropic or OpenAI account
03
Off until you turn it on.
AI assistants stay off for every screen until you switch them on in the builder. Candidates get the agents your organization has keys for, and their start page tells them before they begin what goes to your provider and that reviewers can see their terminal.
04
Watch the terminal live, read-only.
Open the Terminal tab on a candidate's review page while they work and you see their shell as it happens: what they ask the agent, what it runs and what they keep. Reviewers can watch, but they can't type into it.
05
Read every prompt afterwards.
The AI agents tab keeps the conversation in order: each prompt, the tools the agent called and what it answered, with the time and model on every turn. You judge how someone works with an agent, not only the code it left behind.
- Candidates are told up front that the conversation is saved for reviewers
- PromptInvoices with a zero total are still being emailed. Find out why and fix it with a test.Tool · Read
{"file_path":"/workspace/internal/invoice/send.go"}▸ Tool output - AgentThe guard checks inv.Total < 0 instead of inv.Total <= 0, so zero-total invoices get through.Tool · Edit
{"file_path":"/workspace/internal/invoice/send.go","old_string":"if inv.Total < 0 {","new_string":"if inv.Total <= 0 {"}▸ Tool output
06
Switch agents on mid-interview.
In a live interview or pair coding session, the host turns AI agents on or off for everyone from the room's top bar. The running sandbox picks up the change straight away, without a restart.
- Everyone in the room sees whether AI is on
- Same organization keys and gateway as your screens
How it works
Two settings, then it's in their terminal.
Add your provider key
An Anthropic or OpenAI key in Settings → AI assistants, once for the organization. Set a monthly cap there if you want one.
Turn it on for a screen
Switch on AI assistants in the screen builder and give it a token budget per candidate.
Candidates run it
They find claude or codex ready in their sandbox terminal, and you can watch from the review page.
Questions
Who pays for model usage?
Your organization, directly to Anthropic or OpenAI. Kendor doesn't resell tokens or add a margin.
What does connecting without a key mean?
For Anthropic, your Claude Console organization can trust Kendor through workload identity federation, so no long-lived API key is created or shared. OpenAI uses an API key.
Doesn't allowing AI make the score meaningless?
The grade is a verified re-run of what was submitted, so it reflects code that actually works. How the candidate got there — what they asked, what they accepted, what they rewrote — is the part your reviewers judge.
More of Kendor
All features- Coding screensAsync take-homes in a real browser IDE, graded on a clean re-run.Read more
- Live pair codingOne link, one shared sandbox, a verdict from every interviewer.Read more
- InterviewsSelf-booking, invites that update in place, live review.Read more
- Real environmentsSeven pinned languages, PostgreSQL and Valkey, one kendor.yaml.Read more
- Repo importCloud or self-hosted Git, private repos too, one fixed snapshot.Read more
- EnterpriseAshby integrationAshby stages send the screen; results come back as fields.Read more
Run your next screen on Kendor. Read what the code says.
Kendor works with engineering teams in the EU, the UK and the US. Tell us what you're hiring for and we'll set up your organization with you.