Set token budgets for AI agents
Questo articolo non è ancora tradotto, quindi è mostrato in inglese.
Cap what Claude Code and Codex can spend with a token budget per candidate on each screen and a monthly token cap for the whole organization.
There are two limits, and both are optional. A Token budget per candidate on a screen caps what one candidate's agent can spend on that screen. A Monthly token cap under Integrations caps what all candidates together can spend in a calendar month. When either runs out, the agent's next call is refused with a clear message in the terminal. With neither set, spend is bounded only by your vendor account.
What counts as a token
Both limits count what the agent reads and writes: new input sent to the model plus the model's output. Cached context is not counted. Coding agents send the whole conversation again on every turn, and that repeated part is served from the vendor's cache, so it would otherwise eat a budget without reflecting any new work.
Claude Code and Codex draw from the same budget. A candidate who uses both spends from one total.
Budget per candidate on a screen
Editors and owners set this in the screen builder, next to the switch that turns the agents on.
- Turn on AI assistants in the terminal under Options in the settings rail.
- Enter a number in Token budget per candidate. The unit is million tokens, so
1means one million and0.5means half a million. - Leave the field empty for no limit. The rail then reminds you: "No limit set. A candidate's spend is bounded only by your vendor account."
The budget is per candidate attempt and is added up across all the challenges in the screen, so a three-challenge screen with a budget of 2 gives the candidate 2 million tokens in total, not 2 million per challenge. Usage is kept when the sandbox restarts or the candidate comes back later, so restarting doesn't reset it.
You can change the budget while candidates are mid-screen. Running sandboxes pick up the new limit straight away.
What the candidate sees
A small ring in the editor's header fills as the agent works. Hovering it shows, for example, "AI tokens: 400k of 1M used". It turns amber at 80% with "You are close to the AI token budget for this screen", so running out is never the first sign. Once it is used up, the tooltip says "The AI token budget for this screen is used up".
Monthly cap for the organization
Owners set this on the Claude Code or Codex page under Integrations → AI assistants. Both pages show the same field, because the two agents share one cap.
- Enter a number in Monthly token cap, in million tokens.
- Choose Save.
- Leave it empty for No cap.
Under the field you see this month's spend, such as "6.4M of 20M tokens used this month". The month runs on UTC calendar months and starts again at zero on the 1st.
The cap covers every agent call billed to your keys: candidates on screens, live coding interviews, and your own team's sandboxes in the builder.
Tip. If you use identity federation for Claude Code, you can also put a spend limit on the Anthropic workspace in your Claude Console. That limit is enforced by Anthropic, on top of Kendor's caps.
Live coding interviews
Live coding interviews have no per-candidate budget. Agent use in a live session counts only toward your organization's Monthly token cap, so set one if you want a ceiling there.
When a budget runs out
Kendor's gateway checks the budget before forwarding each call, including calls that are already running in parallel. Once a limit is reached, the next call isn't sent to Anthropic or OpenAI, and the agent shows one of these messages in the terminal:
| Limit reached | Message |
|---|---|
| Token budget per candidate | "The AI token budget for this assessment has been used up" |
| Monthly token cap | "The hiring organization's monthly AI token limit has been reached" |
The turn that was already in progress finishes, so the last count can land slightly above the limit. The candidate can keep working without the agent: the editor, terminal, tests and preview are unaffected.
To let a candidate continue with the agent, raise the screen's budget or the monthly cap. Running sandboxes pick up the change without a restart.
Note. Separately from tokens, each sandbox session has a fixed limit on the number of calls the agents can make. When it is reached, the agent shows "This session has reached its AI usage limit".
Where usage shows in review
On a submission's review page, the header shows what the candidate's agent used, such as "AI 1.4M / 2M tokens" with a budget or "AI 1.4M tokens" without one. Hovering it explains that the figure is what the assistant read and wrote during the attempt, with cached context not counted.
Your organization's running total for the month is on the Claude Code and Codex pages under Integrations. For a cost breakdown by model, use your Anthropic or OpenAI console, where the usage is billed.