Live coding interviews assume no AI. Your candidates use it daily.
Onsite is a sandbox where candidates work with an AI agent on real problems, and you evaluate how they direct it.
Book a demoOne real candidate, your loop, we’re on the call.
session · candidate-427 · 41:12LIVE
$ onsite start --task retry-pipeline
sandbox ready · agent attached
candidate pull the retry logic out of the worker, keep the backoff, test the interrupted case
agent proposed 3 files, 84 lines
candidate rejected worker.ts — “swallows the cancel signal”
candidate took over retry.ts manually · 21 lines
12 tests passed · 1 added by hand
→ interviewer sees every prompt, every rejection, every takeover
Leetcode-style screens now measure prompt-recall, not engineering.
Banning AI in interviews tests a job that no longer exists.
Allowing it with no structure gives you no signal either way.
How it works
01
Candidate gets a realistic task in a sandboxed environment with an AI agent.
02
They work the way they would on the job: directing, reviewing, correcting the agent.
03
Interviewer sees what they asked for, what they rejected, and when they took over.
Interviewer view · prompts, rejections, takeoversYou see judgment, taste, delegation, and code review under real conditions, instead of whiteboard recall.
JudgmentWhich problems they hand to the agent, and which they keep.
TasteWhat they accept as good enough, and what they send back.
DelegationHow precisely they scope work before the agent starts.
Code reviewWhat they catch in generated code before it ships.
Inside a session
Real agents in the terminal
Claude Code and Codex run inside the sandbox — candidates use the same agents your engineers do, on Onsite's keys.A full workspace, not a toy editor
Editor, file tree, terminals, a one-click test runner, and an AI pair — the sandbox works like the candidate's own machine.You watch it live
The observer view streams the candidate's screen in real time — read-only, with the option to pause the interview.What you stop doing
Zip filesCandidates open a link. The task, dependencies, and tests are provisioned in an isolated sandbox — nothing for them to install.
Expensed API keysCandidates never bring an API key or pay for the interview. Run on Onsite's key, or connect your org's.
Lossy screen sharesWatch the session live. Afterward, every prompt, file edit, terminal command, and test run replays in order from the transcript.
Off-the-shelf tasksUpload a repo — TypeScript, JavaScript, or Python — and it becomes your task, smoke-tested in a real sandbox before a candidate ever sees it.
What teams ask first
Isn't this just letting them cheat?The agent is part of the task. You are not grading whether they used it, you are grading how they used it: what they asked for, what they rejected, and when they took over.
How is it scored?Every prompt, file edit, terminal command, and test run is captured. Onsite generates a report that scores four competencies — prompting, validation, workflow, and communication — each backed by links to exact moments in the transcript. The interviewer reads it, checks the moments, and records their own decision.
Does it fit our existing loop?It replaces the technical screen or the onsite coding round. In a pilot we run it inside your loop, with your task, and you keep your own rubric.
What do candidates consent to?Everything is upfront: candidates see what’s captured before they start — prompts, file edits, terminal, and a live view of the interview tab, which is the only thing shared. Tab sharing needs a desktop Chromium browser like Chrome or Edge. Details are in the privacy policy.
DEMO
See a full session end to end, then decide if it fits your loop.
Book a demoInterviews that show how an engineer works with an agent, not how well they remember prompts.
