A large benchmark is not the only way to evaluate a coding agent. VS Code's eval team found that repeatedly running one deliberately tiny task exposed changes in reliability, tool use and wasted effort that would be hard to see in an occasional real-world trial. The technique is useful because the task stays still while the model or harness changes around it.
Make the task boring on purpose
The VS Code eval described in the source gives the agent a minimal instruction in an empty workspace: create a file containing a fixed string. The result has an unambiguous pass condition, and the most direct solution is obvious.
That simplicity removes variables. If the prompt, starting state, tools and expected result are fixed, a change in behaviour is more likely to come from the model, agent harness or surrounding infrastructure rather than from a different task.
VS Code says this smoke test accumulated 50,974 runs across 30 models over six months. The important finding was not that models could complete the task. Most could. The useful signal was how they completed it: whether they acted directly, planned unnecessarily, explored an empty workspace, chose a more complicated tool or produced far more output than the task required.
What to record
A pass/fail result is necessary, but it is not enough if you want to detect why an agent has changed. For each run, record at least:
- whether the expected result was produced;
- latency;
- output-token usage or another consistent cost measure;
- the sequence of tools the agent called;
- the failure mode when the task did not complete correctly.
The tool sequence is particularly useful. Four tool calls might be reasonable or wasteful depending on what they were. A trace showing planning, directory listing, search and file creation tells you much more than the number four.
Use the same task before and after a change
The point of a tiny eval is comparison. Run the same task before and after something that could alter agent behaviour: changing model, updating the agent harness, modifying tool descriptions, changing system instructions or altering infrastructure.
Then compare the traces. Did the pass rate move? Did latency increase? Did the agent start making extra tool calls? Did a new failure mode appear? Did output grow substantially even though the requested work did not?
This gives you a cheap regression detector. It does not tell you whether a model is generally better at software development, but it can tell you that something about your setup has changed.
Choose a task from your own workflow
You do not need to copy VS Code's exact file-writing task. Pick the smallest task in your environment that has one clear correct result and can be reset reliably.
For example, you might ask an agent to:
- change one known configuration value and verify it;
- add one deterministic test fixture;
- make a tiny edit to a disposable sample project;
- run one known command and report a fixed property of the result.
Keep the prompt and starting state unchanged between runs. If the task itself keeps moving, you lose the ability to tell whether the agent changed or the problem changed.
Run it often enough to become a baseline
VS Code's team uses its tiny task as a smoke test around wider evaluation work. That is a sensible pattern for ordinary development too. Run the small eval before model onboarding, after harness changes or before a larger test suite. Over time, the accumulated results become a behavioural baseline.
A single run is weak evidence because model behaviour can vary. Repetition is what makes the test useful. You are looking for a pattern, not a dramatic conclusion from one success or failure.
Do not optimise your whole workflow for the tiny test
The source is explicit about the limitation: this is one short-horizon task with one correct answer. Planning and exploration that are wasteful here may be valuable on a long, ambiguous change. VS Code says it continues to use broader benchmark tasks rather than optimising its harness around this one smoke test.
That distinction matters. A tiny eval is good at answering narrow questions such as:
- Is the end-to-end agent path still working?
- Has a model or harness change added obvious overhead?
- Has a new failure mode appeared?
- Does the agent still recognise when a simple task needs a simple solution?
It is not enough to answer whether the agent can safely implement a complex feature, understand your architecture or make good product decisions.
Small, stable and repeatable beats impressive
A useful eval does not need to look like real production work. For regression detection, stability can be more valuable than realism. Start with one task whose correct result is undeniable, log enough detail to explain each run and repeat it whenever the system changes.
You still need broader tests for real coding ability. But a tiny fixed task gives you something those larger tests often do not: a sensitive early warning that the model or harness is behaving differently before you trust it with something harder.