A coding agent can waste a surprising amount of effort before it makes a useful change: broad searches, repeated reads and several nearby hypotheses before it commits to one. VS Code tested whether changing the agent's own instructions could reduce that exploration without making the result materially worse. The useful lesson is not a magic prompt. It is that agent instructions can be treated as something you test rather than something you write once and trust.
What VS Code changed
In a two-week experiment with GPT-5.5 agent traffic, the VS Code team compared its existing system prompt with two alternatives. Both alternatives pushed the agent towards the same basic behaviour: start from a concrete anchor, gather only enough local evidence to form a plausible hypothesis, make a grounded edit sooner and validate it rather than continuing to explore.
The smaller treatment added a compact reminder to search economically. The larger treatment made the edit-and-validate loop more explicit: identify a local hypothesis, choose a cheap check that could disprove it, make a small edit and validate immediately afterwards. The treatments were compared with an evenly sized control group on live traffic.
That distinction matters. The experiment was not asking whether a longer prompt is always better, or whether agents should stop investigating unfamiliar code. It was testing a specific failure mode: unnecessary exploration after there is already enough evidence to take a discriminating action.
The larger instruction changed measurable behaviour
VS Code reported that its larger prompt treatment produced the strongest overall result. Compared with the control, median time to first edit was 5.68% lower, or 3.9 seconds faster, and the 95th-percentile time to first edit was 9.30% lower, or 38.8 seconds faster. The 95th-percentile token total per turn fell by 7.64%, and average tool calls per turn fell by 8.54%.
The quality picture was less one-sided. Commit survival moved up by 0.68%, but that change was not statistically significant. Ten-minute survival moved down by 0.44%, and that small decline was statistically significant at p=0.0493. VS Code therefore treated the quality movement as a trade-off to watch rather than claiming that the efficiency gains were free.
The team shipped the larger treatment as the default GPT-5.5 system prompt in its harness. That is evidence that this particular instruction change helped in that particular model, harness and traffic mix. It is not evidence that copying the same wording into every coding agent will produce the same result.
The part worth copying is the experiment
If your agent repeatedly searches too broadly, rereads unchanged files or delays a small testable edit, write an instruction aimed at that observable behaviour. Keep it narrow enough that you can tell whether it changed anything.
A useful instruction can express a sequence rather than a personality:
- start from the most concrete file, symbol, failing behaviour or command you already have;
- read enough nearby context to form one plausible local hypothesis;
- choose a cheap check that could disprove that hypothesis;
- make the smallest grounded change that lets the check tell you something;
- validate after the edit instead of continuing to search by default;
- do not reread unchanged context unless new evidence makes it relevant.
Those points paraphrase the behaviour VS Code tested. They are more useful than an instruction such as be efficient because each one describes something you can observe in an agent run.
Decide what you will measure before changing the instruction
The VS Code experiment used several measures at once because efficiency on its own is not enough. Fewer tokens are not an improvement if the resulting code is rejected or immediately rewritten.
For your own workflow, choose measures that match the problem you are trying to fix. Depending on what your tools expose, that could include:
- time until the first useful edit;
- number of searches, file reads or other tool calls before that edit;
- tokens used on comparable tasks;
- whether the requested behaviour passes its focused check;
- whether the resulting change survives your review and is kept.
You do not need VS Code's production-scale A/B infrastructure to use the same principle. What matters is making the comparison less subjective. Run a small set of comparable tasks with the current instructions, change one instruction, then run comparable tasks again. If you change the model, task, tools and instructions at the same time, you will not know which change caused the result.
Use instructions to set a decision rule, not to suppress necessary investigation
The strongest part of the VS Code treatment was not simply search less. It gave the agent a point at which searching should stop: once it had a concrete anchor, a falsifiable local hypothesis, a nearby code path and a cheap discriminating check, the next useful action was an edit or probe followed by validation.
That is a better control mechanism than imposing an arbitrary number of file reads. Some tasks genuinely require more investigation. The instruction should help the agent recognise when more exploration is still buying information and when it has become drift.
Keep the quality guardrail
The experiment also gives a reason not to optimise only for speed or token count. The larger treatment improved the strongest efficiency measures while one short-term code-survival measure moved slightly in the wrong direction. If your own instruction makes an agent faster but increases rework, failed checks or rejected changes, the cheaper run is not necessarily the better run.
This is the same principle as testing any other coding-agent efficiency claim: define the behaviour you want to change, measure the current result, alter one controllable part of the workflow and keep a quality measure beside the efficiency measure. The prompt is part of the harness. Treat it with the same suspicion you would treat any other optimisation.
For a broader example of testing claimed agent optimisations rather than accepting the headline number, see Three Coding-Agent Efficiency Claims Were Tested. If you need a small repeatable task to detect behavioural changes over time, Use a Tiny, Stable Eval to Catch Coding-Agent Regressions applies the same idea from another direction.