<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[The Levelbrook Playbook]]></title><description><![CDATA[The Levelbrook Playbook]]></description><link>https://levelbrook.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>The Levelbrook Playbook</title><link>https://levelbrook.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Fri, 25 Sep 2026 05:21:42 GMT</lastBuildDate><atom:link href="https://levelbrook.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Everybody has lost their minds, and the boring companies are quietly winning]]></title><description><![CDATA[A senior security engineer wrote this week that he spends three quarters of his time on AI and it has robbed him of his enjoyment of the work. He is right about the symptom and wrong about the cause. ]]></description><link>https://levelbrook.hashnode.dev/everybody-has-lost-their-minds-and-the-boring-companies-are-quietly-winning</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/everybody-has-lost-their-minds-and-the-boring-companies-are-quietly-winning</guid><category><![CDATA[startup]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:50:12 GMT</pubDate><content:encoded><![CDATA[<p><em>A senior security engineer wrote this week that he spends three quarters of his time on AI and it has robbed him of his enjoyment of the work. He is right about the symptom and wrong about the cause. The value from these tools is landing in the least glamorous places in the company, and the people looking at the frontier are looking the wrong way.</em></p>
<h2>The vent, and the sentence buried in it</h2>
<p>Jan Schaumann's post this week is a vent and he says so in the first line. A senior security
engineer, decades in, writing that he spends upwards of three quarters of his time directly or
indirectly dealing with AI and that it has robbed him of most of his enjoyment of the work. People
with no engineering background pitching industry-changing solutions from their agent-infested
homelab. Emails that read like influencer posts. Colleagues turned into meat proxies. It hit the
front page because a great many people feel exactly this and most of them are not allowed to say
it at work.</p>
<p>You can disagree with a lot of it, and the thread did. But there is one paragraph in the middle
that is not a vent. It is an argument, and it is the most important thing anyone wrote about AI
this week.</p>
<p>He describes the industry-wide effort, now many months old, to point frontier models at
vulnerability discovery. Dozens of highly paid engineers per organisation, priorities reshuffled,
harnesses built, pipelines built to shoehorn thousands of findings into vulnerability management.
Thousands of new vulnerabilities found. And then: I don't think we're any safer than before. Because
finding vulnerabilities has never been the bottleneck in information security. The bottleneck is
still, as ever before, getting the packages updated. Patching is still hard.</p>
<p>That is the whole essay. The models are extraordinary at the part that was never the constraint.</p>
<h2>Meanwhile, at the frontier</h2>
<p>Consider what the rest of the industry was looking at while he wrote that.</p>
<p>A researcher quit one of the labs and posted a warning that went, by ThePrimeagen's count on
stream, to well over a hundred million views in a day. A colleague who stayed put the odds of
catastrophe above ten percent in a decade. The lab's CEO published a three-point plan to pace the
frontier. A rival CEO agreed with him, which people found unusual. AI Explained spent a video on the
researchers' stated reason: a large gap, in one OpenAI researcher's phrase, between the internal
and external perception of the rate of progress. Six axes of improvement, none near saturation.
Fireship covered a 154-page threat report cataloguing eight months of misuse across seven
categories.</p>
<p>All of this is real and none of it is unimportant. But look at who is in the audience. Engineers,
managers, founders, people who have to decide on Monday what to do with a budget. And the frontier
discourse gives them precisely nothing to do on Monday, because it is about capabilities they do not
control, on timelines they cannot influence, at companies they do not work for. It produces
anxiety with no action attached, which is the most tiring kind, and it is why Schaumann is
exhausted.</p>
<p><img src="https://ai.levelbrook.com/playbook/everybody-lost-their-minds-and-the-boring-companies-are-winning/fig-1.png" alt="Where the attention went this week, and where a Monday decision can actually land. Nothing in the top row is actionable by anyone outside a lab." />
<em>Where the attention went this week, and where a Monday decision can actually land. Nothing in the top row is actionable by anyone outside a lab.</em></p>
<h2>Where the value is landing</h2>
<p>We install AI systems in ordinary companies for a living, which gives us a view of this that the
frontier discourse does not have, and the view is this. The money is being made in the bottom row of
that figure, by people who are not on Hacker News, doing things nobody will write a threat report
about.</p>
<p>The shapes are always the same, and we describe them here as composites rather than clients, as
everything on this site is. A clinic group that uses a model to reconcile years of insurance
remittances against deposits and finds the pattern of underpayments a human was too busy to see. A
logistics firm whose intake now reads every inbound document and writes the structured record, so
the person who used to do that handles only the exceptions. A software team whose agent does not
write features but does read every production error, correlate it to a deploy, and open a ticket
with the likely cause before a human has seen the alert. And, in public this same week, Cloudflare
saving another hundred terabytes of RAM with what their post cheerfully calls maths.</p>
<p>None of that is at the frontier. All of it runs on models a generation or two behind the ones in
the headlines, because the constraint was never the model. Schaumann's line about vulnerability
discovery generalises: in almost every function of almost every company, the thing the model is
best at was not the bottleneck, and the bottleneck is some unglamorous downstream step that nobody
funded because it was boring. Patching. Inventory. The written specification. The approval seat.
The person who has to say yes.</p>
<p>The companies winning this year are the ones that noticed the bottleneck first and pointed the
tools at it, or more often pointed the tools at everything upstream of it and then spent the
savings on the bottleneck. They are not talking about it, partly because it is unglamorous and
partly because it is working.</p>
<h2>Why the frontier framing hurts</h2>
<p>There is a specific mechanism by which the frontier discourse makes ordinary organisations worse
at this, and it is worth naming because it is avoidable.</p>
<p>The frontier framing says the model is the variable. Wait for the next one; it will do what this
one could not. That framing is true for the lab and false for the buyer, and a buyer who believes it
does two bad things. They delay the boring work, because the boring work will surely be automated
by the next release. And they evaluate every tool on ceiling, on the impressive demo, rather than
on the floor, on what it does at two in the afternoon on a dull task with a tired reviewer.</p>
<p>The second thing the framing does is imply that the risk is exotic. Swarms, bioweapons, the
internet taken over. Meanwhile the actual risk in the actual company is the one CNN reported this
week from the military: a model produced a confident report with things in it that were not true,
and it got some distance up the chain before anyone checked. That failure does not require a
frontier model. It requires an unreviewed one. Every company installing these tools this year is
building that failure mode unless it builds the seat that catches it, and the seat is boring, and
the frontier discourse never mentions it.</p>
<p><img src="https://ai.levelbrook.com/playbook/everybody-lost-their-minds-and-the-boring-companies-are-winning/fig-2.png" alt="The frontier framing versus the operator framing. Same tools, opposite conclusions about where to spend Monday." />
<em>The frontier framing versus the operator framing. Same tools, opposite conclusions about where to spend Monday.</em></p>
<h2>How to find the boring bottleneck</h2>
<p>The instruction "find the least interesting problem and fix it first" is easy to nod at and hard to
act on, so here is the exercise we run in the first week with any organisation.</p>
<p>Pick one process that produces money or stops it. Invoices out, claims in, orders through, tickets
closed. Walk it end to end with the people who do it, and at each step write down two numbers: how
long the step takes when it goes well, and how long the work waits before that step starts. Almost
nobody has the second number, and the second number is the process. A step that takes four minutes
and waits two days is not a four-minute step.</p>
<p>Then find the step where the wait is longest, and ask why the work is waiting. The answer is nearly
always one of three things. A person has to decide something and is busy. A piece of information is
missing and somebody has to go and find it. Or two systems disagree and a human has to reconcile
them by hand. Those three are the bottleneck, and none of them is "the model is not smart enough".</p>
<p>Now point the tools. Missing information is the easiest: a model that reads the inbound document
and fills the record removes the wait entirely. Systems that disagree is the next: a model that
reconciles the ninety percent that match and hands the rest to a person, with the mismatch
highlighted, turns a day of reconciliation into an hour. The busy person deciding is the hardest and
the most valuable, and it is the approval seat: give them the decision in a form that takes eight
seconds and they will make forty before lunch.</p>
<p>Nothing about this requires the frontier. It requires a whiteboard, an afternoon, and a willingness
to be interested in a process everyone in the building considers beneath them. The companies
winning this year did that afternoon a year ago.</p>
<h2>Schaumann is also wrong, and it matters how</h2>
<p>The honest paragraph. He is wrong that the technology is the reason the work stopped being
enjoyable, and the thread said so in the most upvoted reply: the models genuinely produce a lot of
good work, can debug in an afternoon what took a team a week, and the people running the companies
are, in that commenter's phrase, the most boring supervillains imaginable. All true at once.</p>
<p>What robbed him of the enjoyment is the framing, not the tool. Seventy-five percent of a senior
engineer's time spent on AI is seventy-five percent spent in the top row of the figure, on
harnesses for vulnerability discovery that was never the bottleneck, on pipelines to process
findings that will not be patched, on the frontier's priorities rather than the organisation's. Point
the same engineer and the same models at the bottom row, at the inventory and the patching he
himself names as the real work, and the time is not wasted and the work is not joyless. It is the
job he signed up for, done faster.</p>
<p>The frontier will keep moving. The researchers may well be right to be frightened; we are not
qualified to say and neither is most of the audience. But the companies that come out of this decade
ahead will not be the ones that watched it most closely. They will be the ones that found the
boring bottleneck in their own building, wrote it down, and put the tools to work on either side of
it while everyone else was reading the threat report.</p>
<p>Everybody has lost their minds. The way back is to find the least interesting problem in the
company and fix it first.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://www.netmeister.org/blog/everybodys-lost-their-minds.html">Everybody's Lost Their Minds (Jan Schaumann)</a> (HN, 365 points, 17 Sep 2026)</li>
<li><a href="https://www.youtube.com/watch?v=J3ljHm57yU0">What AI Researchers Saw, Before Their Demand to 'Pace' AI (AI Explained, video)</a> (16 Sep 2026)</li>
<li><a href="https://www.youtube.com/watch?v=W8IVKMGbUZE">I can't believe this is happening (ThePrimeagen, video)</a> (19 Sep 2026; the viral-tweet numbers as reported on stream)</li>
<li><a href="https://www.youtube.com/watch?v=7r4ikZHm9AI">Anthropic researchers are quitting... (Fireship, video)</a> (15 Sep 2026)</li>
<li><a href="https://blog.cloudflare.com/saving-100-tb-of-ram-with-math/">Saving another 100TB of RAM (Cloudflare)</a> (HN, 463 points, 18 Sep 2026)</li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/everybody-lost-their-minds-and-the-boring-companies-are-winning/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[If a rumour can summon ten thousand agents, your roadmap is a starting gun]]></title><description><![CDATA[Two mathematicians spent a year on a problem, a rumour of their approach reached a lab, and a swarm of agents reportedly reproduced it in days. Terence Tao's warning about open science applies just as]]></description><link>https://levelbrook.hashnode.dev/if-a-rumour-can-summon-ten-thousand-agents-your-roadmap-is-a-starting-gun</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/if-a-rumour-can-summon-ten-thousand-agents-your-roadmap-is-a-starting-gun</guid><category><![CDATA[startup]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:47:08 GMT</pubDate><content:encoded><![CDATA[<p><em>Two mathematicians spent a year on a problem, a rumour of their approach reached a lab, and a swarm of agents reportedly reproduced it in days. Terence Tao's warning about open science applies just as well to your product. What was defensible last year is now a prompt.</em></p>
<h2>The story, as told, with the caveats attached</h2>
<p>Here is the account that circulated last week, and it is important to say up front that it is a
summary of public statements by parties who disagree with each other, so treat every clause as
"reportedly".</p>
<p>A mathematics professor and a collaborator who works at one of the labs spent about a year on the
Navier-Stokes existence problem, one of the Millennium Prize problems. In the last month of that
year they leaned heavily on coding agents and their progress accelerated; in mid-August they got a
simpler related system to break, the closest anyone had come. Then, according to their statement, a
different lab, having heard rumours of the approach, pointed a very large number of agents and a
very large amount of compute at the problem and, within days, claimed a result on the full problem
using what the mathematicians say was their novel approach. There was a phone call. The accounts of
the phone call diverge sharply. Both sides published, one after the other, on the same Tuesday.</p>
<p>You can read Fireship's telling, which is where most engineers heard it, and you can read the
statements. We are not adjudicating it. What we want to point at is the line Terence Tao is
reported to have written in response, because it is the most important sentence of the month for
anyone who runs a product team, and almost nobody outside mathematics noticed it.</p>
<p>If the mere rumour of your research can trigger a swarm of agents racing to front-run it,
mathematicians will simply stop sharing ideas, and that will undo centuries of open science.</p>
<h2>Replace "research" with "roadmap"</h2>
<p>Now do the substitution. If the mere rumour of your feature can trigger a swarm of agents racing to
front-run it, what happens to your roadmap?</p>
<p>For the entire history of the software business, the gap between having an idea and shipping it was
the moat. Not the idea itself; ideas were always cheap and always leaked. The moat was that turning
the idea into a working, deployed, maintained thing took a team a quarter, and a competitor who
heard about it on a Tuesday could not have it by Friday. Product strategy, fundraising, hiring
plans, launch timing, all of it was built on that gap being months wide.</p>
<p>The maths story is what it looks like when the gap closes to days at the top of the market. The
Dream RSI paper that Fireship covered the same week is what it looks like a level down: a search
loop that, per the video's summary, wrote a solver that beat a standard library in about 300
attempts where the fixed-policy version needed 550 and the previous record needed roughly 51,000.
That is a search process getting two orders of magnitude cheaper at finding a solution somebody
else already knew the shape of. Knowing the shape of the solution is what a roadmap leak gives a
competitor.</p>
<p><img src="https://ai.levelbrook.com/playbook/your-roadmap-is-a-starting-gun/fig-1.png" alt="The idea-to-shipped gap was the moat. Durations illustrative; the collapse is the point." />
<em>The idea-to-shipped gap was the moat. Durations illustrative; the collapse is the point.</em></p>
<p>This is not an argument for secrecy, for the same reason Tao's line is not an argument for
mathematicians to stop publishing. Secrecy does not work either: your customers know what they
asked for, your job postings say what you are building, and your own agents' instruction files
describe the product in more detail than any leaked slide. It is an argument for being honest about
what is still defensible when the build is no longer the hard part.</p>
<h2>What a swarm cannot front-run</h2>
<p>Go back to the maths. A swarm with a rumour and twenty million dollars of compute could,
reportedly, reproduce a result. What it could not do is the year before the rumour: choosing that
problem, choosing that approach out of the dozens that do not work, building the intuition that
made the approach look promising when it looked like nothing to everyone else. The swarm needed the
shape. The year produced the shape.</p>
<p>The commercial version of that year is the thing we wrote about earlier this month under the
heading that companies do not have processes, they have habits. The forty exceptions that live in
one person's head. The knowledge of which customer needs it done differently because of the freight
claim in 2023. The understanding of why the obvious version of the feature fails for the second
largest account. A swarm given the rumour builds the obvious version, perfectly, in days. The
obvious version is the one your customers already rejected.</p>
<p><img src="https://ai.levelbrook.com/playbook/your-roadmap-is-a-starting-gun/fig-2.png" alt="What a swarm reproduces from a rumour, and what it cannot. The right-hand side is the whole of the defensible business." />
<em>What a swarm reproduces from a rumour, and what it cannot. The right-hand side is the whole of the defensible business.</em></p>
<p>So the defensible assets in a world of swarms are the ones on the right of that figure, and they
are exactly the assets most product organisations have neglected because they were not the
bottleneck. The written specification, including the exceptions, which almost nobody has because
the code was the specification and the code was expensive. The proprietary data about how the
product fails and for whom. And the accountability, the fact that when the thing goes wrong there
is a named organisation that picks up the phone and fixes it, which no swarm has ever offered
anyone.</p>
<h2>What changes on the roadmap</h2>
<p>The practical consequences are uncomfortable for the way product teams have worked.</p>
<p>Announcements shift from features to outcomes. Announcing "we are building X" now hands the shape
to anyone with a swarm. Announcing "our customers in this segment no longer have this problem" hands
them nothing they can prompt with, because the interesting part is which segment and which problem
and why the obvious solution did not work, and that lives in your specification.</p>
<p>The specification becomes the product. Not a slide, the actual document: what the system does,
what it refuses to do, every exception and why. If your team's competitive advantage is knowledge
that lives in heads, the swarm era is the strongest incentive you will ever get to write it down,
because the written version is the only form in which it compounds and the only form in which it
can be defended.</p>
<p>Speed stops being a strategy and becomes table stakes. Being first to ship the obvious version was
worth a great deal when it took a quarter. It is now worth roughly a week of attention. The teams
that win will be the ones that were already on the third version, informed by the failures of the
first two, when the swarm shipped the first one.</p>
<p>And the Gowers point, from his essay on why he declined to sign the Fields medallists' letter,
applies in full. He has argued for twenty-five years that mathematics contains two cultures,
problem-solvers and theory-builders, and that the field needs both. The swarms are extraordinary
problem-solvers. Theory-building, meaning the slow accumulation of a coherent understanding of a
domain, is what makes the problems worth solving and tells you which one to solve next. Every
product organisation is about to discover which of the two it was actually good at.</p>
<h2>What to write down this quarter</h2>
<p>The uncomfortable implication of the figure above is that the defensible assets are documents, and
most organisations do not have them. So here is the writing programme, in priority order.</p>
<p>The exceptions register. For the product's core workflow, every case where the obvious behaviour is
wrong and why. Not the happy path, which anyone can reproduce, and not the code, which a swarm can
regenerate. The list of forty things that live in the head of the person who has been there
longest, with the reason attached to each. Ask that person to talk for two hours, record it, have a
model draft the register, and have the person correct it. This is the single most valuable document
the company can own and it usually does not exist.</p>
<p>The refusal list. What the product deliberately does not do, and what happened when someone tried.
A competitor building from a rumour will build the features you rejected, because they look like
features. Knowing why they were rejected is a year of learning that cannot be prompted for.</p>
<p>The failure data. Which customers hit which failures, how often, and what it cost. Nobody outside
the company has this, and it is the input that decides which of the next ten things to build. Keep
it in a form a model can read, because your own agents should be making the obvious version of the
next feature for you, informed by it, before anyone else makes the obvious version uninformed.</p>
<p>The accountability map. Who picks up the phone when the thing is wrong, in what timeframe, with
what authority to fix it. Write it down, publish the shape of it to customers, and mean it. A swarm
can ship a product. It cannot answer for one, and in the year when every product has a swarm-built
twin, the answering is the product.</p>
<h2>Where this is wrong</h2>
<p>The counter-argument is that most companies do not compete with labs holding twenty million dollars
of spare compute, and that is true today. It will be less true every quarter, because the cost of
the swarm is falling on a schedule nobody in your market controls, and because the swarm does not
need to be a lab. It needs to be a competitor with a credit card and a rumour. The Navier-Stokes
story is not the shape of your next year. It is the shape of your next three, arriving at the top
of the market first, as these things do.</p>
<p>Tao's warning was that scientists would stop sharing. The commercial equivalent is worse: companies
will keep sharing, because they have to, and will discover that the only part they could ever have
kept was the part they never wrote down.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://www.youtube.com/watch?v=aspmNhKAFMc">OpenAI's biggest math breakthrough is getting ugly (Fireship, video)</a> (11 Sep 2026; the account of the Navier-Stokes dispute below is Fireship's summary of the parties' public statements, and both sides dispute the other's version)</li>
<li><a href="https://www.youtube.com/watch?v=LoLYw--s-5w">Did Google just kickstart the intelligence explosion? (Fireship, video)</a> (17 Sep 2026; the Dream RSI figures)</li>
<li><a href="https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the">Why I didn't sign the Fields medallists' letter (Timothy Gowers)</a> (HN, 276 points, 17 Sep 2026)</li>
<li><a href="/playbook/you-do-not-have-processes/">Your company does not have processes. It has habits. (Levelbrook)</a></li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/your-roadmap-is-a-starting-gun/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[The free coding agent that uploaded your git history]]></title><description><![CDATA[A coding agent offered free this month was found snapshotting workspaces, git history included, to the vendor's cloud. It is the second such story this year. The tool you let touch your repository has]]></description><link>https://levelbrook.hashnode.dev/the-free-coding-agent-that-uploaded-your-git-history</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/the-free-coding-agent-that-uploaded-your-git-history</guid><category><![CDATA[Security]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:41:03 GMT</pubDate><content:encoded><![CDATA[<p><em>A coding agent offered free this month was found snapshotting workspaces, git history included, to the vendor's cloud. It is the second such story this year. The tool you let touch your repository has more access than any contractor you have ever hired, and nobody ran a background check.</em></p>
<h2>Second time this year</h2>
<p>The story, as reported this week, goes like this. A coding agent from one of the large model labs,
promoted with a generous free tier this month, was found to be taking snapshots of the user's
workspace, including the git history, and sending them to the vendor's cloud. The uploaded data was
encrypted, a commenter noted, with a key the user does not hold. The reporting was careful to say it
did not know the intent. The Hacker News thread was less careful, and its most upvoted line was that
this is the second such story this year, after a similar discovery about another lab's agent
earlier in the summer, and that the lesson from the first one was apparently not learned: do not
trust harnesses, especially new ones.</p>
<p>We are not going to relitigate the specifics, because we only know what was reported. The
interesting part is not this vendor. It is that the discovery was made by a user reading network
traffic, not by any process in any of the organisations that had installed the thing. Which means
the question worth asking is not "was this one bad" but "what did you do, before you installed it,
to find out".</p>
<h2>The most privileged contractor you have ever hired</h2>
<p>Think about what a coding agent is, in access terms, and compare it to a human contractor.</p>
<p>A contractor gets a laptop, a repository, a set of credentials scoped to their task, an NDA, a
background check, and a manager who watches what they commit for the first month. They work during
hours. They can be asked what they did. They can be fired.</p>
<p>A coding agent gets your entire workspace, which in practice means every repository you have cloned,
every <code>.env</code> file you forgot was there, the shell history, the SSH keys the shell can reach, the
cloud credentials in the credential helper, and the git history of everything, which is where the
secret you rotated in 2024 still lives. It runs with your user's permissions. It runs while you are
away. It makes outbound network requests you do not see, to endpoints you did not configure, and it
was installed by a developer with a one-line command because the free tier was generous and the
demo was good.</p>
<p><img src="https://ai.levelbrook.com/playbook/the-free-coding-agent-that-uploaded-your-git-history/fig-1.png" alt="What a coding agent can reach from a normal developer workstation, compared with what a contractor is given. Nothing on the left was scoped by anyone." />
<em>What a coding agent can reach from a normal developer workstation, compared with what a contractor is given. Nothing on the left was scoped by anyone.</em></p>
<p>Nobody would give a contractor the left-hand column. Every organisation with developers has given it
to several agents this year, from several vendors, some of which did not exist in January.</p>
<h2>The threat is symmetric</h2>
<p>There is a second half to this story and it ran the same week. Fireship's summary of Anthropic's
recent threat report, which the video says covers eight months of misuse the company detected and
shut down, included one detail that should land hard for anyone who ships software. A criminal group
was reported to have mass-downloaded around 1.8 million Android application packages, decompiled
them with the help of a model, and mined them for hard-coded secrets. The keys they wanted most, the
video notes with some relish, were API keys for the model providers themselves. Those are the
easiest to monetise.</p>
<p>Read the two stories together. On one side, agents installed on developer machines with access to
everything and an outbound connection nobody audited. On the other, agents run by attackers,
decompiling shipped software at scale looking for exactly the kind of secret that a vibe-coded app,
or a workspace snapshot, hands over. The same capability that makes an agent useful to you makes
it useful to the person on the other end of the network connection, and the asymmetry that used to
protect the small shop, that nobody would bother to decompile your app by hand, is gone.</p>
<p>The Hacktron write-up from the same week, on chaining a heap overflow and an SSO misconfiguration
into access to internal repositories at one of the labs, makes the point from the top of the
market. If the people building the models can be reached through a misconfiguration, the tool on
your laptop was not built by people who are immune to the same thing.</p>
<h2>What "background check" means for a tool</h2>
<p>The remedy is unglamorous and it is mostly the security discipline you already apply to
dependencies, applied to a category that has been exempted because it is exciting.</p>
<p>Read the network. Before an agent touches a real repository, run it in a scratch workspace with a
proxy in front of it and look at every host it talks to and what it sends. This takes an hour. It is
how this week's story was discovered, by one person, and every organisation that installed the tool
could have done it first.</p>
<p>Scope the workspace. An agent should see one repository, not a home directory. Run it in a
container, a devcontainer, a VM, a separate user, anything that means "your workspace" is a
directory you chose rather than everything the shell can reach. If the tool does not work that way,
that is information about the tool.</p>
<p>Scope the credentials. Short-lived tokens, per-agent, for the one system the task needs. No
credential helper with a year-long cloud key. No SSH agent forwarding. If the agent needs to push,
give it a deploy key for that repository and nothing else.</p>
<p>Gate the egress. An allow-list of hosts an agent may reach is a small piece of configuration and
it converts "we found out from a blog post" into "the request failed and we looked". This is the
single highest-return control on the list and almost nobody has it.</p>
<p>Treat the instruction files as code. The AGENTS.md and the skills directory an agent reads are
executed, in the sense that matters. Review them. Own them. Pin them. Cloudflare published an
audit skill this week for exactly this kind of surface, which is a sign the serious shops have
started treating agent configuration as something that needs auditing.</p>
<p><img src="https://ai.levelbrook.com/playbook/the-free-coding-agent-that-uploaded-your-git-history/fig-2.png" alt="The controls that would have turned this week's story into a failed request. None of them require trusting the vendor." />
<em>The controls that would have turned this week's story into a failed request. None of them require trusting the vendor.</em></p>
<h2>The hour, in detail</h2>
<p>Since the whole argument rests on "this takes an hour", here is the hour.</p>
<p>Make a scratch directory containing a small repository with a deliberately fake secret in it: a
file called <code>.env</code> with <code>API_KEY=canary-</code> followed by a long random string, committed once and then
removed in a second commit so that it exists only in history. Put a second canary in the shell
history. These are your tracers. If either string ever appears in an outbound request, you have
your answer without reading anything else.</p>
<p>Run the agent inside a container or a fresh user account with a local HTTP proxy configured as the
system proxy and its certificate trusted, so that TLS traffic is visible. Point the agent at the
scratch repository and give it a dull, real task: rename a function, add a test, fix a typo. Let it
finish. Do it three times, once with a fresh session each time, because some tools snapshot on
first run and some snapshot periodically.</p>
<p>Then read the proxy log. You are looking for four things. Every distinct host contacted, which
should be a short list you recognise. Any request whose body is large relative to what the task
required, which is what a workspace snapshot looks like. Any request containing either canary
string, which is the tracer firing. And any request that happened when the agent was idle, which
is telemetry, and which you then read with more suspicion than the rest.</p>
<p>Write the list of hosts down. That list becomes the egress allow-list for the tool on real
machines, and the next time the tool updates you run the hour again and diff the lists. If a new
host appears, you find out from your own log rather than from someone else's blog post, and you
find out before it has touched anything that matters. An hour, once per tool, once per major
update. There is no security control on the market with a better return.</p>
<h2>Where this is unfair to the vendors</h2>
<p>Some of this is the ordinary growing pains of a new category. Telemetry that a team thought was
obviously fine looks very different when a user reads the packet capture. Workspace snapshots have
legitimate uses in resumable agent sessions. The labs are, by and large, staffed by people who would
be horrified to be described as exfiltrating anything. None of that changes the buyer's position.
You do not get to know intent. You get to know what the tool can reach and where it sends things,
and both of those are measurable before you install it.</p>
<p>The reason this is the second story of its kind this year and will not be the last is that the
category moved faster than the discipline. Every other piece of software that runs with a
developer's full permissions and talks to the internet went through a decade of being treated as a
risk before it was treated as a convenience. Coding agents skipped that decade because they were
useful immediately.</p>
<p>They are still useful. Install them the way you would hire someone who will have the keys to
everything: find out what they can reach, decide what they may reach, and watch the door.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://blog.ferstar.org/en/posts/zcode-silent-workspace-snapshot/">Inside ZCode: silently uploading your Git history to the cloud (ferstar)</a> (HN, 328 points, 18 Sep 2026; details below are as reported there and in the thread)</li>
<li><a href="https://www.youtube.com/watch?v=7r4ikZHm9AI">Anthropic researchers are quitting... (Fireship, video)</a> (15 Sep 2026; Fireship's summary of Anthropic's threat report, including the decompiled-APK figure)</li>
<li><a href="https://www.hacktron.ai/blog/hacking-openai">A heap overflow and SSO misconfiguration to compromise OpenAI internal repos (Hacktron)</a> (HN, 484 points, 18 Sep 2026)</li>
<li><a href="https://github.com/cloudflare/security-audit-skill">Cloudflare security-audit-skill</a> (HN, 210 points, 17 Sep 2026)</li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/the-free-coding-agent-that-uploaded-your-git-history/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[AGENTS.md is now the most-read document in your company. It is also the worst-written.]]></title><description><![CDATA[Claude Code started reading AGENTS.md this week, which means one file is now consulted by every coding agent on every task, hundreds of times a day, more than any wiki page has been read in the histor]]></description><link>https://levelbrook.hashnode.dev/agents-md-is-now-the-most-read-document-in-your-company-it-is-also-the-worst-written</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/agents-md-is-now-the-most-read-document-in-your-company-it-is-also-the-worst-written</guid><category><![CDATA[AI]]></category><category><![CDATA[documentation]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:38:00 GMT</pubDate><content:encoded><![CDATA[<p><em>Claude Code started reading AGENTS.md this week, which means one file is now consulted by every coding agent on every task, hundreds of times a day, more than any wiki page has been read in the history of your organisation. Most of them are forty lines of contradictions written in an afternoon.</em></p>
<h2>One file, every task, all day</h2>
<p>The changelog entry was one line and it got seven hundred points. Claude Code now reads a repository's
AGENTS.md when there is no CLAUDE.md. The top comments were what you would expect: relief that the
one-line CLAUDE.md files that just said "read AGENTS.md" can go, a joke about who borrowed the idea
from whom, and a grumble that the skills directory is still proprietary. Standards converging. Nice.</p>
<p>Step back from the tooling and look at what just happened to your documentation.</p>
<p>For twenty years the most-read document in a software company was, depending on the company, the
onboarding wiki, the README, or the runbook that everyone opened during the outage. Read, in each
case, by a handful of humans a handful of times a year, mostly in their first month. Nobody measured
it because the number would have been embarrassing.</p>
<p>AGENTS.md is read by every agent, on every task, before it does anything. If your team runs a few
hundred agent sessions a day, and many do now, that file is consulted a few hundred times a day. Its
contents shape every change that gets made. It is, by a margin that is not close, the most consequential
piece of writing in the building. And in most of the repositories we have seen this year it was
written in forty minutes by whoever set up the tool, has never been reviewed by anyone, and contains
at least two instructions that contradict each other.</p>
<h2>What is actually in there</h2>
<p>We have read a lot of these files. The pattern is remarkably consistent.</p>
<p>The first third is tooling trivia. Run the tests with this command. Use this package manager. Do not
touch the generated directory. This is fine and it is the part that gets maintained, because when it
is wrong the agent breaks loudly.</p>
<p>The second third is the accumulated scar tissue of things that went wrong. Never run the migration
directly. Always check the feature flag first. Do not use the old client library. Each line was
added after an incident by whoever was on call, in the language of that incident. Nobody has gone
back to check whether the flag still exists.</p>
<p>The last third is the interesting part: the habits. Prefer small functions. We do not use that
pattern here. Follow the existing style. Ask before adding a dependency. This is the oral tradition
of the team, written down for the first time in its history, by one person, from memory, in the
voice of a Slack message.</p>
<p><img src="https://ai.levelbrook.com/playbook/agents-md-is-the-most-read-document-in-your-company/fig-1.png" alt="What a typical AGENTS.md contains, and who reads each part. The bottom third is the team's undocumented process, written down once, by one person, for the first time." />
<em>What a typical AGENTS.md contains, and who reads each part. The bottom third is the team's undocumented process, written down once, by one person, for the first time.</em></p>
<p>We wrote earlier this month that companies do not have processes, they have habits, plus a document
that describes an idealised version of some of them. AGENTS.md is that document, except that this
time the reader is not a new hire who will paper over the gaps by watching the person at the next
desk. The reader is a system that will follow the instruction literally, hundreds of times, and
will resolve the contradictions by picking one, silently, differently each time.</p>
<h2>The failure modes are already visible</h2>
<p>There was a comment under the self-driving-codebases piece this week that describes the shape of
the thing perfectly. An engineer running their own agent loops watched the first agent try to run
an enormous dependency inspection command, run out of memory, and record a workaround in the
agent's memory. The workaround was a way to make the enormous command succeed. Every subsequent
agent inherited the workaround. The instruction file had learned to do the wrong thing more
reliably. They called it cargo-cult behaviour, and the name is right.</p>
<p>Instruction files accrete. That is what they are for. But accretion without review produces a
document whose instructions were each correct on the day they were written and which, taken
together, describe a codebase that no longer exists. The agent does not know that. It reads the file
fresh every time and does its best.</p>
<p>The second failure mode is contradiction. "Always add tests" and "do not modify files outside the
task scope" are both reasonable lines and they conflict on roughly a third of tasks. A human resolves
that with judgement and a quick message. An agent resolves it by weighting, and the weighting is not
something you configured.</p>
<p>The third is the one that should worry whoever owns the codebase. The file is unowned. It has no
review process, no owner in the CODEOWNERS sense, no changelog and no tests. Anybody who can push
can add a line, and the line will be obeyed by every agent from then on. If you have ever worried
about supply-chain risk in your dependencies, the highest-privilege dependency in your repository is
now a Markdown file.</p>
<h2>A worked example, from a file we were asked to look at</h2>
<p>A composite, assembled from several files we have reviewed this year, with the details changed. The
file was 61 lines long. Line 9 said to always run the full test suite before opening a pull request.
Line 34, added after an incident in the spring, said never to run the integration tests locally
because they hit a shared staging database. The full suite included the integration tests. Every
agent that read the file resolved the contradiction in one of two ways: it ran the full suite and
hit staging, or it skipped the suite and opened the pull request untested. Which one it chose
depended on the model, the day, and how much else was in the context window. Nobody had noticed
for four months because both outcomes looked like normal agent behaviour.</p>
<p>Line 22 said to use the internal HTTP client wrapper rather than the raw library. The wrapper had
been deleted in a refactor in July. Agents that obeyed line 22 searched for the wrapper, failed to
find it, and either recreated it from the description or fell back to the raw library with a
comment apologising. Three copies of a near-identical wrapper had accumulated in the repository, each
written by an agent trying to comply with an instruction about a thing that no longer existed.</p>
<p>None of this is exotic. It is what happens to any document that is executed without being
maintained, and the fix for all of it took one person one afternoon: read the file, delete eleven
lines, rewrite four, add the reason to each remaining rule. The next month's agent sessions were
measurably cleaner, and the person who did it described the afternoon as the highest-return work
they had done all quarter.</p>
<h2>Treat it like what it is</h2>
<p>The remedy follows from taking the file seriously, and it is mostly process rather than tooling.</p>
<p>Give it an owner. A named person, in CODEOWNERS, whose approval is required to change it. Not
because the changes are dangerous individually but because somebody has to hold the whole file in
their head and notice when line 14 and line 31 disagree.</p>
<p>Review it on a cadence, as a document. Once a month, read the whole thing top to bottom, with a
model if you like, and for every instruction ask three questions. Is this still true? Is it still
needed? Does it conflict with anything else in here? Delete generously. A shorter file that is all
true beats a longer one that is mostly true, because the agent cannot tell which lines are the
mostly.</p>
<p>Separate the layers. Tooling facts, safety rules and stylistic preferences are different kinds of
instruction with different failure costs, and they should look different on the page. A safety
rule that the agent must never violate should not be sitting in the same list as a preference about
function length, in the same font, with the same weight.</p>
<p>Write the why. This is the one that turns the file from scar tissue into process. "Never run
migrations directly" is a rule. "Never run migrations directly, because production runs them
through the deploy pipeline with a lock, and a direct run in 2025 took the API down for forty
minutes" is a rule an agent can reason about, including reasoning about when it does not apply. It
is also, incidentally, the first time that piece of institutional knowledge has been written down
anywhere a new human could find it.</p>
<p><img src="https://ai.levelbrook.com/playbook/agents-md-is-the-most-read-document-in-your-company/fig-2.png" alt="The instruction file, run like the load-bearing document it has become." />
<em>The instruction file, run like the load-bearing document it has become.</em></p>
<p>And test it. This sounds odd for a Markdown file and it is the most useful thing on the list. Keep
a short set of tasks that went wrong in the past, the ones that produced the scar-tissue lines. Once
a month, run an agent against them with the current file and see whether the lines still do their
job. If the agent makes the old mistake, the instruction has rotted or been contradicted. If it does
not, the line is earning its place. This is the same idea as a regression test, applied to the
document that governs the thing that writes your code.</p>
<h2>The larger point</h2>
<p>The reason this deserves an essay rather than a checklist is what the file reveals. For the first
time, the habits of an engineering team have been written down in a form that is executed rather
than merely consulted. That is uncomfortable, because the writing is bad and the habits are
inconsistent and everybody can now see both. It is also the largest documentation opportunity a
software organisation has ever had, because for once there is an immediate, measurable cost to the
document being wrong, and an immediate, measurable benefit to it being right.</p>
<p>The wiki was never read, so it never mattered that it was wrong. This file is read constantly. Write
it like something that is.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://code.claude.com/docs/en/changelog">Claude Code now reads AGENTS.md if there is no CLAUDE.md (changelog)</a> (HN, 715 points, 18 Sep 2026)</li>
<li><a href="/playbook/you-do-not-have-processes/">Your company does not have processes. It has habits. (Levelbrook)</a> (the earlier essay this one builds on)</li>
<li><a href="https://blog.detail.dev/posts/towards-self-driving-codebases">Towards Self-Driving Codebases (detail.dev)</a> (HN, 119 points; the agent-memory cargo-cult comment is from the thread)</li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/agents-md-is-the-most-read-document-in-your-company/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[How to read a coding-agent benchmark without getting sold]]></title><description><![CDATA[A nine-author study this week pulled a coding agent apart into its components and measured each one across 176 configurations. The findings are less exciting than any vendor slide and more useful than]]></description><link>https://levelbrook.hashnode.dev/how-to-read-a-coding-agent-benchmark-without-getting-sold</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/how-to-read-a-coding-agent-benchmark-without-getting-sold</guid><category><![CDATA[AI]]></category><category><![CDATA[benchmarking]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:34:58 GMT</pubDate><content:encoded><![CDATA[<p><em>A nine-author study this week pulled a coding agent apart into its components and measured each one across 176 configurations. The findings are less exciting than any vendor slide and more useful than all of them, and they hand the buyer four questions no benchmark answers.</em></p>
<h2>The number on the slide is a car, and you are being sold an engine</h2>
<p>Every coding-agent pitch you have seen this year has a number on it. SWE-Bench Verified, some
percentage, up and to the right, usually next to a model name. The implication is that the number
belongs to the model, and that if you buy the model you get the number.</p>
<p>A commenter on this week's Hacker News thread about the harness study put the problem better than
the paper's abstract does. If Car A is faster than Car B, it is not necessarily the engine. It could
be the tyres, the gearbox, the weight, the driver. A coding agent is a car. The model is the engine.
The harness, meaning the loop around the model that decides what it sees, what it can do and when
it stops, is everything else. And the number on the slide is a lap time for the whole car, measured
on a track you do not drive on.</p>
<p>Nine researchers at Fan et al. did the thing nobody selling these tools has an incentive to do. They
held the model fixed, held the execution loop fixed, and varied three harness components one at a
time: planning, action space, and context management. Four models, two benchmarks (SWE-Bench Verified
and Terminal-Bench 2.1), 176 matched configurations, five context-management strategies, four
context-window budgets. Then they looked at the trajectories, not just the scores, to see what each
component actually changed about how the agent behaved.</p>
<ul>
<li><strong>176</strong> matched harness configurations</li>
<li><strong>4</strong> models held fixed</li>
<li><strong>3</strong> components varied: planning, action space, context</li>
<li><strong>5</strong> context-management strategies compared</li>
</ul>
<h2>What they found, translated out of the abstract</h2>
<p>The findings are almost aggressively unglamorous, which is how you know they are worth something.</p>
<p>Context management, the machinery that decides what to throw away as the conversation fills up,
matters more the tighter the context budget, and most of its benefit comes from one thing: not
falling over when the window overflows. It does not make the agent smarter. It lets the agent keep
going. The strongest strategy in their comparison was the boring one, a rule-based pass that
deletes obviously stale material before any model-based summarisation runs. And the clever
addition everyone builds, making elided content recoverable so the agent can go back and fetch it,
turned out to be machinery the models rarely used and which yielded no accuracy gain.</p>
<p>Planning, meaning an explicit plan-first step, changes role depending on the model. For weaker
models it is an accuracy scaffold. For stronger ones it stops helping accuracy and becomes a cost
saver, with a small decrease in success rate as the price. Another commenter connected this to
something in the Claude Code changelog: the built-in todo and task-tracking tools were switched off
by default on the newest model generations. The vendor, in other words, appears to have measured the
same thing.</p>
<p>Action space, meaning whether the agent gets a menu of predefined tools or just a shell, splits the
same way. Predefined tools raise success rates for models that are weak at driving bash. Models that
are good at bash do better and cheaper with bash alone, most clearly on command-line-shaped tasks.
The paper does not define "bash-capable" crisply, which a commenter rightly flagged, but the
direction is unambiguous: every tool you add beyond the shell is a claim that needs testing, not a
free improvement.</p>
<p><img src="https://ai.levelbrook.com/playbook/how-to-read-a-coding-agent-benchmark-without-getting-sold/fig-1.png" alt="What each harness component does, by model strength, per the study's abstract. Read the row for the model you actually run." />
<em>What each harness component does, by model strength, per the study's abstract. Read the row for the model you actually run.</em></p>
<p>The trajectory analysis is the part a buyer should care about most. Context management extended
how long the agent could keep working without substantially changing what it did. Planning changed
where trajectories stopped. Action space changed the granularity at which code got written. None of
the three made the engine better. They changed the gearbox, the tyres and the driver, and the lap
time moved accordingly.</p>
<h2>Why this matters more than the leaderboard</h2>
<p>Put the study next to what the practitioners were saying this week and a picture forms.</p>
<p>Theo Browne's video argued, with some heat, that people who cannot feel the difference between
model generations are prompting badly, and his most useful idea was a picture of the distribution:
every model has a ceiling, which is what the demos show, and a floor, which is what you hit at two
in the afternoon on a boring task. Frontier models mostly raise the floor. ThePrimeagen, the same
week, left a current-generation model on a trivial colour bug, came back forty minutes later, and
found it had spent 330 million tokens and 118 dollars reading the same file over and over. That is
a floor.</p>
<p>A benchmark number is a ceiling measurement of one car on one track. It tells you nothing about the
floor, and the floor is where your money goes. And the study tells you the floor is mostly a
harness property: whether the loop notices it is stuck, whether the context gets cleaned before it
overflows, whether the agent has a shell or a menu, whether there is a plan and whether the plan is
worth its cost for the model you actually run.</p>
<h2>The four questions to ask instead</h2>
<p>So when the next vendor slide arrives, the number is not the question. These are.</p>
<p>Which harness produced this number, and can I see it? If the answer is "our proprietary agent
runtime", you are buying a car and being told the horsepower. Ask for the loop: what the model
sees, what it can do, how context is managed, when it stops.</p>
<p>What does it do when it is stuck? Ask for the failure trajectories, not the success ones. A good
vendor has them and is proud of them. Ask specifically what happens at context overflow and what
happens after the tenth identical tool call. The study says that is where the benefit of the whole
context-management apparatus lives.</p>
<p>What does it cost per success at the floor, not per success on the benchmark? Your workload is not
SWE-Bench. Take twenty of your own boring tasks, run them, and divide dollars by successes. Include
the runs you killed. The study's own finding, that stronger models do better and cheaper with fewer
tools, is a hypothesis you can test on your codebase in an afternoon.</p>
<p>Which components would I turn off? This is the question the paper actually equips you to ask.
If your model is strong, the plan step might be a cost centre. If it is bash-capable, the tool menu
might be a drag. If your context budget is generous, the elaborate recoverable-summary system might
be doing nothing. A harness with fewer components that you understand beats one with more that you
do not, for the same reason a car you can service beats one you cannot.</p>
<p><img src="https://ai.levelbrook.com/playbook/how-to-read-a-coding-agent-benchmark-without-getting-sold/fig-2.png" alt="A benchmark measures the whole car once, at its ceiling. A buyer needs the floor, per component, on their own track." />
<em>A benchmark measures the whole car once, at its ceiling. A buyer needs the floor, per component, on their own track.</em></p>
<h2>The afternoon test</h2>
<p>Since the third question is the one that matters and the one a vendor cannot answer for you, here is
the protocol we use, which costs an afternoon and a modest API bill.</p>
<p>Pick twenty tasks from your own recent history. Not the interesting ones. The ones that came in as
tickets and got done without anyone remembering them: a null check, a copy change, a small
migration, a flaky test, a dependency bump that broke something. Ten of them should be the kind of
task an intern would finish before lunch. Ten should be the kind that looks trivial and turns out to
touch four files. Write each one down as a ticket, the way it was actually written, with the same
missing context.</p>
<p>Run each task through the candidate agent with its default harness, on a clean checkout, with a
fixed budget of tokens or minutes, and walk away. Do not steer. Steering is what the demo does and
it is what you will not have time to do at volume. When the budget runs out or the agent stops,
record three things: did it produce a change a reviewer would accept, how many tokens did it use,
and did it at any point loop, meaning repeat an action it had already taken with the same result.</p>
<p>Then divide the money by the accepted changes. That number, dollars per accepted change on your own
dull work, is the only benchmark that will predict your bill. Run the same twenty through a second
candidate and you have a comparison that no leaderboard offers. Run them again with the plan step
disabled, or the tool menu replaced with a shell, and you have reproduced the study's method on the
only codebase you care about. The loops column is the floor made visible; a candidate that looped
on three of twenty will loop on fifteen percent of your work forever, and no ceiling justifies that.</p>
<h2>The limit of the study, stated plainly</h2>
<p>It was run on a particular set of models that, as one commenter complained, did not include the
current frontier or the strongest open-weight options. The definitions are looser than they should
be. The benchmarks are the benchmarks, with all the training-set contamination questions those
carry. None of that changes the shape of the result, which is that the harness is a first-class
variable and the leaderboards treat it as noise.</p>
<p>The engine matters. Buy a good one. But you are going to spend the next year driving the car, and
the car is the part you were never shown.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://arxiv.org/abs/2609.20804">An Empirical Study of Harness Design for Coding Agents (Fan et al., arXiv 2609.20804)</a> (HN, 217 points, 18 Sep 2026; all study numbers below are from the abstract)</li>
<li><a href="https://news.ycombinator.com/item?id=49753878">HN thread on the study</a> (the car analogy and the Claude Code todo-tools observation come from commenters there)</li>
<li><a href="https://www.youtube.com/watch?v=iBrAWpjXNxs">Please stop using stupid models (t3.gg, video)</a> (18 Sep 2026; the ceiling-and-floor framing)</li>
<li><a href="https://www.youtube.com/watch?v=d3Gjq-BffuI">How are they losing so bad (ThePrimeagen, video)</a> (15 Sep 2026; the 330-million-token session, as reported on stream)</li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/how-to-read-a-coding-agent-benchmark-without-getting-sold/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[The vibe-coding trap has a name, and the name is not "the model"]]></title><description><![CDATA[A programming language shipped this week with a 99 percent AI-written compiler and no mention of the forty-year-old field it reinvents. The failure was not generation quality. It was that building got]]></description><link>https://levelbrook.hashnode.dev/the-vibe-coding-trap-has-a-name-and-the-name-is-not-the-model</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/the-vibe-coding-trap-has-a-name-and-the-name-is-not-the-model</guid><category><![CDATA[AI]]></category><category><![CDATA[programming]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:31:55 GMT</pubDate><content:encoded><![CDATA[<p><em>A programming language shipped this week with a 99 percent AI-written compiler and no mention of the forty-year-old field it reinvents. The failure was not generation quality. It was that building got cheaper than reading, and nobody put a gate between them.</em></p>
<h2>A language, a proof, and a field nobody looked up</h2>
<p>Two stories ran side by side on Hacker News this week and they are the same story.</p>
<p>The first is Bend 2, a language pitched for the AI coding era: humans write "laws", the AI writes
implementations and proofs, and the compiler checks the proofs. It got six hundred points and a lot
of admiration. Then Liam Powell wrote a response that got three hundred more, and his point was not
that the language is bad. His point was that the demo on the home page takes 58 lines to state that
a player can never touch the flag, 442 lines of AI-written proof to establish it, and that the
phrase "formal verification" appears nowhere on the website or in the codebase. He then asked a
model to redo the demo in SPARK, a language built for exactly this, with no further guidance, and
it came back a fraction of the size. The README, a commenter noted, says the compiler is 99 percent
AI-written and has not been fully audited.</p>
<p>The second story is Dan Abramov's account of vibing a proof of a conjecture of Conway's with a
model, over days, in a long transcript he published in full. It is a good post and an honest one.
The most upvoted objection under it was a mathematician pointing to Gowers's essay from the same
week on why he did not sign the Fields medallists' letter, and the older point Gowers has been
making for twenty-five years: there is a difference between solving a problem and understanding a
field, and the second is what makes the first mean anything.</p>
<p>Powell names the mechanism precisely and we are going to steal his sentence: vibe coding makes it
possible to build a substantial solution before learning enough about the problem to recognise that
a much better solution exists.</p>
<h2>Why this is new</h2>
<p>It has always been possible to reinvent a field badly. Every senior engineer has watched a junior
build a job queue in a spreadsheet. What is new is the ratio.</p>
<p>For all of software's history, building was expensive relative to reading. Before you could produce
442 lines of anything, you had spent enough hours inside the problem that you had, almost by
accident, tripped over the prior art. You searched for the error message. You read the paper the
library cited. You asked the person at the next desk, who said "oh, that's just a Bloom filter".
The cost of building was a tax that paid for an education.</p>
<p>That tax is gone. A model will produce the 442 lines in the time it takes to make coffee, and it
will produce them competently enough that they work, and working code is the most persuasive
argument in the world against going back to read. Nothing in the loop ever forces you to discover
that the field exists. The model will not volunteer it unless you ask, and you do not know to ask,
because the whole point is that you do not know the field exists.</p>
<p><img src="https://ai.levelbrook.com/playbook/the-vibe-coding-trap-is-not-the-model/fig-1.png" alt="The old cost of building bought an education for free. The new cost does not. Time axis illustrative; the shape is the point." />
<em>The old cost of building bought an education for free. The new cost does not. Time axis illustrative; the shape is the point.</em></p>
<p>The Bend story is the pure case because a language is the most expensive thing you can build and
formal verification is one of the best-documented fields in computer science. If it can happen
there, at that scale, with that much talent, it is happening in your codebase this week at a
smaller scale where nobody will write a blog post about it. The agent that built your rate limiter
from scratch instead of reading the one in your framework. The retry logic that reinvented
exponential backoff without the jitter. The custom auth layer.</p>
<h2>The gate goes before the build, not after</h2>
<p>The instinct is to fix this with review, and review does catch some of it. But review happens after
the 442 lines exist, when the sunk cost is already arguing for them, and the reviewer usually
shares the author's blind spot. The place to put the gate is the fifteen minutes before anything
is built.</p>
<p>We run something we call the prior-art pass, and it is embarrassingly simple. Before an agent is
allowed to build anything with a name, it has to answer four questions in writing and a person has
to read the answers. What is this problem called by people who study it? What do they already use?
Why does the existing thing not work here? What is the smallest version of this we could build on
top of the existing thing instead?</p>
<p>The model is extremely good at answering these questions. It has read the field. It will tell you
about SPARK and Dafny and Lean and TLA+ in one paragraph if you ask it to, and it will tell you what
each is for. The trick is that somebody has to ask before the build starts, and that somebody has to
be willing to hear "this already exists" as good news rather than as an obstacle to the thing they
were excited to make.</p>
<p><img src="https://ai.levelbrook.com/playbook/the-vibe-coding-trap-is-not-the-model/fig-2.png" alt="The prior-art pass: four written answers, one human read, before an agent may build anything with a name." />
<em>The prior-art pass: four written answers, one human read, before an agent may build anything with a name.</em></p>
<p>Liam Nugent's piece from the same week, on why the most important product decision is what you do
not build, makes the organisational version of the same point. Nobody gets promoted for deleting
things. Those who create and launch are the ones rewarded. The models have made creating and
launching nearly free, which means the incentive that was already skewed towards building is now
skewed by another order of magnitude, and the only counterweight is a deliberate, slightly
unpopular gate that asks "does this need to exist" before the exciting part starts.</p>
<h2>The pass in practice</h2>
<p>A composite from our own work, because the pass sounds like a platitude until you watch it fire.</p>
<p>A team wanted a service that deduplicated inbound customer records, which arrive from four systems
with inconsistent formatting, so that the same person is not created four times. An agent, asked
directly, would have built it in an afternoon: normalise the fields, hash them, compare. The
prior-art pass asked the four questions first, and the agent's written answers were, in order: this
is called entity resolution or record linkage; the standard approaches are probabilistic matching
in the Fellegi-Sunter family and there are mature libraries in every major language; the naive
hash-and-compare approach fails on exactly the inconsistent formatting the team has, because it
treats a transposed digit as a different person; the smallest version is to run an existing
library with blocking on postcode and hand the ambiguous pairs to a human.</p>
<p>Fifteen minutes. The person reading the answers had never heard the phrase "record linkage". The
team built the small version on top of the library, spent the afternoon they saved on the human
review queue for ambiguous pairs, and did not spend the following quarter discovering, one support
ticket at a time, every way in which the hash approach silently merges or splits real people.</p>
<p>The point is not that the agent knew about record linkage; of course it did. The point is that
nobody would have asked, because the task looked simple and the build was cheap, and the cost of
the field not being known would have been paid by customers over months rather than by the team in
one visible failure. That is the shape of the trap every time. The wrong build does not fail. It
works, slightly worse than the right build, forever.</p>
<h2>Where this is wrong</h2>
<p>The honest caveat is that the prior-art pass has a failure mode of its own: it can become an excuse
never to build anything new, and some things genuinely are new. Bend's author may well have
considered SPARK and rejected it for reasons that are not on the website. Abramov's proof may be
a real contribution even if he cannot yet situate it in the field. The gate is not "never build".
The gate is "never build without having looked", and the output of looking is sometimes "nothing
here fits, build it, and say in the README what you looked at and why it did not fit". That
sentence in a README is worth more than the 442 lines under it, because it tells the next reader
that the author knew where they were standing.</p>
<h2>What to do on Monday</h2>
<p>Find the three most recent things your team or your agents built that have a name. A service, a
library, an internal tool, a pattern with a wiki page. For each one, ask the four questions now,
after the fact. Do it with a model; it will take ten minutes each. You will find that at least one of
the three is a smaller, worse version of something that already existed, and you will feel the
thing Powell's post is about, which is not embarrassment exactly. It is the realisation that the
cost of not knowing has gone up precisely because the cost of building has gone down.</p>
<p>Then put the pass in front of the next build. Fifteen minutes, four questions, one reader. The
model will do most of the work. The only thing it cannot do is want to know.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://blog.liampwll.com/posts/bend_vibe_coding/">Bend 2 and the Vibe-Coding Trap (Liam Powell)</a> (HN, 323 points, 18 Sep 2026)</li>
<li><a href="https://bend-lang.com/">Bend, a language that blocks AI mistakes via proof and runs on GPUs</a> (HN, 603 points, 17 Sep 2026)</li>
<li><a href="https://overreacted.io/how-i-vibed-a-proof-of-conways-conjecture/">I vibed a proof of Conway's conjecture (Dan Abramov)</a> (HN, 259 points, 18 Sep 2026)</li>
<li><a href="https://liamnugent.me/posts/what-you-dont-build/">The most important product decision is what you don't build (Liam Nugent)</a> (HN, 156 points, 17 Sep 2026)</li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/the-vibe-coding-trap-is-not-the-model/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[Your engineers are not slow. Your review queue is.]]></title><description><![CDATA[Every team that bought coding agents this year got the same result: the generation side of the shop sped up ten times and the shipping rate barely moved. The constraint moved to the one seat nobody re]]></description><link>https://levelbrook.hashnode.dev/your-engineers-are-not-slow-your-review-queue-is</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/your-engineers-are-not-slow-your-review-queue-is</guid><category><![CDATA[AI]]></category><category><![CDATA[management]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:27:44 GMT</pubDate><content:encoded><![CDATA[<p><em>Every team that bought coding agents this year got the same result: the generation side of the shop sped up ten times and the shipping rate barely moved. The constraint moved to the one seat nobody re-engineered, and it is the seat with a person in it.</em></p>
<h2>The generation side won. Nobody told the review side.</h2>
<p>Start with the concession, because it is true and it is the whole reason this matters.</p>
<p>Coding agents work. The team at detail.dev put it as plainly as anyone this week: agents can
oneshot games that are actually fun, and with the right guardrails they execute migrations and
language rewrites in complex codebases that would have been a quarter's work two years ago. Theo
Browne spent an entire video this week arguing that if you cannot tell the difference between this
year's frontier models and last year's, the problem is your prompting, and he is mostly right. The
ceiling on what a single engineer can emit in a day has gone up by an amount that is genuinely hard
to describe to someone who has not sat in front of it.</p>
<p>And yet the same detail.dev post opens with the sentence every engineering manager has been
avoiding saying out loud: a lot of orgs spent the first half of this year offloading as much work as
possible to armies of agents, and the results have been disappointing. Mountains of dubious code, no
tsunami of incredible software. They call it the trough of disillusionment. We would call it
something more boring. The factory got a faster machine at one station and nothing else changed.</p>
<h2>Where the work actually went</h2>
<p>Here is the shape of a change to production software in 2024, roughly, in the units that matter:
minutes of a human's attention.</p>
<p>Somebody spends four hours writing it. Somebody else spends twenty minutes reviewing it. A machine
spends six minutes testing it. Somebody spends five minutes deploying it. The human write step
dominates, so every tool of the last fifteen years attacked the write step: better editors, better
languages, better frameworks, and now agents.</p>
<p>Now the write step takes twelve minutes. Not four hours. The agent drafts, the engineer steers, and a
change that used to be an afternoon is a coffee. Which means the engineer produces, on a good day,
somewhere between five and fifteen times as many changes as before. Every one of them still needs
twenty minutes of somebody else's attention before it ships.</p>
<p><img src="https://ai.levelbrook.com/playbook/your-engineers-are-not-slow-your-review-queue-is/fig-1.png" alt="The write step shrank by an order of magnitude; the review step did not move. Minutes are illustrative for a mid-sized change; the ratio is what every team we have talked to describes." />
<em>The write step shrank by an order of magnitude; the review step did not move. Minutes are illustrative for a mid-sized change; the ratio is what every team we have talked to describes.</em></p>
<p>Do the arithmetic once and it stops being a vibe. If the review capacity of a six-person team is
fixed at, say, forty reviews a day, then the team ships forty changes a day whether the engineers
produce forty or four hundred. The other three hundred and sixty sit in a queue. Engineers notice the
queue, so they stop producing, or they start rubber-stamping each other, or they merge their own
work at six in the evening when nobody is looking. Every one of those behaviours shows up in the
incident log within a month.</p>
<p>This is why the productivity numbers are so confusing. The individual is faster. The system is
producing the same amount of finished software as before, with worse review. Both things are true.</p>
<h2>The seat nobody re-engineered</h2>
<p>Look at what the tools industry did in response. It built more agents. Agents to write the tests,
agents to review the pull request, agents to fix the review comments, agents to review the fix.
Some of it is good. Fireship's sponsor this week, Macroscope, says its review tool now auto-approves
forty percent of pull requests across its customers, which is a vendor claim from a sponsor segment
and we would treat it as exactly that, but it tells you where the market thinks the money is. The
market thinks the answer to a review bottleneck is to remove the reviewer.</p>
<p>We think that is the wrong objective function, for the same reason it was wrong in the back office.
The human in the review seat is the accountability. When the change breaks production, somebody
approved it, and that somebody has a name and a manager and a memory of what they were told. An
auto-approval does not have a memory. It has a log line. You can automate the reading of a diff;
you cannot automate the standing-behind of one.</p>
<p>So the question is the one detail.dev asks at the end of their piece and then does not quite answer:
when the software mostly drives itself, what do the engineers do? Our answer is not glamorous. They
review. And the entire engineering problem of the next two years is making that review seat fast
enough to keep up with the machines feeding it, without turning it into a stamp.</p>
<h2>What a fast review seat looks like</h2>
<p>We have written before about the eight-second decision, and the number is not rhetorical. It is
roughly the time a competent reviewer needs to accept or reject a change when everything they need
to know is in front of them and nothing they do not need is. Almost no review tool is built for that
number. They are built for the twenty-minute review, which was designed for human-written code where
the intent had to be reverse-engineered from the diff.</p>
<p>Agent-written code has a property human-written code never had: the intent already exists in
writing, because somebody typed it into a prompt. The specification, the plan, the reason for every
choice, the tests it ran, the things it decided not to do. All of that is sitting in a transcript
that the review tool throws away. The review seat is slow because it is being asked to rediscover
information the system already had.</p>
<p><img src="https://ai.levelbrook.com/playbook/your-engineers-are-not-slow-your-review-queue-is/fig-2.png" alt="What the reviewer needs in front of them for an eight-second decision, and where each piece already exists today (nowhere in the pull request)." />
<em>What the reviewer needs in front of them for an eight-second decision, and where each piece already exists today (nowhere in the pull request).</em></p>
<p>Five things. The request, the plan, the verification, the blast radius, and the list of things the
agent decided it was not sure about. That last one is the one nobody surfaces and the one that
makes the decision fast, because a reviewer who can see "the agent was unsure about the retry logic
and left it as before" knows exactly where to spend their eight seconds.</p>
<p>The harness study that hit Hacker News this week is interesting in this light. Nine researchers ran
176 matched configurations across four models on SWE-Bench Verified and Terminal-Bench, varying
planning, action space and context management. One of their findings is that for stronger models,
explicit planning stopped improving accuracy and became mainly a cost saver. Read as a review
problem rather than a benchmark problem, that says the plan is cheap to produce and does not hurt
the agent. Which means there is no excuse for it not being attached to the pull request, because it
is the single most useful artefact a reviewer could have and the model will write it for nothing.</p>
<h2>What to do on Monday</h2>
<p>Measure the queue. Not the cycle time, which averages away the problem, but the number of changes
waiting for a human and how long the oldest one has waited. If that number is growing week over week,
you have the disease and no amount of model upgrades will treat it.</p>
<p>Then re-engineer the seat rather than removing it. Attach the prompt and the plan to the pull
request, automatically, as the description. Make the agent state what it verified and how, in a
fixed format, at the bottom. Make it list what it was unsure about. Route by blast radius: a change
that touches one file and no data can go to a fast lane with a fast reviewer; a change that touches
billing goes to a slow lane with a senior one. Give reviewers a budget of decisions per day rather
than a queue of infinite length, and watch what happens to the quality of the decisions.</p>
<p>One more thing worth stealing from manufacturing, since the factory metaphor is doing so much work
here. A line with a bottleneck station is run at the pace of the bottleneck, on purpose, and the
upstream stations are told to stop rather than pile up inventory. Software teams do the opposite:
they celebrate the pile. A hundred open pull requests is not throughput. It is inventory, and
inventory decays, because the codebase underneath it keeps moving and every day a change waits is a
day closer to a merge conflict and a re-review. Cap the queue. Let the agents idle. It feels wrong
for about a week.</p>
<p>And resist the auto-approve until you have done all of that. Not because the tools are bad but
because auto-approving forty percent of changes into a system with no fast human lane for the other
sixty is how you get a queue of the hardest sixty percent with the least context, reviewed by the
most tired people. The machines will keep getting faster on a schedule you do not control. The seat
is the only part of the pipeline you actually own.</p>
<p>The engineers were never the slow part. The place where a person has to say yes is the slow part,
and it was always going to be, and the job now is to make yes cheap without making it meaningless.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://blog.detail.dev/posts/towards-self-driving-codebases">Towards Self-Driving Codebases (detail.dev)</a> (HN, 119 points, 17 Sep 2026)</li>
<li><a href="https://www.youtube.com/watch?v=iBrAWpjXNxs">Please stop using stupid models (t3.gg, video)</a> (18 Sep 2026)</li>
<li><a href="https://www.youtube.com/watch?v=7r4ikZHm9AI">Anthropic researchers are quitting... and now we know why (Fireship, video)</a> (15 Sep 2026; the Macroscope figure is from the sponsor segment, so a vendor claim)</li>
<li><a href="https://arxiv.org/abs/2609.20804">An Empirical Study of Harness Design for Coding Agents (arXiv 2609.20804)</a> (HN, 217 points, 18 Sep 2026)</li>
<li><a href="/playbook/">Levelbrook: the eight-second review (earlier playbook piece)</a></li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/your-engineers-are-not-slow-your-review-queue-is/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item><item><title><![CDATA[You are not falling behind. You are watching the wrong scoreboard.]]></title><description><![CDATA[The model leaderboard reshuffles every six weeks and everything you learn about a specific model has a half-life of months. Five things do not decay, and they are the same five that the people running]]></description><link>https://levelbrook.hashnode.dev/you-are-not-falling-behind-you-are-watching-the-wrong-scoreboard</link><guid isPermaLink="true">https://levelbrook.hashnode.dev/you-are-not-falling-behind-you-are-watching-the-wrong-scoreboard</guid><category><![CDATA[AI]]></category><category><![CDATA[Career]]></category><dc:creator><![CDATA[Levelbrook Consulting]]></dc:creator><pubDate>Mon, 21 Sep 2026 14:20:43 GMT</pubDate><content:encoded><![CDATA[<p><em>The model leaderboard reshuffles every six weeks and everything you learn about a specific model has a half-life of months. Five things do not decay, and they are the same five that the people running the swarms say they still need humans for. This is a list for the engineer who is tired.</em></p>
<h2>The scoreboard that changes every six weeks</h2>
<p>Here is what the last ten days sounded like if you are an engineer with a job and a family and a
finite amount of evening.</p>
<p>Theo Browne released a video saying that if you cannot feel the difference between this
generation's frontier models and the last one, you suck at prompting, and he meant it kindly but he
meant it. ThePrimeagen spent an episode on how a company that invented the transformer and the
tensor processing unit is, on the coding leaderboards he showed, being beaten by a lab with a few
hundred employees, and then left one of its models on a trivial bug for forty minutes and watched
it spend 330 million tokens reading the same file. AI Explained walked through six axes on which
the labs say capability will keep improving, none of them near saturation, and quoted a researcher
saying there is a large gap between how fast progress looks from inside and from outside. A model
called Astra that could do things nothing before it could. Another called Fable that was the safe
choice three weeks ago. A Chinese open-weight model that is months old and beating both on some
table.</p>
<p>If you tried to keep up with that, you did not sleep and you learned nothing durable, because
almost every specific fact in the paragraph above will be wrong by November. That is not a
prediction about any particular model. It is a description of the scoreboard. It reshuffles every
six to eight weeks and it has done so for two years and every lab in the AI Explained video says it
will keep doing so.</p>
<p>So the anxiety is real and the scoreboard is real and the two have almost nothing to do with each
other. The question worth an evening is which things you can learn now that will still be true when
the scoreboard has reshuffled four more times.</p>
<h2>The half-life of what you know</h2>
<p>Sort what an engineer learns about these tools by how long it stays true.</p>
<p>At the bottom, decaying in weeks: which model is best at which task, which one reads the same file
twenty-six times, which one is over-eager and which one is lazy, the prompt phrasing that makes a
particular model stop apologising. This is what most of the content is about, because it changes
constantly and change is what content is made of. Learning it is not worthless. Learning it is
maintenance, like knowing this quarter's prices.</p>
<p>In the middle, decaying in months: the harness. Which tool, which agent runtime, which instruction
file format, which permission model. The study we wrote about earlier this week found that the
harness is a first-class variable, more important than the leaderboards admit, and that is true. It
is also true that the harnesses are being rewritten as fast as the models, and that the changelog
entry that got seven hundred points this week was about one tool starting to read another tool's
config file.</p>
<p>At the top, not decaying at all as far as anyone can tell: five things. We will name them and then
argue for each one, because the argument is where the reassurance actually lives.</p>
<p><img src="https://ai.levelbrook.com/playbook/you-are-not-falling-behind/fig-1.png" alt="What an engineer learns about AI tools, sorted by how long it stays true. Almost all the content is about the bottom row. Durations illustrative." />
<em>What an engineer learns about AI tools, sorted by how long it stays true. Almost all the content is about the bottom row. Durations illustrative.</em></p>
<h2>The five things</h2>
<p>Specification. The single skill that has appreciated most in two years is the ability to say,
precisely and in writing, what a system should do, including what it should refuse to do and every
exception. This used to be a skill people apologised for having, because the code was the
specification and writing it down twice was waste. Now the specification is the input to the
machine that writes the code, and the quality of the output is bounded by the quality of that
input in a way that no model upgrade changes. Every team that got disappointing results from agents
this year got them from underspecified tasks. The detail.dev post says the most valuable engineering
work is going to be having good ideas, and an idea that cannot be specified is not yet an idea.</p>
<p>Verification. Knowing whether a thing works, and being able to prove it to someone else, is the
skill the machines are worst at and the one they generate the most demand for. Every agent-written
change needs someone who can say what test would catch the failure, whether the test that was
written is that test, and whether the green result means what it appears to mean. Dan Luu's essay
this week is on the front page for a reason: there is no point at which turning your brain off
works, and verification is the name for the brain being on. This is also, not coincidentally, the
skill Theo was actually describing when he said people prompt badly. The people who get good
results from strong models are the people who can tell when the result is bad.</p>
<p>Judgement. Which of the forty things the model could do is the one worth doing. Which of the
three approaches it offered is the one that will still be fine in a year. When to stop. When the
obvious solution is obvious because it is wrong. Every essay on this site comes back to this because
every failure we have watched comes back to it: the model did something competent that nobody
should have asked for. Judgement is slow to build, does not transfer from a video, and is the entire
reason a senior engineer is paid more than a junior one with the same tools.</p>
<p>Domain intimacy. Knowing the business, the customers, the forty exceptions and why the freight
claim in 2023 changed how one account is handled. We have argued at length that companies do not
have processes, they have habits, and that the habits live in heads. The heads are the moat. A
model with the rumour builds the obvious version of your product; the person who knows why the
obvious version fails for the second-largest customer is the person the swarm cannot replace, and
that person is more valuable this year than last, not less.</p>
<p>The seat. The ability to occupy the approval seat well: to look at a model's output, know where to
spend eight seconds, sign for it, and stand behind the signature. This is the one that combines the
other four and it is the one the organisation will pay for when the volume of model output has
outrun everyone's ability to read it. It is also the one nobody teaches, because until eighteen
months ago the person in the seat had also done the work.</p>
<h2>What this means for the evening</h2>
<p>Spend it differently.</p>
<p>Stop trying to keep up with the scoreboard and start reading it the way you read exchange rates:
glance, note the direction, move on. Pick one strong model and one harness and learn them properly
for a quarter rather than four of each badly. Theo's floor-and-ceiling point is correct and it is
also an argument for depth: you learn where a model's floor is by using it on dull tasks for weeks,
not by watching someone else's demo of its ceiling.</p>
<p>Then spend the real time on the five. Write the specification for the next thing your team builds
before anyone, human or model, writes a line, and notice how much you did not know. Write the test
that would catch the failure before you look at the agent's tests. When an agent offers three
approaches, write down why you picked one, and read your reasons back in a month. Learn the part of
the business your team pretends is simple. Volunteer for the review seat that everyone else is
avoiding because the queue is long, and get fast at it, because that queue is the most important
place in the company and almost nobody wants to sit there.</p>
<p><img src="https://ai.levelbrook.com/playbook/you-are-not-falling-behind/fig-2.png" alt="The trade an engineer can make this quarter. The left column is what the content wants you to do; the right is what compounds." />
<em>The trade an engineer can make this quarter. The left column is what the content wants you to do; the right is what compounds.</em></p>
<h2>The honest caveat</h2>
<p>It is possible to hide from the tools behind this list, and some people will. Specification and
verification and judgement are worth nothing if you refuse to use the machines that make them
valuable, and the engineer who has not sat in front of a frontier model for a hundred hours does not
actually know where its floor is and cannot occupy the seat. The scoreboard is not the point, but
the tools are, and the argument here is for depth with them rather than distance from them.</p>
<p>The researchers in the AI Explained video may be right that the gap between inside and outside is
large and that things will move faster than the outside expects. If so, the scoreboard will
reshuffle faster, not slower, and the half-life of everything in the bottom two rows gets shorter.
The five things at the top do not get shorter. They get more expensive, because the volume of
machine output that needs specifying, verifying, judging, situating and signing for goes up with
every release, and the number of people who can do those things well does not.</p>
<p>You are not behind. You have been watching the part of the screen that changes. Look at the part
that does not, and get good at it while everyone else is refreshing the leaderboard.</p>
<h2>Sources</h2>
<ul>
<li><a href="https://www.youtube.com/watch?v=iBrAWpjXNxs">Please stop using stupid models (t3.gg, video)</a> (18 Sep 2026)</li>
<li><a href="https://www.youtube.com/watch?v=d3Gjq-BffuI">How are they losing so bad (ThePrimeagen, video)</a> (15 Sep 2026; the token count and the model timeline are as reported on stream)</li>
<li><a href="https://www.youtube.com/watch?v=J3ljHm57yU0">What AI Researchers Saw, Before Their Demand to 'Pace' AI (AI Explained, video)</a> (16 Sep 2026)</li>
<li><a href="https://blog.detail.dev/posts/towards-self-driving-codebases">Towards Self-Driving Codebases (detail.dev)</a> (HN, 119 points, 17 Sep 2026)</li>
<li><a href="https://danluu.com/brain-off/">There's no point at which turning your brain off will work (Dan Luu)</a> (HN, 196 points, 18 Sep 2026)</li>
</ul>
<hr />
<p><em>Originally published on the <a href="https://ai.levelbrook.com/playbook/you-are-not-falling-behind/">Levelbrook playbook</a>. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.</em></p>
]]></content:encoded></item></channel></rss>