Source: https://www.youtube.com/watch?v=zcLPGC-tvgk (Robert C. Martin × Matt Pocock, 56:39, 2026-08-19)
Extracted from the transcript in transcript.md.
Agents suffer from messy code the same way humans do — just at a higher threshold. Bob's December/January experience: he let an agent pile change on change without cleaning up, and watched it thrash — fix one thing, inadvertently break another, fix that, break another, going in circles. One agent eventually gave up outright. Mess compounds until the agent stalls.
The response is not more instructions. It is deterministic tooling wrapped around the agent.
Rules written into a long prompt (CLAUDE.md / AGENTS.md style) get treated as "more like guidelines" — the Pirates of the Caribbean problem. The mechanism is lost in the middle: as the context window grows, the beginning and the end retain prominence, the middle is effectively gone. A long rules document buries its own contents.
Deterministic tools don't decay that way. If you want a rule enforced, put it in a script, not in a paragraph.
Only the first few sentences reliably retain priority — everything past that gets shoved into the middle as the context grows. Trim the initial prompt to its absolute minimum, then enforce everything else after the fact with tooling.
The pattern: "You must change the code until this tool says it's OK." The agent then loops — add tests, cut cyclomatic complexity, split functions — until it conforms.
Bob's two tools, both ~2000-era ideas that were impractical for humans because they were too slow and too boring:
- CRAP score — combines test coverage with cyclomatic complexity into a single "how crappy is this function" number. Bob originally ran it in the early 2000s, found plenty of crappy functions, and shelved it because fixing them all by hand was unaffordable.
- Mutation testing — a program walks the source flipping operators (
<→>,==→!=, sign flips) and runs the full suite for each mutation, expecting failure. A mutation that survives means a gap in the tests: "a surviving mutant, and it must be killed." Bob's year-2000 run took an overnight batch; agents do comparable work in ~30 minutes and then plug the holes.
The unlock is not that the ideas got better. It's that agents are fast and don't care how boring the work is.
Bob's phrasing: "It's probably a mistake to impose a human discipline on an agent. It is not a mistake to impose human values on the agent, but there may be thresholds that we need to change."
- Thresholds shift. Agents have huge, perfectly accurate short-term memory, so they tolerate more complexity per function. Bob keeps CRAP below 4 for humans, has set 6 for agents, and is considering 8. (With 100% coverage, a CRAP score of 6 means six fully-tested pathways through the function.) He hasn't found the true ceiling.
- Strict TDD doesn't transfer. Line-of-test-then-line-of-production is a discipline that exists because of how humans are wired — it works around limited short-term memory. Bob is a longtime TDD advocate and still won't enforce it on agents: they revert to writing a function then its test (Ousterhout's style) even when instructed otherwise, and he considers that fine.
Bob's gauntlet:
- Specifier — takes a human-written document, produces a Gherkin acceptance test (given/when/then) and a QA procedure, both written from a human operator's point of view: "You are a human. You are operating this system at the UI. You must prove that the system works."
- Coder — writes unit tests and the implementation for the story, gets the Gherkin passing.
- Cleaner — runs CRAP analysis and general code review, cleans up the mess the coder left behind.
- Hardener — runs mutation testing, "absolutely merciless," drives toward 100% coverage. Takes a good long time.
- QA agent — turns the written QA document into an executable script that drives the system and produces a deterministic result.
Advantages: stages can run in parallel; each agent's context stays small and focused, so lost-in-the-middle bites less and a few more rules survive at the top; each agent starts with a clean context.
Costs: ~10–15s startup per agent, plus re-deriving context from scratch; heavy communication overhead between stages.
Measured result: a task a single agent finishes in 5 minutes with questionable output takes roughly an hour through the gauntlet — against about half a day for a human. Call it 4–5× productivity, at quality "much more than the human would ever put into it."
Pocock's framing, which Bob endorsed: once a session leans in a direction, everything afterward follows that direction, and the only way to clear it is to clear the context.
Bob's illustration: you're having a pleasant conversation with a model about coffee, someone walks past talking about a soap opera, that lands in the context — and from then on every coffee answer is soap-opera flavored. The model can't differentiate what belongs. Keep everything in the window consistent and the direction unconfused, and you avoid a lot of the hallucination people complain about.
Narrow single-purpose agents are partly a trajectory-control mechanism: an implementer trying to make something work has a gentler trajectory than a hardener chasing 100% coverage.
Bob's manual loop (up until roughly a month before the interview): have the agents build the thing, then interrogate them about it — what are the modules, how do they interrelate, how do they talk to each other. "Then I would get scared to death, because the answers were horribly frightening." Then design the module structure himself and hand the agents an implementation plan.
Two tools he built to support this:
- Architecture viewer — a UML-ish diagram of module structure and dependency direction; click a module to see its submodules, click down again to reach the code. Drill to any level.
- Dependency-rule spec file + checker — declares which modules may depend on which and how dependencies should flow. Agents cannot violate it; a checker runs at the end, and violations must be fixed, usually by inverting a dependency, inserting an interface, or splitting a module.
He is trying to automate the design step itself and, as of the interview, "having not a lot of luck."
Ousterhout's shallow-vs-deep distinction: shallow modules have wide interfaces and little hidden behind them; deep modules have narrow interfaces hiding a lot of implementation.
Models pay attention to interface names and structure, and to tests — they read tests to understand what the system does. That lets them work without reading the implementation underneath, which is "both a danger and an advantage": fine as long as the code is consistent with its interface.
Same argument as clean-vs-dirty code. A well-partitioned module with disciplined interfaces is graspable because humans compartmentalize — and so do the models, at some threshold. Load a module with everything under the sun and the agent won't know what it's doing in there.
Bob tried heavy upfront planning, including the week of the interview, and called it "always a disaster" with always the same outcome: you make the plans, then as the agents run you realize they can't follow the plan because you didn't think of everything and they aren't as wise as you are. They run half-cocked, you stop them, back up, rewrite the plan, restart.
He explicitly connects this to the 1970s temptation that produced waterfall, and to agile as the answer to it. His current experiment is the agile approach: run a story or two, look at the architecture at the end, get involved manually where needed, sort things out, then a few more stories. He isn't sure the manual organizing step can ever be eliminated.
On plan-maxing specifically: "The agents love to write plans. Oh my goodness, they love it. And they will embellish the plans and the plans will be gorgeous and beautiful and spell out all kinds of details" — and then fall apart at the end. His read on the industry's move toward spec-driven development is that it's probably not going to work.
The cost-of-change argument. If every change to a house cost $1 — including laying the foundation and the roof — would you pay an architect thousands for a perfect plan the contractor then builds in one shot? Or would you tell the contractor to put the foundation here, look at it, move the kitchen, move the stairs, fix the traffic pattern? The cost of change has dropped about as close to zero as it's going to get. So why pay for expensive upfront planning instead of iterating?
Bob keeps no list of specifications in the repo. They change constantly and then go away. There's no longer an equivalent to source code as "the final specification," because humans aren't the ones writing the source anymore, and a lot of people feel that absence and want a human artifact that defines everything upfront.
His inversion: instead of writing a specification that defines what he wants, he looks at the end result and treats that as the specification.
Corollary, on his own published tools (CRAP runners for Clojure/Java/Go, the mutation tester, the agent harness): "Don't download those. I wrote them for me. Point your agents at them, have the agents look at them, and then build one for you." Specify the essence, then customize to your need.
Send an agent a large specification and it will actually read it. Send a human the same document and you might get a 20% hit rate — "5% maybe."
The inverse also holds, and is the more uncomfortable half: the things agents write, humans don't read. They're producing output on the assumption we'll read all of it, and we don't.
Framing borrowed from Ousterhout: tactical programming is the sergeant on the ground fighting the battle; strategic is the general directing the course of the war. Agents are very good at tactical and bad at strategic — and tactical work is exactly what they've eaten, which is where newcomers used to learn.
Bob is explicit that he doesn't have a confident answer, but his sketch:
- Write code by hand first — a year, maybe, however long it takes — so you know what the agents are dealing with.
- Treat juniors like agents. Whoever is running the agent fleet strategically should hand the new hire the same tasks the agents get, and subject them to the same deterministic tools. Several months of being horribly unproductive and learning a hell of a lot. After that gauntlet, maybe they can be trusted to run an agent of their own. Pocock's summary: "You become a sub-agent of the agent."
- Learn to recognize the thrash, not the bad code. Bob's own diagnostic in December wasn't spotting dog-do in the diff — that wasn't the important part. It was watching the agent struggle, and recognizing the struggle because he'd been through it himself. A novice wouldn't recognize it.
- Read the old books. DeMarco, Yourdon, The Pragmatic Programmer — "the old books, the ones that nobody reads because they're old." Filter out the archaic parts; the lessons were learned in the 70s and 80s and the books are where they're written down. Then you still have to learn it by feeling it.
- The assembly-language analogy. Bob's advice from ten years ago: if you've never written assembly, spend a weekend on it, because if all you do is write Java you live in a fantasy world and there's magic you don't understand. The educational path he sketches runs binary → assembly → something like C → something like Python → agent work with deterministic tools → strategically supervising agents.
Pocock's addition: the feedback loop on strategic programming has always been brutally long — your mistakes surface nine months out, and people who change jobs every six months may never see their own. Agents compress that loop, so the feedback arrives sooner.
Bob's answer, attributing the premise to Dijkstra: software is the most complicated thing humans have ever attempted. The fundamentals are how we organize that complexity into a form that can be conceived — not just by humans, but by the models too, since the models are modeled after humans.
On the people who say fundamentals no longer matter: "They will learn, and they will learn that the hard way, and it won't take very long." It might take longer than he expects, because the agents are good — but he's watched them hit the wall and knows the wall is there.
The abstraction-layer parallel (closing exchange): every step up the abstraction ladder — binary, assembly, compilers, high-level languages, and now the model — drew the same complaint from the layer below, that it would ruin everything and eliminate the jobs. It didn't, and the same rules applied each time.
Bob's closing line, which is the whole talk compressed:
"The rules you throw away are the ones you're going to pick up off the floor in a year and dust off and remember why you need them."
Pocock's coda: people have made the abstraction complaint since Plato, who argued writing would make people stupider.