In March I wrote a post called "The Framework" that laid out how my agentic development process actually works. Five phases, a library of skills, a supervisor/worker orchestrator, and a philosophy built around one adversary: context drift. That post described version 3 of the framework. It had been through three iterations in about three weeks, and I said at the time that it wasn't the only way to do this, just the way I'd arrived at.
Version 6 merged at the end of September. The March post is still the right place to start if you're new here, because the bones haven't moved. But three generations of changes have landed on top of those bones, and the interesting part to me isn't the list of what got added, it's why each version happened.
None of it came from the model getting smarter. Every version is the process catching itself failing at something and writing the fix down.
What Didn't Change
First, what didn't change, which is most of it.
Phase 1 is still Otter.ai and talking. Phase 2 is still the convergence loop in Claude Web, still producing a PRD, a Screen Inventory, and a Decisions Log. Phase 3 is still a skill that decomposes the PRD into a milestoned backlog of vertical-slice stories on a GitHub Project board. The engineering standards repo is still the single source of truth that every session pulls before it does anything. The orchestrator still spawns a fresh subagent for every stage so no one agent carries baggage from the last ticket. Milestones are still hard stops. Circuit breakers still exist. Before any ticket starts, main still gets tagged as a rollback point.
If you read the March post and then looked at the repo today, you'd recognize it. What you'd notice is that the pipeline got longer, the rules got sharper, and there's a whole layer sitting on top that didn't exist in March.
v4: The World Got Bigger
The March post's pipeline ended at integration tests. By late spring it didn't.
v4 added a UI test stage: Playwright tests, written on their own PR after the implementation and integration tests merged. It added infrastructure to the security review, so the same agent that runs the OWASP Top 10 pass now auto-detects Bicep, Docker, and GitHub Actions files and reviews those too. It added two entire phases that the March post didn't have: deployment and promotion (auto-deploy to Dev on merge, manual promotion to Staging and Production through GitHub Environments), and ongoing project health for when you come back to a repo after time away.
The v4 addition that mattered most was small and boring: a background CI watcher. After every merge, the orchestrator spawns an agent whose only job is to watch the CI run that merge triggered. If it goes green, the watcher reports and disappears. If it goes red, a fix agent spins up, pulls the logs, diagnoses the failure, opens a repair PR on its own branch, and merges it, all in parallel with whatever ticket the orchestrator is already working on. Before this, a CI break was silent. Three tickets would merge on top of it before anyone noticed. Now the orchestrator finds out within minutes of the merge that caused it.
v4 also added triage, conformance auditing, and backlog reconciliation for when a PRD changes mid-project. None of it was dramatic. It was the framework growing to cover the parts of a real project that the March version had left to me.
v5, Part One: The Pipeline Was Making Too Much
In mid-June I asked Claude to assess the test suite on FormIt, the largest project running on the framework. The question was simple: we love tests, we subscribe to the testing pyramid, and we seem to be writing way too many. Assess.
The numbers came back at 4,245 tests. Sixty-six of them were pure unit tests. Nearly two thousand were integration tests against a real database. Seven hundred and fifty-seven were Playwright. The assessment called it an ice-cream cone, not a pyramid, and it was already costing real money: four of the five most recent commits were Playwright stabilization.
The useful part wasn't the count, it was the diagnosis, which pointed at the shape of the pipeline instead of any one test. Test-writing was purely additive across three skills and three PRs, each running from fresh context, blind to what the others had written. Every tier carried its own mandate of "happy path plus edge plus error, for every method," so the same behavior got verified three times, weighted toward the most expensive tier. The testing standard itself prescribed writing both a unit test and an integration test for every endpoint. And the engineering review was a one-way ratchet: it could reject a PR for missing coverage, up to three iterations at three stages, but it had no rule for flagging a redundant test. Every feedback loop in the system pushed the count up. Nothing ever pushed it down.
v5's answer was to define the tiers properly and enforce them from both directions. The standard now names four tiers (Unit, Contract, Integration, UI) and says what each one owns and what never goes there. The governing rule sits above the table: test each behavior once, at the cheapest tier that can catch its failure mode. The refine stage assigns every behavior to exactly one tier before implementation starts. The reviewer checks the boundary both ways, flagging business logic that landed in the Contract tier and contract assertions that got repeated in Integration, and it always cuts the more expensive duplicate.
Then the rules got IDs. There are eleven Test Rules now, TR-1 through TR-11, in one numbered block in the testing standard. Skills never restate a rule; they cite it. Every PR carries a small grid with the applicable TR rows down the side and two columns, one the producer fills in and one the reviewer fills in, and every cell needs evidence: a line reference or pasted output, never a bare checkmark. The producer fills it in before saying done; the reviewer fills it in trying to prove the producer wrong. I'd noticed subagents follow rules a lot better when handed an explicit checklist to fill out rather than prose to remember, and the grid is that observation turned into a mechanism.
Two other v5 rules. "Fix the code, not the test": a failing test is a signal the code is wrong, and editing the test to make it pass is reserved for a test that was written badly, never for a test that's inconvenient. And the critical-path cap: UI tests covering a critical user journey get tagged, and the repo-wide count of tagged journeys has a floor of three and a hard ceiling of ten. Ten is a count, not a percentage, on purpose. A percentage of a suite that grows to thousands doesn't mean anything. Ten is ten.
That ceiling got tested for real on a repo sitting at 602 UI tests when the rule went in. That's a separate post.
v5, Part Two: The Orchestrator Forgot
The second v5 story started with a measurement I'd been avoiding. The orchestrator was reliable for about three tickets. Somewhere between five and ten, it started forgetting workflow steps. Not failing loudly, just skipping things, the same context drift I'd built the whole subagent setup to prevent, now showing up in the one agent that didn't get a fresh context, because it was the one spawning everyone else.
The obvious fixes didn't work, and I checked. Subagents can't spawn subagents, so you can't put an overseer above a per-ticket orchestrator above the stage agents. An agent can't clear its own context mid-run. The only supported way to get fresh context per unit of work is to start a new process.
So that's what v5 does. The orchestrator became stateless. On every launch it reads its state from a per-run tracking issue on the board and picks a mode. WORKING means there are ready tickets: process up to N of them (N defaults to three, the measured reliability threshold), checkpoint each one's state to the tracking issue as it finishes, and exit. CLEANUP means there's nothing left: run the full UI suite, compute the pyramid ratio and file a drift ticket if it's crossed a threshold, audit that every ticket really did get its reviews and its grids, inject any fix tickets that fell out of all that, and exit.
Relaunching it is the job of what the plan literally calls the dumb loop. It's a bash script. It knows three things: time out a session that hangs, restart a session that crashes, and stop after a maximum number of iterations so a runaway can't go forever. It relaunches the orchestrator until the run reaches a fixpoint, meaning no ready tickets and a CLEANUP pass that injected nothing new, at which point the orchestrator marks the run complete and the loop breaks. There's no intelligence in the loop at all. That lives in the orchestrator, the state lives in the tracking issue, and the loop is just a heartbeat.
The tracking issue mattered beyond crash recovery. Its body is the live state of the run. Its comments are an append-only event log. There's an operator slot on it where I can leave a message that the next relaunch will read, which is how I steer a headless run without stopping it. And a separate read-only monitor skill narrates the run off that issue in plain English, so I can check on it from my phone without parsing a board.
The tracking issue has already paid for itself. When a run went off the rails one night this summer and the injected-ticket count started climbing instead of falling, I could see the whole run from the issue, and I caught it that night instead of a week later. v5 didn't prevent that one, but it's the reason I saw it.
The Layer That Didn't Exist: Engineering Discipline
Everything so far is about what to build and how to test it. The last piece of v5, finished at the end of June, is about how to think, and it exists because of two specific mistakes.
In the first, a refinement spec asserted that the UI captured a value "from the POST response handler." It read as fact. It was wrong. The POST payload type never carried that field; it only existed on a different GET response. A second pass that grepped the actual type collapsed the claim in about two minutes. In the second, a triage diagnosed a 500 error as "missing blob storage," which would have sent someone off to edit infrastructure that was already correct, while the real exception sat in the runtime logs the entire time. Pulling the stack trace pointed straight at the render code and ended the theory.
Same failure both times. The model saw a symptom, came up with a story that explained it, and wrote the story down as fact without checking the source. Confident and wrong is worse than "I don't know," because somebody goes and acts on it.
The fix is a new standard with five rules, ED-1 through ED-5, that every skill now cites the same way it cites the TR rules. Grounded claims: any statement about how the code behaves gets traced to a file and line before it's written as fact. Adversarial self-review: treat your own just-produced output as a suspect to cross-examine, not a product to defend. Hypothesis labeling: anything you can't ground right now gets written as a hypothesis with the evidence needed to settle it. Scope-fork detection: if the approach you picked changes the blast radius, say so. And the cold read, ED-5, which is the one I'd been doing by hand for months.
Regular readers will recognize the cold read. It's the Joseph Maneuver, the trick my friend Mike taught me of asking "what am I missing?" until the answers turn trivial. ED-5 makes it mandatory and makes it terminal: once a ticket is saved, the skill hands it to a fresh subagent that gets the ticket text and the repo but not the author's reasoning, and asks one question, what's missing. You withhold the reasoning on purpose. Hand a reader your reasoning and they inherit your blind spots and hand them right back. The findings come back, the author confirms or rejects each one against source, and only then does it report done. It's positioned as the last thing the skill does, because I'd watched instances run the silent self-review, update the issue, say "done," and let the internal pass stand in for the one I actually wanted.
One more line from that standard, because there's an incident behind it: "Memory is a claim, not an authority." An instance had saved a note that two orchestration loop processes running at once meant something was wrong and it should halt. The driver had since changed. The note was stale. The instance acted on it anyway and killed a healthy run. So now the rule is that a recalled memory is background, not instruction. Before acting on one, re-ground it against the current source. If the memory and the code disagree, the code wins.
v6: More Than One Machine
v6 merged at the end of September. It lets several orchestrators run against one repo at the same time, each in its own clone, normally one per computer. That's all it does.
The design is dumb on purpose, same as the loop. There's no shared pool of tickets and no coordination between runs. A run is identified by its tracking issue number, given explicitly. You either start one with a ticket list or continue one with a run number; the orchestrator never goes looking for work on the board, and the old "no arguments means do everything in Up Next" mode is gone on purpose. A small read-only skill takes a ticket list and a number of machines and hands back the batches, following a few rules: a dependency never crosses batches, a milestone goes last in the batch that holds all its stories, tickets that will touch the same files stay together. I review the split before anything launches. The only shared-state problem the design had to solve was a red main, since both runs merge to it, and the answer there is that whoever has an open fix PR is the fixer and everyone else waits.
Getting there took seven weeks, which is its own story, because the version I designed first was a shared pool with claim handshakes and worker liveness and a reaper for dead workers, and it went through three cold reads that each came back with more open questions than the one before. I'm the king of solving problems I don't have.
Where It Stands
Six months ago the framework was five phases and a pipeline. Today the phases are the same, the pipeline is longer, and there are three sets of numbered rules layered over it that didn't exist in March. The standards cover what to build, the Test Rules cover how to test it, and the ED rules cover how to reason about it. Each rule has an ID, and the skills cite the ID instead of restating the rule, because restating is how drift starts.
Every one of those rules has a specific failure behind it, and you just read them. None of this was designed up front. It got discovered one incident at a time and written down before I forgot it.
In March I wrote that most of the failure modes people experience in agentic development aren't limitations of the AI, they're architecture problems. I still believe that. What I didn't see then was that the architecture I'd keep having to fix was the process, not the app. The model is more capable than it was six months ago, and that has barely mattered. What mattered was that every time the process failed, I had a place to write down why, and a mechanism to make sure the next session read it.
So that's what the framework is now. Less a pipeline than a place to write down what went wrong, and a mechanism that makes the next session read it. Given who keeps causing the failures, I need that more than I need a smarter model.
Kevin Phifer is the founder of Theoretically Impossible Solutions LLC, specializing in agentic AI development and consulting. You can reach him at kevin.phifer@theoreticallyimpossible.org.