Kratos Labs · Field note
Never Let an AI Coding Agent Loop Over Its Own Reasoning
Long AI coding sessions can drift when one agent plans, builds, rewrites the tests, and judges its own work. This workflow separates those roles with fresh context, frozen contracts, and read-only evaluation.
This article expands a post first published on X.
Long AI coding sessions often fail for a simple reason: one agent gets too many jobs.
The same context plans the task, chooses the architecture, writes the code, edits the tests, and decides whether the work succeeded. Every stage inherits the conclusions, assumptions, and blind spots of the stage before it.
Planning can look excellent because it begins with clean context. By implementation, that context contains discarded directions, stale assumptions, and a growing defense of its own earlier choices.
The problem is not that AI agents loop. The problem is that the same agent loops over its own reasoning.
A long context is not an independent reviewer
An agent can change a decision halfway through a session and make the change sound reasonable. It can adjust the architecture to fit the code it already wrote. It can weaken a test that blocks completion. It can quietly redefine done.
Each move may look defensible in isolation. Together, they move the goalposts.
The final review is then performed by the same context that created the implementation and remembers why every compromise felt necessary. That is not independent evaluation. It is self-approval with a longer transcript.
Better prompting does not fix the structure. Separation of duties does.
The loop should persist through artifacts, not conversation
The reliable loop is:
observe-> plan-> freeze-> build-> falsify-> acceptEach role starts in fresh context. Each handoff happens through inspectable artifacts: a repository snapshot, a frozen task contract, protected tests, a focused diff, a running system, and an evaluation receipt.
Continuity still exists. It moves out of chat history and into files, tests, and commits.
1. Snapshot the repository before anyone touches it
Start by recording the state the task inherits:
- current architecture
- current failures
- current Git state
- verified build and test commands
- known assumptions and unknowns
Without that baseline, the agent may treat an old failure as part of its assignment. It may rewrite architecture it never understood. It may report a clean result because it never recorded what clean meant before the work began.
A compact baseline can be enough:
task: add export commandcurrent architecture: cli -> service -> databaseexisting failures: integration/export-timeout.test.tsverified commands: npm test, npm run typecheck, npm run cli -- --helpprotected surfaces: cli exit codes, export format, database schemaopen assumption: large exports may exceed the current request timeoutThe snapshot is not bureaucracy. It prevents the implementation from rewriting the starting line.
2. Freeze a decision contract after planning
Planning should end in a small contract that defines the work:
- objective
- non-goals
- architecture constraints
- protected tests and user-facing behavior
- definition of done
- verification commands
- known assumptions
Once the contract is frozen, the builder cannot change it.
If implementation reveals that the contract is wrong, stop the build and reopen planning. Produce a new contract as a deliberate revision. Do not let the builder edit the target while trying to hit it.
This boundary matters because implementation pressure is real. The closer an agent gets to a difficult edge case, the more attractive a smaller definition of success becomes.
3. Protect the tests that define success
Builders should be able to add tests. They should not be able to rewrite the core tests that define acceptable behavior.
Otherwise a test failure becomes an invitation to change the test instead of the code.
Sometimes a requirement genuinely changes. That still belongs in a new planning decision, with the contract and protected test set revised together. It should never happen as an unreviewed side effect of implementation.
The important distinction is simple:
- adding a regression test strengthens the boundary
- weakening an acceptance test moves the boundary
Only the first belongs to the builder.
4. Give every role fresh context
Fresh context does not mean throwing away useful information. It means deciding what information each role is allowed to inherit.
The planner receives the repository snapshot and the task.
The builder receives the frozen contract and the repository. It does not need the entire planning conversation or every abandoned direction.
The evaluator receives the contract, the diff, and the running system. It does not receive the builder's rationale.
That last omission is deliberate. The evaluator should inspect what the system does, not be persuaded by the story of how it was built.
Role isolation removes a subtle source of bias. A fresh evaluator has no identity invested in the implementation and no memory of the compromises that produced it.
5. The builder produces a candidate, not a verdict
The builder should never declare victory. It can only produce a candidate.
A separate evaluator decides whether that candidate satisfies the frozen contract. The evaluator should be read-only:
- it cannot fix the code
- it cannot weaken the tests
- it cannot redefine done
- it can only return
ACCEPTorREJECTwith exact reproduction steps
When the evaluator rejects a candidate, the rejection becomes a new input to the build. It does not become an invitation for the evaluator to patch the code and then approve its own patch.
This keeps authorship and judgment separate.
6. Falsify the real user-facing surface
Lint, types, and unit tests are necessary. They are not the whole product.
The evaluator should try to break the surface a user touches:
- run the real CLI and inspect exit codes, stdout, and stderr
- call the real API with valid, invalid, and unexpected inputs
- use the browser and exercise actual state transitions
- verify persistence after reloads, retries, and interrupted operations
- test the boundaries between components, not only each component alone
Green test suites can hide broken exit codes, newline injection, stale browser state, and inputs nobody thought to encode in a fixture.
The evaluator's job is not to confirm that the implementation looks sensible. Its job is to find the cheapest reproducible reason to reject it.
7. Turn every rejection into a ratchet
A rejection should improve more than the current patch.
bug escaped -> add a regression testarchitecture drifted -> add a mechanical boundary checkinvalid state appeared -> encode an invariantmanual mistake repeated -> add an automated gateFixing one instance is rarely enough. The workflow should become less able to repeat the same class of mistake.
This does not mean adding a rule for every imaginable risk. Let real failures earn their controls. A narrow check attached to an observed failure is more valuable than a giant speculative framework nobody trusts.
8. End every accepted slice in one atomic checkpoint
An accepted slice should contain:
1 behavior1 frozen contract1 focused diff1 evaluation receipt1 commitThen kill the context.
The next slice starts from the committed repository, not from a conversation that has been running for hours. If something later fails, the system has a clear rollback point and an exact record of what was accepted.
Small checkpoints also make drift visible. A commit that tries to change three behaviors, two architectural decisions, and the test strategy is not an atomic slice. It is a new planning problem pretending to be implementation.
What fresh context does and does not solve
This workflow does not make an AI model trustworthy.
It does not guarantee that the frozen contract is correct. A bad contract can produce a consistently wrong result. It does not help if the evaluator cannot run the real system. It does not replace human judgment for product, security, legal, or business decisions.
What it does is make failure easier to detect and harder to hide.
The builder cannot move the goalposts. The evaluator cannot repair the candidate it is judging. Tests cannot quietly shrink to fit the implementation. Every accepted change has a contract, a diff, a receipt, and a checkpoint.
The goal is not perfect autonomy. It is bounded failure with evidence.
A compact operating checklist
Before letting an AI coding workflow run unattended, verify that:
- the repository baseline is recorded
- the task contract is frozen
- the tests that define success are protected
- planning, building, and evaluation run in fresh contexts
- the builder can produce only a candidate
- the evaluator is read-only and tests the real surface
- every rejection creates a regression test, invariant, or gate when the failure class warrants it
- every accepted slice ends in one focused commit
If one of these boundaries is missing, the loop can still run. It simply cannot prove that it stayed on task.
Trust is the wrong target
You cannot prompt an agent into becoming a reliable judge of its own work.
You can build a system where it is never allowed to be its own final authority.
That is how long AI coding sessions become operable: not by extending one conversation forever, but by splitting the work into independent roles connected by evidence.
The broader reliability layer, including history, narrow checks, and human approval, is covered in Make the company workspace reliable.