Kratos Labs · Field note
Make the company workspace reliable: checks, history, time, and human approval
Make your AI company OS reliable with narrow checks, full-history Git gates, result receipts, multi-host discovery, bounded workers, and exact-version human approval.
Part 3 of 3. Part 1 gave every fresh AI session a home. Part 2 gave one repeated job a recipe. This is the layer that keeps that job trustworthy as the system grows.
It is Monday morning. The weekly update for Example Client is supposed to be ready.
The schedule says the job ran at 09:00. The current client folder looks clean. The draft reads well.
Then the owner notices the gaps.
An earlier saved revision said Example Client had moved into delivery, even though that label was retired. A later revision removed it, so the current folder looks fine. The scheduler recorded that something started, but nobody can prove that the right sources were read or that a valid update was produced. A second AI host could not find the weekly-update recipe. The owner remembers approving the draft, but nobody can show which saved version that approval covered.
Nothing here requires a bad model. A workflow can look successful while its evidence is missing.
Part 1 gave every model a shared company workspace. Part 2 turned the Example Client update into a repeatable job with one official status page, dated evidence, a written recipe, a predictable draft, and owner review.
Part 3 closes the loop. It asks the system to show what it used, what it produced, what it checked, what remains uncertain, and where human authority stops it.
Reliability is evidence, not confidence
For a company workspace, reliable does not mean flawless. It means a run can answer five ordinary questions:
- What sources did it use?
- What did it produce?
- What check ran?
- Did the check pass, find a problem, or fail to run?
- Which exact version is waiting for human approval?
The shape is small:
request -> sources -> draft -> check -> saved version -> owner approvalThat chain turns a polished answer into a traceable work product. Keep Monday's Example Client update in view. The owner wants it prepared on a schedule. A second AI host should find the same recipe. A research worker and a writing worker may work at the same time. Several saved revisions may exist before anything leaves the machine.
Each capability adds a new way to look successful while being wrong. The right response is not a giant process. It is one small proof for one failure that already happened.
Start with the failure a check can catch
The first failure is the retired client stage.
An old Example Client draft called the stage delivery. The current status page says that label is no longer valid. Correcting the sentence in one chat is not enough. A new session could find the old wording in another saved surface and repeat it.
Build a narrow, read-only check for this one problem. Let it inspect only the official current surfaces for the weekly update: the Example Client status page and the weekly-update recipe. The dated work note is historical evidence for a reporting period. It may legitimately preserve an old label, so it is outside this narrow current-surface check. The checker must never rewrite either current source.
The check needs three outcomes:
- Clean: the retired stage is absent from the surfaces it owns.
- Blocked: the retired stage is present.
- Unknown: the check itself could not complete, so it has no result.
Unknown is not a softer clean result. If the checker cannot read a required file or its own rules are damaged, the workflow must stop and say verification is unknown. A broken measuring instrument cannot report a safe room.
Build it from the observed failure:
- Name the retired stage precisely.
- Name the exact current surfaces the check may inspect.
- Create a disposable known-bad example containing the retired stage.
- Confirm the check blocks the bad example and stays clean on the clean one.
- Confirm neither run changes the official source files.
For Example Client, the pass condition is concrete: the bad example turns red, the clean example passes, and the status page remains unchanged. The fail condition is also concrete: the old stage passes through, or the checker breaks and the workflow calls that clean.
The limit matters. A text check can catch a retired label in named files. It cannot decide whether the update is strategically useful, whether an open risk is described well, or whether the owner should send it. A narrow check earns trust by saying exactly what it cannot judge.
That boundary also tells you where to start. Do not scan every file because a universal checker sounds impressive. Start with the named current surfaces where the retired stage could return. Widen the boundary only when another real failure proves the first one was too small. Save the check result in the weekly update's receipt, so a later reviewer can see what was measured instead of seeing only a green label with no scope.
Put two doors around Git history
The retired stage creates a second failure. Imagine two saved revisions:
Revision A: Example Client is in delivery.Revision B: The retired stage is removed.Revision B makes the current folder look clean. Sharing the history can still share Revision A. The bad statement remains part of what leaves the machine.
Think of two doors. The first checks the next saved snapshot before it enters history. The second checks every new saved revision before that history is shared outside the machine.
Now the Git hook names make sense. The first door is a pre-commit check. It examines the staged snapshot about to be saved. The second is a pre-push check. It examines every commit that would be transferred, not only the latest folder.
That difference is the whole point. If Revision A contains the retired stage and Revision B deletes it, a tip-only check sees B and misses the unsafe history underneath it.
Use this acceptance test:
Create a disposable history where Revision A contains the retired stage and Revision B removes it. The current folder is clean. The share gate passes only if it still catches Revision A.
There is one more proof question. A tracked configuration file can show that a guard is intended to exist. It cannot prove that this computer is using it. Run one known-bad red test locally. The guard is active only when the machine blocks the bad case.
These doors do not decide what is true about Example Client. The status page still owns current truth. Git shows how the page, recipe, and draft changed. Current truth and historical evidence have different jobs.
That distinction prevents a common cleanup mistake. Removing a sentence from the current folder can correct today's view without changing what an earlier revision contains. The history gate is not asking you to pretend the earlier mistake never happened. It is asking you to know whether that revision is still part of the history you are about to share. A correction becomes trustworthy when the current source is fixed, the unsafe revision is detected, and the next saved version carries the corrected rule.
Make Monday produce a result, not a timestamp
Return to 09:00. A schedule record proves that something started. It does not prove that the Example Client status page was read, that the reporting week's evidence was found, or that a usable draft exists.
Keep dates beside the fact they govern:
**Due:** 2026-09-15 - review the experiment and record the outcome. **Stale after:** 2026-09-30 - verify this research again before relying on it.The first marker says work needs attention. The second says a claim may be old enough to require a fresh check. A generated clock can read these markers across owner files and show what is due, stale, or waiting for review. It is a view over the source files, not a second place where the team maintains the same commitments.
The Monday job needs a result receipt. Its proof should name the request, the sources, the output, the checks, and the result:
requested jobsources usedoutput createdchecks completedresult: ready for owner review, blocked, or unknownFor Example Client, that receipt might name the current status page, the dated work note, the draft path, and the retired-stage check. The dated work note can be a source used by the job, but it is not an input to the retired-stage check. The proof marker advances only after the result is verified. A timer firing is never its own proof.
Test the distinction deliberately. Create a past-due marker and confirm the clock makes it visible. Then make a disposable weekly-update attempt with a required source missing or no output created. Pass means the receipt reports blocked or unknown, and the proof marker does not advance. Fail means the schedule says the update is ready because a timer fired.
The receipt should also make recovery ordinary. If the status page was unreadable, the next action is to restore access and rerun the check. If the dated note was missing, the result stays blocked until the evidence exists. If the draft was written but validation failed, the owner sees a draft waiting for correction rather than a false ready signal. A useful receipt narrows the next question after a bad Monday.
The clock has a limit. It can aggregate structured markers. It cannot infer every important date buried in prose, and it cannot decide whether a deadline was wise. Structure makes time visible. It does not replace judgment.
Give every AI host a signpost to one recipe
The next Monday failure is quieter. Claude finds the weekly-update recipe. A second file-aware AI host sees the same folder but misses it. That host improvises from whichever files it notices first.
Keep one canonical recipe. Give each host a tiny signpost:
Trigger: weekly client updateFull procedure: skills/weekly-client-update/SKILL.mdThe signpost helps a host find the job. The recipe owns how the job works. A script, when needed, performs a fixed check or transformation. Do not copy the complete recipe into every host's instruction file. Copies drift, and a new stop condition can reach one host while another keeps the old procedure.
For Example Client, the acceptance test is the same request in two fresh hosts:
Both hosts must reach the same recipe, read the same official sources, save within the same output boundary, and stop at owner review.
Matching signposts prove the intended setup. Only the live two-host test proves that each local host actually discovered and followed it. The second host is useful because it exposes assumptions the first host allowed to stay hidden.
Bound parallel workers and review from fresh context
The owner now wants a stronger Monday update. A research worker should inspect the week's evidence while a writing worker prepares the draft.
If both workers receive the broad instruction help with the client update, they can edit the same draft, overwrite unfinished work, or stage files owned by the other worker. More workers do not create more capacity when ownership is unclear.
Give each worker a brief with one result and one territory:
Objective: one concrete resultOwned file or responsibility: one exact, non-overlapping areaSources: the minimum official filesDo not touch: every other worker's filesDone means: one checkable resultReturn: artifact path, checks run, and unresolved questionsFor Example Client, the research worker can own an evidence note listing supported completed-work claims and open questions. The writing worker can own the weekly-update draft and use only that evidence note, the official status page, and the recipe. Neither worker edits the other's file. The main session reconciles the two artifacts before they enter history.
This is a procedural boundary because the workers still share a working tree. File ownership, explicit paths, and exact staging keep the territory visible. The system does not claim magical isolation.
Independent review is a separate job. Give the reviewer the draft and the official sources in a fresh context. Do not give it the writer's defense of every choice. The reviewer must reconstruct the claims from evidence.
Use a known omission as the acceptance test:
Plant one deliberate omission in a disposable draft. The reviewer passes only if it finds the omission without being told where it is.
For Example Client, omit one open risk from the draft. A useful reviewer compares the draft with the status page and names the gap. It does not merely repeat the writer's rationale. The main session still decides which objection applies. Review is evidence, not automatic permission to rewrite everything.
The same separation applies to the reviewer's scope. It can check whether a completed-work claim has dated support, whether the open risk survived into the draft, and whether the saved artifact is the one under discussion. It does not become the client owner, invent a missing fact, or authorize a send. Independent review is valuable because it adds another reconstruction of the evidence, not because it transfers authority to another model.
Keep quality separate from authority
The Example Client update may now have the right sources, a clean retired-stage check, a result receipt, a history gate, and a fresh-context review with no unresolved objection.
It still must not send itself.
Approval belongs to one exact saved version. If the owner approves a draft and someone changes one sentence later, that is a new version and needs a new approval decision. The record should name the artifact and its saved revision, not only a reusable filename.
The authority boundary is simple:
- Reading and diagnosis can happen inside the requested task.
- Reversible workspace work can happen inside the agreed boundary.
- External communication requires explicit approval for the exact artifact.
- Financial or irreversible action requires stronger verification and explicit approval.
A clean verification result is not send authority.
Checks can show what they measured. A reviewer can find a missing claim. Neither can decide whether the timing, relationship, or commercial judgment behind a message is right. The owner approves the version that leaves the company.
Be honest about what is live
The technical receipt supports a live Kratos layer with an operating contract, owner files, canonical skills, Claude and Codex discovery surfaces, decision, error, and tombstone ledgers, structured time markers, a non-mutating repository checker, commit and push gates, bounded workers, and mandatory fresh-context review for selected high-risk artifacts.
Some parts are deliberately narrow. Mechanical checks cover named surfaces, not every semantic error. Git can compare tracked configuration, but it cannot fully prove local hook activation without a red test. Parallel workers still share a working tree, so ownership remains partly procedural. Schedules are narrow, and a start record is not a result. The full planned negative-fixture pack and one composite quality command are not complete.
Other layers are not built by choice or by current need. There is no knowledge graph, generalized workflow engine, web dashboard, or claim that every business lane runs unattended. The Monday Example Client updater in this story is an implementation pattern with acceptance tests. It is not a claim that every scheduled workflow already succeeds without the owner.
Those limits are part of the proof. A mechanism is live only when the receipt supports it. A partial mechanism says what it covers and what it cannot prove. An unbuilt mechanism stays out of the present tense.
Monday, replayed
The owner asks for this week's Example Client update in ordinary language. The mature loop is:
ordinary request-> official sources-> one canonical recipe-> bounded worker-> predictable draft-> narrow checks-> independent review when risk earns it-> exact saved version-> human approval-> Git history-> correction improves the next runThe host reads the operating contract and official client sources. One recipe defines the work. The exact saved version waits for the owner. Only then can an approved artifact leave the company.
Part 1 made the company readable by a fresh model. Part 2 made one useful job repeatable. Part 3 makes the growing system honest about what ran, what it checked, what remains uncertain, and where a human must decide.
The model is replaceable. The company layer is cumulative.
Reliability does not mean removing the human. It means the system knows what it can do, what it can prove, and where it must stop.
Take the Part 2 workflow and add one red test for one failure that has already happened. Do not build the rest until the work earns it.
Follow for future receipts from real use: the files that changed, the checks that caught something, the procedures that failed between models, and the complexity that was not worth keeping.
Previous: Part 2 - build one repeatable AI workflow.