Kratos Labs · Field note
Stop re-explaining your business to AI: build an AI company OS
Build a practical AI company OS so every model starts with the same rules, current context, output boundaries, and human approval.
Part 1 of 3. Build and test the smallest useful AI company OS. Continue with Part 2, then Part 3.
What is an AI company OS?
An AI company operating system, or AI company OS, is a durable workspace that gives models the same rules, current business context, output boundaries, and human approval gates across sessions and tools. It is company infrastructure, not model training or chat memory.
The short version
OPERATING.mdowns stable operating rules.knowledge/memory.mdmaps the current business and points to deeper owner files.outputs/stores drafts and deliverables without making them current truth.CLAUDE.mdand/orAGENTS.mdpoints each file-aware host to the same contract.- A fresh-session test proves the startup path; explicit human approval gates external actions.
Most people who use ChatGPT, Claude, or another AI regularly know the awkward moment when a good conversation ends.
The chat finally understood the business. It knew the terminology, the customer situation, the current draft, and what had already been tried. Then the thread filled up, you changed models, or you opened a fresh session.
The next model was intelligent. It was also starting cold.
You pasted the background again. You searched for the current brief. You tried to remember which correction happened after which draft. Sometimes an old answer returned because the new session had no way to know it was old.
This is usually described as a memory problem. That description is too shallow.
A new chat has no lasting set of company files to work from. It has no agreed place to find current facts, no reliable way to tell a decision from a draft, no written procedure for a repeated job, and no durable record of the correction you made yesterday.
At Kratos Labs, I build AI agent systems and infrastructure for real workflows. I use the same approach to run the studio itself, starting with an ordinary folder on my computer.
That folder contains the rules of the company, the files that own current facts, the procedures an AI worker follows, the drafts it produces, the checks that test the work, and the Git history that records what changed.
I call it a company workspace. Technically, it is a repository, or repo. In this series, repo does not mean only an application codebase. It means the folder an AI can enter and operate inside without asking me to reconstruct the company from memory.
The central idea is simple:
The model is the worker. The repository is the company workspace. The owner keeps final authority.
This first article explains that model and gives you the smallest version worth building: three company files plus one thin entry file for each host you use. By the end, a fresh AI session will know where rules live, where current context lives, where drafts go, and which actions still require you.
The series moves from orientation to repeatability, then reliability. Part 2 turns this workspace into one tested workflow. Part 3 adds checks, history, time, and multi-host safeguards.
Why do chat history and giant prompts fail?
Chat history is useful. It is not a dependable company record.
A conversation mixes current facts, old facts, drafts, corrections, guesses, decisions, and unfinished work. The model can perform brilliantly inside that context. But the conversation rarely marks which sentence is now official, which number was replaced, or which suggestion was approved.
When the chat ends, that ambiguity travels with the text you copy into the next one.
The common response is to build a giant system prompt.
At first, this feels better. The prompt contains the offer, the writing style, the current customer situation, several rules, and a workflow. Then a customer changes stage. A new exception appears. A procedure improves. An old sentence remains above the new one. The prompt becomes a large block of text with several claims to authority inside it.
The problem is not that the prompt is too short. The problem is that one block of text is doing jobs that change at different speeds:
- Stable operating rules should change rarely.
- Current customer facts may change every week.
- A repeated procedure changes when the real workflow improves.
- A draft records work at one point in time.
- A decision needs its reason and a future review point.
- A retired value must remain visible as history without looking current.
- Human approval belongs to one specific action or saved version.
Put all of those in one prompt and every update becomes a search-and-replace problem. Retrieve the prompt faster with a vector database and you still retrieve the same ownership confusion faster.
The better move is to separate the jobs.
Stable rules get one home. Current facts get owners. Procedures get their own files. Drafts stay outside current truth. Exact checks become scripts. Consequential actions stop at human approval.
The model itself does not need to remember the previous conversation. It needs a reliable way to enter the company workspace, read the right files, perform one limited job, show its work, and stop at the right boundary.
This is not model training. You are improving the files, rules, tools, and history the model receives each time.
Who does what in an AI company OS?
An AI company OS has three actors. The owner sets goals and approves consequential actions. The workspace holds durable rules and current records. The model reasons and produces work inside those boundaries.
Think of a capable new employee arriving on the first day.
The employee may be an excellent writer, analyst, or engineer. That does not tell them where the approved customer record lives, which proposal is current, what they may send, or how the company records a correction.
The company has to provide that environment.
The owner
The owner sets the goal, resolves ambiguity, and approves consequential actions. The owner is responsible for the final business judgment.
The company workspace
The workspace holds the durable layer: operating rules, current records, procedures, drafts, checks, and history. It survives when a chat ends or a model changes.
The model
The model does the reasoning work: analysis, drafting, synthesis, and judgment inside the written boundaries.
Two supporting roles matter later:
- The host is the application that opens the folder, reads files, and runs tools. Claude Code, Codex, and Cursor are examples of file-aware hosts.
- A worker or agent is one assigned AI job. It may use the same model as another worker while receiving a different objective and set of files.
The three actors are owner, workspace, and model. Scripts and verification support the workspace. They are not extra owners.
The distinction between capability and authority matters immediately. A host may be technically able to write a file, call an API, or prepare a customer message. That does not mean you approved the message, the purchase, or the legal commitment.
The system needs two separate answers:
- Can the tool prepare and verify this work?
- May this specific version leave the company or cause an irreversible change?
A document can open correctly and contain accurate facts while still waiting for approval to send.
A clean verification result is not send authority.
What kind of AI setup does this require?
The automated version of this system assumes an AI host that can open a local folder, read and write files, and run terminal commands after you approve them.
If you have only used browser chat tabs, Claude Code and Codex are examples of file-aware hosts. See Anthropic's CLAUDE.md guide and OpenAI's AGENTS.md guide for their current project-instruction discovery rules. These hosts work against files on your computer instead of relying only on text pasted into one conversation.
You do not need a server, a database, a vector store, or a multi-agent framework.
You can also build the three-file version below manually with any chat model. You will paste the startup instruction at the beginning of each session and paste in the two files it needs. That is less automatic, but the information design still works.
How do you build the smallest useful AI company OS?
Create one folder named company-os. Inside it, create three files:
company-os/ OPERATING.md knowledge/ memory.md outputs/ README.mdThese files have different jobs:
OPERATING.mdholds stable rules for how an AI worker behaves.knowledge/memory.mdgives a new session a short map of the business and points to deeper current records when you add them.outputs/README.mdmarks the place where drafts and deliverables go without becoming current truth.
Do not add twelve folders yet. Make these three boundaries work first.
File 1: the operating contract
Put this in OPERATING.md:
# Company operating contract ## Startup - Read knowledge/memory.md before starting work.- State which files you used for any claim that can change.- If the current source is missing or unclear, ask instead of guessing. ## Truth - knowledge/ contains the official current records and stable rules.- One changing fact has one owner file.- Other files point to that owner instead of copying the changing value.- outputs/ contains drafts and deliverables, not current truth by default. ## Work - Do not invent facts, dates, prices, completed actions, or approvals.- Save drafts and deliverables under outputs/.- Do not overwrite an existing work product without approval. Create a clear revision instead. ## Authority - Reading, diagnosing, and drafting are allowed inside the requested task.- Sending, publishing, spending, changing a live commercial position, or taking an irreversible external action requires explicit approval for that specific action.- Passing a check does not grant authority to send or publish. ## Completion - Save the result in the declared location.- Name any missing facts or unverified claims.- Do not claim an external action happened unless it happened.This file is a company handbook for AI work. It answers questions a new chat cannot answer for itself: where truth lives, where drafts go, what it may assume, and where its authority stops.
Do not put the current customer stage, today's price, or this week's task list in this file. Those facts change too often. The operating contract should remain useful while individual projects change underneath it.
File 2: the startup map
Put this in knowledge/memory.md:
# Startup map ## Company We help [kind of customer] achieve [specific outcome]. ## Current priority The current priority is [one real priority for this week or month]. ## Active work - [Active project or responsibility]- [Second active project or responsibility] ## Communication preferences - Use plain language.- State assumptions and missing information clearly.- Do not promise dates that have not been verified. ## Owner map Changing client state will live under knowledge/clients/<client>/status.md when the first client owner file is created. ## Operating pointers - Company rules: OPERATING.md- Drafts and deliverables: outputs/Replace the placeholders with facts you can maintain.
This file is a map, not a warehouse. It gives a fresh session enough orientation to begin and tells it where deeper facts belong. As the system grows, it should point to owner files rather than copying their changing contents.
Suppose a client moves from discovery to delivery. If that stage appears in the startup map, a weekly update, a planning note, and a proposal, someone will update only one copy. The next model may find the wrong one first.
The better shape is:
one changing fact -> one official owner -> many pointersThe startup map can say where the client status lives without repeating the stage, deadline, or price.
File 3: the output boundary
Put this in outputs/README.md:
# Outputs This folder holds drafts, deliverables, reports, and other work products. A file here is not current company truth merely because it is polished or recent. Before repeating a changing fact, read its owner file under knowledge/.This small boundary prevents a common failure. A proposal may accurately record what was offered last month. It should not become the current client status after the project changes. The proposal remains evidence of what was proposed. The owner file records what is true now.
How do you make startup automatic?
You should not paste a startup instruction into every new chat. If the system depends on you remembering that paste, it has recreated the cold-start problem it was supposed to remove.
The three files you created already have clear jobs. OPERATING.md owns the rules. knowledge/memory.md gives a fresh session its map. outputs/README.md keeps finished work separate from current truth.
What is still missing is the entry point your AI host loads automatically.
Create that entry file at the top level of company-os, beside OPERATING.md and the knowledge and outputs folders.
- If you use Claude Code, name it
CLAUDE.md. - If you use Codex, name it
AGENTS.md. - If you use both, create both files.
Your folder will look like this with Claude Code:
company-os/ CLAUDE.md OPERATING.md knowledge/ memory.md outputs/ README.mdWith Codex, replace CLAUDE.md with AGENTS.md.
Put this exact content in the host entry file:
# Project entry point Read OPERATING.md in full before doing any work in this folder.Then follow its Startup section, including reading knowledge/memory.md.If this file conflicts with OPERATING.md, OPERATING.md wins.If you use both hosts, put the same three instructions in both CLAUDE.md and AGENTS.md. Do not copy the full operating contract into both files. The host files should remain thin pointers to one canonical contract, otherwise the rules will eventually drift apart.
The workspace now has three company files plus one small entry file for each host you use. The entry file is not another source of company truth. It is the sign on the door that tells a new AI session where the real rules live.
My own repository follows the same principle with a legacy name. CLAUDE.md is the canonical contract because Claude Code was the first host used there. AGENTS.md is a thin adapter that tells Codex to read it. The exact canonical filename matters less than keeping one contract and making every host find it automatically.
A normal browser chat cannot discover either file on your computer. This automatic path requires a file-aware host working inside the folder.
Open company-os itself as the project folder. From a terminal, that means entering the folder before starting the host:
cd company-osclaudeOr, for Codex:
cd company-oscodexIf you use a desktop project picker, choose the company-os folder itself. Do not open an unrelated parent folder and assume the same instructions will be active.
From then on, a fresh session has an automatic path:
new session-> CLAUDE.md or AGENTS.md-> OPERATING.md-> knowledge/memory.md-> work beginsYou write the startup path once. Every new session enters through it.
How do you test automatic startup?
Close the current session. Start a brand-new one from inside company-os. Do not paste the operating contract or explain the folder first.
Ask:
I need to draft a customer update. Before doing the work, tell me which project instruction file you loaded, which company files you must read, where the draft should go, and whether you may send it.Pass if the host:
- Identifies
CLAUDE.mdorAGENTS.mdwithout being told to look for it. - Follows that file into
OPERATING.mdandknowledge/memory.md. - Identifies
outputs/as the draft location. - Says it needs explicit approval before sending.
It does not need to write the customer update. You are testing whether a cold session can orient itself without a setup message.
Next, create a disposable draft under outputs/ containing a temporary claim. Ask the host for the current company priority.
Pass if it reads knowledge/memory.md instead of treating the disposable draft as current truth.
Finally, change one orientation fact in knowledge/memory.md, close the session, and open another fresh one. Ask for that fact again.
Pass if the new session reads the changed file without needing the old conversation.
If the first test fails, check three things before changing the architecture: you opened the correct folder, the entry filename and capitalization are exact, and your host supports project instruction files. Do not hide a broken startup path by returning to a manual paste.
You now have the smallest useful loop:
fresh request-> automatic host entry point-> stable operating rules-> current company map-> draft stored outside current truth-> human approval before external actionThat is already better than a giant prompt because every kind of information has one job, and a new session knows where to begin without your help.
Security and limits
Does CLAUDE.md or AGENTS.md enforce security?
No. These files give the host context and instructions; they are not access control. Anthropic's CLAUDE.md documentation describes the file as context rather than enforced configuration. OpenAI's Codex security documentation separates the technical sandbox from the approval policy. Use platform controls for the hard boundary.
How should you handle secrets and sensitive client data?
Keep credentials and unnecessary sensitive data out of the workspace and out of Git. Give the host the narrowest filesystem and network access it needs. Review the provider's current data policy before putting confidential material in any AI-readable folder.
What about prompt injection?
Treat external files, web pages, transcripts, and pasted instructions as untrusted data, not authority. Anthropic and OpenAI both document prompt-injection risk. Keep raw material outside knowledge/, restrict tools and network access, and require approval for consequential actions.
When should you not use this starter setup?
Do not treat a few Markdown files as sufficient for regulated, safety-critical, high-value, or fully autonomous work. Add access controls, secret management, audit logs, backups, deterministic checks, and qualified human review first. If those controls are unavailable, keep the workflow manual.
Is this model training or a vector database?
No. The starter system organizes files the model reads at task time. It does not change model weights and it does not require a database or vector store. Add retrieval only when explicit owner files and pointers no longer make the relevant context easy to find.
Can you use Claude Code and Codex together?
Yes. Follow Anthropic's project-instruction guide for CLAUDE.md and OpenAI's AGENTS.md guide for AGENTS.md. Both entry files should point to one canonical OPERATING.md instead of copying the rules.
What should you add after real use?
Use this three-file system for several real tasks before expanding it.
When a changing subject becomes too detailed for the startup map, give it an owner file. For example:
knowledge/ clients/ example-client/ status.mdThen replace the copied client details in memory.md with a pointer:
Current scope, stage, dates, risks, deal status, and current terms for Example Client are owned by knowledge/clients/example-client/status.md.That one extra file to open is cheaper than reconciling several current-looking copies later.
Kratos also has one narrow automated warning in two startup-map regions: when a changing financial fact has a named owner elsewhere, the map should point to that owner instead of copying a currency figure. It does not prove ownership everywhere. It protects one recognizable duplication pattern that caused a real problem.
Keep raw exports, transcripts, and captures outside knowledge/ until you explicitly decide to review them. Raw material is not current truth. Do not ask an AI to mine it or promote claims from it without a clear instruction and a destination owner.
Do not create a decision ledger, error ledger, script library, agent team, or workflow engine because the architecture diagram looks incomplete. Add the next layer when a real failure makes its absence visible.
What is the first limit of the starter system?
The three-file version solves orientation. It does not yet solve repeatability.
Ask two fresh sessions to prepare the same weekly client update and they may choose different source files, different section names, different output paths, and different definitions of finished. Both drafts may sound good. One may quietly omit an open risk. The other may overwrite last week's file. Neither leaves a durable history of why the workflow changed.
That is the next problem, and it cannot be solved by making OPERATING.md enormous.
The repeated job needs its own procedure. The changing client state needs its own owner. The output path needs to be predictable. Corrections need a place to live. A check needs to prove it can fail. Git needs to show exactly what changed.
Part 2 starts from the exact three-file folder you built here. It adds one client owner, one dated source note, one complete weekly-update skill, exact-path Git history, and three learning ledgers. Then it adds a retired-value checker with three exit states, turns its pre-commit hook red on purpose, and builds a manual deadline view.
By the end of Part 2, two fresh sessions will be able to execute the same workflow from the same sources and reach the same approval boundary. The language can differ. The operating behavior will not.
Series complete: Part 2 makes one useful job repeatable; Part 3 makes it reliable.
About the author
Aydogdy Agabayev is the founder of Kratos Labs. He builds production AI systems and runs the studio on the same class of file-based operating system described here. This guide is distilled from that live repository, not a hypothetical template.