Kratos Labs · Field note
Inside the Git Repository That Runs Most of Our AI Agency
A practical look at the file-based operating system behind Kratos Labs, from lead research and planning to website operations, coding, client documents, and correction.
A Git repository runs most of Kratos Labs.
That sounds like a claim about automation. It is really a claim about where the agency's operating context lives.
On X, we put it more bluntly:
this git repo runs 90% of my ai agency
The percentage describes coverage of the operating layer. It does not mean an agent owns every decision or clicks every button. Human judgment, relationship actions, and public approval remain human. The repository owns the context, workflow, checks, and record around them.
That difference is what turns a collection of AI tools into an operating system.
The repository is the control plane
Most AI stacks are organized around apps. Ours is organized around state and responsibility.
The repository gives each layer an explicit job:
knowledge/contains curated business truth and names the owner of each live fact.raw/quarantines unverified material until it has been checked.outputs/holds research, briefs, proposals, invoices, and other deliverables.ai/records decisions, errors, and retired claims that must not quietly return.skills/andscripts/turn recurring work into procedures that different agents can execute and verify.
If you want the foundation underneath this structure, start with how we built an AI company OS. The useful distinction here is what happens after the folders exist: the repository starts running the agency's recurring work.
What it runs in a normal week
1. It turns Reddit noise into usable demand signals
The pain-signal pipeline scans Reddit across the niches we care about, applies a mechanical prefilter, removes duplicates, and routes the surviving posts through qualitative review.
One verified full scan looked at 340 posts. Ninety-one passed the mechanical filter, and 23 survived the evidence gate as real pain signals. A later daily delta scanned another 330 posts, found nine new candidates, and kept one. The volume changes. The standard does not.
The point is not to produce a large spreadsheet. It is to preserve the original buyer language, separate firsthand pain from vendor promotion, and give offer research a stream of evidence that can be challenged later.
2. It sources and qualifies 25 Sales Navigator leads per workday
The standing workday target is 25 screened Sales Navigator leads. The system sources candidates, checks decision-maker status and onboarding relevance, removes duplicates, and writes the finalized batch to the tracker.
The human boundary is explicit: the repository can prepare and record the batch, but we still perform every LinkedIn connection action. Bookkeeping is not send authority.
That boundary matters because a useful operating system should make the next action easier without pretending that every relationship action belongs to an agent. The same principle appears throughout our repeatable workflow design.
3. It maintains the agency website
Every Monday, the SEO control plane checks the site, Search Console, and GA4. It looks for crawl failures, indexing problems, broken measurement, and technical regressions.
Verified technical fixes can be implemented, deployed, and checked in production inside that scope. Material positioning and new public editorial copy remain approval-gated. The system can own the machinery without silently becoming the publisher.
4. It prepares us for booked calls
A new booking notice is a trigger, not another inbox item to remember. The system detects the booking, puts it on the agenda, and assembles a prospect brief from the person's company, role, public activity, likely priorities, and relevant conversation angles.
The result is a prepared call with traceable research, not a last-minute search across five tabs.
5. It mirrors the daily plan to Google Calendar
The daily ledger remains the source of truth. Once the plan is agreed, each task becomes a time block in Google Calendar and is classified by motion, such as client work, outbound, content, or operations.
The calendar is a projection of the operating plan, not a second task database. That one-way rule prevents the plan, the calendar, and the accountability record from drifting apart.
6. It runs a six-phase research doctrine
Important research does not begin with an open browser and a vague request to find everything. The repository requires a declared brief, evidence gathering, synthesis, adversarial challenge, a decision draft, and market testing when the decision needs it.
Claims receive confidence labels. Sources are judged for directness, freshness, scope, and independence. Contradictory evidence is preserved instead of averaged away.
This produces slower research than a quick answer and much faster decisions than repeatedly reopening the same question.
7. It keeps long coding sessions from collapsing into one context window
Long implementation sessions are divided into scoped work, verification, fresh-context review, and explicit handoff artifacts. The next agent receives the repository state and the evidence, not a compressed story about what the previous agent thinks happened.
That protocol lets work continue across days while reducing the chance that one mistaken assumption contaminates every later pass. We documented the detailed pattern in our fresh-context AI coding loop.
8. It turns corrections into memory
When we catch a factual error or a stale statement, the fix is not complete when one sentence changes. The system records what failed, identifies the fact's real owner, removes stale copies, and adds a durable control when the failure can be checked mechanically.
A retired number can become a tombstone. A recurring structural mistake can become a repository check. A judgment error can enter the decision or error ledger with a review condition.
the same mistake does not get to quietly come back.
This correction loop is the practical core of making a company workspace reliable. Reliability is not the absence of mistakes. It is the ability to make mistakes visible, repair them, and raise the floor afterward.
9. It generates and tracks client documents
Proposals, invoices, RFP responses, and related client documents are generated from branded templates and live owner files. Validation checks the finished artifact, not only the source text.
The corresponding records track what was issued and what happened next, including acceptance, rejection, and payment when those facts are observed. A document does not disappear into a downloads folder after it is sent.
What makes this an operating system
The interesting part is not that an agent can draft a proposal or scrape a website. Individual tools can do both.
The operating-system behavior comes from three contracts:
- Every live fact has one owner. Other files point to it instead of copying it.
- Every consequential action has an authority boundary. Preparing, approving, sending, publishing, and paying are different permissions.
- Every repeatable failure leaves a durable artifact, such as a rule, test, tombstone, or review date.
Those contracts let many agents work on the same business without treating the latest chat as truth. They also make the system inspectable. We can see what it believed, what it produced, who was allowed to act, and which check passed before the result moved forward.
What the open-source version includes
We open-sourced the reusable system behind our private production instance as Operator OS.
It includes the file contracts, workflow patterns, skills, ledgers, checks, and fictional worked examples needed to adapt the approach to another business. It does not publish our client data, credentials, or private operating state. You connect your own email, calendar, analytics, research sources, and sales tools, then set the approval boundaries that fit your work.
Operator OS is free and MIT licensed.
You can get Operator OS on GitHub.
Start with one workflow
Do not try to automate nine departments on the first day.
Write down the business facts one workflow needs. Give each fact one owner. Define the output, the approval boundary, and the check that proves the run succeeded. Then execute it repeatedly until the record is trustworthy.
Add the next workflow only after the first one can survive a new agent, a new context window, and a correction.
Kratos Labs builds governed AI systems for organizations moving from pilots into daily operations. Operator OS is the public, reusable starting point for teams that want to build the operating layer themselves.