Local-first agent harness
Agents, end to end.
Tell it what you need done. It carries the follow-through — across your models, your apps and your machine — and hands you a receipt that shows it actually finished.
It asks before a consequential action leaves your machine, and it asks for that one exact action, not a blanket yes.
The outcome lifecycle. Each segment’s output is the next one’s input.
- 01declaredA contract exists. Nothing has been agreed.
- 02acceptedYou agreed to this exact version.
- 03runningWork is dispatched, owned, deduplicated.
- 04needs‑youOne durable question. Blocks only the task that needs it.
- 05verifyingA separate session reads the delivered bytes, and nothing else.
- 06completedReceipt written, bound to this contract version.
Scroll horizontally to see all six states.All six states are shown in order below.
failed cancelled rolled back The other three ways it ends. Each one says why.
The difference
Permission is the easy part. Completion is the product.
The current generation of agent tooling is astonishingly permissive. It will let you point a model at your inbox, your repository, your browser and your filesystem — and then it hands the last thirty per cent back to you: the assembly, the checking, and the job of working out whether what it produced is what you asked for.
Agent Centipede is built from the other end. The unit is not a prompt or a task or an agent run — it is a thing you wanted done, carried through to a finished result. And the run is not over until something other than the author has read what came out and held it against the definition you agreed to.
The contracts, the gates and the refusals below are not the pitch. They are why the finishing can be believed.
It allows you to do almost anything, which is awesome. It just doesn’t do the work end to end well.
On the tool this replaces
What it does
Between the ask and the receipt
The harness runs on your own computer and drives agent CLIs you have already installed and logged into. Where the work has to reach outside that, reaching outside is a decision you make at the moment it happens.
-
It writes the contract first
An objective, success criteria you can read, constraints, a deadline, a budget in real money and units, an authority mode, what evidence counts, where it escalates, and what happens on cancel or failure. Anything it inferred, it says it inferred.
-
It works across engines
The harness drives the
claude,codexandgrokCLIs on your machine, over the logins you already have. An engine that is missing reports as unavailable instead of taking the fleet down with it. -
It uses a computer and your apps
An isolated cloud Linux desktop, a local VM, or — on the platforms where the safety boundary is certified, and only after you opt in — the machine in front of you. Connected apps run through Composio, on your own key or a managed connection. These are the parts that are not local, and you choose them deliberately.
-
It stops at the exact action
Five classes of thing stop and ask: send, spend, commit, publish, credential. Local reads, local writes and drafts nobody receives are not among them, so it invents no ceremony around work it is already allowed to do.
-
It keeps one card, not chatter
A batch is one durable card with real per-lane states — queued, running, completed, failed, canceled. Hidden prompts, tool transcripts and raw errors stay inspectable and stay out of the conversation.
-
It revises instead of duplicating
New evidence updates the outcome that already exists. A note that arrives twice does not open a second plan, and a choice you settled last week does not get reopened because a date moved.
Why the finishing holds
Six ways a plan never starts
Finishing reliably means not starting things that cannot finish. Refusal happens before work is dispatched and before a budget is touched, in language that names what went wrong.
- budget below cost The plan estimates more than the budget allows, so it stops at the check rather than spending toward the wall.
- deadline already passed The date it was written against is behind us. It says so instead of starting.
- owner mismatch The plan belongs to somebody else. Not started here.
- not a first version Only a first version is ever proposed or started, and the refusal names the version that was pressed.
- standing grant A yes given now would cover actions nobody has seen yet. Authority is deny-by-default and there is no setting that changes that.
- answer off the menu A question with fixed choices, answered with something else, is refused rather than quietly recorded — and you are told your answer did not land.
And one it refuses on principle
The session that produced the work never grades it. The check runs in a separate process and a separate session, which is handed the delivered bytes and the success criteria and nothing else — not the plan, not the transcript, not how any of it was made. Where no second reviewer is available, the product does not offer verification at all, and says prepared — never done.
Proof
A receipt, not a status message
When an outcome closes, what you get is a replayable record bound to the exact agreement it was checked against — readable on its own, without opening the payload it describes.
- outcomeId
- which outcome this closes
- contractId
- which agreement it closes against
- contractVersion
- and which version of it — a verdict covers one version and no other
- criteriaHash
- sha-256 over the success criteria, so a verdict cannot be moved onto a different bar
- workIds
- every piece of work this contract owned
- evidenceRefs
- the attributed, time-bound observations it was judged on
- artifactRefs
- the files that were actually delivered
- verifiedAt
- when the independent read happened
- usage
- what it cost, in cents and units, against the budget you set
The work is read as untrusted data
Everything the run produced reaches the reviewer inside a fence, labelled as untrusted output. A document that announces it has already been approved is not a shortcut — attempted direction is treated as a defect in the work and weighed against it. Absence of proof is a fail: if the artifacts do not demonstrate a criterion, that criterion is not met.
The check refusing and the checker breaking are different facts
A verifier that read the work and rejected it returns a verdict with a reason. A verifier that could not run returns an outage. One is a judgment about your deliverable, the other is a broken tool, and most systems collapse both into the same red light. This one does not — and a rejection meant for another contract version can never terminate yours.
An approval binds to one action, and is spent
Every consequential action is put in front of you with its account, operation, target, payload, expiry and reversibility. Your yes is consumed exactly once. A retry against the same approval fails closed. A payload that changed after you approved it fails closed too, rather than proceeding with a warning.
How it is held to this
44 journeys it has to survive
Each one is a real request with a written definition of success, run against the product rather than against a description of it. A sample of what they insist on:
- Builder output cannot mark itself verified.
- A send stops at the point of action instead of just happening.
- Saying no stops the work, not the app — and it is reported as your decision, not a failure.
- New evidence revises the existing work instead of duplicating it.
- Partial indexing is used honestly, not reported as complete.
- Unavailable, not configured, and available are three different answers.
- Two sources disagree on an amount, and it gets flagged rather than averaged.
- HTTP 200 is not verification.
- Bounded retry with a changed strategy, then an honest blocker.
- The last inch does not get handed back.
Run it
What actually runs locally
Agent Centipede is a desktop application with the harness embedded.
One small server on 127.0.0.1 owns every agent process.
Your transcripts, your keys and the event journal live in a directory
in your home folder. There is no account to make and no service of
ours in the middle.
Some things do leave, and it is better to say which. Model calls go to whichever provider your engine is logged into. A cloud desktop is a cloud desktop, if you choose one over the local VM. Connected apps reach their providers through Composio. None of that is hidden from you, and the consequential half of it stops and asks first.
Requirements
- DesktopmacOS · Windows · Ubuntu (beta)
- RuntimeNode 24+
- Enginesat least one of
claude,codex,grok, already logged in - Verificationa second available reviewer, or it will not claim done