Picture a small factory's scheduling agent. A rush order lands, and the agent does what agents do: it produces a fluent, confident plan to take the job. The plan is locally reasonable. It is also wrong, because the factory's protected medical order, its two "important" orders, and its best-effort work already consume most of the agent-work budget and the exact production-cell interval the rush job wants. The agent didn't lie. It just answered a smaller question than the one that mattered.

That gap is where I decided to spend my hackathon: not making an agent that plans better, but making one that is structurally unable to promise what the whole system can't keep — and that can prove, afterward, that the one thing it was allowed to do is the one thing that happened. The result is FlakeBrake — a commitment firewall for humans and agents — wrapped around a TrueForge mission.

Scarce capacity turns confidence into damage

Autonomous action is at its most dangerous precisely when capacity is scarce, because that's when every commitment is a trade against something already promised. An agent that overcommits under abundance wastes time. An agent that overcommits under scarcity reassigns loss — quietly, to whichever existing obligation was least defended. And the usual failure amplifiers — stale approvals, renamed retries, success declared off a 200 — are all downstream of the yes.

So FlakeBrake treats promise acceptance as an admission decision over the complete versioned portfolio, not a model assertion — and infeasible proposals get a bounded, portfolio-wide replan. In the demo the smallest safe change trims one best-effort order from quantity 10 to 8 while protected work stays untouchable.

Why TrueForge, specifically

I could have hand-rolled an agent loop, and it would have produced exactly the kind of demo I didn't want: a bespoke while-loop where "session," "approval," and "resume" mean whatever my code says they mean that day. The interesting properties of FlakeBrake — pausing a turn for a human, surviving a disconnect, keeping subagent work isolated — are only worth demonstrating if they run on infrastructure I don't control.

TrueForge gave me that substrate. The mission runs inside the pinned 0.1.4 server through its 0.1.3 SDK, and every load-bearing element is a TrueForge primitive: a durable session and turn graph; a root "obligation commander" that alone talks to the owner; three dynamic subagents — portfolio, capacity, and assurance — investigating through four MCP services; a sandbox for the assurance agent's generated analysis code, which never makes decisions; native approval pauses on the four consequential tools; and reconnect that replays durable state instead of re-running the mission.

The chain, every run

  1. Session
  2. Three subagents
  3. Four MCP services
  4. Local sandbox
  5. Native approval pause
  6. Human decision
  7. Reconnect
  8. Verified effect

One honesty note: the credential-free judge mode swaps in a bounded local deterministic model provider so anyone can run the mission reproducibly, with no OpenAI or Daytona accounts. Everything else is the real thing — real TrueForge sessions, turns, subagents, MCP connectors, sandbox execution, native approval events, reconnect. The deterministic profile proves the orchestration and the safety mechanics; it deliberately doesn't claim to prove an external model's judgment. Credential-gated tests that exercise an OpenAI provider and a Daytona sandbox exist in the repo as a separate, optional gate.

How agents keep each other honest

Multi-agent systems have a quieter failure mode than a bad plan: agents relying on each other's claims. TrueForge coordinates the specialist agents, tool calls, sandbox execution, approval pauses, and reconnect; FlakeBrake sits behind the consequential MCP boundaries and checks capacity, authorization, denied-effect equivalence, execution identity, factory results, and replay before another agent — or a human — relies on the consequential outcome.

Agents can propose anything; they cannot make it true. Every consequential effect must pass current-state checks, authorization, fenced execution, and independent verification.

That guarantee has two deliberately distinct layers. The recorded recommendation: TrueForge preserves each specialist thread and its prose in durable history, and FlakeBrake makes no claim to understand or certify arbitrary sentences — and nothing verifies the semantic truth of what a specialist wrote. The governed effect: specialists recommend, and before a recommendation can become a consequential effect, the root's change-control call and FlakeBrake's stores independently re-evaluate the exact action against current authoritative state. Only the root invokes the consequential tools, effect arguments are built from authoritative SQLite state, and raw prose cannot authorize, mutate, verify, or replay anything.

Specialist recommendation → TrueForge root turn → FlakeBrake MCP check → Human approval when required → Fenced execution → Independent read-back → Evidence-bound result

Concretely, in the current build: a capacity recommendation is recomputed against the authoritative versioned basis; a stale basis cannot be treated as current; a differently worded denied action remains denied; an executor cannot declare success merely because it returned a receipt; and a restarted agent cannot repeat an already completed effect. An unsupported, stale, conflicting, duplicated, or unverified consequential claim is mechanically unusable. Routine inter-agent checks are automatic; a person is interrupted only at the policy-defined consequential boundaries.

Two kinds of "no"

The demo's centerpiece is a distinction I now think every agent system needs to make explicit. When the mission requests the rush order's preferred 09:10–09:40 production slot, the owner denies it — a human reading an exact request (mission, tool, expected effect, SHA-256 action digest) and choosing "Deny action." That's the first kind of no: judgment.

Then the planner does the sneaky thing, by design: it submits the same material effect through a different MCP adapter, submit_schedule_change, with a different action name and schema. And here the second kind of no fires: the denial recorded in the ledger is a typed, normalized effect constraint, so the equivalent representation is matched and blocked mechanically — no second interruption, no mutation.

An owner's denial that evaporates on the next rephrase isn't governance. It's a suggestion box.

Keeping those two denials distinct also keeps the human budget honest: the whole mission takes exactly four owner calls, and the mechanical block isn't one of them. The genuinely different 09:40–10:10 alternative is a new effect with its own digest — so it earns a real decision, and gets approved.

A mutation is not a success

The most quietly radical rule in FlakeBrake is that the mutation's own result is treated as a claim, not a conclusion. The approved alternative executes exactly once — allowance claimed, attempt fenced, one synthetic factory mutation, one receipt. And then it reads the factory back through an independent path, appends the actual consumption facts (6 agent work units, 30 production-cell minutes) to the ledger, and only after the read-back matches does it record the terminal verified event that lets the root TrueForge mission complete.

Exactly-once is a recovery philosophy, not a flag

Everything above has to survive the boring disasters: a refresh mid-approval, a lost response, a process restart. My rule was that recovery must replay decisions and effects, never repeat them. The mission, session, turn graph, approvals, tool results, and terminal projection are all durable; a browser refresh re-projects that state rather than re-running anything; a lost approval response reconciles to its existing successor turn; idempotent retries return the original durable result — and the client's monotonic request generations mean a stale poll can never regress a verified outcome. The end-of-demo party trick is deliberately anticlimactic: refresh the browser, and every count stays exactly where it was.

What repeated Qodo review actually changed

Every implementation PR went through the Qodo review app, and its best findings lived in the gaps between components I'd tested separately. On the M4 PR it flagged a live-owner bypass path (High/Security) and a non-atomic acceptance/grant commit (High/Correctness). On the SQLite PR it caught my contention test reporting readiness before contention was actually proven, which would have made the test pass vacuously forever. On the M5 PR it pushed monotonic polling and malformed-request containment further than my first version. The rhythm that emerged — finding, focused regression test, fix, re-review at the exact head, human adjudication — improved the codebase, but it improved my process more: I stopped merging anything whose safety property didn't have a test with its name on it.

Limits, honestly

FlakeBrake v0.1 is a bounded demonstration. It's a synthetic microfactory, not a production integration; capacity assumptions are declared and versioned, not inferred; the judge UI is loopback-only with no multi-user auth story; and deterministic mode proves mechanics, not model quality. The most transferable lesson: put authority in code and evidence in ledgers, and let the model be brilliant inside those walls.

Where this points

Agents are going to be deployed against scarce, contended, real-world capacity, and trustworthy deployment there needs exactly the properties this small system demonstrates end to end: admission before commitment, authorization bound to exact effects, denials that survive renames, one fenced mutation, independent verification, and recovery that never repeats a human's decision. TrueForge supplied the durable spine that made those properties demonstrable rather than aspirational. The rest was mostly discipline — and a willingness to let "not yet" be the agent's most impressive answer.