FlakeBrakeBuilt on TrueForge

A rush order. One safe promise. Proof all the way down.

FlakeBrake is a commitment firewall for humans and agents.

A human-governed TrueForge agent that replans constrained work, pauses for exact authorization, mechanically blocks equivalent denied actions, executes one fenced mutation, verifies the result independently, and reconnects without repeating a single owner decision or effect.

TrueForge action chain every station below is a real TrueForge 0.1.4 interface, exercised on every judge run

  1. Sessiondurable session & turn graph
  2. Three subagentsportfolio · capacity · assurance
  3. Four MCP servicesorders · capacity · simulator · change control
  4. Local sandboxmechanical assurance code
  5. Native approval pausethe turn holds for a person
  6. Human decisionapprove or deny, exactly as displayed
  7. Reconnectdurable replay, no repeats
  8. Verified effectindependent read-back first

Honest by construction: the judge path runs genuine TrueForge 0.1.4 sessions, turns, dynamic subagents, MCP connectors, sandbox execution, native approval events, and reconnect behavior — with a bounded local deterministic model provider, so judging needs no OpenAI or Daytona credentials and produces the same evidence every time.

01The operator problem

A fluent plan is not an authorization.

Agent systems are good at producing a plausible next action. They are far less reliable at deciding whether a new commitment stays feasible next to every promise already accepted — the human review budget, the agent work budget, the production window that protected orders already occupy. The dangerous failure mode is quiet: the agent says yes first and discovers the conflict mid-execution, when capacity is already spent.

Most agent demos never touch this. They show a happy path where the model's confidence is treated as authority, and they skip the questions an operator would actually ask:

Skipped question № 1

Who authorized that exact effect?

A chat transcript that says "sounds good" binds nothing. FlakeBrake pauses the TrueForge turn and shows the owner the exact tool, arguments, expected effect, and a SHA-256 action digest before anything is allowed to happen.

Skipped question № 2

What happens when the agent asks again, differently?

Renaming an action or switching to another tool adapter is the oldest trick in the book. FlakeBrake's denials are typed, durable constraints that recognize the same material effect through equivalent representations — and block it without another human call.

Skipped question № 3

Did the effect actually happen — once?

A completed tool response is not success. FlakeBrake fences a single execution attempt, stores one receipt, reads the factory back independently, and only then exposes a verified terminal state. Refresh the browser and the counts don't move.

02Built on TrueForge

TrueForge owns the agent loop. FlakeBrake owns the policy.

The mission runs inside the pinned TrueForge 0.1.4 server through the pinned 0.1.3 SDK — not a lookalike shim. TrueForge owns the session and turn graph, the dynamic subagent threads, the MCP connector routing, sandbox execution, the native approval pauses, and reconnection. FlakeBrake supplies what TrueForge deliberately doesn't decide: the admission kernel, exact authorization semantics, denial equivalence, fencing, and independent verification.

Runtime
TrueForge 0.1.4 · SDK 0.1.3pinned in package.json, surfaced live in the judge UI
Root agent
flakebrake-root-obligation-commanderthe only role that talks to the owner
Subagents
Three dynamic rolesportfolio & order analyst · capacity & schedule analyst · assurance & simulation engineer
MCP services
Four Streamable HTTP connectorsfactory-orders · factory-capacity · factory-simulator · factory-change-control
Approval gate
Four gated tools, native pausesevery consequential change-control tool requires an owner decision
Sandbox
Enabled, downloads offthe assurance subagent runs generated analysis code in an isolated TrueForge sandbox
Provider profile
flakebrake-deterministic/m4-missiona bounded local model endpoint registered with the real TrueForge server
Reconnect
Durable session replayrefresh and restart re-project the same session — never re-run it
TrueForge harness strip from the judge UI after a verified run: the action chain marks Mission and Verified result as VERIFIED, with 3 specialist agents, 4 factory tools, Sandbox check, Human pause, and Same-session resume all OBSERVED, above the sentence 'TrueForge coordinated 3 specialist agents, connected 4 factory tools, ran a sandbox check, paused for your decisions, and resumed the same durable session.' Session, runtime, and MCP identifiers sit behind the collapsed technical-evidence disclosure.
Observed, not asserted: the harness ribbon after a verified deterministic run. Before the mission starts every chain station reads Configured; afterwards the stations read Observed and the mission and result read Verified — and the disclosure's facts flip from "4 services configured / Configured / Dynamic · configured" to 4/4 services reached, 1 executed, 3 threads evidenced, Native · 4 owner calls. Runtime evidence from that run, captured at the current build.

Configured

What the root agent spec declares before any run: four MCP services, sandbox enabled, dynamic subagents enabled, four approval-gated tools. Setup facts — never presented as runtime proof.

4 services configured · Dynamic · configured

Observed

What this run actually exercised, from durable runtime evidence: services reached, sandbox executions, subagent threads, owner calls, model requests.

4/4 services reached · 3 threads evidenced

Verified

The only state presented as success: an independent read-back of authoritative state confirming the one approved effect, recorded as a terminal event.

Terminal verified success

What the deterministic profile does and doesn't prove. The judge run proves the orchestration and safety mechanics through TrueForge's genuine public interfaces; it does not evaluate an external model's judgment, because the provider is a bounded deterministic endpoint built for reproducible judging. Separately, the repository carries credential-gated test coverage — skipped without an external M0 configuration — that exercises an OpenAI model provider and a Daytona sandbox provider through the same boundaries. That coverage is an optional assurance gate, never part of the credential-free judge path.

03The commitment firewall

How agents keep each other honest

TrueForge coordinates specialist agents, tool calls, sandbox execution, approval pauses, and reconnect. FlakeBrake sits behind the consequential MCP boundaries and checks capacity, authorization, denied-effect equivalence, execution identity, factory results, and replay — before another agent, or a human, relies on the consequential outcome.

Agents can propose anything; they cannot make it true. Every consequential effect must pass current-state checks, authorization, fenced execution, and independent verification.

Specialists recommend. Before a recommendation can become a consequential effect, the root's change-control call and FlakeBrake's stores independently re-evaluate the exact action against current authoritative state. Only the root invokes the consequential change-control tools, and in the deterministic fixture the effect arguments are constructed from authoritative SQLite state — raw prose cannot authorize, mutate, verify, or replay an effect. That precision matters, so the firewall's claim comes in two distinct layers:

Layer 1 · Recorded recommendation

Preserved, not certified

TrueForge keeps each specialist thread and its output in durable history — context the root works from, evidence a judge can read. FlakeBrake makes no claim to understand or certify arbitrary prose, and this isn't deception detection: another agent can consume the authoritative domain outcome, but nothing verifies the semantic truth of a specialist sentence.

Layer 2 · Governed effect

Authoritatively enforced

Capacity, authorization, equivalent denial, execution identity, read-back, and replay are enforced against current authoritative M1–M4 state before another agent or human relies on the consequential outcome. An unsupported, stale, conflicting, duplicated, or unverified consequential claim is mechanically unusable.

  • An agent claims"This replan fits."
    The firewall

    Capacity recommendations are recomputed against the authoritative versioned basis — the deterministic kernel re-evaluates the complete portfolio at the MCP boundary, and the assurance engineer independently recomputes demand, ranking, and protected-order preservation in the sandbox.

    M1 kernel · versioned Promise Basis
  • An agent claims"The earlier evaluation is still good."
    The firewall

    A stale basis cannot be treated as current. Acceptance atomically compares the complete evaluated version tuple; once portfolio v2 is durable, the v1 REPLAN can never be accepted.

    M2 ledger · version compare-and-swap
  • An agent claims"This is a different action now."
    The firewall

    A differently worded denied action remains denied. Typed effect normalization matches the material effect across tool names and schemas, so the active denial blocks the equivalent representation — no new owner call, no mutation.

    M2 ledger · typed denial equivalence
  • An agent claims"The tool returned a receipt — done."
    The firewall

    An executor cannot declare success merely because it returned a receipt. Independent read-back of authoritative factory state and a terminal verification event stand between the mutation and the word "success".

    M3 boundary · independent read-back
  • An agent claims"I restarted, so I'll run it again."
    The firewall

    A restarted agent cannot repeat an already completed effect. Attempts are fenced and idempotent: replay converges on the original durable result and creates no second mutation.

    M2/M3 · fenced idempotent attempts

Routine inter-agent checks are automatic — nobody is interrupted while the firewall recomputes, matches versions, or converges a replay. Humans enter exactly where policy says a consequence needs an owner: the four approval-gated change-control tools that pause the TrueForge turn.

Architecture sequence the amber station is the only one that costs a human interruption — and only when policy requires it

  1. Specialist recommendationthree subagents investigate & propose
  2. TrueForge root turnonly the root acts on a recommendation
  3. FlakeBrake MCP checkbasis, authorization, equivalence, identity
  4. Human approval when requiredthe four gated tools pause the turn
  5. Fenced executionone claimed, idempotent attempt
  6. Independent read-backauthoritative state, read separately
  7. Evidence-bound resultdurable, replayable, verified

04The deterministic hero scenario

Ten steps from overload to proven, in about three minutes.

A rush order arrives at a synthetic microfactory whose protected, important, and best-effort orders already consume finite human-review, agent-work, and production-cell capacity. Every run of the story below is the same, end to end, because the basis is canonical and the code is deterministic — which is exactly what makes it checkable.

  1. The direct plan comes back REPLANReplan

    Evaluation is side-effect free. The complete portfolio is checked against declared capacity: agent work would go over by 2 and owner decisions over by 1. Nothing is dropped, nothing mutates.

  2. Candidates are compared as complete portfolios

    Replan candidates are whole portfolio states, not local patches. The winner modifies one best-effort display order from quantity 10 to 8. Protected work is a hard constraint and stays untouched.

  3. The owner approves the exact modificationOwner call 1

    TrueForge pauses the turn natively. The owner sees the mission, tool, expected effect, and the SHA-256 action digest — and authorizes only that.

  4. The fresh promise is accepted atomicallyOwner call 2

    Portfolio v2 is durable before readmission. The fresh ADMITTABLE evaluation and its exact authorization grant commit in one transaction — the stale v1 REPLAN can never be accepted.

  5. The owner denies the primary reservationOwner call 3

    The 09:10–09:40 slot conflicts with protected production commitments. The denial is not a chat message; it becomes an active, typed constraint in the ledger.

  6. The equivalent retry is blocked mechanicallyAuto-blocked

    The planner submits the same material effect through a different MCP adapter, submit_schedule_change. Typed effect normalization recognizes it, and the active denial blocks it — no extra owner call, no mutation.

  7. The distinct alternative is approvedOwner call 4

    09:40–10:10 is a genuinely different effect, so it earns a real decision of its own. The owner approves it against its own digest.

  8. Exactly one fenced mutation executesExactly once

    The grant's allowance is claimed, the attempt is fenced and idempotent, one synthetic factory mutation runs, and one receipt is stored. A retry returns the original durable result instead of acting twice.

  9. The factory is read back independentlyRead-back

    Verification doesn't trust the mutation's own response. Authoritative state is read separately, and actual consumption is appended to the ledger: 6 agent work units, 30 production-cell minutes.

  10. Only then: verified successVerified

    The terminal verified event is recorded, and the root TrueForge mission is allowed to complete. The UI's outcome flips to "Verified success" — and it means it.

The judge UI's external owner boundary paused on the first decision. The panel reads 'Your decision is required', names the action 'Select Portfolio Modification — Modify order/best-effort-display: quantity 10 → 8', offers a collapsed 'Durable action identity' disclosure, and presents Deny action and Approve action buttons with a recommendation note.
Owner call 1, as the judge sees it. TrueForge holds the turn; the UI never auto-approves. The response can authorize only the exact digest and arguments displayed — missing, stale, malformed, or replayed-with-different-arguments responses fail closed.

05The owner boundary

Denied once by a human. Denied again by the ledger.

These are two different safety properties, and FlakeBrake demonstrates both in the same run. Conflating them is how agent demos oversell — a human "no" that evaporates on the next rephrase isn't governance.

Human judgment

Owner denial

A person looks at the exact 09:10–09:40 reservation — mission, predecessor turn, tool, expected effect, digest — and chooses Deny action. That choice is recorded durably, with its rationale, as an active constraint.

Mechanical enforcement

Equivalent-action denial

When the same material effect arrives through a different tool name and schema, typed effect normalization matches it to the active denial and blocks it automatically. No second interruption, no mutation, no chance to relitigate by rewording.

Judge UI policy panel showing two entries: 'Owner denied primary interval' with the rationale that the primary interval conflicts with protected production commitments, and 'Blocked automatically — same denied action' recording that the equivalent reservation for cell-alpha 09:10–09:40 cannot bypass the active denial.
Both denials, side by side, from the record. The mechanical block consumed no owner call — the four-call budget is itself an asserted safety property of the run.

About that digest: every consequential action is bound to a SHA-256 content digest of its exact identity and arguments. It's a precise fingerprint for matching decisions to actions — a digest, not a cryptographic signature, and FlakeBrake never pretends otherwise.

06Exactly once, then proven

The whole run fits in eight numbers.

4owner calls
1mechanical block
1acceptance
1attempt
1factory change
1change record
2measured facts
0duplicates

Counts rendered live by the judge UI and asserted by the deterministic test gates — including zero duplicate approvals and zero unauthorized mutations. Refreshing the browser changes none of them.

  1. Rung 1

    Factory change record

    This record proves the one-time-locked factory command committed. By itself, it is not verified success — and the UI says so in exactly those words.

  2. Rung 2

    Independent read-back

    Authoritative factory state is read separately from the mutation result, and actual consumption facts are appended to the immutable ledger.

  3. Rung 3 · the only "success"

    Verified terminal event

    Only after the read-back matches is a terminal verified event recorded and the root mission allowed to finish.

07Reconnect without regret

Refresh the browser. Nothing happens twice.

Mission identity, the TrueForge session and turn graph, approvals, tool results, and the terminal projection are all durable. A refresh calls the read-only state API and replays that projection — it does not restart the mission, reopen an approval, or repeat an effect. The harness ribbon labels the state plainly: Durable session replayed.

The same discipline runs deeper than the browser: a lost approval response reconciles to its existing successor turn, same-mission retries converge on the original durable result, and the client's monotonic request generations and durable revisions mean an old poll can never overwrite a newer decision or regress a verified outcome.

08Operator proof center

Safety and impact, from the record — not from the narrator.

Judges shouldn't have to reconstruct the causal chain from a timeline. The Operator Proof Center sits at the top of the judge UI and answers the four operator questions at a glance — direct plan, safe winner, owner boundary, durable outcome — then backs each answer with disclosure drawers built from the same canonical projection the ledger holds: exact control decisions, capacity before/after impact, durable proof and replay, and optional technical identities.

The Operator Proof Center in its Verified record state. Summary cards read: Direct plan 'Doesn’t fit yet' (REPLAN) with agent work over by 2 and human decisions over by 1; Safe winner 10 → 8 for the best-effort display order with protected work unchanged; Human decisions 3 allowed · 1 denied with 1 equivalent action mechanically blocked; Durable outcome 1 factory change · verified with 1 change record, 1 verified completion, 2 measured facts. Below are disclosure rows to show decision history, exact capacity math, the durable ledger, and technical identifiers, plus a 'What FlakeBrake prevented' summary naming the denied 09:10–09:40 interval and the approved 09:40–10:10 mutation.
The verified record, one screen. Even the counterfactual — "What FlakeBrake prevented" — is composed only of facts derivable from the deterministic fixture: the overload amounts, the protected order kept at 10, the bounded 10 → 8 change, and which interval did and didn't become the factory result.

Architecture

Five deliberate layers, one trust boundary at a time.

The model can propose; deterministic code evaluates and authorizes; the owner decides; the ledger remembers; verification reads back independently. Each layer exists so no other layer has to be trusted for something it can't prove.

  1. M1

    Deterministic admission kernel

    Pure portfolio feasibility, candidate ranking, Promise Basis versioning, and typed effect comparison. Returns ADMITTABLE, REPLAN, or REJECT — and never mutates anything.

  2. M2

    Immutable SQLite ledger

    Append-only admissions, owner decisions, denials, grants, shared allowances, fences, attempts, receipts, and actual-consumption facts. Acceptance and grant issuance commit atomically; database-incarnation identity keeps replaced files from impersonating state.

  3. M3

    Synthetic factory & MCP boundary

    An invocation-owned factory behind four Streamable HTTP MCP services. Exact-once mutation, independent read-back, and idempotent replay live here — the browser never touches SQLite.

  4. M4

    TrueForge mission orchestration

    The root obligation commander, three dynamic subagents, sandbox execution, four MCP connectors, native approval pauses, and durable session/turn recovery — TrueForge 0.1.4 running the loop for real.

  5. M5

    Judge UI & control boundary

    A loopback-only service projecting canonical backend state, with strict request validation, monotonic polling, bounded shutdown, and the Operator Proof Center on top.

The full trust-boundary and data-flow write-up, including the end-to-end mermaid diagram, lives in docs/architecture.md; the frozen normative contract is PRODUCT_SPEC_v0.1.md.

Code quality

Reviewed in public, remediated with regressions.

Every implementation PR was reviewed by the GitHub Qodo application, remediated with focused regression tests, re-reviewed at the exact head, and merged only after separate human adjudication. The trail is public — and it's presented here as evidence of process, not as a claim that automated review replaces testing or human approval.

PR #5 · M4 orchestration

Qodo caught the owner boundary leaking.

  • High · Security — a live-owner bypass path
  • High · Correctness — non-atomic acceptance/grant
  • High · Reliability — rollback overwriting concurrent writes
Review trail →
PR #6 · SQLite contention

Even the test protocol got reviewed.

  • High · Test reliability — readiness reported before contention was actually proven
  • Lifecycle — startup exit handling in the child process
Review trail →
PR #7 · M5 judge UI

The UI's replay discipline was hardened here.

  • Medium · Reliability — monotonic polling
  • High · Correctness — durable failed-mission recovery
  • High · Reliability — malformed-target containment
Review trail →

FlakeBrake went through 16 Qodo-reviewed pull requests, resolving more than 50 findings, with zero unresolved findings on the merged release. Earlier milestone trails live in merged PR #2, #3, and #4; later rounds — submission readiness, M5 polish, and the Judge Clarity corrections — in #8, #9, and #14. Each preserves initial findings, remediation rounds, exact-head re-reviews, and the final human adjudication.

Three-minute demo

Watch the whole story in one take.

Run it yourself

One command to the control room.

Deterministic judge mode · Node 22+
git clone https://github.com/dvellon/flakebrake.git
cd flakebrake
npm ci
npm run m5:ui
# open http://127.0.0.1:4173 · choose "Start hero mission"
# no OpenAI or Daytona credentials required

The judge service binds to loopback only, creates its stores in an invocation-owned temporary root, and cleans up on Ctrl+C. The full credential-free verification ladder — typecheck, deterministic test suites, and the Firefox browser gate — is in the README.