Harness Engineering Platform for Coding Agents

Roll out better coding-agent harnesses with confidence

Version, compare and roll out coding-agent setups using real pull request, review and verification outcomes, so the team default is chosen by evidence, not vibes

Works with whatever deploys your config today: GitOps, MDM, or an internal platform. Fullbeam only tells you what happened afterwards.

  • GitHub App
  • Claude Code
  • Codex
  • OpenCode

One GitHub App for outcomes; a lightweight agent connection captures the runtime.

Fullbeam · acme/backend

harness@v13 vs harness@v12

Canary · 20%

Release hypothesis

Reduce post-review revisions by at least 20% without increasing CI failures by more than 2 percentage points.

Met · −54% / +1pp

What changed in harness@v13

  • ~verify-before-commit skillv4 → v5
  • ~code-search MCPv3 → v4
  • +pre-PR verification hookadded

Primary outcome · post-review revisions

0.6 ↓ 54%

harness@v12: 1.3

Correctness guardrail · CI failure rate

11% +1pp

harness@v12: 10% · policy: ≤ 2pp

Evidence

  • 53 PRs entered the window · 91% verified coverage
  • 46 included: 23 candidate · 23 matched default
  • 7 excluded: 3 mixed model/runtime · 2 weak session links · 2 emergency hotfixes

Recommendation

Expand canary 20% → 50%

keep harness@v12 as the default

Illustrative rollout decision · not customer data

Expand harness@v13 to 50%. Do not make it the default yet.

Across 46 comparable pull requests, 23 on the candidate and 23 matched on the default, work covered by harness@v13 needed 54% fewer post-review revisions than harness@v12. CI failures were 1 percentage point higher, within the predeclared guardrail. With 91% verified coverage, the rollout policy recommends expanding the canary from 20% to 50%, not making it the team default yet.

No magic quality score. Every recommendation includes the sample size, exclusions, coverage, observed improvements, possible regressions and the policy that produced it.

harness@v13

Assigned release

verify skill v5 · code-search MCP v4 · pre-PR hook enabled

Effective runtime

48 exact · 3 partial · 2 unknown

what was actually active during each covered session

2 sessions still used code-search MCP v3

Session → commit

51 verified or strong · 2 weak

every link retains its evidence confidence

Pull-request comparison

46 included · 7 excluded

the exact reviewable PR revision associated with each session

Outcome

review · verification · rework · merge · revert

The evidence chain

The assigned release is not always the harness that actually ran.

Fullbeam records which versioned instructions, skills, hooks, MCP servers, permissions, memory, model and runtime were actually active during each covered session. It then links the session to the exact commit, pull-request revision and engineering outcome, preserving evidence confidence at every step.

Missing runtime data and weak session links remain partial or unknown. They are never silently folded into a clean comparison.

The contents of your instructions, skills and configurations stay in your source system. Fullbeam only needs their versioned identities for comparison.

From aggregate to evidence

Trace every result back to the PR.

Open any comparison metric and inspect the exact pull-request revision, agent plan, decisions, verification, reviewer feedback and structured outcome behind it. Reviewers see how the change was produced; Platform teams can audit the evidence behind the rollout decision.

Part of the same invented example, not customer data or a live repository.

acme/backend · #482harness@v13

Add team-scoped SSO policy

Agent plan

  1. 01Add policy schema✓ verified
  2. 02Protect existing account migrations✓ verified
  3. 03Verify tenant isolation flowstest missing

Reviewer feedback

“Please add a tenant-isolation regression test before merge.”

Review outcome

Another revision requested

Structured reason: incomplete verification before PR creation

↳ counted in post-review revisions for harness@v13

How it works

Stop tuning AGENTS.md, skills, and MCP servers by vibes.

  1. 01Version

    Treat each shared setup change as a Harness Release: versioned instructions, skills, tools, permissions, memory, model defaults and workflows.

  2. 02Observe

    Record the effective harness and runtime that actually ran, preserving exact, partial and unknown coverage instead of assuming that the assigned release was active.

  3. 03Link

    Connect covered sessions to commits, pull-request revisions, review, CI, merge, rework and revert outcomes.

  4. 04Decide

    Compare the candidate with the current default and apply the rollout policy: expand, continue the canary, modify and retest, roll back, or collect more evidence.

Who it’s for

Built for teams with a real harness rollout to make.

repeated agent usage

Meaningful coding-agent volume

Enough covered sessions and comparable pull requests to evaluate a shared change, not one developer experimenting with a prompt.

shared harness

A shared harness that changes

Centrally managed instructions, skills, hooks, MCP servers, permissions or workflows released across multiple engineers or repositories.

candidate vs default

A candidate and a current default

A real decision to expand, continue, modify or roll back a Harness Release.

Fullbeam is not for teams merely tracking coding-agent adoption. It is for teams changing a shared setup and needing to know whether the change earned a wider rollout.

Which version should everyone be on?

Bring one candidate Harness Release and the current default. Fullbeam captures what actually runs, compares the resulting pull requests and shows the evidence behind the rollout decision.