Meaningful coding-agent volume
Enough covered sessions and comparable pull requests to evaluate a shared change, not one developer experimenting with a prompt.
Harness Engineering Platform for Coding Agents
Version, compare and roll out coding-agent setups using real pull request, review and verification outcomes, so the team default is chosen by evidence, not vibes
Works with whatever deploys your config today: GitOps, MDM, or an internal platform. Fullbeam only tells you what happened afterwards.
One GitHub App for outcomes; a lightweight agent connection captures the runtime.
Release hypothesis
Reduce post-review revisions by at least 20% without increasing CI failures by more than 2 percentage points.
Met · −54% / +1pp
What changed in harness@v13
Primary outcome · post-review revisions
0.6 ↓ 54%
harness@v12: 1.3
Correctness guardrail · CI failure rate
11% +1pp
harness@v12: 10% · policy: ≤ 2pp
Evidence
Recommendation
Expand canary 20% → 50%
keep harness@v12 as the default
Illustrative rollout decision · not customer data
Across 46 comparable pull requests, 23 on the candidate and 23 matched on the default, work covered by harness@v13 needed 54% fewer post-review revisions than harness@v12. CI failures were 1 percentage point higher, within the predeclared guardrail. With 91% verified coverage, the rollout policy recommends expanding the canary from 20% to 50%, not making it the team default yet.
No magic quality score. Every recommendation includes the sample size, exclusions, coverage, observed improvements, possible regressions and the policy that produced it.
harness@v13
Assigned release
verify skill v5 · code-search MCP v4 · pre-PR hook enabled
Effective runtime
48 exact · 3 partial · 2 unknown
what was actually active during each covered session
2 sessions still used code-search MCP v3
Session → commit
51 verified or strong · 2 weak
every link retains its evidence confidence
Pull-request comparison
46 included · 7 excluded
the exact reviewable PR revision associated with each session
Outcome
review · verification · rework · merge · revert
The evidence chain
Fullbeam records which versioned instructions, skills, hooks, MCP servers, permissions, memory, model and runtime were actually active during each covered session. It then links the session to the exact commit, pull-request revision and engineering outcome, preserving evidence confidence at every step.
Missing runtime data and weak session links remain partial or unknown. They are never silently folded into a clean comparison.
The contents of your instructions, skills and configurations stay in your source system. Fullbeam only needs their versioned identities for comparison.
From aggregate to evidence
Open any comparison metric and inspect the exact pull-request revision, agent plan, decisions, verification, reviewer feedback and structured outcome behind it. Reviewers see how the change was produced; Platform teams can audit the evidence behind the rollout decision.
Part of the same invented example, not customer data or a live repository.
Agent plan
Reviewer feedback
“Please add a tenant-isolation regression test before merge.”
Review outcome
Another revision requested
Structured reason: incomplete verification before PR creation
↳ counted in post-review revisions for harness@v13
How it works
Treat each shared setup change as a Harness Release: versioned instructions, skills, tools, permissions, memory, model defaults and workflows.
Record the effective harness and runtime that actually ran, preserving exact, partial and unknown coverage instead of assuming that the assigned release was active.
Connect covered sessions to commits, pull-request revisions, review, CI, merge, rework and revert outcomes.
Compare the candidate with the current default and apply the rollout policy: expand, continue the canary, modify and retest, roll back, or collect more evidence.
Who it’s for
Enough covered sessions and comparable pull requests to evaluate a shared change, not one developer experimenting with a prompt.
Centrally managed instructions, skills, hooks, MCP servers, permissions or workflows released across multiple engineers or repositories.
A real decision to expand, continue, modify or roll back a Harness Release.
Fullbeam is not for teams merely tracking coding-agent adoption. It is for teams changing a shared setup and needing to know whether the change earned a wider rollout.
Bring one candidate Harness Release and the current default. Fullbeam captures what actually runs, compares the resulting pull requests and shows the evidence behind the rollout decision.