Private preview · Managed benchmarksRequest access →
Fullbeam
Resources

How to evaluate coding-agent stacks

A useful coding-agent evaluation has to match the system and work you may release. These notes explain how Fullbeam defines that system, controls the comparison, counts engineering effort, and handles missing runtime evidence.

What the release gate checks

The evidence behind the decision

Each result stays tied to the Stack Release, workload, repository state, and coverage that produced it. If evidence is missing, the gap remains visible in the release call.

01

Name the release

Learn why a model name is too small a unit and how a Stack Release captures the harness, instructions, tools, permissions, routing, workflow, and runtime.

02

Build a fair comparison

Use representative repository work, equivalent starting states, repeated critical cases, explicit acceptance checks, and visible evidence coverage.

03

Turn results into a rollout boundary

Read private qualification, controlled canaries, workload restrictions, rollback evidence, and regression cases as one release process that starts before exposure and keeps learning afterward.

04

Keep evidence inside a clear boundary

Separate declared releases from effective runtime, send bounded evidence, exclude sensitive bodies from retained evidence, and qualify releases without hidden reasoning traces.

Bring us the next stack change

Start with the Core Thesis, then use the benchmark, cost, and runtime notes to plan a private comparison your team can defend.

One current releaseOne candidatePrivate repository workA workload-level decision