Qualification economics and evidence
Read cost per accepted task, workload decisions, runtime coverage, and engineering guardrails.
Agent Stack Qualification connects immutable baseline-versus-candidate attempts to acceptance, cost, runtime, and downstream engineering evidence.
Questions it answers
- Which current and candidate Stack Releases were frozen, and what actually ran?
- What did each arm cost across all valid attempts, including failures?
- How many tasks satisfied the frozen repetition-aggregation and acceptance policy?
- Which workloads have candidate-only regressions or insufficient evidence?
- What CI, reviewer-intervention, rework, merge, and abandonment outcomes followed?
Use Cost qualification to inspect overall and workload-level cost per accepted-task equivalent, pass evidence, candidate regressions, cost coverage, assigned/effective drift, and the frozen decision scope. Accepted-task equivalents sum each task's frozen repetition-aggregated pass rate; mean aggregation can therefore produce a fractional denominator.
Raw scheduled-work cost and cost per accepted-task equivalent answer different questions. A cheaper run is not an economic win if it produces fewer accepted-task equivalents. Failed valid attempts stay in the cost numerator; missing cost is reported as missing coverage.
The current immutable Release Policy still applies its economics threshold to raw scheduled-work savings. The acceptance-adjusted metric is decision context, not an independent verdict. Read the policy disposition and both economic measures together.
Read comparisons carefully
Fullbeam reports observed associations, sample sizes, and source coverage. It does not claim that a prompt, skill, model, agent, or engineer caused an outcome.
For example, Fullbeam may report that a candidate passed the frozen private policy at a lower cost per accepted-task equivalent and that no review-revision increase was observed in a bounded canary. That is evidence for a rollout decision, not universal proof that the candidate caused an improvement.
Unavailable data is not the same as zero. Fullbeam labels missing or partial sources so that absent telemetry is not presented as no activity.
Decide by workload
Use repeated, like-for-like evidence to allocate workloads and record a decision such as:
- promote;
- promote for selected workloads;
- continue evaluation;
- block;
- roll back a production canary;
- insufficient evidence.
Apply the rollout through the system that owns deployment. Fullbeam records the decision and later observes the external action; it does not execute it.
Offline attempt economics and production outcome confirmation are distinct evidence layers. Do not claim that an offline saving survived production until production cost is attributed to the effective Stack Release with sufficient coverage.