Problem
Fixtures can score planted findings, but there is no versioned baseline showing whether a changed skill improves recall, increases unsupported findings, or behaves inconsistently across runs.
User outcome
Reviewers can evaluate a skill change against transparent historical evidence rather than intuition.
Scope
- Define per-fixture metrics for expected findings, location accuracy, unsupported findings, finding volume, and run completion.
- Store baselines by skill, version, provider, and model.
- Compare changes with confidence intervals or explicit run-level variance.
- Require human review of baseline changes.
Non-goals
- Declaring one model or skill universally best.
- Blocking changes on a single stochastic run.
Acceptance criteria
Validation
- Introduce a known regression and verify the comparison detects it.
- Run clean fixtures to measure false-positive pressure.
Relationships
- Depends on reproducible cross-provider evaluation.
- Informs official catalog quality evidence.
Problem
Fixtures can score planted findings, but there is no versioned baseline showing whether a changed skill improves recall, increases unsupported findings, or behaves inconsistently across runs.
User outcome
Reviewers can evaluate a skill change against transparent historical evidence rather than intuition.
Scope
Non-goals
Acceptance criteria
Validation
Relationships