Skip to content
Agent Engineering Lab

LAB NOTE · LAB-003

The baseline got the right answer and emitted one dispatchable action in thirty runs

Lab-002's sixty committed run records, scored twice. Once on the answer and once on the path. The two metrics rank the two systems in opposite directions, and the single assertion that explains it is whether the proposed remediation could actually be executed.

the address was right; the slot was the wrong shape
the address was right; the slot was the wrong shape
1 of 30
The system that looked twice as good on the answer emitted one machine-executable action in thirty runs

Lab-002 measured whether two architectures reached the right answer. It did not measure whether they got there acceptably. An outcome metric cannot tell a system that is safe from one that got lucky, and this Lab puts a number on that gap using runs that already existed.

It runs no models. It reads the sixty run records Lab-002 committed and applies deterministic assertions to each, which makes it free, instant, and exactly reproducible from data already in the repository.

What it found

SystemOutcome passTrajectory passPassed outcome, failed trajectory
baseline, static single pass20.0%3.3%6 of 6
contender, workflow graph with a gate10.0%33.3%2

The two metrics rank the systems in opposite directions. On the answer the baseline looks twice as good. On the path the contender is ten times better. Every single baseline run that passed the outcome check failed the trajectory check.

One assertion explains it.

Names the correct service and versionEmits a dispatchable action
baseline19 / 301 / 30
contender28 / 3022 / 30

The baseline identifies the right remediation in most incidents and almost never expresses it in a form anything can execute. It writes “Roll back deployment of ‘svc-search’ to its previous version (‘v1.8.2’) immediately.” where the executor needs rollback_deploy:svc-search@v1.8.2.

A keyword-overlap outcome score sees a good answer. An executor sees nothing it can dispatch.

Scope and limits

  • Sixty run records from one Lab, two systems, one incident corpus. The rates are properties of those runs.
  • The trajectory assertions are deterministic checks written after the fact. They encode one view of what an acceptable path looks like, and a different view would produce different numbers.
  • Scoring existing records cannot measure anything the records did not capture.

Reproduce

python -m labs.lab_003.score          # writes RESULTS.md and results.json
python -m labs.lab_003.score --print  # stdout only

No API keys, no Ollama, no network. Inputs are labs/lab_002/results.json and labs/lab_002/incidents.json, both tracked in git.