Build agents that survive contact with reality.
Most agents are workflows in costume. Reliability lives in the harness, the loop, and the gate, not the model. This lab publishes the evidence.
Latest
The baseline got the right answer and emitted one dispatchable action in thirty runs
Lab-002's sixty committed run records, scored twice. Once on the answer and once on the path. The two metrics rank the two systems in opposite directions, and the single assertion that explains it is whether the proposed remediation could actually be executed.
-
ARCH-002 · BlueprintUpdated Sep 21 · 13 minDetector Architecture
Four detection techniques, a sourcing decision people mistake for a fifth, one interface contract, and the calibration problem that breaks every ensemble built by averaging scores that do not mean the same thing.
-
LAB-004 · LabSep 21 · 10 minPrompt hardening moved three models three different ways. The gate moved all of them to zero.
Fifty support tickets, thirty carrying injected instructions, run through three guard configurations on three local models. Compliance and reachability are measured separately, because they are not the same number and only one of them is under your control.
-
LAB-005 · LabSep 21 · 9 minThe poison was stored, retrieved and cited every time. What changed was what it was allowed to cause.
An agent reads an incident ticket carrying a false causal claim, remembers it, and meets a matching incident later with a clean ticket. Four configurations, measured on three separate things: whether the claim is stored, whether it is cited, and whether it reaches a forbidden action.
-
LAB-006 · LabSep 21 · 9 minA second attempt was worth six points. Deciding when to take one was worth one.
Lab-001 changed two things at once and reported a seventeen point gap. This varies them independently across five arms, reproduces both original numbers exactly, and finds that the metric Lab-001 called routing quality was measuring agreement with a keyword map.
-
LAB-007 · LabSep 21 · 12 minAsking a model to write a probability throws away what it already knows
One forward pass, two readouts. Read as a word the model said No to all fifty tickets and ranked at chance. Read as a distribution over that same token it ranked all fifty correctly. Plus the corpus arithmetic that decides whether any calibration number you read this year means anything.
"Most agents are workflows in costume. Earn the complexity."Read the manifesto →