07Agent systems
← The index
Evals with a deploy gate
739 cases that can fail a build.
- Made for
- Newfold Digital
- Years
- 2024 – 2026
- Status
- In production
739 regression cases, ~2,700 assertions, ~250 scoring metrics, wired into Jenkins as a blocking gate. Deterministic assertions can fail a build; model-judged ones report and stay out of the way.
Nearly every take-home test ever set is the second version. It looks rigorous and it cannot separate anybody.
Can the marking tell them apart?
Two people hand in the same task. One understood it; one wrote something that looks right. If the marking cannot separate them, the marking is worthless.
Fig. 1