Harness EngineeringAgent Evaluation Variance: Did A Really Beat B?
Agent evaluation variance can make one prompt look better by chance. Design paired repeated trials, report critical failures and inspect discordant outcomes.
Rod Rivera

3 entries filed under this tag.
Harness EngineeringAgent evaluation variance can make one prompt look better by chance. Design paired repeated trials, report critical failures and inspect discordant outcomes.
Rod Rivera
A bounded LLM repair loop needs separate attempt, deadline and work limits. Test late responses, exhausted repairs and forbidden writes with a fake clock.
An AI agent retry policy must treat timeouts, auth errors, schema faults, and policy refusals as different problems, not one loop.
Learn
Save your progress with a free account.
Checking sign-in…
Cookie duty
The one he did not eat is a measurement cookie. Say yes and the site loads Google Analytics (GA4), which stores an identifier in your browser and lets us count which lessons get read.
Change your mind any time with Privacy choices in the footer. Read the Privacy Policy
Optional · newsletter
That is saved and settled. Separately, and this is marketing rather than measurement: Rod writes a weekly email — one lesson, one field note — sent through Kit. It has nothing to do with cookies, and skipping it costs you nothing.
We will send a confirmation email. Click the link in it and the next lesson finds you.