
Why Agents Fail: Harness Engineering for AI Projects
The demo worked in the meeting. Three months later the project is dead. Investigate both the model's capability and the system around it before deciding what failed. This course is about the gap in between: if you fund or greenlight AI projects, five plain-language questions that expose a weak one before you commit budget, and what a project actually worth backing looks like. If you build them, the shared vocabulary and failure taxonomy behind harness engineering, worked through against four real architectures — research agents, workflow agents, decision-support agents, and a named case study on Rasa's own policy-enforced CALM system. The capstone is a nineteen-item failure catalogue and a five-hour exercise: name your own project's archetype and build a working harness skeleton for it, a manifest schema, a deterministic verifier, a human-in-the-loop gate, a real defense against one catalogued failure. The five hours are a suggested timebox for a first attempt. Judge the result against your test cases and recorded failures; elapsed time alone does not demonstrate competence. Migrated from a funder-facing keynote and a three-hour practitioner workshop.
MethodThis course teaches a discipline. The examples use today's tools; the method is meant to outlast them.

Rod Rivera
Professor
Desk

First lesson is free to preview.
What You'll Learn
- How model limitations and operational failures can each contribute to the demo-to-production gap, and why both need evaluation
- Five plain-language questions that expose a weak AI project before you approve the budget
- What a well-run AI agent project's staged-autonomy rollout actually looks like
- The shared vocabulary and six-category failure taxonomy behind harness engineering
- The five-component architecture behind research and synthesis agents, with a real manifest schema and verifier
- The two-halves split and session-isolation discipline behind process and workflow agents
- Why a confidence score in a decision-support agent is a legal promise, and how calibration drift is caught before a regulator finds it
- The architecture behind policy-enforced conversational agents, credited to Rasa's own CALM system by name
- How to name your own project's archetype and use a five-hour timebox to start a harness skeleton, then test what it demonstrates
Prerequisites
- No programming experience required for lessons 01-03 (the executive's diagnostic)
- Comfort reading YAML and simple Python for lessons 04-09 (the engineer's path)
- No prior AI agent building experience required — every architecture is worked from first principles
Syllabus

Ready to keep going?
Lesson one is open. Sign the register with a social account to enter the rest of Why Agents Fail: Harness Engineering for AI Projects and keep your place on any device.