Skip to content

Felt vs. Measured Transformation

Twelve weeks from now, someone on your team is going to say "this automation push is really working." Your job in this lesson is to get suspicious of that sentence before you ever hear it.

The trial that should worry you

Start with a result that has nothing to do with agencies or twelve-week programs. In a randomized controlled trial run by METR, experienced developers used AI tools on real issues from their own repositories. They took 19% longer to finish those tasks. Afterward, they believed the tools had made them about 20% faster.

Read the two numbers side by side. Not "roughly accurate, give or take." Wrong in the flattering direction, by close to forty points, in people who were good enough at their jobs to know their own codebases cold. If experienced developers cannot feel their way to an accurate read on a single task with a clear start and end, a felt sense of "this is going well" is not a small risk. It is the default failure mode.

What this trial does and doesn't prove

It measured one setting: early-2025 tooling, experienced open-source developers, tasks in codebases they already knew. It does not prove every AI-assisted task runs slower. What it does establish, and what does not expire, is narrower and more useful: feeling faster and being faster are two different measurements, and the gap between them was large enough to matter.

Why a twelve-week program is the harder case, not the easier one

Here's the part most people get backward. A single task at least has two moments to compare: before you started, and when you stopped. The developers in that trial had a clock running the whole time, and they still misjudged it by 39 points.

A twelve-week transformation doesn't give you even that much structure. There is no single before-and-after. There's a slow accumulation of Tuesdays: a client call that went a little smoother, a report that took twenty minutes instead of an hour, a week where everyone was slammed anyway for reasons that had nothing to do with automation. None of those moments naturally lines up against a "before" to compare to. By week eight you're not comparing week eight to week one. You're comparing this week to your memory of last week, which is itself already a comparison to the week before that. The reference point drifts along with you, and a drifting reference point can't tell you anything.

One comparison point versus none

That's the whole shape of the problem. A trial with one comparison point still fooled experienced developers by 39 points. A twelve-week program that never fixes a comparison point at all has no floor under it. "Things seem better" is not a measurement. It's a mood, and moods drift with whatever happened Tuesday afternoon.

The fix: one fixed baseline, five numbers, once a week

The fix this whole course teaches is almost insultingly simple to state, which is exactly why it works: fix Week 1 as the baseline and never let it move. Every week after that, you log five metrics against that same fixed point, not against last week's log, not against your memory. The five metrics get a full lesson of their own later (Lesson 4), so they're only named here: an Automation Index, a Time Liberation Score, a Revenue Efficiency Multiple, a Client Capacity Score, and a recurring-revenue percentage.

The point of naming them now isn't to explain them. It's to show you what they replace. Every one of those five numbers exists to answer a question your feelings cannot answer honestly: not "did this week feel better," but "is this week measurably different from Week 1, and by how much."

You won't be doing this logging by hand in a spreadsheet you'll abandon by week three. This course ships a real, tested Claude Code skill that runs the weekly log-read-adjust rhythm for you. It's not a hypothetical or a "here's how you'd build this" exercise. It's working software, and Lesson 2 is entirely about installing it. Nothing about how it works matters yet. What matters right now is only this: by the end of Week 1, you will have a fixed number to hold every future week accountable to.

The order below is the one detail worth fixing in your head before Lesson 2, because getting it backward is what lets the drift back in:

Fix the baseline once

Week 1 gets logged and then locked. It is never re-measured, never "updated to reflect current context," never touched again until graduation.

Log every week after against that same fixed point

Not against last week's log. Against Week 1, every single time, for all twelve weeks.

Let the comparison replace the impression

The report tells you whether Week 7 beat Week 1. Whether Week 7 felt better than Week 6 stops being the question that matters.

The mistake this course exists to prevent

Twelve weeks in, a program can feel completely transformed and be running flat, for the same reason the developers in that trial felt 20% faster while running 19% slower: a felt sense of progress is not evidence of progress, and the longer the program runs without a fixed comparison point, the less that feeling is worth trusting.

A hypothetical, not a case study

Picture two agencies, purely as an illustration, not a real outcome this program has produced. Agency A reports the program "feels transformative": the team is excited, the workflows feel modern, morale is up. Agency B reports the program "feels the same as always," a little underwhelming even. Before reading on, guess which agency's numbers are actually stronger against a fixed Week 1 baseline.

Quick check — A twelve-week automation program feels like it's going well. What does that feeling actually prove?

What you're actually signing up for

Twelve weeks. Five metrics. One number that doesn't move: your Week 1 baseline. The next lesson gets you off the theory and into the terminal, installing the actual skill and pointing it at a real project.

Continue to Lesson 02

Install the real, tested Claude Code skill this course runs on, and see what it refuses to do before it ever logs a number.

Have a question about this lesson?

Reply here and it goes straight to Rod. Same as replying to one of his emails.