Student edition · 90 minutes of dedicated work · 2026-09-09
This is one of two practical units for Chapter 15. Unit A constructs and connects the mechanism; Unit B investigates a controlled failure, repairs it and transfers the invariant. Each is a complete ninety-minute session, with its own setup and required conceptual introductions. Basic Python variables, conditions, loops, functions, lists and dictionaries are the starting knowledge. Libraries and specialized concepts used here are introduced below before the main task.
By the end you should be able to:
- Explain the chapter's mechanism using a prediction and an observed intermediate result.
- Repair the failure: Ignoring reservations makes the supposedly simpler baseline wrong while the agent can still be right. Feed the same Case through evaluate with OfflineShopModel and compare authored expected quantities to both baseline and real tool observations.
- Solve build an oracle that cannot copy the expected answer using changed inputs and an independent expectation.
- Retain your implementation, failed/corrected observations, causal explanation and limits.
| Minutes | Dedicated activity | Evidence you produce |
|---|---|---|
| 0–5 | State the problem and make a prediction | Initial prediction in your own words |
| 5–25 | Foundations and library examples | Values, explanations, revised predictions |
| 25–35 | Trace setup and the main interface | Input → learner function → observation |
| 35–60 | Reproduce, diagnose and repair | Source, visible checks and runtime evidence |
| 60–80 | Implement and challenge the transfer task | Function and a new counterexample |
| 80–90 | Retrieve, explain and save | Exit ticket and retained submission |
Installation is preparation time. These are planning estimates, not measured completion times. Use the reference primers when a term is unfamiliar; in Unit B retrieve an explanation before re-reading it. Run All checks that the artifact executes. Unfinished student functions deliberately produce NEEDS_WORK. Keep your first attempt before opening answers.
This notebook belongs to the nineteen-chapter edition. Its supplied teaching runtime is embedded, so it can run without the textbook or another notebook. Where code uses REFERENCE_LESSON, that is the frozen runtime exercise identifier; the reader-facing chapter and saved unit identifiers use the current edition. Building against a supplied runtime is not proof that you have constructed all of its dependencies.
Run the self-contained setup
Use a Python 3.14 Jupyter kernel and Pydantic 2. If needed, run
%pip install "pydantic==2.13.4" once in a separate cell and restart the kernel.
Package installation needs internet; the lesson itself needs no repository download, API key
or prior notebook. The complete source runtime uses Python 3.14, so this edition does not claim
compatibility with a hosted notebook service's default interpreter.
The collapsed cell contains 89 frozen teaching files. Base85 represents compressed bytes
as text; zlib decompresses them; SHA-256 checks that the decoded files match this edition.
These are supplied packaging operations, not learner algorithms. tempfile creates an isolated
working copy; Path handles file locations; sys.path tells Python where the supplied modules
live. The code is available for inspection below and performs no package installation itself.
The subsequent lesson teaches the libraries used by the mechanisms you will implement.
Run setup on every fresh kernel. It writes scratch runtime files separately from your retained
practical-work/ch15-b folder. Rerunning setup restores the frozen support files and keeps
your saved work. Restarting a kernel clears variables, not saved submission files. Source basis:
Sovereign Agent 444c5f6. Some tasks use reviewed local subprocesses; they are not an OS sandbox.
Supplied offline setup and teaching files
The complete supplied setup cell is in the downloadable notebook. Run it there before attempting the cells below. This web edition omits only that compressed setup payload.
Commit to a prediction before the examples
prediction_notes = {
"prediction": "Write the expected behavior before running the worked example.",
"reason": "Name the input and rule behind that prediction.",
"falsifier": "Name an observation that would prove the explanation wrong.",
"revision": "After execution, explain what changed in your understanding.",
}Reading the Python vocabulary used in this notebook
You need basic assignments, if, loops, functions, lists and dictionaries. The less familiar
features used by the supplied code are introduced here. A library is reusable code that
Python can import. The standard library ships with Python; Pydantic is an additional package.
An import makes a name available, but does not mean that you have completed the exercise.
JSON is text for exchanging structured values. A Python dictionary is an in-memory object;
the JSON representation is a string. Use json.dumps to encode and json.loads to decode.
Decoding proves that text has valid JSON syntax, not that its fields match our business contract.
Predict which of the following two decoded objects could describe a stock count.
import json
intro_data = {"sku": "MANGO", "count": 3}
intro_text = json.dumps(intro_data, sort_keys=True)
print(type(intro_data).__name__, type(intro_text).__name__, intro_text)
print(json.loads(intro_text))
print("Also valid JSON:", json.loads('["not", "a", "stock", "record"]'))
assert json.loads(intro_text) == intro_dataThe first result is a dictionary; the second is a list. Before indexing a decoded object,
check the shape that your function promises to accept. An exception interrupts the normal
path. raise ValueError(...) refuses an invalid value; try/except lets a caller inspect that
expected refusal. Catch the expected class, rather than turning every programming error into
apparent success. finally runs cleanup even when an earlier operation raises.
An annotation, such as count: int, documents the expected type. It does not by itself
enforce the type at runtime. A class defines a kind of object; an instance holds one object's
data. @dataclass asks Python to generate routine construction and comparison methods from
annotated fields. frozen=True prevents ordinary reassignment of the instance's fields; it does
not make every object nested inside those fields immutable. A method is a function attached
to a class; self refers to the instance receiving the call.
from dataclasses import dataclass
@dataclass(frozen=True)
class IntroObservation:
operation: str
count: int
intro_observation = IntroObservation("count-mango", 3)
print(intro_observation.operation, intro_observation.count)
assert intro_observation == IntroObservation("count-mango", 3)A callback is a function passed to another function. This is how the classroom harness
invokes your implementation. The argument candidate below is a function object; parentheses
perform the call. Predict the two answers before execution, then trace the result to the callback.
def intro_apply(candidate, value):
return {"input": value, "observed": candidate(value)}
def intro_double(value):
return value * 2
print(intro_apply(intro_double, 3))
print(intro_apply(lambda value: value + 2, 3))
assert intro_apply(intro_double, 3)["observed"] == 6lambda value: value + 2 is a small anonymous function. A closure is a function that retains
access to values from its surrounding scope. It can bind a tool to a shop snapshot. A shallow
copy duplicates only the outer container; copy.deepcopy also copies nested containers used
in these fixtures. A set stores distinct values; required <= allowed asks whether every
required item is allowed. frozenset is the corresponding immutable set. A tuple groups ordered
values; (value,) is a one-item tuple, including the comma.
Paths and cleanup. Path represents a filesystem location. path / "file.json" constructs
a child path; read_text and write_text read and write text. A context manager, used with
with, manages entry and exit. A temporary-directory context removes its contents on exit.
Save your submission outside temporary runtime directories. Reopening a file is different from
reusing a Python variable: the former tests retained bytes, while the latter only tests this kernel.
from pathlib import Path
from tempfile import TemporaryDirectory
with TemporaryDirectory() as intro_folder:
intro_path = Path(intro_folder) / "observation.json"
intro_path.write_text(json.dumps(intro_data), encoding="utf-8")
intro_reopened = json.loads(intro_path.read_text(encoding="utf-8"))
assert intro_reopened == intro_data
print("Read from a file:", intro_reopened)Retrieval check: explain JSON versus a dictionary, annotation versus validation, class versus instance, and defining a callback versus invoking it. Change the callback above so an incorrect implementation visibly changes the observed output. This distinction will matter when grading your connected work. Reference: Python's JSON, dataclasses, and pathlib documentation.
Pydantic: turn an input dictionary into a checked object
Pydantic is an additional Python library for validating data. Its model is a class describing
fields, not a neural network. Inherit from BaseModel, declare annotated fields, then call
model_validate on incoming data. A field without a default is required. A field with a default
can be omitted. The result is an instance whose values you read with dot notation.
from pydantic import BaseModel, ConfigDict, Field, ValidationError
class IntroCourseRequest(BaseModel):
model_config = ConfigDict(strict=True, extra="forbid")
name: str = Field(min_length=1)
quantity: int = Field(gt=0, le=1000)
note: str = ""
intro_request = IntroCourseRequest.model_validate({"name": "mango", "quantity": 4})
print(intro_request.name, intro_request.quantity, repr(intro_request.note))
assert intro_request.note == ""The annotation says the field's type. Field supplies constraints: gt=0 means greater than
zero, le=1000 means at most 1000, and min_length=1 excludes an empty name. ConfigDict sets
model-wide behavior. strict=True rejects conversions for this integer field, including "4",
4.0 and True; extra="forbid" rejects undeclared keys. Pydantic can otherwise convert some
compatible inputs, so choose this boundary deliberately rather than assuming every accepted
input arrived in the expected type.
Predict which rule refuses each payload. ValidationError reports a failed contract. Its
errors() entries contain loc, the field location, and type, the failure category. Catching
that expected exception lets the notebook inspect the failure and continue.
intro_bad_requests = [
{"name": "mango", "quantity": "4"},
{"name": "mango", "quantity": True},
{"name": "", "quantity": 4},
{"name": "mango", "quantity": 0},
{"name": "mango", "quantity": 4, "approved": True},
{"quantity": 4},
]
for intro_bad_request in intro_bad_requests:
try:
IntroCourseRequest.model_validate(intro_bad_request)
except ValidationError as intro_error:
print([(item["loc"], item["type"]) for item in intro_error.errors(include_input=False)])
else:
raise AssertionError("An invalid input crossed the declared contract")Use model_dump() for a Python dictionary, model_dump_json() for JSON text, and
model_validate_json() to parse and validate JSON. model_json_schema() describes the contract;
it is neither an instance's current values nor an invocation of the business handler.
intro_serialized = intro_request.model_dump_json()
intro_schema = IntroCourseRequest.model_json_schema()
assert IntroCourseRequest.model_validate_json(intro_serialized) == intro_request
print("Actual values:", intro_request.model_dump())
print("Quantity contract:", intro_schema["properties"]["quantity"])
assert intro_schema["properties"]["quantity"]["exclusiveMinimum"] == 0Four is valid input to this schema even if the shop needs six. Pydantic checks the declared shape and constraints; the handler still needs authoritative stock, price and permission. Ordinary assignments to an existing instance are not automatically revalidated unless configured for assignment validation. This lesson validates new input at the boundary and uses the resulting values. Explain these limits before relying on a model object in a transaction or tool call.
Chapter 2's full introduction expands this pattern with a separate data-repair checkpoint. This notebook contains the required pattern here so prior Pydantic experience is not needed. References: models, fields, and strict mode.
SQLite from the first row to an atomic change
A dictionary disappears when its process ends. A database can retain records so a later process
can resume from evidence. SQLite is an embedded database: Python's sqlite3 library opens
a local database file without starting a separate database server. SQL is the language used
to define, select and change its records. A table has named columns and rows. A primary
key identifies a row; a query asks for rows satisfying a condition.
Start with a deliberately small preference table. Read the SQL as instructions: create the
table, insert one named value, then select the value for one session. ? is a parameter
placeholder; the values are passed separately so they are data, not SQL instructions.
fetchone() returns one row or None; it does not guarantee that a matching row exists.
import sqlite3
from pathlib import Path
from tempfile import TemporaryDirectory
with TemporaryDirectory() as intro_sql_folder:
intro_db_path = Path(intro_sql_folder) / "example.sqlite"
intro_db = sqlite3.connect(intro_db_path, autocommit=True)
intro_db.execute("CREATE TABLE preference (id INTEGER PRIMARY KEY, session TEXT, value TEXT)")
intro_db.execute("INSERT INTO preference (session,value) VALUES (?,?)", ("lucy", "09:00"))
intro_row = intro_db.execute(
"SELECT value FROM preference WHERE session=?", ("lucy",)
).fetchone()
print("Matching row:", intro_row)
assert intro_row == ("09:00",)
assert (
intro_db.execute(
"SELECT value FROM preference WHERE session=?", ("another-session",)
).fetchone()
is None
)
intro_db.close()
intro_reopen = sqlite3.connect(intro_db_path, autocommit=True)
assert intro_reopen.execute("SELECT count(*) FROM preference").fetchone()[0] == 1
intro_reopen.close()
print("The row survived closing and reopening its connection.")Index [0] selects the first column of a returned tuple. sqlite3.Row is an alternative row
factory that also permits named-column access. dict(row) then produces an ordinary dictionary.
The book's Database wrapper supplies that configuration and its schema; the wrapper is course
code, while sqlite3 is the standard library. You will use the public connection and transaction
methods explained at the exercise boundary, rather than needing to reconstruct the wrapper.
Now consider a budget. Moving five pence from reserved to spent requires two values to change
together. A transaction makes a group of local changes commit together or roll back together.
The example explicitly controls SQL transactions with autocommit=True and SQL statements.
BEGIN IMMEDIATE starts a write transaction; COMMIT keeps its changes; ROLLBACK discards them.
Predict the row after the deliberately raised exception. Catching an error alone would not undo
the first update; the rollback is the operation that restores the prior state.
intro_ledger = sqlite3.connect(":memory:", autocommit=True)
intro_ledger.execute(
"CREATE TABLE budget (id INTEGER PRIMARY KEY, reserved INTEGER, spent INTEGER)"
)
intro_ledger.execute("INSERT INTO budget VALUES (1,5,0)")
intro_ledger.execute("BEGIN IMMEDIATE")
try:
intro_ledger.execute("UPDATE budget SET reserved=0 WHERE id=1")
raise ValueError("injected failure before the matching spend update")
except ValueError:
intro_ledger.execute("ROLLBACK")
assert intro_ledger.execute("SELECT reserved,spent FROM budget").fetchone() == (5, 0)
intro_ledger.execute("BEGIN IMMEDIATE")
intro_ledger.execute("UPDATE budget SET reserved=0,spent=5 WHERE id=1")
intro_ledger.execute("COMMIT")
print(
"After a complete change:", intro_ledger.execute("SELECT reserved,spent FROM budget").fetchone()
)
intro_ledger.close()The literal :memory: creates a temporary database inside this connection; it is useful for the
small experiment, but the earlier file example establishes persistence. Neither example proves
that a remote supplier rolls back when the local transaction rolls back. An external operation
has its own state and evidence.
SQL you will meet later. UPDATE ... SET ... WHERE ... changes selected rows. AND combines
conditions. ORDER BY makes an ordering explicit; absent that clause, do not rely on row order.
count(*) counts rows; sum(amount) totals a column and can be NULL on an empty input;
coalesce(sum(amount),0) uses zero for that empty aggregate. A UNIQUE constraint rejects
duplicate identities. GROUP BY status computes one aggregate per status.
An invariant is a condition that must remain true across operations, such as nonnegative
reserved money. A snapshot is a consistent view at one point; two separate reads can describe
different moments unless their transaction contract binds them. In the book, with db.immediate()
groups related writes. It is a course-defined context manager with the commit/rollback purpose
you just observed. Do not assume that an arbitrary with connection has identical behavior under
every SQLite autocommit setting.
Your prediction: two workers both read ten remaining pence outside a transaction and each approve seven. Why can both believe the next order fits? Explain what must be checked together with the write. Then change the example's initial reserved amount and repeat the failure. Reference: Python's SQLite tutorial and transaction control.
Evaluation asks a precise question of independently specified evidence
An agent can produce fluent text and still order the wrong amount. An evaluation case fixes inputs, a task and expected business outcomes. An oracle defines those outcomes independently of the candidate. A baseline is a simpler method used for comparison. A metric summarizes observations; it does not replace the per-case evidence that produced it.
For Lucy, relevant facts include product identity, physical stock, reserved stock, target and price. Sellable stock is physical stock minus reserved stock. If physical stock is six, four are reserved and the target is five, the deficit is three. A baseline that ignores reservations will report no deficit and can make a correct agent look wrong.
Calculate a case before writing its oracle
Predict the needed quantities for the rows below, including the exact-threshold row. The expected dictionary is written independently, before the baseline calculation. Do not replace it with the candidate's own output when the comparison fails.
intro_case_rows = [
{"sku": "MANGO", "physical": 6, "reserved": 4, "target": 5},
{"sku": "COCOA", "physical": 7, "reserved": 2, "target": 5},
]
intro_expected = {"MANGO": 3}
intro_baseline = {}
for intro_row in intro_case_rows:
intro_available = intro_row["physical"] - intro_row["reserved"]
intro_need = max(0, intro_row["target"] - intro_available)
if intro_need:
intro_baseline[intro_row["sku"]] = intro_need
print(intro_baseline)
assert intro_baseline == intro_expectedThe baseline omits products with no deficit. Including zero entries would change its declared output contract. Empty input should produce an empty dictionary. Renaming the SKU must change the identity in the output without changing arithmetic. These are useful checks against a candidate that simply looks up the visible example.
Keep the denominator visible
Suppose one system solves nine of ten easy cases and none of two difficult cases. “90%” describes only the easy subset. A reported aggregate needs its population and missingness. Unrun cases are not successes or observed failures; retain their status separately. The same discipline applies to cost and latency: specify which attempts the total includes.
intro_case_results = [
{"case": "ordinary", "status": "PASS", "pence": 2},
{"case": "reserved", "status": "FAIL", "pence": 3},
{"case": "provider-down", "status": "NOT_RUN", "pence": None},
]
intro_observed = [row for row in intro_case_results if row["status"] != "NOT_RUN"]
intro_passes = sum(row["status"] == "PASS" for row in intro_observed)
print("Observed pass fraction:", intro_passes, "/", len(intro_observed))
print("Unrun:", len(intro_case_results) - len(intro_observed))
print("Known recorded cost:", sum(row["pence"] for row in intro_observed))
assert intro_passes == 1 and len(intro_observed) == 2The fraction is one of two observed cases, with one unrun case. A mean alone can hide an important failed case, so the course retains each observation. Holdouts are additional cases withheld from the visible feedback. They test transfer, but a finite holdout suite is still not proof for every possible input or every natural-language claim.
A saved report needs a subject and provenance
Record the evaluated configuration, scenario identities and exact evidence. A digest can detect
changed report bytes; it cannot establish that the report's oracle was correct. A supplied
Case.expected is the reference for grading, not an input the baseline may read to generate its
answer. A candidate returning that field would pass by copying the answer rather than calculating.
The actual baseline function is connected to evaluate, which also invokes the offline shop
model and actual tools. The construction task implements the deterministic calculation from
stock rows. The failure task removes reserved stock from the oracle's arithmetic. Inspect both
baseline and tool observations to determine which side is wrong; do not automatically trust the
component named “baseline.”
Before coding, write cases for no products, exact threshold, a renamed product and reservations larger than physical stock. Add a fluent but wrong explanation while holding correct tool calls fixed. Explain why substring checks cannot establish arbitrary prose faithfulness and why an honest review-required outcome is preferable to a fabricated pass.
Choose an explicit starting point for this independent notebook
This Unit B runs without Unit A. By default it prepares a supplied reference starting point
and labels its provenance. It is not evidence that you built Unit A. To investigate your own
successful implementation, set LEARNER_HANDOFF to its saved path before running the cell.
An invalid selected file refuses; it is never silently replaced with the reference.
SourceTask supplies copied-source execution and handoff validation; RuntimeLab supplies the
controlled failure experiment. Their public operations are introduced beside the main exercise.
The artifact stores identity and observations; no variables from another kernel are required.
LEARNER_HANDOFF = NonePrepare and validate the supplied starting artifact
import json
import runpy
import shutil
import textwrap
from pathlib import Path
COURSE_INPUT = COURSE_WORK / "ch15-unit-a-handoff-v1.json"
if LEARNER_HANDOFF is not None:
learner_input = Path(LEARNER_HANDOFF).expanduser().resolve()
if not learner_input.is_file():
raise FileNotFoundError("The selected learner handoff does not exist")
if learner_input != COURSE_INPUT.resolve():
shutil.copy2(learner_input, COURSE_INPUT)
HANDOFF_ORIGIN = "LEARNER_SELECTED"
else:
source_task_class = runpy.run_path(
str(COURSE_ROOT / "book/always_on/exercises/source_tasks_v1.py")
)["SourceTask"]
reference_task = source_task_class(COURSE_ROOT, 12)
try:
reference_task.install(textwrap.dedent(reference_task.fragment))
reference_observation = reference_task.visible("SUPPLIED_REFERENCE_START")
if reference_observation["status"] != "PASS":
raise RuntimeError("The supplied starting point did not pass its connection check")
reference_task.save(COURSE_INPUT, reference_observation)
finally:
reference_task.close()
HANDOFF_ORIGIN = "SUPPLIED_REFERENCE"
print("Starting evidence:", HANDOFF_ORIGIN)
print("The core task below validates the selected artifact before using it.")Understand the supplied execution interface
The course runtime is provided so your implementation can be connected to real callers and
storage. SourceTask(ROOT, chapter) makes a private copy. install(source) replaces only the
declared function; visible() invokes the real chapter probe; save(path, result) retains a
successful implementation and its evidence. load(path) checks the saved identities and hashes.
inject_failure() changes the declared boundary; repair(fragment) replaces that broken fragment.
close() removes the scratch copy after you retain evidence. These methods are supplied harness
operations, not additional packages you must discover or install.
RuntimeLab provides the same copied-source failure experiment without the complete-function
construction layer. Its run method records exit status, observations and the compared expectation.
A subprocess log from an unfinished learner implementation is feedback about that implementation;
it is not a successful connection. A syntax error in the notebook cell itself is a separate issue
to fix. The task below names which interface it uses.
For direct-function units, the visible driver calls your callback without installing a source string. In either case, trace where your code is invoked. Supplied fixtures, database wrappers and replay models are labelled infrastructure; your own implementation and changed-case explanation are the evidence of learning.
Main practical: construct, connect and challenge
Ignoring reservations makes the supposedly simpler baseline wrong while the agent can still be right. This time you begin with your Unit A implementation and its saved evidence. Lucy can adopt a faulty scripted baseline as her standard unless independently authored answers challenge it.
Verify the handoff
The starting-point cell has selected the Unit A artifact explicitly. A selected learner handoff must validate; the default reference start is labelled separately. Run the setup and keep the runtime and implementation hashes in your submission.
import json
import os
import runpy
from pathlib import Path
ROOT = COURSE_ROOT
SourceTask = runpy.run_path(str(ROOT / "book/always_on/exercises/source_tasks_v1.py"))["SourceTask"]
REFERENCE_LESSON = 12
HANDOFF = Path("ch15-unit-a-handoff-v1.json")
handoff_status = "MISSING"
if HANDOFF.is_file():
task = SourceTask(ROOT, REFERENCE_LESSON)
try:
handoff = task.load(HANDOFF)
handoff_status = "VERIFIED"
print("IMPLEMENTATION", handoff["implementation_sha256"])
finally:
task.close()
print("UNIT_A_HANDOFF", handoff_status)Reproduce and diagnose
Predict the consequence of this injected boundary before executing it:
(sku, threshold - stock)The controlled mutation changes the same implementation you submitted. It refuses if the declared mutation boundary no longer occurs exactly once; inspect an alternative implementation with the instructor before adapting the experiment.
baseline = broken = None
if handoff_status == "VERIFIED":
task = SourceTask(ROOT, REFERENCE_LESSON)
try:
task.load(HANDOFF)
baseline = task.visible("YOUR_BASELINE")
if baseline["status"] != "PASS":
raise ValueError("Saved Unit A code no longer satisfies the visible contract")
task.inject_failure()
broken = task.run("INJECTED_FAILURE", expected=task.spec["expected_broken"])
print("BEFORE", baseline["observation"])
print("AFTER", broken["observation"])
finally:
task.close()
else:
print("HANDOFF_REQUIRED: complete Unit A before performing Unit B")State a diagnosis using those two observations. Name a test that would prove your diagnosis wrong. Feed the same Case through evaluate with OfflineShopModel and compare authored expected quantities to both baseline and real tool observations.
Repair the boundary
Return the complete replacement for the injected fragment. Do not edit the oracle or print a desired observation. Repair the actual source. The starter keeps the defect so the learner outcome remains incomplete.
def repair_fragment():
return "(sku, threshold - stock)"Hint 1 — the consequence
Lucy can adopt a faulty scripted baseline as her standard unless independently authored answers challenge it.
Hint 2 — the evidence
Compare the two observations, then trace the changed field to baseline in src/reference_organizations/store/evaluation.py. Distinguish a schema refusal from a business-rule or authority refusal.
Hint 3 — the design
For each stock row compare available stock with threshold; calculate needed from the same quantities; do not read case.expected to produce an answer.
def connect_repair(fragment):
task = SourceTask(ROOT, REFERENCE_LESSON)
try:
task.load(HANDOFF)
task.inject_failure()
task.repair(fragment)
return task.visible("YOUR_REPAIR")
finally:
task.close()
repair_result = None
if handoff_status == "VERIFIED":
repair_result = connect_repair(repair_fragment())
print("REPAIR", repair_result["status"], repair_result["observation"])
else:
print("REPAIR_NOT_ATTEMPTED: missing Unit A evidence")Transfer under a changed constraint
Use an empty catalog, exact threshold, renamed products and high reservations. A visible-case lookup and an implementation returning case.expected must fail.
Create a fresh task, load your handoff, inject the defect and apply your repair. Then change only the copied probe to exercise the new condition. Keep the actual observation and a prediction written beforehand. Explain why a visible-case lookup or a blanket refusal could pass the original example but fail this transfer.
The instructor's holdout applies your repair to a new copied runtime and checks both the positive case and the missing protection. An exact exception or changed state must cause a failure; no broad error is accepted as successful refusal.
Exit ticket
Submit the original handoff, baseline and broken observations, repair, transfer probe and results. State what Lucy would experience before and after the fix. Identify the guarantee that still requires separate evidence: A passing finite scenario suite is not proof of arbitrary language faithfulness.
passed = repair_result is not None and repair_result["status"] == "PASS"
exercise_report = {
"unit": "ch15-b",
"attempted": int(repair_result is not None),
"completed": int(passed),
"failed": int(repair_result is not None and not passed),
"skipped": int(repair_result is None),
"connection": "PASS" if passed else "NOT_READY",
"handoff": handoff_status,
}
print("EXERCISE_REPORT=" + json.dumps(exercise_report, sort_keys=True))Changed-constraint construction: Build an oracle that cannot copy the expected answer
Allow twenty minutes. Spend three minutes predicting, ten implementing and tracing, five on a new case of your own, and two explaining the surviving limitation. This is dedicated work, not an invitation to run a supplied answer. Both units revisit the same invariant after different core experiences; in Unit B, attempt this task from memory before consulting Unit A.
Implement transfer_check(rows). Each row has sku, physical, reserved and target nonnegative integer fields. Return a dictionary of only positive deficits, calculated from physical minus reserved. There is deliberately no expected field in this interface. Inputs have unique identities and must remain unchanged.
Write your expected values before running the table. Keep one accepted case and one refusal. Your function is passed directly into the driver below. The driver copies inputs and checks they remain unchanged; it does not replace your implementation with the reference answer.
Hint 1 — identify the authoritative inputs Name the source field for each output value. Which input changes while the rule remains the same?
Hint 2 — choose the boundary cases Start with exact empty, exact equality and one value on each side of the boundary where valid. Do not add a special case for a visible product name or operation identity.
def transfer_check(rows):
raise NotImplementedError("Calculate independent deficits from stock inputs")import copy
import json
TRANSFER_CASES = [
(
"reserved deficit",
[[{"sku": "MANGO", "physical": 6, "reserved": 4, "target": 5}]],
{"MANGO": 3},
),
("exact threshold", [[{"sku": "MANGO", "physical": 7, "reserved": 2, "target": 5}]], {}),
("empty", [[]], {}),
(
"renamed identity",
[[{"sku": "PEAR", "physical": 1, "reserved": 0, "target": 4}]],
{"PEAR": 3},
),
]
def same_transfer_value(actual, expected):
if type(actual) is not type(expected):
return False
if isinstance(expected, dict):
return actual.keys() == expected.keys() and all(
same_transfer_value(actual[key], value) for key, value in expected.items()
)
if isinstance(expected, list):
return len(actual) == len(expected) and all(
same_transfer_value(a, e) for a, e in zip(actual, expected, strict=True)
)
return actual == expected
def run_transfer(candidate, cases):
observations = []
for label, arguments, expected in cases:
supplied = copy.deepcopy(arguments)
before = copy.deepcopy(supplied)
raised = None
try:
actual = candidate(*supplied)
except NotImplementedError:
raised = "NotImplementedError"
actual = {"unfinished": True}
except Exception as error:
raised = type(error).__name__
actual = {"raises": raised}
expects_error = isinstance(expected, dict) and set(expected) == {"raises"}
correct = (
raised == expected["raises"]
if expects_error
else (raised is None and same_transfer_value(actual, expected))
)
passed = correct and same_transfer_value(supplied, before)
observations.append(
{"case": label, "expected": expected, "observed": actual, "passed": passed}
)
print("PASS" if passed else "NEEDS_WORK", label, "expected", expected, "observed", actual)
return observations
transfer_observations = run_transfer(transfer_check, TRANSFER_CASES)
TRANSFER_PASSED = all(row["passed"] for row in transfer_observations)
print("TRANSFER_STATUS", "PASS" if TRANSFER_PASSED else "NEEDS_WORK")Design a counterexample and retrieve the mechanism
Add one new case with an independently calculated expected outcome to TRANSFER_CASES and rerun
the driver. Change one condition at a time. Then deliberately replace your candidate with a
constant answer in a temporary copy and show a case that rejects it. Restore your implementation.
Explain why that counterexample is stronger than repeating the original example with a new name.
Without viewing the worked example, write the invariant in words and trace one observed value back to its input. Identify which part is a local fixture result and which claim would need a live provider, host or external-system observation. Keep a first attempt even if you used a hint.
Save your evidence and explain the result
Fill the prediction notes and your explanation before saving. Include the exact observed value, the input or retained row that caused it, your code's invocation point, one failed hypothesis, and the strongest claim the evidence still cannot support. A completed code cell alone does not earn explanation credit. Do not label reference-start behavior as your own Unit A construction.
Keep this edited notebook, the Markdown if used for notes, saved handoff files, and the JSON record below. Your work folder survives scratch cleanup and can be reopened in a new kernel. An instructor can ask for an unseen case after the visible checks; keep your implementation general.
explanation_notes = {
"causal_trace": "Explain the input, learner invocation and observed result.",
"failed_hypothesis": "Describe a prediction the evidence changed.",
"remaining_limit": "Name the guarantee not established by this experiment.",
}
course_submission = {
"unit": "ch15-b",
"planned_minutes": 90,
"starting_evidence": globals().get("HANDOFF_ORIGIN", "INDEPENDENT_UNIT_A"),
"prediction": prediction_notes,
"explanation": explanation_notes,
"core_report": exercise_report,
"transfer": transfer_observations,
"explanation_review": "HUMAN_REVIEW_REQUIRED",
}
submission_path = COURSE_WORK / "ch15-b-submission-v1.json"
submission_path.write_text(
json.dumps(course_submission, indent=2, sort_keys=True), encoding="utf-8"
)
print("Saved evidence:", submission_path)
print(
"COURSE_REPORT="
+ json.dumps(
{
"unit": "ch15-b",
"transfer_passed": TRANSFER_PASSED,
"starting_evidence": course_submission["starting_evidence"],
"edition": "student",
},
sort_keys=True,
)
)Keep building with Prof Rod
Found this material through a colleague, classroom or shared download? Get the complete book at profrod.ai/book and join the Prof Rod learner community. Bring one result, one question or one failure you learned from. Share this resource with another learner and keep its source links with it so they can find the full course and future updates.