Harness EngineeringAI Approval Stale Data: Why Snapshots Expire
A human approval binds to one data snapshot. Learn why AI approval stale data checks must recheck inputs before acting, not just before asking.
Rod Rivera

41 entries filed under this tag.
Harness EngineeringA human approval binds to one data snapshot. Learn why AI approval stale data checks must recheck inputs before acting, not just before asking.
Rod Rivera
AI tool error handling must distinguish an empty catalog from unavailable data. Use separate result types and a mutation test that catches a swallowed timeout.
AI agent tool design should expose effects and permission boundaries. Compare an overloaded inventory function with explicit read and draft tools.
AI agent latency measurement means tracing the whole user-visible path, not just timing the model call that feels easiest to optimize.
Agent evaluation variance can make one prompt look better by chance. Design paired repeated trials, report critical failures and inspect discordant outcomes.
A bounded LLM repair loop needs separate attempt, deadline and work limits. Test late responses, exhausted repairs and forbidden writes with a fake clock.
AI cache authorization needs current permissions, revision-aware keys and a defined release point. Test revoked access before and during an answer request.
A retrieved note claims approval for a purchase. See why the agent prompt injection boundary must still deny the write.
Design LLM validation error feedback that identifies malformed fields, limits repair attempts and keeps budget decisions under application control.
A dark software factory can ask agents to improve its tests. Learn how to preserve acceptance authority, inspect proposed changes and challenge the checker with a faulty candidate.
AI agent trusted configuration keeps budgets and prices outside generated proposals. Test top-level and nested overrides against a closed Python schema.
Build a dark software factory around a concrete change request. Separate implementation from acceptance, preserve evidence and decide which decisions remain human.
Turn a working notebook cell into a callable contract with declared inputs, rejected cases and stable outputs before another service depends on it.
An AI agent approval boundary means accepted_draft carries no supplier receipt. Learn to separate validation from execution.
Validate an MCP token for the receiving service as well as its issuer, expiry and scope. Test wrong-audience denial before protected work executes.
An agent evaluation environment must freeze inputs, tools and an independent oracle, or the same message silently grades against two different answers.
AI judge calibration means checking a trace evaluator against blind human labels before routing any case on its verdict alone.
A dictionary of order lines can hide duplicate rows. Check the list for repeats before conversion, not after, to preserve evidence.
RAG access control means filtering by permission before retrieval, not after the answer is written. A worked two-user case shows where checks belong.
AI reviewer context isolation starts with the actual request payload. Build a tested allowlist and separate primary evidence from prior verdicts.
Why a date-only release rule needs one clock across build, request and cache before a scheduled publication timezone decision is trustworthy.
Prompt caching measurement means separating context size, billed usage and latency instead of reading one counter as proof of all three.
Faster drafting can raise total cost if acceptance drops. Work the break-even arithmetic before you trust a lower per-call price.
Learn Pydantic strict validation to decide which conversions a quantity field should accept, and which ones deserve a clear rejection.
An AI agent retry policy must treat timeouts, auth errors, schema faults, and policy refusals as different problems, not one loop.
A timed-out tool call leaves you unsure a write happened. Choose an idempotency key, and reconcile the result before retrying anything.
Build an AI workflow baseline that counts setup, review and repair, separates elapsed time from human effort, and keeps the task mix comparable.
A checkpoint that saves state is not the same as a receipt from the outside world. Learn what must persist before an actor is replaced.
A changed source revision should retire a stored fact without deleting its history. Here is how to tell which memories must go.
Evaluate model routes on ordinary tasks, missing data and policy conflicts, then set acceptance thresholds before comparing cost savings.
AI agent per user credentials means separating who invokes an agent from which account its tools use, so one caller never inherits another's access.
Keep MCP job state durable while connections change. Check authenticated identity and job permissions on each protected request before resuming work.
A new hostname doesn't mean new authority. Learn why AI agent preview deployment needs credential isolation, not just a branch deploy.
A coding agent benchmark evaluation measures one task distribution. Learn to check whether it matches your actual release decision.
Learn to turn one failed agent run into an agent regression dataset that tests recovery without pretending to know the original cause.
Check an LLM proposal against trusted stock, prices and budget after its fields validate, with explicit errors for wrong quantities and stale data.
Learn to write an AI agent handoff checklist that lets a new engineer verify state and evidence instead of trusting a summary.
Anthropic measured it: users approve about 93% of permission prompts. A gate that opens 93% of the time is a doorbell. Here is what the permission system actually enforces, what it cannot, and the standing rules worth adopting before you delegate anything that matters.
Your lunch break has a token price. The model remembers nothing between turns, so every request re-sends the whole conversation — which means cost scales with context carried, not work requested. Here is the cost function, the dials that matter, and the ledger line worth keeping.
Choose an agent task you can verify, bound its effects, and plan the handover. A practical delegation grid, six operating modes, and checks that can miss the real failure.
Make agent work resumable from records: capture decisions, source revisions, test evidence and unresolved effects, then check whether another session can continue safely.
Learn
Save your progress with a free account.
Checking sign-in…
Cookie duty
The one he did not eat is a measurement cookie. Say yes and the site loads Google Analytics (GA4), which stores an identifier in your browser and lets us count which lessons get read.
Change your mind any time with Privacy choices in the footer. Read the Privacy Policy
Optional · newsletter
That is saved and settled. Separately, and this is marketing rather than measurement: Rod writes a weekly email — one lesson, one field note — sent through Kit. It has nothing to do with cookies, and skipping it costs you nothing.
We will send a confirmation email. Click the link in it and the next lesson finds you.