AI agents can implement a feature in minutes. The hard part arrives after: shipping it to real users when you haven't read the code - and, increasingly, when nobody has.
On September 30, Kent C. Dodds spent ninety minutes on a MEGA livestream showing exactly how he does that on Kody, his own SaaS. Not a demo repo. His product, his delivery loop, his gates. This Drop is that session turned into something you can carry to your own stack: the recording, and the fifteen principles underneath it that don't depend on any particular tool.
#The loop, in one picture
What Kent runs is a chain, and each link exists to make the next one trustworthy without a human reading the diff:
- Intent - a goal, discussed with the agent until the context is shared
- Primitives - the documented building blocks every change is classified against
- Implementation - by an agent, in its own context
- Review - by different agents and CI, in their own contexts
- Gates - tests, preview environments, risk grading
- Flags - feature flags and experiment audiences so shipping is reversible
- Production feedback - observability, ship notifications with cost attached
- Cleanup - scheduled agents that delete, migrate and enforce principles after the implementer is gone
Reading is nowhere in that chain. Verification is in six of the eight links.
◆ Bonus for this drop
verification-stack-checklist.md
One-page Markdown checklist: the gates, reviewers, flags and cleanup jobs Kent runs before anything reaches users - adapt it to your repo.
FILE3 KB
verification-stack-checklist.md // gates, reviewers, flags, cleanupFree — we'll email you a one-click link. You'll also get each new Drop. Unsubscribe anytime.
#The fifteen principles
#Trust
1. You trust code because you can verify it, not because you wrote it. You never trusted code because you typed it. You trusted it because you could check it. Agents writing faster than you can read doesn't change the principle - it raises the stakes on verification, gates and rollback, and lowers the value of authorship to nearly zero.
2. You cannot ship good software if you don't understand what it's for. Process without product context is empty. That used to be delegable: a PM understood users, you took tickets. Agents take tickets now. The human job moved to understanding users, pain and intent.
3. In user feedback, hunt the core problem, not the suggested solution. The solution a user proposes is sometimes useful and often noise. Find the pain underneath. Early on, think hard about every piece of feedback without committing to implement it - and hold on to the users who report real problems in the space you're solving.
#Context
4. Align shared context before you ask for implementation. Load the agent with what you know - including asking questions you already have answers to. If it explains something wrong, you've caught the misalignment before a line of code exists. Lock intent first; then "go build it."
5. Keep climbing the abstraction ladder. Tab-complete → chat while watching the code → chat that owns the edit → agents managing agents. The useful move is always upward. The question to ask constantly: is the thing I'm doing right now something an agent could do? If yes, hand it off. If it can't do it reliably yet, make it reliable.
6. Durable knowledge lives in the repo, not in one agent's memory. A brand-new agent with no chat history should produce good work. Vision, invariants, decision records, contribution rules, system structure - all in the repository. The goal is agent-agnosticism: switch models or harnesses without losing the product's brain.
7. Make the project self-healing with a friction log. Agents quietly work around pain - slow starts, broken scripts, missing docs - and never complain. Instruct them to log friction instead. A scheduled cleanup agent fixes the underlying issue, and every future run gets cheaper.
#Primitives
8. Primitives are how you stop agents inventing slop. Document the building blocks - auth, UI surfaces, storage, billing - and force every change to classify itself: reuse, combine, expand, or (rarely) add a new primitive. Prefer progressive discovery over dumping all docs at once; every new agent is a new hire, and overload confuses them. Grade work low/medium/high risk by which primitives it touches, and agents gravitate to the smallest safe change.
#Shipping
9. Prefer two-way doors; use feature flags when you can't. Ask how easy a decision is to back out of. Shipping a public primitive users depend on is a one-way door. Flags and "experiments, may break" audiences turn many one-way doors into two-way ones - ship, learn, unship, no apology email.
10. Agents don't delete code. Make cleanup intentional. Agents add fallbacks and leave legacy paths forever; nobody pays them to delete. Track migrations with cleanup issues that state how you'll know removal is safe. Instrument unused paths. Run a scheduled "legacy reaper" that deletes what's safe and instruments what isn't - and the same for tests and docs.
11. Do it the hard way first, then automate, then make it deterministic. Do painful work manually long enough to understand the real problem. Then: human work → agent work → deterministic software for the repeatable parts. Agents don't feel your pain, so they won't push work into durable software unless you require it. Software any agent can call beats automations trapped in one vendor's walled garden.
#Verification
12. Separate implementers from reviewers - different context windows. Never let the agent that built the change be the only reviewer. Separate AI reviewers and CI, each in its own context, catch what the author's context rationalizes away. Tell the implementer to treat review comments as suggestions to validate, not orders. Earn trust in reviewers by using them manually before automating the loop.
13. Agent-written tests are not a sufficient gate. Same agent wrote the feature and the tests? Green CI is a weak signal. Comfort shipping without reading the diff comes from layers: strong AI review, CI, preview environments the agent can exercise, feature flags, and disaster recovery you have actually tested. Tests matter. They aren't enough alone.
#Economics
14. Judge spend by value created, not sticker price. Parallel agents cost real money. Cut cost first: friction log, better docs, deterministic packages. If you still spend, make sure what ships is worth more than it cost - and put cost and difficulty into ship notifications so you stay honest.
15. Your unique value is no longer implementing code. Leverage is the argument for all of this - not "ship more, forever." Offload implementation so attention goes to judgment, product direction, hard trade-offs, and the long conversations that don't fit in a ticket. Implementing code stopped being a unique human value proposition. Imagining, judging and caring still are.
#Where this goes in MEGA
This session is Week 3 - Product - in miniature. Kent's week takes the loop above and makes you build it on your own repo: primitives documented, reviewers separated, gates that fail on an empty diff before they pass on your change, flags, and the cleanup agents that keep it honest. Week 2 with Angie builds the agents the loop runs on; Week 4 with John turns it into a factory.
#Further reading
- I was wrong about MCPs - Kody's capability layer, which is what principle 8 looks like as architecture
- Don't make the agent remember how to do its job - principle 11 in practice: procedures become packages
- Towards Autonomous Product Development - the same loop, run with coordinators and workers in the cloud
- Kody, open source: github.com/kentcdodds/kody