Before you leave...
Take 20% off your first order
20% off
Enter the code below at checkout to get 20% off your first order
Discover summer reading lists for all ages & interests!
Find Your Next Read

The new unit test, for AI that thinks.
Your large language model crushed the benchmark. Then it shipped, and the support tickets started. Sound familiar? A nondeterministic system can be brilliant in the demo and wrong at scale, and you cannot regression-test it the way you test real code. The number you were steering by was measuring someone else's problem.
Done Is a Function You Write is the field manual for LLM evaluation, the discipline that quietly became the center of AI engineering. Its core move is a reframe: "done" is not a vibe or a leaderboard rank, it is a function you write. Once you can write that check, you can delegate to the machine everything that passes it.
Eval-driven development, from your first failing check to a suite that guards every deploy:
This is hands-on AI engineering, not theory. Read it and you will write the checks that decide what a thinking machine is allowed to do, delegate exactly as much as those checks prove safe, and ship AI you can stand behind. That is the scarce skill no model upgrade erases.
For engineers, data scientists, and applied-AI teams building with large language models and agentic AI, past "can the model do it" and stuck on "can I trust it enough to ship." Part of the Build Agents You Can Trust series, in The Verifier's Library.
Thanks for subscribing!
This email has been registered!
Take 20% off your first order
Enter the code below at checkout to get 20% off your first order