Evals, Observability & Quality
Measure AI behavior, catch regressions, trace cost and latency, and keep workflows improving.
12 stories (5 articles · 7 videos)
Start here
A few good first pieces before you browse the full feed.
10 min readEvals for non-engineers: know if your AI workflow is getting better or worse
A practical introduction to evaluating an AI workflow with representative examples, explicit rubrics, qualified labels, release rules, and a spreadsheet when that is sufficient.
Intermediate
13 min readBuilding evals that actually catch regressions
Most eval suites look impressive but miss real regressions. Building evals that catch what matters requires careful dataset construction, sensitive metrics, judge calibration, and a culture of trust. The patterns from teams that get this right.
Advanced
12 min readObservability for LLM apps: tracing, costs, latency, quality drift
Extend ordinary observability with multi-step traces, attributable cost, prompt and model versions, evaluated quality signals, privacy controls, and workload-derived alerts.
AdvancedMore in this topic
69 minutesThe Agent Landscape - Lessons Learned Putting Agents Into Production
MLOps.community. Prosus's VP of AI and an AI engineer report what actually broke when they deployed agents across the group's portfolio companies: prompt-injection pen-testing before launch, an unsafe write when a Jira agent choked on human shorthand, stale context handled by making agents surface their assumptions, and fallback design that merged or killed agents once they added cognitive load. It reads like the article's failure-mode register replayed as a live postmortem.
Advanced
10 min readProduction AI failure modes: what breaks after the demo
Build a failure-mode register for hallucination, stale context, prompt injection, unsafe tool use, schema drift, weak fallback, and observability gaps.
Advanced
11 min readChunking, reranking, and hybrid search: make RAG actually work
Most RAG implementations work poorly because they get three things wrong. A practical guide to chunking documents, reranking results, and combining keyword with semantic search — without becoming a search engineer.
Intermediate
107 minutesWhy AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Lenny's Podcast. Hamel Husain and Shreya Shankar walk through the entire eval workflow on a real property-management AI assistant — looking at traces, open and axial coding of errors, deciding when to stop, building an LLM-as-judge, and validating it against human judgment. This is the rare long-form conversation that is genuinely aimed at PMs and team leads rather than ML engineers, and it covers the same "30 minutes a week after setup" rhythm the article recommends.
Intermediate
3 minutesEvaluate prompts in the Anthropic Console
Anthropic. A three-minute Anthropic walkthrough of running a real eval inside the Workbench — auto-generating realistic test cases, grading outputs, tweaking the prompt, and re-running the same suite side-by-side. The view count sits below the usual bar, but for "how do I actually do this without writing code" this is the cleanest official demo and slots neatly under the more strategic Husain/Shankar conversation.
Intermediate
55 minutesHow to Systematically Setup LLM Evals (Metrics, Unit Tests, LLM-as-a-Judge)
Dave Ebbelaar. A working AI engineer walking through his actual eval ladder — assert-style unit tests, reference-free metrics, LLM-as-judge alignment with humans, and the analyze/measure/improve loop. The structure is the closest match on video to the article's argument that evals are a regression-catching system, not a leaderboard.
Advanced
19 minutesHow to Construct Domain Specific LLM Evaluation Systems: Hamel Husain and Emil Sedgh
AI Engineer. Hamel Husain and Rechat's CTO walk through the eval system behind a real AI product: why generic off-the-shelf evals fail, a layered setup of assertions, logged traces with human review, and LLM judges kept aligned with a domain expert, and how a working eval system unlocks data curation and fine-tuning. It is the production case study for the article's argument that evals are regression-catching machinery, not a leaderboard.
Advanced
9 minutesLangSmith in 10 Minutes
LangChain. A guided tour of an LLM trace, project, and dataset by LangChain's co-founder — token cost, latency, error rate, feedback aggregation, drilling into a single retrieval-step span. It's the closest visual analogue to what the article describes when it talks about "every call is a span" and why structured traces beat print logging.
Advanced
154 minutesInstrumenting & Evaluating LLMs
Hamel Husain. Hamel Husain, Eugene Yan, Brian Bischof, Harrison Chase, and Shreya Shankar working through tracing, log analysis, LLM-as-judge, and the workflow around looking at real production data. Sit with it the same way you would a long podcast — it is the single best deep treatment of the article's "look at your traces" thesis on YouTube.
Advanced