9 minutesLangSmith in 10 Minutes
You can navigate traces, projects and datasets in LangSmith and read off token cost, latency, error rate and per-span detail.
AI Engineer. Hamel Husain and Rechat's CTO walk through the eval system behind a real AI product: why generic off-the-shelf evals fail, a layered setup of assertions, logged traces with human review, and LLM judges kept aligned with a domain expert, and how a working eval system unlocks data curation and fine-tuning. It is the production case study for the article's argument that evals are regression-catching machinery, not a leaderboard.
Recorded mid-2024; specific tools have moved on but the system design has not. Treat it as an architecture reference, not a vendor guide, and keep the domain expert in the loop as the talk insists.
See how a production team layers assertions, human review and LLM judges so regressions surface before release.
Comfortable with basic LLM evaluation terminology and the article's eval-ladder framing.
Last reviewed: Jul 15, 2026
Continue through the same learning path with the next curated companion videos.
9 minutesYou can navigate traces, projects and datasets in LangSmith and read off token cost, latency, error rate and per-span detail.
75 minutesBuild a first MCP server and understand how tools, schemas, prompts, resources and transports fit together.
19 minutesImprove tool design so agents select the right action with the right parameters.
Hand-picked external courses that go deeper on this topic.
Emory University Goizueta Business School faculty
The deeper, university-level counterpart to our beginner HubSpot marketing pick — Emory's business-school treatment goes past 'how to prompt' into training generative models for brand-specific output, the purchase-funnel economics of AI-generated content, and a full module of genuine skepticism about when generative AI is and isn't worth using in marketing.
IBM AI Academy
The genAI-era answer to the executive-strategy question. Three short courses aimed squarely at business leaders — no technical background required — on where generative AI creates value, how to govern it responsibly, and how to turn a vague "we should use AI" into a concrete, defensible use case. Rated 4.6 across ~700 reviews. Best taken before your next AI budget decision.