9 minutesLangSmith in 10 Minutes
You can navigate traces, projects and datasets in LangSmith and read off token cost, latency, error rate and per-span detail.
Dave Ebbelaar. A working AI engineer walking through his actual eval ladder — assert-style unit tests, reference-free metrics, LLM-as-judge alignment with humans, and the analyze/measure/improve loop. The structure is the closest match on video to the article's argument that evals are a regression-catching system, not a leaderboard.
Some tool choices will age, but the ladder is sound: deterministic checks first, then model-graded checks validated against humans. Do not skip calibration just because an LLM judge is easy to add.
Design an eval ladder that catches regressions before prompt or model changes reach users.
Experience shipping or maintaining an AI workflow with known failure examples.
Last reviewed: May 18, 2026
Continue through the same learning path with the next curated companion videos.
9 minutesYou can navigate traces, projects and datasets in LangSmith and read off token cost, latency, error rate and per-span detail.
75 minutesBuild a first MCP server and understand how tools, schemas, prompts, resources and transports fit together.
19 minutesImprove tool design so agents select the right action with the right parameters.
Hand-picked external courses that go deeper on this topic.
Emory University Goizueta Business School faculty
The deeper, university-level counterpart to our beginner HubSpot marketing pick — Emory's business-school treatment goes past 'how to prompt' into training generative models for brand-specific output, the purchase-funnel economics of AI-generated content, and a full module of genuine skepticism about when generative AI is and isn't worth using in marketing.
IBM AI Academy
The genAI-era answer to the executive-strategy question. Three short courses aimed squarely at business leaders — no technical background required — on where generative AI creates value, how to govern it responsibly, and how to turn a vague "we should use AI" into a concrete, defensible use case. Rated 4.6 across ~700 reviews. Best taken before your next AI budget decision.