107 minutesWhy AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Learn the product-builder eval loop: inspect traces, label failures, define criteria, test changes and compare against human judgment.
Tina Huang. A clean tour of the current model landscape grouped by tier — flagships, lite models, mid-tier, specialized — with concrete picks for what each tier is actually good for. This is the "know your options before you route" half of the article, and Huang frames cost-vs-capability the same way the article does without leaning on benchmark hype.
Model names, pricing and capabilities change quickly. Use this for the decision pattern, then verify current model behavior before adopting it.
Compare flagship, lite, mid-tier and specialized models so routing decisions are based on task fit, cost and latency instead of brand preference.
None beyond the article — it's the know-your-options half before any routing work.
Last reviewed: May 18, 2026
Continue through the same learning path with the next curated companion videos.
107 minutesLearn the product-builder eval loop: inspect traces, label failures, define criteria, test changes and compare against human judgment.
211 minutesUnderstand the modern LLM stack well enough to reason about tokens, training, tools and failures.
18 minutesLearn why schema-first LLM calls need typed objects, validators, retries and explicit handling for malformed or hallucinated fields.
Hand-picked external courses that go deeper on this topic.
Anthropic Academy
MCP is the protocol that's quietly replacing one-off tool integrations across the AI tooling ecosystem. Learn it from the source. By the end you'll have built and deployed your own MCP server, connected an LLM client to it, and understood why this standard is the closest thing the field has to USB-C.
João Moura (Founder, CrewAI)
Doubles as our sales and customer-support vertical pick and a genuinely practical agent-building course: you build an agentic sales pipeline (lead scoring, personalized outreach) and a customer-support data-insights pipeline as two of the five hands-on projects, taught by CrewAI's own founder. Requires basic Python, so it sits with our other builder-track courses rather than the no-code picks.
Hugging Face
The clearest open-source treatment of agentic systems available. Anchored in the three frameworks engineers actually evaluate (smolagents, LlamaIndex, LangGraph) rather than one vendor's stack. Concludes with a benchmark assignment and public leaderboard — accountability your team can verify.