32 minutesFast LLM Serving with vLLM and PagedAttention
Understand why serving engines, batching and KV-cache memory dominate self-hosted inference economics.
Prompt Engineering. Walks through Anthropic's prompt caching against Gemini's context caching with concrete latency-and-cost reductions per use case (long-document chat, few-shot, multi-turn). The breakdown of cache-write surcharge vs. cache-read discount is exactly what the article assumes when it talks about when caching pays off.
Recorded in August 2024, so the model names and per-token prices are dated. The cache-write-surcharge vs. cache-read-discount mechanic still holds; check current provider pricing pages before running the numbers for your workload.
You can estimate when prompt caching pays off by weighing cache-write surcharges against read savings for your real workloads.
Working knowledge of LLM API pricing and long-prompt workloads.
Last reviewed: May 18, 2026
Continue through the same learning path with the next curated companion videos.
32 minutesUnderstand why serving engines, batching and KV-cache memory dominate self-hosted inference economics.
34 minutesEvaluate AI product pricing and specialization around measurable outcomes rather than seat counts.
6 minutesRecognize the core architecture of a voice agent and the failure points that affect customer trust in real calls.
Hand-picked external courses that go deeper on this topic.
Emory University Goizueta Business School faculty
The deeper, university-level counterpart to our beginner HubSpot marketing pick — Emory's business-school treatment goes past 'how to prompt' into training generative models for brand-specific output, the purchase-funnel economics of AI-generated content, and a full module of genuine skepticism about when generative AI is and isn't worth using in marketing.
IBM AI Academy
The genAI-era answer to the executive-strategy question. Three short courses aimed squarely at business leaders — no technical background required — on where generative AI creates value, how to govern it responsibly, and how to turn a vague "we should use AI" into a concrete, defensible use case. Rated 4.6 across ~700 reviews. Best taken before your next AI budget decision.