34 minutesHow AI is Reinventing Software Business Models ft. Bret Taylor of Sierra
Evaluate AI product pricing and specialization around measurable outcomes rather than seat counts.
Anyscale. Walks through why naive LLM serving wastes 60–80% of GPU memory, how PagedAttention borrows OS-style paging to fix that, and why continuous batching produces the 24× throughput numbers the article uses in its math. After this, the article's "you'll be lucky to hit 50% utilisation" line stops feeling abstract.
The PagedAttention mental model remains important, but throughput numbers and serving-engine defaults age quickly. Use this for fundamentals, then benchmark your own model, hardware, batch shape and uptime requirements before making a hosting decision.
Understand why serving engines, batching and KV-cache memory dominate self-hosted inference economics.
Basic knowledge of transformer inference, GPU memory limits and API-vs-self-hosting tradeoffs.
Last reviewed: May 18, 2026
Continue through the same learning path with the next curated companion videos.
34 minutesEvaluate AI product pricing and specialization around measurable outcomes rather than seat counts.
6 minutesRecognize the core architecture of a voice agent and the failure points that affect customer trust in real calls.
37 minutesEvaluate private AI as an infrastructure and governance decision instead of defaulting to either SaaS or self-hosting by instinct.