Companion videos

Picking the right model for the job — companion videos

The article is a decision cheat sheet: which provider, which tier, which "thinking" toggle for which task. The pick below shows someone actually moving between those models on screen so you can see why a working practitioner would reach for one over another in the moment.

Primary pick

2:11:12
How I use LLMs

Andrej Karpathy

Karpathy spends explicit chapters on "Be aware of the model you're using, pricing tiers" and "Thinking models and when to use them," then keeps switching between ChatGPT, Claude, Gemini, Grok, and Perplexity throughout the rest of the walkthrough. It is the closest thing to watching the article's cheat sheet applied live by someone with strong opinions about when each tier earns its keep.

What you should get from this: Learn to route work across fast, cheap, deep-reasoning and source-grounded tools instead of using one model for everything.

Watch or know first: Comfort using at least one mainstream chatbot and comparing outputs on the same task.

AI Expert note: The model-picker habits are valuable; exact model names, prices, context limits and rankings are not stable. Re-run the comparison on your own tasks before standardizing a team recommendation.

Open video page

Also worth watching

16:50
The New, Smartest AI: Claude 3 – Tested vs Gemini 1.5 + GPT-4

AI Explained

Older than the article (March 2024), but the methodology is what's useful: a single careful reviewer running the same hard tasks — OCR, theory of mind, instruction following, math — through three frontier models side by side and showing exactly where each one cracks. The model names are dated, the framework for comparing models is not.

What you should get from this: See a structured comparison method you can reuse when deciding which model is good enough for a task.

Watch or know first: Know that Claude 3, Gemini 1.5 and GPT-4 are no longer the frontier baseline.

AI Expert note: Treat this as a testing-method video, not a current ranking. The concrete winners are historical; the useful part is running the same examples, checking failure modes and resisting vague "best model" claims.

Open video page