Voice & Audio
Voice mode, audio generation, translation, cloning risks, and voice agents.
8 stories (3 articles · 5 videos)
Start here
A few good first pieces before you browse the full feed.
7 min readChatGPT Voice mode: talking to AI like a friend
Talking to AI feels strange for about ninety seconds, then it often becomes the easiest interface for thinking out loud. A practical guide to voice mode — what it is great at, what it is bad at, and how to actually use it.
New to AI
7 min readAI voice and audio: from cloning to podcasts to translation
An orientation map for voice cloning, narration, transcription, and translation, with consent, disclosure, privacy, and verification boundaries.
BeginnerMore in this topic
6 minutesAI Voice Agents: How They Actually Work & Why They Sound So Human
CX Foundation. Breaks voice agents into the practical pipeline: speech recognition, language model, business-system APIs, text-to-speech and interruption handling. That gives the article's rollout framework a concrete technical foundation before readers choose Twilio, Retell, Vapi, LiveKit or another platform.
Advanced
9 min readVoice agents for customer flows: where they work and where they fail
Voice agents are useful when the flow is bounded, the data is available, and the fallback is clean. A practical decision framework for Twilio/Retell-style systems, disclosure, handoff, testing, and rollout.
Advanced
11 minutesHow to Clone Your Voice with AI - Realistic AI Voice Clones (Full Tutorial)
ElevenLabs. Official walkthrough that contrasts Instant Voice Cloning (a minute of audio, results in seconds) with Professional Voice Cloning (30 minutes to several hours of audio, much higher fidelity). The recording-quality guidance — mic, room, levels, pre-processing — is the part that's hardest to find elsewhere and matters most for getting a clone you'll actually use.
Beginner
16 minutesHow to Use ElevenLabs - Best Text to Speech AI Voices (FULL GUIDE)
Alec Wilcock. Tour of the platform that anchors most of the article's examples — text-to-speech, speech-to-speech, voice design, and voice cloning, all in one screen-recorded walkthrough. Goes through the free-vs-paid limits and the controls that actually matter (stability, similarity, style) without overselling.
Beginner
26 minutesIntroducing GPT-4o
OpenAI. This is the May 2024 keynote where ChatGPT's real-time voice mode was first demoed. Mark Chen does the breathing-exercise demo, Barrett Zoph does the math-tutor demo, and then they switch to real-time Italian–English translation. Twenty-six minutes of "ah, that's what they mean by talking to AI like a friend." The model and capabilities shown have only improved since.
New to AI
3 minutesTwo GPT-4os interacting and singing
OpenAI. Two instances of voice mode talking to each other, one of which has camera access to describe the room. Three minutes long and the most efficient way to internalise what makes voice mode different from old "press the microphone, wait, listen" interfaces — interruption, tone, music, real-time vision, all in one clip.
New to AI