A text interface lets both parties pause without explanation. Speech is less forgiving. Every delay sounds like hesitation or absence. Talking over someone feels different from producing a long paragraph. A response that cannot be interrupted becomes a small captivity. When an agent speaks, timing becomes visible behavior.
Modern systems expose capabilities needed for a more conversational loop. OpenAI describes its Realtime API as direct speech-to-speech processing designed for low-latency interaction. Google’s Live API documentation exposes voice activity detection, interruption handling, turn-coverage settings, and non-blocking tool behavior. These are platform capabilities, not proof that a particular experience feels natural. They reveal the design surface.
Our thesis is that voice-first AI should be judged as a participant in turn-taking, not as a text model with audio output. The voice itself matters, but the social mechanics around it matter more.