← Back to Geekish
Developer Tools • September 17, 2026

Gemini's new audio models are aimed at real-time voice apps

Google is pushing Gemini further into voice-native software with models built for live dialogue and transcription workflows.

Illustration of audio waveforms flowing through developer windows and AI model blocks
Image: Geekish-generated illustration based on Google's Gemini audio developer update.
VOICE AIGeekish sourced quick read

Original Geekish context based on the source linked below.

The short version

Google says developers can now build with Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, and Gemini 3.5 Transcribe. The stated focus is real-time voice applications, live dialogue, and audio transcription.

Why this matters

Voice AI is moving from demo mode into product plumbing. Real-time models are useful for assistants, call tools, tutoring apps, accessibility features, and any interface where typing is the slow part.

The developer signal

The names tell the story: Live for conversation, Extended Thinking for harder dialogue tasks, and Transcribe for turning spoken audio into text. Google is packaging audio as a developer surface, not just a consumer assistant trick.

Geekish take

The next wave of AI apps may feel less like chat boxes and more like tiny operators listening, responding, summarizing, and taking notes in real time. The hard part will be latency, reliability, privacy, and making voice interfaces feel useful instead of pushy.

Want more tech without boring tech-site energy?

Follow Geekish for sourced quick reads, AI, gadgets, apps, creator tools, and internet culture.

Get the tech drop