Crypt0's NewsCrypt0's News

AI

Gemini 3.8 Live Speaks 97 Languages and Thinks Out Loud While It Works

Google DeepMind announced two new voice AI models on September 15, and the pitch is impressive, with real time visual processing, 97 languages with seamless mid conversation switching, and a thinking mode that reasons out loud while it works. The models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, are rolling out through the Gemini API, Google AI Studio, Search Live for consumers, and enterprise previews in Workspace and Gemini Enterprise. The launch follows Google's rapid release cadence this year, with Gemini 3.8 Flash arriving just two weeks earlier.

The two models are clearly differentiated by use case. The standard 3.8 Live model is built for scale and cost efficiency, aimed at fluid dialogue and visual grounding. The Extended Thinking variant targets complex, layered tasks that need deeper reasoning. Both detect and switch among 97 languages automatically, even mid sentence. Both process visual input in near real time, so you can point a camera at something and discuss it while talking. Both execute tool calls and API requests in the background while the conversation keeps flowing, which means you keep talking while the model keeps working.

The benchmarks Google chose to feature are genuinely strong. Extended Thinking takes the top spot on Artificial Analysis' Speech to Speech Quality Index at 82.6, ahead of GPT Live 1 Astra (Medium) at 81.5 and Grok Voice Think Fast 2.0 (High) at 81.3. It leads on the tau Voice agentic benchmark at 68.6 percent, ahead of GPT Live 1 Astra at 67.9 percent and Grok Voice at 56.5 percent. It scores 97.7 percent on Big Bench Audio reasoning and leads Sierra's tau banking leaderboard, which tests how well a voice agent resolves customer service style banking tasks, at 35.1 percent. Google's demo videos show the model narrating a chess game, converting a whiteboard sketch into working React components, and guiding live troubleshooting.

The pricing is aggressive too. Audio input runs about 3 dollars per million tokens, roughly half a cent per minute, and audio output about 12 dollars per million tokens, under two cents per minute. Yet the feature that deserves the most attention costs Google nothing in benchmarks and everything in trust. Extended Thinking narrates its own reasoning while it works, filling the silence with phrases like let me check that while background tasks keep running. That verbalized reasoning solves the oldest awkwardness of voice assistants, the dead air that makes people wonder whether the call dropped. And every audio output carries SynthID watermarking, so generated speech can be identified as machine made. Google is framing transparency as a product feature, and its rivals are still catching up to that move. The next time a voice assistant fills a pause with let me check that, it may be thinking in your language, at your speed, and telling you about it as it goes.

Quick answers

What is this story about?

Google DeepMind announced two new voice AI models on September 15, and the pitch is impressive, with real time visual processing, 97 languages with seamless mid conversation switching, and a thinking mode that reasons out loud while it works. The models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, are rolling out through the Gemini API, Google AI Studio, Search Live for consumers, and enterprise previews in Workspace and Gemini Enterprise. The launch follows Google's rapid release cadence this year, with Gemini 3.8 Flash arriving just two weeks earlier.

Why does this story matter?

The pricing is aggressive too. Audio input runs about 3 dollars per million tokens, roughly half a cent per minute, and audio output about 12 dollars per million tokens, under two cents per minute. Yet the feature that deserves the most attention costs Google nothing in benchmarks and everything in trust. Extended Thinking narrates its own reasoning while it works, filling the silence with phrases like let me check that while background tasks keep running. That verbalized reasoning solves the oldest awkwardness of voice assistants, the dead air that makes people wonder whether the call dropped. And every audio output carries SynthID watermarking, so generated speech can be identified as machine made. Google is framing transparency as a product feature, and its rivals are still catching up to that move. The next time a voice assistant fills a pause with let me check that, it may be thinking in your language, at your speed, and telling you about it as it goes.

Sources

New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.

← Back to Crypt0's News