Skip to content
Android

Gemini 3.5 Transcribe Takes On OpenAI’s GPT-Live Voice AI

Voice-to-text AI is reshaping how people interact with their devices, as rapid improvements in dictation and transcription models make it possible to operate phones and computers with minimal typing. While physical smartphone keyboards have seen a modest revival, a parallel trend is pushing toward...

Gemini 3.5 Transcribe Takes On OpenAI's GPT-Live Voice AI - voice-to-text AI
Voice-to-text AI is reshaping how people interact with their devices, as rapid improvements in dictation and transcription models make it possible to operate phones and computers with minimal typing. While physical smart

Voice-to-text AI is reshaping how people interact with their devices, as rapid improvements in dictation and transcription models make it possible to operate phones and computers with minimal typing. While physical smartphone keyboards have seen a modest revival, a parallel trend is pushing toward a future where users barely touch their screens at all.

Purpose-built dictation apps have surged in popularity. Options on mobile include Wispr Flow, Yaps, and Typeless, while Aqua Voice and Monologue serve Mac users. Most of these tools share the same core function and follow a freemium pricing model, offering limited free usage alongside premium tiers that run roughly EUR 7 to EUR 10 per month.

OpenAI and Google Push Real-Time Voice Models

Last month, OpenAI introduced GPT-Live, a model built for near real-time, full-duplex communication that layers on top of a multimodal large language model capable of delegating tasks. It marked the latest step in an ongoing effort to improve the reliability of voice input across laptops, desktops, and phones with a microphone.

Google has now countered with Gemini 3.5 Transcribe, which the company says is faster and more accurate than GPT-Live. For Pixel 11 owners, the model is already built into the Rambler tool, which replicates functionality found in many standalone dictation apps.

Voice Input Meets Dynamic Action

Gemini 3.5 Transcribe matches GPT-Live’s capabilities while promising quicker performance, and Google’s reach gives it a significant advantage. The company wants users to speak directly to Google Docs and Sheets, as well as Gemini Live and Spark on Mac and mobile. Through the Gemini API, developers can also integrate voice input into their own applications.

Voice input itself is not new, but this emerging approach pairs accurate transcription with dynamic action. Models can intelligently strip out verbal tics and reformat spoken words into tables or charts. A user could dictate an email, instruct the system to insert a specific photo from a camera roll, and embed a map with directions, all within a single voice command.

The combination of speed and adaptability is appealing, though it comes with tradeoffs. Typing offers a tactile, contemplative connection to a thought that voice cannot fully replicate. Even so, a growing group of users, particularly those with fine motor challenges, stands to benefit from these advances. For Pixel 11 owners, Gemini 3.5 Transcribe is available now through the Rambler tool.

Source
Image: 9to5google.com

The US tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your US list is ready.

Shop Amazon Tech Deals Shop Amazon Tech Deals