HeadlinesBriefing favicon HeadlinesBriefing.com

Gemini 3.5 Transcribe: Precision Speech-to-Text

Hacker News •
×

Google introduces Gemini 3.5 Transcribe, its most precise speech-to-text model designed for intelligent voice interactions. Unlike conventional models, it handles background noise, complex jargon, and disfluencies, converting raw audio into accurate, polished text.

Developers can integrate this model via the Gemini API, Google AI Studio, and Gemini Enterprise Agent Platform. It powers real-time streaming with sub-second latency and pre-recorded audio processing with speaker attribution and word-level timestamps.

The model achieves a Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming, a 70% improvement in time-to-final transcription over its predecessor, Chirp 3. It supports over 85 languages and custom vocabulary.

Available in the Gemini app on Android and macOS, it enhances everyday tools like Gboard and Chrome. Features include smart transcription, function calling, and multi-speaker identification, making it a versatile solution for voice-driven applications.