Google has newly announced "Gemini 3.5 Transcribe," an advanced audio transcription technology. Going beyond conventional speech-to-text conversion, this technology features the ability to detect verbal stumbles, mistakes, and fillers in real-time, automatically correcting and transforming them into organized, highly readable text.
Gemini 3.5 Transcribe supports audio in 85 languages worldwide. Its standout feature is the ability to remove unnecessary noise from conversational speech while maintaining context for accurate transcription. By outputting the speaker's intended message as natural prose, it is expected to dramatically boost efficiency in meeting minutes creation and content production.
While conventional speech recognition engines focused on "transcribing heard audio verbatim," Gemini 3.5 leverages the advanced reasoning capabilities of large language models to interpret the speaker's intent and proofread the text. This compensates for the inherent imperfections of spoken language, significantly reducing the effort required for post-editing.
Google plans to progressively integrate this technology into various platforms and services. Armed with multi-language support and real-time correction capabilities, the company aims to enhance productivity in global communication.