Google has released Gemini 3.5 Transcribe, a specialized speech-to-text model with 4% word error rate for live audio and 2.6% for recorded audio. The model supports 85 languages, auto-detects language, distinguishes up to 3 speakers, and offers custom vocabulary features. Now available in Gemini app (macOS), via API, and coming to Chrome for form-filling—directly applicable for Thai SMEs automating voice documentation and multilingual customer interactions.
← Back to all articles