01 What happened

Google detailed developer features of Gemini 3.5 Transcribe in September. The model had been released the previous month; the update explains custom vocabulary, live transcription, and file processing.

02 Key details

  • Supports real-time streaming in over 85 languages with automatic language detection and switching.
  • Allows the use of custom_vocabulary to prioritize up to 1,000 domain-specific terms, jargon, or names.
  • Processes audio files up to one hour long with speaker labeling, timestamps, and smart transcript cleaning.
  • Google reports a 4.0% average word error rate for streaming and 2.6% for non-streaming in internal evaluations.

03 Why it matters

Developers can improve transcription accuracy for specific use cases by integrating custom vocabularies and automated formatting features.

04 Who it matters to

Software developers and engineers working with voice interfaces.

Original sourceGoogle