Google is making voice transcription sound a lot more polished.
The company has introduced Gemini 3.5 Transcribe, a new speech-to-text model designed to turn natural speech into clean, readable text. Instead of simply writing down every sound a person makes, Gemini AI transcription can remove filler words such as “um” and “ah,” handle spoken corrections and automatically format the resulting text.
Google announced Gemini 3.5 Transcribe on August 26, describing it as its most precise speech-to-text model yet. The technology is designed for everything from voice notes and meetings to real-time dictation and other voice-powered applications.
Gemini AI transcription cleans up the way people actually speak
Anyone who has read a word-for-word transcript of a conversation knows that spoken language can look surprisingly messy on the page.
People pause. They repeat themselves. They change their minds halfway through sentences. And words such as “um,” “uh” and “ah” appear far more often than most speakers realize.
Gemini 3.5 Transcribe attempts to solve that problem with what Google calls Smart transcription.
When enabled, the system can remove filler words, reduce false starts and resolve corrections made while someone is speaking.
For example, someone might say they want to schedule something for Tuesday and then immediately correct themselves to Wednesday. Rather than reproducing the entire correction, Smart transcription can place the intended Wednesday date directly into the finished text.
The result is closer to edited written language than a traditional raw transcript.
Google is not forcing every transcript to be cleaned up
An important detail is that Gemini 3.5 Transcribe can also produce traditional word-for-word transcripts.
Google’s developer documentation describes two transcription modes.
Verbatim mode preserves what was actually spoken, including filler words, repetitions, pauses and false starts. Smart mode, meanwhile, is designed to improve readability by removing those speech patterns and cleaning up grammar and formatting.
That distinction matters.
A polished transcript can be useful for writing emails, notes or documents, but there are situations where retaining someone’s exact words is important. Interviews, research projects and certain professional records may require a more literal transcript.
By offering both approaches, Google is allowing developers and users to choose between accuracy to the original speech and a cleaner final document.
Gemini AI transcription can understand corrections
Removing “ums” and “ahs” is only part of what the new system does.
Gemini AI transcription is also designed to understand the intention behind spoken corrections.
People rarely dictate perfectly formed sentences. Someone might begin a thought, stop, replace a word or correct a date without restarting the entire sentence.
Gemini 3.5 Transcribe can recognize some of these changes and incorporate the corrected information into the finished text.
Smart mode can also automatically organize speech into paragraphs, numbered lists and bullet points. Dates, numbers and currencies can be formatted rather than simply reproduced exactly as they were spoken.
This could make voice typing considerably more practical for longer pieces of writing.
More than 85 languages are supported
Google says Gemini 3.5 Transcribe can automatically detect and transcribe more than 85 languages.
The model is also designed to handle different accents and dialects, an important challenge for speech-recognition technology used across multiple countries.
Users and developers can provide custom vocabulary as well. This can help the system recognize unusual names, industry terminology, company-specific phrases or specialized technical language that ordinary speech recognition might misunderstand.
That could make the technology especially useful for professionals working in areas where specialized terminology appears frequently.
Google claims major transcription accuracy improvements
Google is also emphasizing accuracy rather than simply the AI editing features.
According to the company, Gemini 3.5 Transcribe achieved an average word error rate of 4.0% for streaming transcription and 2.6% for non-streaming transcription in testing measured by Artificial Analysis.
Google says the model is designed to remain accurate in noisy real-world environments and can better recognize difficult alphanumeric information such as postal codes and order numbers.
As with any benchmark, real-world performance will vary depending on factors such as microphone quality, background noise, accent and the complexity of what is being discussed.
It can identify different speakers
Recorded conversations create another problem for transcription systems: knowing who said what.
Gemini 3.5 Transcribe includes speaker identification for prerecorded audio. Google says it can attribute speech and provide timestamps for conversations involving up to three speakers.
Support for more than three speakers is currently considered experimental.
That capability could make the model useful for transcribing interviews, meetings, customer-support calls and group discussions.
Instead of receiving a large block of text, users can potentially see which speaker made each statement.
Gemini on Mac already uses the technology
Some consumers have already started encountering Google’s smarter approach to dictation.
The Gemini app for macOS allows users to speak into almost any window by holding the Fn key. Google’s intelligent dictation system then converts the speech into polished text and inserts it at the cursor.
It can automatically remove “ums” and “ahs” while also recognizing mid-sentence corrections.
Google says Gemini 3.5 Transcribe is also being used for voice capabilities such as Rambler on Android.
The broader goal is clear: Google wants speaking to a computer to feel less like traditional voice typing and more like having an editor clean up your words while you talk.
Gemini AI transcription could change voice typing
Traditional dictation has always presented users with a trade-off.
Speaking can be considerably faster than typing, but the resulting text often requires cleanup. Fillers, repeated phrases, punctuation errors and verbal corrections can leave users spending extra time editing what the software produced.
Gemini AI transcription attempts to perform some of that editing automatically.
That could make voice input more attractive for writing emails, drafting documents, recording ideas, creating meeting notes or composing messages while working on other tasks.
It also represents a broader shift in speech recognition.
Older transcription systems primarily tried to answer one question: “What words did the person say?”
AI-powered systems increasingly try to answer another: “What was the person trying to say?”
There are still reasons to check AI-generated transcripts
The convenience comes with an obvious limitation.
Once software begins cleaning up what a person says rather than reproducing it exactly, users need to pay attention to whether the system has interpreted their intended meaning correctly.
Removing an “um” is unlikely to change much. Resolving a self-correction or restructuring an entire sentence involves more interpretation.
That makes reviewing important transcripts worthwhile, particularly when dates, figures, names or important statements are involved.
Google’s decision to keep a Verbatim mode alongside Smart transcription provides an alternative when preserving the original wording matters.
Google is turning transcription into an editing tool
Gemini 3.5 Transcribe shows how quickly voice technology is moving beyond basic speech recognition.
Google is no longer simply trying to produce a transcript of everything someone says. With Gemini AI transcription, the company wants AI to clean up speech as it happens, remove distracting filler words, understand corrections and organize spoken thoughts into text that is immediately easier to read.
For anyone who frequently dictates messages or spends time cleaning up transcripts, that could be a meaningful improvement.
And while “ums” and “ahs” are small details, removing them automatically highlights a much bigger change: AI transcription is becoming less about recording speech word for word and more about turning spoken ideas into finished writing.







