Somebody in your business is typing up a recording right now. It might be the office manager writing out what was agreed in Tuesday's meeting. It might be a salesperson reconstructing call notes from memory at 5 p.m., three calls after the call actually happened. It might be you, scrubbing back and forth through a voice memo hunting for the one number the customer said out loud and nobody wrote down.
That job just got a lot smaller. Google has released a speech-to-text model that does not simply write down every sound a person made — it cleans the transcript up on the way out.
Here is what actually changed, what the accuracy figure really means, where it still falls over, and how to try it this week without signing anything.
What Google announced
On August 26, 2026, Google introduced Gemini 3.5 Transcribe, the transcription model behind dictation in Gboard, Chrome, and the Gemini app.
The headline is not that it hears words accurately. Computers have been reasonably good at that for years. The headline is what it does with them afterward. In its own words, the model "seamlessly handles self-corrections (like 'let's meet Tuesday—no, Wednesday'), removes filler words ('ums' and 'ahs'), auto-formats your text."
Alongside that, Google says the model:
- Labels who said what. It identifies multiple speakers and stamps them with timestamps — with an important limit, covered below.
- Detects the language by itself. Google says it automatically detects and transcribes over 85 languages, so nobody has to tell it in advance which one is being spoken.
- Learns your vocabulary. It supports custom vocabulary — the part numbers, drug names, street names, and surnames that generic transcription reliably mangles.
- Works live or after the fact. There are two versions: one for real-time streaming as someone speaks, and one for audio you already recorded.
Why "clean" is the part that matters
If you have ever used automatic transcription and quietly gone back to typing notes by hand, this is why.
Old transcription was faithful in the worst way. It wrote down everything: every "um," every false start, every time somebody said Tuesday and then corrected themselves to Wednesday. The result was technically an accurate record of the noise in the room and practically unusable. Reading it took nearly as long as listening to the recording, and cleaning it up took longer than typing notes from scratch. So the transcript sat in a folder and the afternoon job stayed an afternoon job.
Stripping fillers and resolving self-corrections is a small-sounding change with an outsized effect: it turns a raw dump into something a human will actually read. Speaker labels do the same thing for meetings and calls, because "who committed to that" is usually the only question the recording is being consulted to answer.
What the accuracy number actually means
Google reports the model's word error rate — the share of words a transcript gets wrong, whether misheard, dropped, or invented. The figures in the announcement are 2.6 percent for pre-recorded audio and 4.0 percent for real-time streaming. Google credits Artificial Analysis, an outside benchmarking outfit, with the measurement.
Translate that into something you can picture. A 2.6 percent error rate means roughly 26 wrong words in every 1,000. That is a couple of sentences' worth of mistakes scattered across a long transcript — fine for getting the gist, not good enough to paste into a contract without reading it.
Two honest caveats before you plan around that number:
- It is a benchmark, not your audio. Benchmark recordings are not a speakerphone on a conference table, a cab with the window down, or a customer calling from a job site. Your real-world error rate will be worse than the lab figure, and how much worse depends entirely on your recordings.
- Live is worse than recorded. The 4.0 percent streaming figure is meaningfully higher than the 2.6 percent for recorded audio. If accuracy matters more than immediacy, record first and transcribe after.
The limit nobody will mention in the demo
Speaker labeling is the feature that makes this useful for meetings, and it is also the one with a hard ceiling. Google states that the model provides timestamps for up to three speakers, and that support for more than three is experimental.
Three is a sales call. Three is a one-on-one plus a note-taker. Three is not your Monday morning team meeting with seven people around a table, and it is not a board meeting. Beyond three voices you are relying on a feature Google itself labels experimental, which is a polite way of saying do not build a process on it yet.
That limit should shape where you point this first. Two- and three-person conversations — sales calls, support calls, client intake, site walkthroughs, your own dictated notes — are where it is ready. Large meetings are where you still check the output line by line.
Where a business would actually use it
Google names the uses it built this for: voice agents, real-time captioning, post-call analytics, voice-driven interfaces, and dictation. Translated out of vendor language, here is what that looks like in an ordinary business:
- Call notes that write themselves. Recorded sales and support calls become a readable, speaker-labeled record without anyone retyping. What got promised is searchable instead of remembered.
- Dictation instead of paperwork. A technician or nurse or inspector talks through a report in the truck on the way back, and the formatted text is waiting. Custom vocabulary is what makes this survive contact with your industry's jargon.
- Captions for anything you record. Training videos, webinars, and customer-facing clips get captions that do not read like nonsense.
- Accessibility that is not an afterthought. Real-time captioning means a deaf or hard-of-hearing employee or customer is in the conversation as it happens, not reading a cleaned-up summary afterward.
- Multilingual customers without a scramble. Automatic language detection across 85-plus languages means a call in Spanish or Vietnamese produces a transcript without anyone configuring anything first.
What it costs
This is where honesty is more useful than a number. Google's announcement does not state a price. Anyone quoting you a per-minute rate for this model is not reading it off that post.
What the announcement does tell you is how you reach it. Developers get it through two interfaces — a streaming version for live audio and a separate one for pre-recorded files — available through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. Those are developer tools. There is no "buy transcription" button for a business owner.
The free path is the one most people will actually use: the same model already powers dictation in Gboard, Chrome, and the Gemini app. If you want to feel the difference, that costs nothing.
How to try it this week
No budget, no procurement, no consultant. About twenty minutes.
- Dictate one thing you would normally type. A follow-up email, a quote, a set of notes. Use dictation in Chrome, Gboard, or the Gemini app and see how much cleanup the result actually needs.
- Run one recording you already have. Pick a real call with background noise and your industry's vocabulary in it — not a clean test recording. That is your true error rate, not the benchmark.
- Read the output against the audio, once. Count the mistakes yourself. That one exercise tells you more about whether to trust it than any vendor figure.
- Decide what still needs a human. Anything that becomes a contract, a medical note, a legal record, or a number you will invoice against gets read by a person. Everything else can probably ship as-is.
- Check your consent rules before you record anything. Recording laws vary by state and country, and a transcript is a record you now have to store, protect, and eventually delete. Better transcription makes recording more tempting; it does not make it automatically legal.
The pattern worth noticing
The interesting thing here is not that a machine got better at hearing. It is that the machine got better at judgment — deciding that "um" was not meant, that Tuesday was retracted in favor of Wednesday, that this sentence ends here. Those used to be the parts only a person could do, and they were exactly the parts that made transcription a job rather than a button.
That is how this generation of AI tends to reach a small business. Not as a product you evaluate and buy, but as a quiet upgrade inside something you already use, which turns a task that used to eat an afternoon into a task that eats a minute. The upside arrives for free. The limits arrive silently too — the three-speaker ceiling and the gap between benchmark accuracy and your noisy conference room are real, and neither one appears on screen while you are reading the transcript.
We break down developments like this one for people who run businesses, not for engineers. More plain-language explainers are in the AIWAS knowledge base, and short takes go out on @AIWASai.