transcription

The AI put your words in someone else's mouth

By Jack Jay · 24 August 2026 · 9 min read

The recap email lands eleven minutes after the call ends. You skim it on your phone in the hallway. Under Action items, second bullet:

Jack to have the migration plan ready by Friday.

You never said that. You said Friday looked tight. Someone else — you think it was Priya, though the call had nine people on it — said she'd try for Friday anyway, and the conversation moved on. Now there's a line with your name on it in eleven inboxes, and by the time you're back at a laptop the VP has replied "great, thanks Jack."

There is a worse version, and if you use these tools long enough you'll get it. Not a commitment you didn't make, but an opinion you don't hold. Your colleague pushes back hard on a vendor. The notes file that objection under your name. Now you're writing the email explaining you did not say the thing everyone read you saying, a genre nobody comes out of looking good.

A wrong word is noise. A wrong name is a record.

Transcription gets words wrong constantly and it mostly doesn't matter. Your product ships as "Cadence" and the transcript says "Cadance." Someone says "SOC 2" and you get "sock two." A reader repairs those without noticing, because the error announces itself. Broken text looks broken.

Attribution errors don't behave that way. "Sarah agreed to Friday" parses perfectly. It's grammatical, it's specific, it has a name in it — and names are the part of a summary people quote and act on. The failure is invisible at the exact moment it matters most, when someone forwards it to a person who wasn't in the room.

That forward is the whole problem. A meeting summary carries the authority of a record even though nobody signed it and the thing that produced it was guessing. A commenter on Hacker News put the failure mode plainly: "The worst is when someone uses these, blindly trusts the output, and then sends out an email after the meeting with a list of hallucinations that everyone supposedly 'agreed to'." Someone else in the same thread was blunter: "We have multiple people across our projects using these and they all fucking suck. They mix things up all the time. Wrong todos, wrong people assigned to todos."

The tool is a guest, and guests only hear the room

Here's the part most comparison posts skip, and it decides whether any of this ever gets fixed.

Almost every AI notetaker joins your call the way a person would: as a participant. It dials in, subscribes to the audio, and what reaches it is the meeting already mixed down — every microphone summed into one stream, the same one you hear. Pulling that apart into "who said what" is diarization, and diarization is inference. The model listens to voice characteristics — pitch, resonance, cadence, the shape of someone's vowels — clusters the segments that sound alike, then labels each cluster a speaker.

Under good conditions it works well. Two people, two decent headsets, clean turns. That's the demo, and the demo is honest about the demo.

Now switch to the conditions your actual meetings run under. In testing, clean one-on-one accuracy ran around 88 to 90 percent for both Otter and Fireflies and fell to roughly 75 percent in group calls with five or more people, with Granola's attribution getting rough past four participants. Overlapping voices collapse into each other — with several people on a call, Otter "can struggle to distinguish between voices. It might mislabel speakers or treat one person as two." Phone codecs strip the detail the model needs, which is why speaker identification "defaults to 'Speaker 1' whenever our freelance writers join from their phones." A Fireflies user gives the plain version: "It doesn't record the right person...it will say Aaron and then actually be Mike."

A two-star App Store review maps the boundary almost exactly: not the app for you "if your meetings are heated and people are talking over each other or if anyone mumbles or speaks too fast, or has an accent. In the 56 meetings I recorded and used this app to transcribe, all but one had to be edited, line by line."

This isn't one vendor being sloppy. When Hyprnote launched on Hacker News, two users hit the same wall the same day — one tried it "in a 30-minute meeting with 20 people and it put the entire conversation under a single speaker, in a single paragraph," and another called the inability to tell who said what "a show stopper". The founder's reply is the most useful sentence in the thread: "finally, speaker identification is a big challenge for us too."

The conference room isn't a hard case. It's an impossible one.

Six people around a table sharing one omnidirectional mic. That case doesn't strain the architecture. It breaks it.

To separate those six, the model has to recover six sources from one channel. Two colleagues sitting a similar distance from the same microphone, with similar voices, produce a nearly identical signal — and no metadata rides along to break the tie, because there's only one connection. The information needed to split that room was never captured. It isn't missing because the model is small. A bigger model gives a better guess. It stays a guess.

The same thread has this case too: when "some people are remote and others are in the office with a shared mic...it's a total dealbreaker for us." You don't have to take a user's word for it. Otter's own documentation says that "if multiple people speak into the same mic, Otter can struggle to separate speakers when voice quality is poor in a large room," and that when people talk at once, it "becomes less clear to whom Otter should correctly attribute a speaker." The vendor is telling you where the wall is.

What 2026 actually changed

Something real shipped this year, so it's worth saying.

Microsoft moved on enrollment. Express voice enrollment hit targeted release in mid-June 2026 and general availability between the end of June and the end of August 2026, on by default for enterprise tenants, letting people register a voice profile by speaking during a meeting instead of sitting through setup. Otter has an equivalent it calls Speaker Learning — stored voice data used "to help automatically identify and label speakers across conversations," with an admin switch to turn it off.

Enrollment helps, genuinely. It converts an open-ended clustering problem into a matching problem against known voiceprints, which is far easier. It also inherits every limit of the thing underneath it. Teams voice and face enrollment is restricted by app language and unavailable in GCCH or DoD environments. And it does not repair the shared-mic case, which an MSP admin found out the expensive way: "We have Teams Rooms set up yet I cannot for the life of me work out how to stop it transcribing participants as 'Speaker 1', 'Speaker 2' etc. I have a Copilot license, I have my face and voice enrolled, yet it makes no difference."

Microsoft owns the client, the room hardware, the transcription stack, your identity directory. If the guess could be engineered away, it would have been by now.

The second failure: notes that invent things

Now put a summarizer on top of a transcript that has the labels wrong.

Summarization is compression, and compression is where invention happens. The model has read an enormous number of meeting summaries, so it knows the shape one is supposed to have: a decision, an owner, a date. Hand it forty minutes of hedging and circling and it produces that shape anyway. A salesperson describes the mechanism exactly: those summaries "confidently invent action items nobody agreed to. The model smooths a messy call into tidy 'next steps' and half of them were never said."

The tidiness is the tell. Real meetings are not tidy. They reverse themselves, and a summarizer that resolves the reversal in the wrong direction produces something confidently backwards — one user reports having "the note taker say essentially the opposite of what we decided on because we flip-flopped multiple times during the meeting." A consultant names the deeper cost: the tool "forces the notes to look like every other meeting it knows, which is not helpful in understanding what actually happened."

Stack the two failures and you get the bullet from the top of this post. A wrong speaker label, then a summarizer that turns a hedge into a commitment and staples a name to it.

What you can do on Monday

None of this fixes the architecture. All of it lowers your exposure.

SituationWhat helpsWhat it costs
Recurring calls, same peopleVoice enrollment where it exists — Teams, Otter Speaker LearningVoice data held on a vendor's servers
Five or more peopleSay names out loud before speaking and on handoffFeels stilted for four minutes
Room with one shared micHave in-room people join on their own laptops, headsets onSetup friction, echo if anyone forgets
Long recordingsManual speaker taggingOne user reported an hour of tagging for a ten-hour meeting
Product names, jargonCustom vocabulary where supported — Google Meet users keep asking without success, including for misspelled namesOnly as good as the list you keep

Then the rule that does more work than all of it combined: never forward an AI summary you haven't read against your own memory of the meeting. Not skimmed. Read, holding the question "did anyone actually say this" against every line with a name and a date in it. Forward it unread and you have converted a machine's guess into your assertion. Everyone downstream will treat it as yours.

A different way to build it

The guess exists because of where transcription happens. Every error above traces back to one decision: transcribe the mixed room, then reconstruct who was in it.

Team Meet inverts that. Each participant's microphone is transcribed on their own machine, so the words arrive already attached to the connection they came from. The speaker isn't a prediction from timbre. It's the channel. Nine people, a heavy accent, someone talking over someone else, a phone codec chewing the high frequencies — none of it touches attribution, because attribution was never an acoustic question here. Cross-talk stops being a labeling catastrophe and becomes what it is: two people talking at once, on two separate streams, both legible.

The second rule matters as much, and it addresses the invented action item. Every claim in a Team Meet summary carries a verbatim quote from the transcript, checked as a literal substring — not a paraphrase, not a semantic match, an exact string present in the record. If the quote isn't found, the item doesn't render as fact. It renders as unverified, visibly, so a reader can tell what was said from what the model thought fit the shape of the meeting. A note that invents a constraint is worse than one that's merely wrong, because a group will plan around it. Wrong notes get corrected. Invented constraints get built on.

Two honest limits. Two people sharing one laptop are still one connection, and we label them as one rather than inventing a split we can't support — the same physics that defeats everyone else defeats us, and the difference is that we say so. And this applies to meetings held on Team Meet. A bot joining someone else's call gets the mixed stream like every other bot, and inherits every failure in this post.

The rest of the market is working on a better guess. That work is real and the models keep improving. But a meeting record is one of the few documents where being confidently wrong costs more than being incomplete, and "we know who spoke because we know which mic it came from" is a different kind of answer than "the model is fairly sure."

teammeet.io