An interactive explainer

Who Said That?

Why does one meeting transcript say Jessica and another Speaker 2?

An accurate transcript, with notes and action items waiting the moment a call ends, has become almost mundane. Speech-to-text got good enough that most people stopped wondering what happens underneath it.

PaulCan your team deploy this by October?
JessicaOctober will be difficult. November is more realistic.

Few months back, I spent time with the team responsible for how our conversation intelligence (CI) product integrates with meeting providers like Zoom, Teams and Google Meet. I knew how audio is captured as streams and turned into text.

But what I had never thought about was how that text gets split up by speakers at all, how the audio gets split where the voice changes, and the pieces grouped by whoever they sound like.

That part has a name: diarization, derived from diary: a record of what happened, in order, with the time attached.

But diarization gets as far as this:

Speaker 1Can your team deploy this by October?
Speaker 2October will be difficult. November is more realistic.

It still does not know that Speaker 2 is Jessica. That is a whole different problem, and trying to understand it led me to a fascinating fact: quite often the answer to this problem depends less on the speech model than on where the audio came from in the first place.

Three separate systems hide inside the word transcription:

  • What was said?: Speech recognition
  • Who spoke when?: Diarization
  • Who is that person?: Identity resolution

Where did the audio come from?

To understand how a system separates and identifies speakers, we will have to go one step back and learn what audio did the CI tool receive in the first place.

During a Zoom, Meet or Teams call, each person sends their own microphone audio to the meeting platform. Paul's voice and Jessica's voice therefore begin as separate audio streams. The meeting app then sends the other participants' audio to each participant's device, where their voices are played together as one conversation.

A conversation-intelligence product has to get access to that audio. Whether it receives one mixed track or participant-level audio determines how much diarization work is needed later. If it is captured after the mix, Paul and Jessica are combined, so the system has to work out who spoke from the sound alone.

Capturing closer to the meeting platform may preserve more structure: separate participant audio, participant IDs, or metadata about who was speaking. For example, in Zoom's Real-Time Media Streams, audio can arrive already associated with a participant ID and name.

Common ways to get the audio

  1. Meeting bot. Joins the call as a participant and records the audio and meeting media available to it. What it receives depends on how the bot integrates with the meeting platform: some bots work with mixed audio, while integrations using richer media APIs may also receive speaker metadata or participant-level audio. It also has to be let in: audio from before it was admitted does not exist anywhere downstream.

    Good to know: This makes bot join time and waiting-room time useful operational metrics when evaluating a conversation-intelligence product.

    Infrastructure providers such as Recall.ai handle much of this bot lifecycle for developers: scheduling bots in advance, tracking states such as joining, waiting room and recording, and providing controls for how long bots should wait before leaving.

  2. Desktop capture. Records audio directly from a participant's computer, typically using the microphone and system audio. The trade-off is that remote participants may arrive through the computer's system audio rather than as separate participant streams. The product may therefore need additional speaker information or diarization to work out who said what.

    Good to know: Products such as Granola and Wispr Flow use this bot-free approach, so they can work across meeting platforms (Teams, Zoom, etc.) without joining the call as a participant.

  3. Native media API. Receives media directly from the meeting platform, like Zoom RTMS. Where the platform supplies participant attribution, identity becomes a lookup rather than an acoustic guess.
Figure 1 - Where did you tap into the call?
What the application receives at the selected capture point

Figure scrolls sideways →

Tap point
Capture architecture decides what the models still have to guess. The ribbon shows who actually spoke; the track beneath it shows what the recording itself can tell apart. Nothing downstream can add information the tap point discarded.

Diarization: separating speakers in mixed audio

What happens next depends on what survived during capture. If the system receives participant-specific audio and metadata, some of the speaker information may already be available.

But a very common case is that the system receives one mixed recording. Paul and Jessica are now part of the same audio track, and the system has to separate their voices from the sound itself.

This is where diarization begins.

A recording is a long list of numbers, each describing how hard the air pushed at one instant, thousands of times a second. There is no field for a person, and nothing marks where one speaker stops. A human in the room separates the voices easily, using direction and eyesight - neither of which is available in the mixed audio.

Diarization answers one question - who spoke when. It broadly runs in four steps.

  1. Find the speech. A voice-activity detector marks where speech exists and discards silence, keyboard clatter and the dog next door. It finds speech, not speakers.
  2. Cut it up. What remains is divided into short regions. The cuts do not fall neatly at speaker changes, because where the speaker changes is exactly what is not yet known.
  3. Represent each region. A speaker-embedding model turns each region into a numerical representation of how the voice sounds. Segments from the same speaker tend to produce more similar representations than segments from different speakers.
  4. Group them. The system groups the similar ones together. The number of groups is how many people it believes were talking.

The result is not Paul and Jessica. It is SPEAKER_00 and SPEAKER_01.

Figure 2 - Two voices, still no names
Diarization stepped through: raw audio, speech detection, embeddings, clustering

Figure scrolls sideways →

This is a simplified 2D view of speaker embeddings. In reality, they have many more dimensions. Segments from the same voice tend to sit closer together. The overlap sits between the two groups because it contains both speakers.

Overlap in speaking is where a single mixed track causes maximum errors. Clustering gives each region one label, so when two people talk at once it has to pick a winner, and the other person's words vanish or land under the wrong name.

Does quality of diarization still matter?

Different transcription providers make different trade-offs here. Deepgram, AssemblyAI and others each run their own diarization models, and the differences are real: how much speech they need before they can confidently separate a speaker, how well they handle short interjections, overlapping speech, noisy calls or phone-quality audio, and how quickly speaker labels stabilise in a live transcript.

So diarization is not just a checkbox in a transcription API. If speaker attribution matters to your product, the quality of this layer is something worth evaluating.

Names are not in the audio

Suppose diarization works perfectly. The system has successfully separated the conversation into Speaker 1 and Speaker 2.

There is still one problem: it does not know that Speaker 2 is Jessica. Nothing in Jessica's voice contains the string "Jessica." A name cannot be recovered from the waveform alone.

The cleanest way is by using participant metadata. If the meeting platform says user_id: 42 is Jessica, and the audio is associated with that ID, the system can simply map the audio to her name. No voice recognition is needed.

If the audio is mixed, the system may instead use active-speaker metadata from the meeting platform and align those timestamps with the diarized turns. A turn is a stretch of the conversation assigned to each anonymous speaker.

A third option is voice identification: comparing an unknown voice with a stored reference recording of a known person. This requires prior enrolment. Zoom's Smart Name Tags is one example: users can enrol a voice recording so Zoom can recognise them as speakers later. This is different from diarization, because the system is matching the voice against a known person rather than merely grouping similar speech together.

So turning Speaker 2 into Jessica can be a lookup, an alignment, or a voice match - depending on what information survived when the meeting was captured.

Figure 3 - Turn Speaker 2 into Jessica
Evidence used by the selected route to resolve Speaker 2 to Jessica

Figure scrolls sideways →

Identity source
The words never changed. The evidence available for identity did. With participant metadata, Jessica is a lookup. Without it, the system may be left with an anonymous speaker.

What speech-to-text does, and does not do

At this point, two questions are separate: who spoke when, and who that speaker actually is. Speech recognition turns audio into words and timestamps. It solves a different problem from diarization and speaker identity.

It also predates generative AI by decades. Modern systems evolved from older statistical speech-recognition methods to neural and Transformer-based models, with systems like Whisper making high-quality multilingual transcription much more accessible. Generative AI did not invent transcription. It mainly changed what products could do with the transcript afterwards.

A wrong name produces a fluent wrong answer

For many downstream CI applications, an LLM reads the transcript and generates something a user may act on: a summary, an answer, a next step, or a recommendation.

Jessica: November is more realistic.

→ Timeline risk: the customer now expects a November deployment.

If Jessica's sentence is incorrectly attributed to Paul, the output can still sound completely reasonable. The business meaning, however, may be reversed: the vendor now appears to be proposing November instead of the customer.

A missing sentence is easy to notice, but a sentence attributed to the wrong person is harder to catch.

Every downstream insight inherits the mistakes made upstream. A better LLM cannot recover audio that was never captured, and perfect speech recognition cannot fix the wrong speaker attached to a sentence.

The whole chain, from microphone to named transcript
The six stages between Jessica speaking and a speaker-attributed transcript line

Diagram scrolls sideways →

Every arrow is a place information can be lost, and nothing later in the chain can put it back. Capture is the stage that decides how much the two after it still have to work out.

When the transcript is wrong, where to look?

That brings us back to the original question: why does one transcript say Jessica while another says Speaker 2?

The speech model may not be the problem at all. The issue could be how the audio was captured, whether speaker information survived, how diarization grouped the voices, or how identity was resolved.

This matters when debugging or improving a conversation-intelligence product. Before trying to improve a model, the product team needs to understand which part of the pipeline actually failed.

Sometimes the fix is better transcription. Sometimes it is better diarization. And sometimes the real problem happened much earlier, when the audio or speaker metadata was captured.

To improve the output and accuracy of any CI product, a product manager first needs to know which layer to fix. And this explainer is a map of all those layers - a simpler place to start.