One VTU logoOne VTU
All posts
EngineeringAIAudio

An audio file cannot tell you who is speaking

Our two-speaker audio overview highlights the line being spoken as it plays. The finished recording contains nothing that says where a turn begins — so the timings had to be measured while the audio was still being made.

8 September 20264 min read

A two-speaker audio overview is the feature people ask about first. You point it at your notes, it writes a conversation about them, and two voices read it back — one asking, one explaining.

It plays like a podcast, so people expect it to behave like one: the transcript sits on screen, and the line being spoken is highlighted as you hear it.

That highlight is the part that turned out to be hard, and the reason is not obvious.

The recording does not contain the answer

The obvious way to build it is to work backwards from the audio: analyse the finished file, find where the speaker changes, and map those boundaries onto the transcript.

That works on a real podcast, because a real podcast is edited. Speakers overlap and pause and leave gaps, and those gaps are exactly what a boundary detector looks for.

Our recording has no gaps. It is assembled from one request per segment, head to tail, with nothing in between:

segment 1 PCM | segment 2 PCM | segment 3 PCM | ...

Not a millisecond of silence at any seam. So there is nothing in the file to detect. Structurally, a finished recording of two people talking is indistinguishable from one person talking continuously.

The answer exists, but only for a moment

Here is the thing that changed the design: while the audio is being assembled, we know exactly where every segment begins.

A segment is one request to a speech model, and the response comes back as raw PCM. Its byte length, divided by the format's bytes per second, is its exact duration — 24 kHz, mono, 16-bit, so this is arithmetic rather than estimation.

Those numbers are true only while the segments are still in hand. They cannot be recovered afterwards, because the only thing that survives is the joined file, and the joins are invisible. So the timings are measured during assembly and stored beside the recording.

It is a strange sentence to write in a code comment: if you are looking at a finished recording and wondering where a speaker starts, that answer was knowable a few minutes ago and is not knowable now.

The error that cannot accumulate

There is a second-order detail worth being honest about.

Segments are not one turn each. Requests are batched to keep the number of calls down — which is also why a retry re-pays for seventy seconds rather than five minutes — so a single segment usually holds several turns from both speakers. We know exactly where that segment starts. We do not know exactly where the fourth line inside it starts.

So inside a segment the position is apportioned by character count: the share of the segment's spoken text that comes before this line. That is an estimate, and it can be a second or two out.

What matters is that the error is bounded. The next segment boundary is exact again, so the estimate is slightly wrong for a line or two and then snaps back to true. It cannot drift, because it is never more than one segment away from a point we measured.

An estimate that resets every seventy seconds is a very different thing from one that accumulates. The first is invisible in use. The second is the bug where the highlight is perfect for a minute and then a line and a half behind for the rest of the recording.

What that meant for recordings made earlier

They are not highlighted.

A recording made before this existed has no timings stored, and they cannot be pulled back out of the file — which is the whole point of the section above. The same is true of a run that was interrupted at the last moment: the audio file survived, the timings did not.

Those play with the transcript visible and nothing marked. The alternative was guessing from the audio, and a highlight that is confidently wrong is worse than no highlight at all: it sends you to the wrong paragraph, and you believe it.

The rebuild that would have made it stutter

One more detail, because it decides whether a feature feels finished.

Playback position updates five times a second. Rebuilding the transcript on every one of those ticks would throw away and rebuild every line in the list five times a second — usually to change nothing at all, because the position has not crossed a line yet.

So the transcript subscribes to the position stream itself, reduces it to one event per spoken line, and rebuilds only when that line actually changes. Beside it, the transport controls subscribe to the same stream and care about something else entirely. Two subscribers, one stream, and no shared rebuild.

Why it was worth the trouble

Because a two-speaker overview is a revision tool, not a podcast. You listen to it on the way to college and then want to find the part you half-remembered. Seeing the line while you hear it is the difference between a recording you listen to once and one you can actually work from.

Keep reading

Related reading

EngineeringAIProduct

Every rule that makes a timetable a timetable is English text in a prompt

We went looking for the scheduler in our AI study timetable and there is not one. No overlap detection, no chronological sort, no feasibility check — the plan is whatever the model returns, and the rules exist only as sentences asking it to behave.

12 September 20266 min read