Writing

Streaming speech: where the latency actually goes

In short: in a voice pipeline that waits for each stage to finish, the first sound waits for the whole answer to be written and then spoken. Streaming doesn't make any stage faster. It lets the stages…

published
read time
5 min
words
900
lang
en
filed under
Engineering

In short: in a voice pipeline that waits for each stage to finish, the first sound waits for the whole answer to be written and then spoken. Streaming doesn't make any stage faster. It lets the stages overlap, so the first sound waits only for the first sentence.

A few weeks ago I wrote about a voice assistant I built in a weekend: text in, a language model, text-to-speech, audio out over a WebSocket. It works, and for long answers it feels slow. This post is about why, and about how to see where the time goes in your own pipeline instead of guessing.

Two clocks, and only one of them matters

Every voice reply has two times:

  • Time to first sound: from the moment the user stops to the moment they hear something.
  • Total time: until the last word has played.

People forgive a long answer. They don't forgive silence. If the assistant starts talking quickly, a reply that takes a while to finish feels like conversation. If it is silent for the same total time and then plays everything at once, it feels broken. So the number to chase is the first one, and the first one is all about what each stage waits for.

The waterfall

buffered LLM TTS send play streamed by sentence LLM TTS play first sound first sound time
The bars are not to scale and not measured. The point is where the dots land: same stages, very different silence.

In the buffered version, text-to-speech can't start until the model has finished the whole answer, the audio can't be sent until the whole file exists, and playback can't start until the whole message has arrived. Every stage adds its full length to the silence.

In the streamed version, the model's tokens are cut into sentences as they arrive. The first sentence goes to text-to-speech while the model is still writing the second. The first chunk of audio starts playing while the rest is still being made. The total time barely changes. The silence shrinks to roughly one sentence of writing plus one sentence of speech.

What each stage waits for

StageBuffered: waits forStreamed: waits forWatch out for
Language modelThe full answerThe first sentenceA long first sentence
Text to speechThe full textOne sentenceOdd pauses between sentences
TransportThe full fileOne audio chunkBase64 inside JSON
PlaybackThe full messageThe first chunkFormats that can't play in pieces

One trap is worth calling out. In one project, the WebSocket handler for speech looked like streaming. It generated the speech, then sent it in small chunks, each with an index and an isLast flag. The client received many messages, which felt like streaming. But the file was finished before the first chunk left the server. Chunking a finished file is not streaming. The first sound still waits for the whole thing.

Measure it, don't guess

The fix starts with seeing it. I stamp each request with a few marks from the moment it arrives, and log them together. The sentence splitter sits between the model's token stream and the speech calls.

import re
import time

SENTENCE_END = re.compile(r"[.!?]\s")

class Marks:
    def __init__(self):
        self.t0 = time.perf_counter()
        self.ms = {}

    def mark(self, name):
        # keep only the first time each mark is hit
        self.ms.setdefault(name, round((time.perf_counter() - self.t0) * 1000))

async def sentences(stream, marks, min_chars=20):
    buf = ""
    async for chunk in stream:
        if not chunk.choices:
            continue
        delta = chunk.choices[0].delta.content or ""
        if delta:
            marks.mark("first_token")
        buf += delta
        while (m := SENTENCE_END.search(buf, min_chars)):
            sentence, buf = buf[: m.end()].strip(), buf[m.end():]
            marks.mark("first_sentence")
            yield sentence
    if buf.strip():
        yield buf.strip()

async def reply(ws, client, messages, tts):
    marks = Marks()
    stream = await client.chat.completions.create(
        model="gpt-4o-mini", messages=messages, stream=True
    )
    async for sentence in sentences(stream, marks):
        audio = await tts(sentence)
        marks.mark("first_audio_ready")
        await ws.send_bytes(audio)
        marks.mark("first_audio_sent")
    marks.mark("done")
    return marks.ms

A few notes on this:

  • min_chars stops the splitter from cutting at "Dr." or "e.g." at the very start of a reply, and from sending a one-word sentence that sounds clipped.
  • This version still runs speech for sentence two only after sentence one is sent. A small queue, with one task writing sentences and another speaking them, overlaps those too.
  • The server can't hear the speaker. Add one more mark on the client, when playback actually starts, and send it back. That is the real first sound.
TipLog the marks for a short answer and a long one side by side. In a buffered pipeline, first sound grows with answer length. In a streamed one it should stay roughly flat. If it doesn't, something is still waiting for "complete".

What streaming costs you

It isn't free. Speaking sentence by sentence can lose the flow between sentences, so the voice sometimes resets its tone at each full stop. Errors get harder: if speech fails on sentence three, the user has already heard one and two. And the client has to handle audio arriving in pieces, in order, and know when the reply is over. All of that is worth it for a voice product. For a tool that just saves an mp3 to disk, it isn't.

Do this next

Add five marks to your pipeline: request received, first token, first sentence, first audio sent, and first sound on the client. Log them for every request for a day. Then find the first stage whose mark grows with the length of the answer. That stage is waiting for "complete", and it's the one to change first.

related

Keep reading