Writing

Real-time transcription on a laptop: what Whisper gets wrong

In short: live on a laptop, Whisper invents text in silence, repeats itself, transcribes your speakers, cuts words at chunk edges and mangles names. Most of the fixes sit around the model, not in it:…

published
read time
4 min
words
839
lang
en
filed under
Engineering

In short: live on a laptop, Whisper invents text in silence, repeats itself, transcribes your speakers, cuts words at chunk edges and mangles names. Most of the fixes sit around the model, not in it: gate the silence, overlap the chunks, stop feeding it its own past text, use a headset, and tell it the words to expect.

Earlier this year I put a small tool on GitHub called Transcriber. It turns an audio or video file into a text file with Whisper. On a clean recording it's very good, good enough for notes and first drafts.

Then I tried to make it live. Speak into the laptop, see the text appear a few seconds later, no cloud. That is a different problem. The model is the same. Almost everything around it changes, and the mistakes change with it.

Why live is harder than files

With a file, Whisper sees the whole recording. It works through it in windows of about thirty seconds and has plenty of context on both sides of every word.

Live, you don't have that. To keep the delay down you feed it short chunks of a few seconds as they arrive. Each chunk has little context. Some chunks are pure silence. Some start or end in the middle of a word. On a laptop CPU, every pass also has to finish before the next chunk is ready, or you fall behind.

time chunk 1 chunk 2 chunk 3 a word lands here overlap
Each chunk carries the last second of the one before, so a word cut at a boundary appears whole in the next pass.

What it gets wrong

These are the failure patterns I ran into, and the usual ones anyone building live transcription on Whisper will meet. The fixes are the ones that helped.

ConditionWhat goes wrongWhat helped
Silence or quiet roomFluent text that nobody said, often a stock phraseDrop quiet chunks before the model
Steady noise: fan, caféDropped words and low-confidence guessesMic closer to the mouth, cut quiet chunks
Echo from laptop speakersThe other side of a call transcribed as youA headset, or capture the mic only
AccentsNames and technical terms replaced by common wordsA short glossary in the initial prompt
Short chunksWrong language picked for a chunkSet the language, don't detect it
Chunk boundariesWords cut in half or droppedOverlap chunks by about a second
Feeding back past textThe same line repeated again and againTurn off conditioning on previous text
CarefulWhisper will write confident, grammatical sentences for audio with no speech in it. If silence reaches the model, invented text reaches your transcript, and nothing in the output marks it as invented.

The accent row matters to me personally. My first languages are Azerbaijani and Persian, and my English carries that. Whisper handles accented speech far better than older systems I've used. Where it still slips is on words it hasn't seen much: names, project terms, library names. A few of those in the initial prompt goes a long way.

The loop that runs on a laptop

Here is the shape of a live loop with the fixes in place. It uses sounddevice to read the microphone in a callback, so no audio is lost while the model is busy, and the open source whisper package for the model.

import queue
import numpy as np
import sounddevice as sd
import whisper

SR = 16_000
CHUNK = 5 * SR      # seconds of new audio per pass
OVERLAP = 1 * SR    # carried over so edge words aren't cut

model = whisper.load_model("base.en")
q = queue.Queue()

def on_audio(indata, frames, time, status):
    q.put(indata[:, 0].copy())

def loud_enough(x, rms=0.01):
    return np.sqrt(np.mean(x ** 2)) > rms

buf = np.zeros(0, dtype=np.float32)
tail = np.zeros(0, dtype=np.float32)

with sd.InputStream(samplerate=SR, channels=1, dtype="float32", callback=on_audio):
    while True:
        buf = np.concatenate([buf, q.get()])
        if len(buf) < CHUNK:
            continue
        chunk, buf = buf[:CHUNK], buf[CHUNK:]
        audio, tail = np.concatenate([tail, chunk]), chunk[-OVERLAP:]
        if not loud_enough(chunk):
            continue                         # silence never reaches the model
        out = model.transcribe(
            audio,
            language="en",                     # don't guess per chunk
            fp16=False,                        # running on CPU
            condition_on_previous_text=False,  # stops repeat loops
            initial_prompt="Glossary: PyTorch, Whisper, Montreal.",
        )
        print(out["text"].strip())

A few notes on it. The loudness gate is crude. A proper voice activity detector is better at telling speech from a loud fan, and it's the first thing I'd swap in. The overlap means the seam between chunks can produce the same word twice, so compare the start of each new line with the end of the last one and drop the repeat. And the model size is a trade. The small English-only models keep up on a laptop CPU. The larger ones are more accurate and fall behind.

Before you build your own

Record five minutes of your real setup: your room, your microphone, your voice, a stretch of silence and a stretch with someone talking on speaker. Run it through Whisper as a file first. Every mistake you see there will be worse live, and now you know which rows of the table above you need to fix first.

related

Keep reading