July 28, 2026/5 min readEmotion detection that must never slow the voiceIn short: the emotion classifier in the supervised care companion runs beside the reply, never in front of it. It starts when the user's words arrive, nothing waits for it, and when it finishes it…
July 14, 2026/5 min readBarge-in is the whole productIn short: a voice companion you cannot interrupt is a recording with extra steps. Barge-in, the user talking over the reply and the reply stopping, is what makes it feel like a conversation, and it…
June 30, 2026/5 min readWhy the service speaks English to everyoneIn short: the supervised care companion speaks English to every participant because every line it says in a hard moment has to be signed off by a clinician, and a translation is a new text that nobody…
November 18, 2025/5 min readRewriting a Next.js backend in Python without users noticingIn short: I moved an app's API out of Next.js routes and into a FastAPI service one endpoint at a time, keeping every request and response shape the same, and only pointed the frontend at the new…
November 4, 2025/5 min readStreaming speech: where the latency actually goesIn short: in a voice pipeline that waits for each stage to finish, the first sound waits for the whole answer to be written and then spoken. Streaming doesn't make any stage faster. It lets the stages…
October 21, 2025/4 min readA voice assistant in a weekend: LLM to speech over WebSocketsIn short: a working voice assistant is one WebSocket, one language model call and one text-to-speech call, in that order. I built one in a weekend with FastAPI and OpenAI's APIs. It works, it's…
September 9, 2025/5 min readSpeech to facial animation: a week with Audio2FaceIn short: NVIDIA's Audio2Face turns speech into a stream of ARKit blendshape weights, with emotion folded in. It gives a talking character believable lips with no animator, but it is one stage in a…
November 5, 2024/4 min readReal-time transcription on a laptop: what Whisper gets wrongIn short: live on a laptop, Whisper invents text in silence, repeats itself, transcribes your speakers, cuts words at chunk edges and mangles names. Most of the fixes sit around the model, not in it:…