A voice assistant in a weekend: LLM to speech over WebSockets
In short: a working voice assistant is one WebSocket, one language model call and one text-to-speech call, in that order. I built one in a weekend with FastAPI and OpenAI's APIs. It works, it's…
- published
- read time
- 4 min
- words
- 848
- lang
- en
- filed under
- Engineering
In short: a working voice assistant is one WebSocket, one language model call and one text-to-speech call, in that order. I built one in a weekend with FastAPI and OpenAI's APIs. It works, it's public, and its latency is honest about every shortcut I took.
I wanted a small, clean version of something I build over and over: text goes in, a character answers, and the answer comes back as speech. No frontend, no database, no accounts. Just the pipe. The code is public as EArts on my GitHub, and this post walks through how it's put together and where the time goes when you use it.
The contract: one message in, one message out
The client opens a WebSocket and sends JSON. Only two fields are required: type, which is always "text", and content, between 1 and 10,000 characters. You can also override the model, the voice (alloy, onyx or nova) and the temperature. The defaults are gpt-4, nova and 0.7.
The server answers with one JSON message: type: "audio", the format (mp3), the audio as base64, and the text the model wrote. If something fails, it answers with type: "error" and one of three codes: INVALID_REQUEST, OPENAI_API_ERROR or INTERNAL_ERROR. Three codes is enough for a client to decide whether to fix its input, retry, or give up.
The shape
The project has four files that matter: main.py starts the FastAPI app, utils/websocket_handler.py holds the connection handler and the Pydantic models, utils/openai_service.py wraps the two API clients, and utils/config.py loads settings. Secrets live in .env and nothing else does. Models, defaults and the system prompt live in config/config.yaml, so changing the character's personality is an edit to a YAML file, not a code change.
The handler
Stripped down, the loop looks like this. Validation happens in the Pydantic model, so a bad request never reaches the paid APIs.
import base64
from typing import Literal
from fastapi import WebSocket, WebSocketDisconnect
from openai import AsyncOpenAI
from pydantic import BaseModel, Field, ValidationError
client = AsyncOpenAI()
SYSTEM_PROMPT = "You are a friendly character. Keep answers short."
class TextRequest(BaseModel):
type: Literal["text"]
content: str = Field(min_length=1, max_length=10_000)
model: str = "gpt-4"
voice: Literal["alloy", "onyx", "nova"] = "nova"
temperature: float = Field(0.7, ge=0.0, le=2.0)
async def handle(ws: WebSocket):
await ws.accept()
try:
while True:
raw = await ws.receive_json()
try:
req = TextRequest(**raw)
except ValidationError as e:
await ws.send_json({"type": "error", "error": "INVALID_REQUEST", "message": str(e)})
continue
chat = await client.chat.completions.create(
model=req.model,
temperature=req.temperature,
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": req.content},
],
)
text = chat.choices[0].message.content
speech = await client.audio.speech.create(
model="tts-1", voice=req.voice, input=text, response_format="mp3"
)
await ws.send_json({
"type": "audio",
"format": "mp3",
"data": base64.b64encode(speech.content).decode(),
"text": text,
})
except WebSocketDisconnect:
pass
The real version also catches API errors and maps them to OPENAI_API_ERROR, and anything else to INTERNAL_ERROR, so the socket stays open after a failure instead of dying with it.
Where the time goes
I didn't put a stopwatch on each stage for this post, so there are no milliseconds here. But you don't need them to see the problem. Look at what each stage waits for.
| Stage | Waits for | Grows with | Delays first sound? |
|---|---|---|---|
| Receive and validate | One small JSON message | Nothing that matters | Barely |
| LLM reply | The complete answer | Answer length | Yes |
| Text to speech | The complete answer text | Answer length | Yes |
| Encode and send | The complete mp3 | Audio length | Yes |
| Decode and play | The complete message | Audio length | Yes |
Every row except the first says "complete". The user hears nothing until the model has finished writing, the voice has finished speaking into a file, and that file has crossed the network in one piece. Base64 also makes the audio about a third bigger on the wire. A short answer feels fine. A long one feels like the assistant went to make coffee.
What I'd keep, and what I'd change
Keep:
- The strict request model. Bad input fails fast, for free, with a clear error.
- Secrets in
.env, everything else in YAML. I can share a config without leaking a key. - The test client.
test_client.py --save-audiowrites the reply tooutput.mp3, so I can listen to exactly what a user would hear.
Change:
- Stream the model's tokens and cut them at sentence boundaries.
- Send each sentence to TTS as soon as it's complete, and play the first one while the rest are still being written.
- Send audio as binary frames instead of base64 inside JSON.
- Add conversation history. The request schema has no field for it, so every message starts from zero.
Run it yourself
Clone EArts, copy .env.example to .env and add your OpenAI key, then start the server with python main.py. In a second terminal run python test_client.py --text "Tell me a joke" --save-audio. Then ask for a long story and compare how long you wait for each. That difference is the next thing to fix, and it's what my next post is about.
related