October 6, 2026/4 min readTelling users the truth about ingest timeIn short: a tool that runs local models on a laptop is slow at bulk work, and a large photo import can run for hours, often overnight. I put that sentence near the top of the setup guide, before the…
August 11, 2026/5 min readWhat happens when the AI goes down mid-sessionIn short: when the model provider fails mid-sentence, the user sees whatever the companion had already said, followed by a calm, pre-written pause line. Their unanswered question is held, and the next…
July 14, 2026/5 min readBarge-in is the whole productIn short: a voice companion you cannot interrupt is a recording with extra steps. Barge-in, the user talking over the reply and the reply stopping, is what makes it feel like a conversation, and it…
December 30, 2025/4 min readBuilding a product alone: what I automated, what I did by handIn short: running a company of one, I automated anything that repeated and had a clear right answer, like deploys, billing and checking the retrieval pipeline. I kept doing by hand anything where the…
November 18, 2025/5 min readRewriting a Next.js backend in Python without users noticingIn short: I moved an app's API out of Next.js routes and into a FastAPI service one endpoint at a time, keeping every request and response shape the same, and only pointed the frontend at the new…
October 7, 2025/5 min readLeaving a job to build: the first two weeksIn short: I've left my job at a medical AI company to build my own product, alone. I gave myself two weeks of setup before any product code: a one-page brief, accounts with budget alarms, a deploy…
May 20, 2025/5 min readMost AI never leaves the notebookIn short: most models die between the demo and the first real user, and rarely because the model was bad. They die because nobody owns them, the real data looks different, the metric doesn't match a…
March 25, 2025/5 min readCloud Run or EC2: how I chose for a model that listensIn short: if the service takes a finished recording, scores it and answers, Cloud Run is usually the simpler and cheaper home. If it holds a live audio stream open, needs a GPU, or has to stay warm…
February 11, 2025/4 min readWhat "in production" means in a hospitalIn short: in a hospital, a model is in production when a clinician can see its output during real care, a named group has approved that, a named person owns it when it breaks, and the ward runs fine…
December 3, 2024/4 min readRunning AI on your own machinesIn short: when data is not allowed to leave the building, the smallest useful setup is two machines on the same network: one that asks, one with the GPU that answers. It takes about thirty lines of…