Building a product alone: what I automated, what I did by hand
In short: running a company of one, I automated anything that repeated and had a clear right answer, like deploys, billing and checking the retrieval pipeline. I kept doing by hand anything where the…
- published
- read time
- 4 min
- words
- 889
- lang
- en
- filed under
- Product
In short: running a company of one, I automated anything that repeated and had a clear right answer, like deploys, billing and checking the retrieval pipeline. I kept doing by hand anything where the output was a decision: reading answers, talking to users, and choosing what not to build.
It is the last week of the year, so here is a look back. I started a company on my own: one person, every technical decision, building an AI product. I am not going to describe the product itself here. This post is about how one person keeps all of it running.
The rule I ended up with
I did not start with a rule. I started by doing everything by hand and getting tired. The rule came out of that:
- If I do it more than twice and there is a correct answer, a script does it.
- If the output is a judgment about what the product should be, I do it, every time, even when it is slow.
The trap is the middle. Plenty of tasks feel automatable because they are boring, but they are boring judgment, not boring procedure. Reading fifty answers from the model is boring. It is also the only way I know if the product works.
Where everything landed
Automated
- Builds and deploys, from a push to a running service
- Billing and subscriptions
- Sign-in
- A fixed set of questions run against the retrieval pipeline on every change
- Tracing of every model call
- Ingesting new documents into the corpus
By hand
- Reading real answers, start to finish
- Checking that citations point at the right passage
- Every conversation with a user
- Deciding which documents belong in the corpus
- Writing the prompts
- Saying no to features
What I automated, and why it was cheap
Most of the left column is not clever. It is the stuff every product needs and nobody pays you for. Payments, auth and deploys are solved problems. The common mistake in solo projects is building these yourself because it feels like progress. I bought or scripted all of it in the first weeks and then stopped thinking about it.
Two items on that list matter more than the rest.
The question set
Retrieval breaks quietly. You change the chunk size, or the prompt, or the embedding model, and nothing crashes. The answers just get a bit worse. So I keep a fixed list of questions with what a good answer should cite, and the pipeline runs them on every change. It does not tell me the product is good. It tells me when I made it worse, which is the thing I cannot notice alone at midnight.
Tracing
Every model call is logged with its inputs, the retrieved passages and the output. When a user says "it told me something wrong", I can find the exact call and see whether retrieval missed or the model ignored what it was given. Without that, every bug report is a guess.
What I kept doing by hand
The right column is the actual job. A few notes on it.
Reading answers. No metric I have tried replaces reading. A citation can point to a real document and still be the wrong paragraph. An answer can be correct and useless. I read a batch after every meaningful change.
Choosing the corpus. Ingestion is automated. Deciding what to ingest is not. A bad source in the corpus poisons answers far from where you would look for it.
Users. I answer every message myself. Partly because there is nobody else. Mostly because it is the fastest way to find out which feature people actually use, which is rarely the one I spent the most time on.
Saying no. Alone, every feature you add is a feature you maintain, forever, on your own. The cheapest code is the code I talked myself out of.
The release loop
Put together, shipping a change looks like this. The middle steps run themselves. The first and the last are mine.
- Make the changeOne change at a time, so I know what moved.
- Run the question setAutomatic. If answers got worse, stop here.
- DeployAutomatic. A push becomes a running service.
- Read the tracesBy hand. A batch of real calls, start to finish.
- DecideKeep it, roll it back, or write down why it is not worth it.
If you are about to build alone
A short list, in the order I would do it:
- Buy payments, auth and hosting in week one. Do not build them.
- Get deploys down to one command, then to zero.
- Write twenty questions with known good answers before you tune anything. Run them on every change.
- Log every model call with its inputs. You will need it the first time a user complains.
- Block an hour a week to read raw outputs. Put it in the calendar like a meeting, because it is the first thing you will skip.
Then keep a list of everything you did by hand this week. Anything that shows up three weeks in a row and has a right answer is your next script.
related