Writing

Grounded answers or honest refusals: a rule for retrieval over personal files

In short: when a system answers questions over someone's own files, there is one rule I would not trade for anything: every answer quotes the person's own sources verbatim, or it is an honest refusal.…

published
read time
5 min
words
999
lang
en
filed under
Engineering

In short: when a system answers questions over someone's own files, there is one rule I would not trade for anything: every answer quotes the person's own sources verbatim, or it is an honest refusal. Here is how I put together a small local system around that rule, and why.

What it is for

Most of what I know about my own life is scattered: PDFs in a downloads folder, years of photos, voice memos I never listened to again, chat exports. I wanted one place that reads all of it, links it into something navigable, and answers questions about it without sending any of it to anyone.

Two requirements shaped everything else. It has to run on a laptop, and it is not allowed to invent things. A second brain that confidently remembers a dinner you never had is worse than no second brain at all.

The architecture

Everything runs locally in containers, except the models, which run natively on the host so they can use the machine's GPU.

your files ingest local models Postgres + pgvector source of truth graph rebuildable answer quote or refuse
One database holds the truth. The graph is a view of it that can be thrown away and rebuilt.
  • Postgres with pgvector is the source of truth. Your content, its vectors and its metadata live there. If it is not in Postgres, it does not exist.
  • The graph is a view of the source store. It is what makes your files navigable as a web of linked things rather than a list. It is also rebuildable from Postgres at any time, which means it can never be the only copy of anything, and a bad graph is a rebuild rather than a data loss.
  • Models run through a local model runner on the host. Or, if you prefer, your own key for a hosted model, stored encrypted at rest. Embeddings run locally either way.
  • Every model call is traced locally, so I can see exactly what was asked and what came back, and the trace never leaves the machine.

The reason for one source of truth and one projection is backup and sanity. There is one thing to back up and restore. Everything else is derived.

Models by role, not one model

The system does not use one global model. Work is routed to a small number of model roles, and each role can be pointed at a different model.

RoleUsed forDefaultIf you have memory to spare
Quick stepsLight extraction and classification during ingestOne small multimodal modelA larger text model, or a hosted key
AnsweringThe main model that answers your questionsThe same small modelA larger text model, or a hosted key
VisionUnderstanding photos during ingestThe same small modelA small vision specialist
EmbeddingTurning text into vectors for retrievalA small embedding modelNothing. It is frozen

The default covers three roles with one model because a small general model that is also multimodal is good enough for all three. One download covers quick steps, answering and vision on a normal laptop, and the embedder is a small addition on top.

# Models run on the host, not in a container, so they get the GPU.
roles:
  quick:  small-multimodal   # one model covers three roles
  answer: small-multimodal
  vision: small-multimodal
  embed:  small-embedder     # frozen once anything is ingested

Splitting roles has one trap. In most model families the vision variant is a separate model, not a mode flag, so pointing the vision role at the plain text model of the same family does not work. And the smallest vision models in some families are heavy next to everything else in the stack.

On a weak machine, a small text-only model is a legitimate choice. Photos then degrade to a text-only understanding from the filename, the metadata and any text available, instead of failing outright. Degrading is better than failing.

CarefulDo not change the embedding model casually. Every stored vector lives in that model's space, so switching silently breaks retrieval for everything already ingested until a full re-embed runs. The system has a guard that blocks a silent embedder change for exactly this reason. It is the one role where a careless change does real damage.

Refusing to make things up

The contract with the user is short. An answer is grounded only in your own files and comes with verbatim citations, or the system says it does not know.

Verbatim matters. A paraphrased citation can drift from the source and still look cited, and you would have to open the file to notice. A quoted span is something you can check against the file directly. If the system cannot find words in your files that support an answer, it should not have an answer.

The refusal matters just as much. With a small local model, the temptation is to let it fill gaps from what it already knows about the world. For a second brain that is exactly the wrong behaviour. "I don't have anything about that in your files" is a correct answer. A plausible guess about your own life is not.

What you need to run something like this

  • A laptop with a capable GPU for the default local setup. On a CPU-only machine a single photo caption takes minutes instead of seconds, which makes local vision too slow to be practical. There, use a hosted key for vision, or accept text-only photos.
  • Memory. The containers need a comfortable amount of memory, the tracing stack adds a little more, and the local model needs its own share on top. A mid-range laptop's worth of memory is a realistic floor for running everything at once.
  • Patience on first import. Every photo pays a vision call. That is a post of its own.

If you are building anything that answers questions over personal data, write down the refusal sentence before you write the prompt. Then try to make it say something your files do not contain. If you can, fix that before you add a single new feature.

related

Keep reading