What happens when the AI goes down mid-session
In short: when the model provider fails mid-sentence, the user sees whatever the companion had already said, followed by a calm, pre-written pause line. Their unanswered question is held, and the next…
- published
- read time
- 5 min
- words
- 1,083
- lang
- en
- filed under
- Engineering
In short: when the model provider fails mid-sentence, the user sees whatever the companion had already said, followed by a calm, pre-written pause line. Their unanswered question is held, and the next turn answers it. No stack trace, no silence, no error screen.
Three things a site will meet anyway
The supervised care companion is built for adults in a supported care programme, and it often runs on shared machines at programme sites. A site does not care whether a failure is my fault or a provider's. It cares what a vulnerable person sees. So I wrote down the three things a site will meet whether or not the code is ready for them, and built for each one:
- A model outage in the middle of a conversation.
- A group session where several people start within a minute.
- A database backup that nobody has ever restored.
The design goal for all three is the same: the participant should never be the one who discovers the failure.
An outage mid-sentence
The provider will fail while a user is reading a reply. Before this work, what happened next lived in three places. Now it lives in one module, and a failing turn goes through the same steps every time:
- The call failsA timeout, a dropped connection, a 429 or a 5xx is worth another attempt. A wrong key or a malformed request is not, because another attempt would fail the same way.
- Wait, with jitterRetries back off exponentially with full jitter, so a room full of devices does not come back in lockstep.
- Stop once words have landedThe moment any part of the reply has reached the user, retrying is over. See the warning below.
- Pause, not failWhen the attempts are spent, the user keeps what was already said, and a drafted pause line is shown after it.
- Record the whole storyThe turn is stored and marked as paused. That keeps the half answer out of every read of the conversation, and leaves an operator the full picture.
- Come back to the questionThe user's unanswered message is held for a short window. The next turn opens with a drafted "I am back" line and answers what they asked before the gap, not instead of it.
That last step matters more than it looks. A device can lose its own copy of the conversation when it reconnects. Because the held message lives on the server, the user still resumes where they were.
A group session starting together
Picture a group session where several people start within a minute, and a few of them send their first message in the same second. Without a ceiling, that is a burst of simultaneous calls to a provider with rate limits, and a few unlucky participants waiting on connections nobody will answer.
So turns go through a queue: a ceiling on turns in flight, with a waiting room in front of it.
A queued user is told, in the character's drafted words, that lots of people are talking to him and he is coming in a moment. A user who cannot be served inside the wait ceiling is told to try again shortly. An operations endpoint reports queue depth and the current settings, so a support person can see the queue without reading logs.
The defaults were measured rather than picked. In my own load test, a burst of concurrent session starts stayed well under a second. The test fires many sessions and turns at once with the provider stubbed, so it costs nothing to run again.
Failure, mapped to behaviour
| What fails | What the service does | What the user sees |
|---|---|---|
| Timeout, dropped connection, 429, 5xx | Retries with jittered backoff | A slightly slower reply |
| Wrong key, malformed request | Does not retry | The pause line |
| Provider fails after words were sent | Stops retrying, stores a paused turn | What was said, then the pause line |
| User writes again after a pause | Answers the held question first | "I am back", then an answer to what they asked |
| Too many turns at once | Queues the turn | A line saying lots of people are talking |
| Queue wait runs out | Refuses the turn cleanly | A line asking them to try again shortly |
| Persona file missing or broken | Refuses to start at all | The app's drafted service-unavailable message, never a generic assistant |
The last row is the odd one out, and on purpose. A service that cannot load the character's persona does not start, because a generic assistant talking to a person in distress is a worse outcome than an outage.
The backup nobody has restored
Everyone has backups. Few people have restored one. A scripted drill dumps the database, restores it into a scratch copy, and compares the two table by table, plus the migration ledger, the views, the constraints and the indexes. It never writes to the source, and it drops what it created. It is meant to run quarterly, and it has a runbook.
One thing I learned running it: the drill needs a direct database connection rather than a pooled one, because it creates a database, and a connection pooler can ignore the database name in a connection string. The run finished in well under a minute end to end, on an empty database. That proves the procedure and the permissions. It says nothing yet about restoring a large, real set of conversations, and I would rather say that out loud than let a fast run imply otherwise.
Try it on your own demo
Start a conversation in your product, and while a reply is streaming, cut the network to your model provider. Watch the screen as a user would. If what you see is a spinner that never ends, a red error, or a sentence that rewrites itself, you have found your next week of work. Then restore last night's backup into a scratch database and diff it. Both are far cheaper to find on a Tuesday afternoon than in a live session.
related