Running AI on your own machines
In short: when data is not allowed to leave the building, the smallest useful setup is two machines on the same network: one that asks, one with the GPU that answers. It takes about thirty lines of…
- published
- read time
- 4 min
- words
- 849
- lang
- en
- filed under
- Engineering
In short: when data is not allowed to leave the building, the smallest useful setup is two machines on the same network: one that asks, one with the GPU that answers. It takes about thirty lines of Python, and it grows into containers and a scheduler once more than one project shares the hardware.
Why the cloud is often off the table
A lot of my work in the last few years has been in healthcare. At a hospital network, and in a research consortium across several countries, the first question about any model was never "which GPU?" It was "where does the data go?" For a lot of clinical data the honest answer has to be "nowhere". It stays on machines the institution owns, on a network it controls.
The consortium went one step further with federated learning: the model travelled between sites and the data stayed home. That is the heavy version. Most teams need something much simpler first. They have one decent machine with a GPU sitting under a desk, and several people on ordinary laptops who want to use it.
I put the simple version on GitHub as on-Premise-AI. It is deliberately small. Computer 1 sends a request over the local network. Computer 2 runs the model and sends the answer back.
The smallest setup that works
Two packages, flask and requests. The server loads the model once at start-up, which matters more than it looks, and exposes one endpoint.
# server.py, on computer 2 (the one with the GPU)
from flask import Flask, request, jsonify
app = Flask(__name__)
def load_model():
# stand-in: load your real model here, once
return lambda inputs: [len(str(x)) for x in inputs]
model = load_model()
@app.post("/predict")
def predict():
inputs = request.get_json()["inputs"]
return jsonify({"outputs": model(inputs)})
if __name__ == "__main__":
app.run(host="0.0.0.0", port=5000)
# client.py, on computer 1
import requests
r = requests.post(
"http://192.168.1.20:5000/predict", # computer 2's address on the LAN
json={"inputs": ["first case", "second case"]},
timeout=60,
)
r.raise_for_status()
print(r.json()["outputs"])
Find computer 2's address with ipconfig on Windows or ip addr on Linux, start the server, run the client. If it hangs, it is almost always the firewall on computer 2 blocking the port.
host="0.0.0.0" means anyone who can reach that machine can call your model, with no password. "Inside the hospital" is a lot of people. Allow only known client addresses in the firewall, put a shared token in a header, never turn on Flask's debug mode on a shared network, and do not log the request bodies if they hold patient data.What breaks first
The demo works on day one. These are the things that show up in the first few weeks, in the order I would expect them:
- Timeouts. A large input takes longer than the client is willing to wait. Set a real
timeouton the client and return an error the user can read, not a stack trace. - One request at a time. Flask's built-in server is for development. Put the app behind
gunicornorwaitress, and decide how many model copies fit in GPU memory before you add workers. - "Which version answered?" Return the model version in every response. When a result looks odd two weeks later, that field is the first thing you will want.
- The machine reboots. Someone installs updates on Friday night. Run the server as a service so it comes back by itself.
What it grows into
The second project is where the simple setup starts to hurt. Two models want the same GPU, they need different library versions, and one of them is loaded all day while being used twice. The README lists the next two steps, and they are the right ones:
- Docker, so several projects run side by side. Each model gets its own image with its own dependencies. Upgrading one never breaks the other.
- Orchestration and resource management for the shared hardware. Something has to decide which container gets the GPU, queue requests when it is busy, and unload models nobody is using.
At that point you are running a small private cloud, and the same tools apply: a reverse proxy in front, a queue, health checks, and a dashboard that shows GPU memory per container. None of it needs the data to leave the building.
When this is the wrong answer
Owning machines is not free. Someone patches them, someone replaces the disk, and someone gets the call when the GPU fan dies. If your data is allowed in a managed cloud with the right agreements in place, that is often cheaper in staff time. The on-premise route earns its keep when the rules say so, or when the same machine is busy most of the day.
If you want to try it, take two laptops on the same Wi-Fi, run server.py on one and client.py on the other, and swap the stand-in for the smallest real model you have. Then add the version field and the token before you show it to anyone.
related