We do analysis for clients whose data is, by nature, sensitive. That shapes how we build our tooling. One of the most consequential choices we made was to run our AI inference on hardware we own, inside our own network, rather than through third-party APIs. This post explains what that actually means, what it costs us, and when it's the right call for you.
The default path, and its quiet cost
The easy default for AI work is to send your data to a third-party model API. It's fast to start: no hardware to buy, no models to download, no serving stack to maintain. You get a URL, you send a request, you get an answer.
But there are costs that don't show up on the invoice. Your data leaves your building and lands on someone else's infrastructure. The provider's retention policies apply to what you send, and you're typically not the one who sets them. Usage is metered, so heavy analysis work racks up a bill that scales with the work. And the provider can change the model, the pricing, or the terms under you — the endpoint you built your workflow on can quietly become a different model or a different price.
For a consultancy that handles client material, that's not a theoretical risk. It's a real risk to manage, and managing it usually means avoiding the default path.
What on-premises actually means
Concretely, we run open-weight models on hardware we own and control. Our lab has about 96GB of GPU memory, and we serve models locally through OpenAI-compatible endpoints. When we run an analysis, the data stays on our network. Nothing leaves it.
The stack is unglamorous. We serve models with local inference engines — vLLM for the bigger workloads, llama.cpp for the rest — running quantized open-weight models. Quantization means the model weights are stored at reduced precision, which trades a little quality for a large reduction in memory, so a capable model fits in the hardware we have. Because the models are open-weight, we choose which one to run and pin the exact version. And because we serve it ourselves, there's no per-token billing: the cost of a run is electricity, not a meter.
One thing worth noting: we test both dense models and mixture-of-experts models, because the architecture affects real-world speed a lot more than spec sheets suggest. Two models with similar parameter counts can have very different throughput on the same hardware.
A typical serving command looks like:
# serve a quantized open-weight model locally
llama-server -m /models/qwen-exl3 --port 8000 --ctx-size 32768
After that, any OpenAI-compatible client can point at the local endpoint and work as if it were talking to a cloud API.
The hidden work (and why it matters to clients)
We want to be honest about what this involves, because it's not free. Running inference well is operations work. The hardware runs hot, draws real power, and is loud — heat, power, and noise all need tuning, not just a one-time setup. Quantization trades a little quality for memory, so the choice of quantization level is a real decision, not a default. Long context windows eat memory, because the model's working cache grows with the context, and a large context can push a model off the GPU entirely.
Performance has to be measured, not assumed. Published benchmark numbers are from other people's hardware with other people's settings. We benchmark our own setups and tune power limits on our own hardware, rather than trusting published numbers, because the difference between an assumed throughput and a measured one is the difference between a plan that works and one that doesn't.
This work matters to clients because it's the difference between "we can run this" and "we can run this reliably, at a known cost, with the data staying in the room."
When on-premises makes sense — and when it doesn't
On-premises inference is not automatically the right answer. It's the right answer for specific situations:
- Sensitive or confidential data. If the data shouldn't leave your building — client records, internal documents, anything under NDA — on-prem is the only option that keeps it there.
- Repetitive, high-volume analysis. If you're running the same kind of analysis many times, the absence of per-token billing means the marginal cost of each run is just power. The hardware pays for itself.
- Reproducibility and model version control. If you need the same model, at the same version, producing the same behavior across runs, owning the model and the serving stack gives you that control.
And there are situations where the cloud still wins:
- Frontier capability for hard one-off reasoning. If you need the very best model at the frontier for a single difficult problem, the cloud has it and you don't.
- Burst workloads. If the work is spiky — a big project that comes and goes — renting capacity is cheaper than owning hardware that sits idle most of the time.
Where this fits for our clients
On-prem inference is not the product. It's infrastructure. It's what lets us take on analysis we otherwise couldn't: financial records, transaction graphs, internal documents, anything that should never leave the room. The models do the heavy lifting, but the value is that the work can happen at all, without the data going anywhere.
If your data can't leave the building, that doesn't mean the analysis can't happen.
On-prem AI is less about running a model and more about making analysis possible where the data has to stay.