There's a growing scene of people running large language models on their own hardware — no API, no cloud, no per-token billing. For a lot of use cases it's not just a hobby; it's a genuinely better fit. Here's why, and where to start.
Why run it locally?
Privacy. Your data never leaves the machine. For sensitive work — research, internal documents, anything you wouldn't paste into a public API — that's a real advantage.
Cost. After the hardware is paid for, inference is free. Heavy, repetitive use that would rack up a cloud bill becomes just an electricity cost.
Control. You pick the model, the quantization, the context length, and the serving stack. No rate limits, no provider outages, no terms-of-service changes under you.
Iteration. Experimenting with prompts, models, and settings is fast when there's no network round-trip and no meter running.
The practical starting point
The most common stack today is llama.cpp (or a server built on it) serving an OpenAI-compatible API. The basic shape:
- Pick a model that fits your memory. Bigger is better, but only up to the point where it fits comfortably.
- Quantize it. Full-precision models are huge. Quantized versions (e.g. EXL3, GGUF quants) trade a little quality for a large reduction in memory.
- Serve it. Point a client (an app, a script, an agent) at the local endpoint.
A minimal serving command looks like:
llama-server \
-m /path/to/model-exl3 \
--port 8000 \
--ctx-size 32768
Then any OpenAI-compatible client can talk to http://localhost:8000/v1.
What to watch
- VRAM is the constraint. The model weights plus the KV cache for your context window all have to fit. If it doesn't, you spill to CPU and things get slow.
- Quantization is a tradeoff. Aggressive quants save memory but can degrade quality. Find the sweet spot for your use case.
- Context length costs memory. The KV cache grows with context. A 128K context window is not free.
- Benchmark your own setup. Numbers vary wildly by hardware. Measure tokens/second on your machine with your model before you plan around it.
Where this goes next
We'll be writing more about the local AI scene — specific model choices, quantization tradeoffs, and the real-world problems that come up when you run these things day to day.
Running a model locally is less about beating the frontier and more about having a capable, private, always-available tool that's yours.