← All posts

Running Local LLMs: Why It Matters and How to Start

Privacy, cost, and control — why running large language models on your own hardware is worth the effort, and a practical starting point.

There's a growing scene of people running large language models on their own hardware — no API, no cloud, no per-token billing. For a lot of use cases it's not just a hobby; it's a genuinely better fit. Here's why, and where to start.

Why run it locally?

Privacy. Your data never leaves the machine. For sensitive work — research, internal documents, anything you wouldn't paste into a public API — that's a real advantage.

Cost. After the hardware is paid for, inference is free. Heavy, repetitive use that would rack up a cloud bill becomes just an electricity cost.

Control. You pick the model, the quantization, the context length, and the serving stack. No rate limits, no provider outages, no terms-of-service changes under you.

Iteration. Experimenting with prompts, models, and settings is fast when there's no network round-trip and no meter running.

The practical starting point

The most common stack today is llama.cpp (or a server built on it) serving an OpenAI-compatible API. The basic shape:

  1. Pick a model that fits your memory. Bigger is better, but only up to the point where it fits comfortably.
  2. Quantize it. Full-precision models are huge. Quantized versions (e.g. EXL3, GGUF quants) trade a little quality for a large reduction in memory.
  3. Serve it. Point a client (an app, a script, an agent) at the local endpoint.

A minimal serving command looks like:

llama-server \
  -m /path/to/model-exl3 \
  --port 8000 \
  --ctx-size 32768

Then any OpenAI-compatible client can talk to http://localhost:8000/v1.

What to watch

  • VRAM is the constraint. The model weights plus the KV cache for your context window all have to fit. If it doesn't, you spill to CPU and things get slow.
  • Quantization is a tradeoff. Aggressive quants save memory but can degrade quality. Find the sweet spot for your use case.
  • Context length costs memory. The KV cache grows with context. A 128K context window is not free.
  • Benchmark your own setup. Numbers vary wildly by hardware. Measure tokens/second on your machine with your model before you plan around it.

Where this goes next

We'll be writing more about the local AI scene — specific model choices, quantization tradeoffs, and the real-world problems that come up when you run these things day to day.

Running a model locally is less about beating the frontier and more about having a capable, private, always-available tool that's yours.

Working on a network problem?Tell us what you’re trying to understand.

Start a conversation →