Open-weight models have become good enough that running an LLM on your own hardware is a realistic option, not a research project. Families such as Llama, Qwen, Mistral, Gemma and DeepSeek offer models from a few billion to hundreds of billions of parameters, many under licences that allow commercial use. The question for most companies is no longer "can we?" but "should we, and how?". This guide covers the decision, the tooling, hardware sizing and the operational work that vendors' blog posts tend to skip.

When self-hosting makes sense

Self-hosting is justified by one of four reasons:

  1. Data cannot leave your infrastructure. Regulated data, client contracts that forbid third-party processing, classified or defence-related work. This is the most common and strongest reason; see privacy for LLM apps.
  2. High, steady volume of a narrow task. Classification, extraction or embedding of millions of items per day, where a small fine-tuned model on owned GPUs beats per-token pricing. See fine-tuning vs RAG vs prompting.
  3. Latency or offline requirements. Edge deployments, factories, on-premises products without reliable internet.
  4. Control. Pinning an exact model version for years, customising decoding, or avoiding vendor policy changes.

Self-hosting is usually not justified for general assistants, complex agents and coding tasks, where frontier API models remain substantially stronger, or for low and spiky volumes, where idle GPUs cost more than API calls. Many teams end up hybrid: frontier APIs for hard tasks, self-hosted models for sensitive or high-volume ones.

Choosing an inference server

Tool Best for Notes
vLLM Production serving on GPUs PagedAttention, continuous batching, OpenAI-compatible API, tensor parallelism, structured output
SGLang High-throughput serving, complex prompting programs RadixAttention prefix caching (Zheng et al., 2023)
Hugging Face TGI Production serving in the HF ecosystem Mature, good integration with HF Hub
Ollama Developer machines, small internal tools One-command setup, model library, built on llama.cpp
llama.cpp CPU, Apple Silicon, edge devices GGUF quantized models, minimal dependencies

The key innovation behind vLLM is PagedAttention, which manages the attention key-value cache in pages like virtual memory, reducing fragmentation and enabling much larger batches; the paper reports 2–4x higher throughput than prior systems at the same latency (Kwon et al., 2023). Combined with continuous batching — adding new requests to a running batch instead of waiting — this is why a dedicated server serves many more users per GPU than a naive script.

Our default: Ollama for local development, vLLM for production. Both expose OpenAI-compatible endpoints, so application code stays the same.

# Production-style vLLM server with an OpenAI-compatible API
vllm serve Qwen/Qwen2.5-14B-Instruct \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.90 \
  --enable-prefix-caching \
  --api-key "$VLLM_API_KEY"

Sizing GPU memory

GPU memory is the binding constraint. Two components dominate:

Model weights: parameters × bytes per parameter.

Model size FP16/BF16 INT8 4-bit
8B ~16 GB ~8 GB ~5 GB
14B ~28 GB ~14 GB ~9 GB
32B ~64 GB ~32 GB ~18 GB
70B ~140 GB ~70 GB ~40 GB

KV cache: grows with context length × concurrent sequences. For long contexts and many users, the KV cache can exceed the weights. Grouped-query attention in modern models reduces it; prefix caching shares it across requests with the same system prompt.

Rules of thumb: leave 20–30% of GPU memory for KV cache and activations beyond the weights; a single 24 GB GPU comfortably serves a quantized 8–14B model to a small team; 32B models fit on one 80 GB GPU in FP16 or one 48 GB GPU quantized; 70B-class models need multiple GPUs or aggressive quantization.

Quantization

Quantization stores weights in fewer bits, cutting memory and often increasing speed:

  • GPTQ and AWQ are post-training methods for 4-bit GPU inference with small quality losses (Frantar et al., 2022; Lin et al., 2023).
  • FP8 is well supported on recent data-centre GPUs and is close to lossless for many models.
  • GGUF quantization levels (Q4_K_M, Q5_K_M, Q8_0…) are the llama.cpp/Ollama format for CPU and Apple Silicon.

Quality loss depends on the model and task — reasoning and non-English text can suffer more than English chat. Always evaluate the quantized model on your own eval set, including Ukrainian if relevant; see LLM evals.

Deployment on Kubernetes

For production, run inference servers as Kubernetes workloads on GPU nodes:

  • GPU node pool with the NVIDIA device plugin or GPU operator; taints so only inference pods land there.
  • Model storage: pre-download weights to a persistent volume or node-local cache; pulling 30+ GB at pod start makes scaling painfully slow.
  • Readiness probes on the model endpoint — loading a large model takes minutes.
  • Autoscaling on queue length or GPU utilisation rather than CPU; scale-to-zero only if cold starts are acceptable.
  • A gateway in front (an OpenAI-compatible router or LiteLLM-style proxy) for authentication, rate limits, routing between models and fallback to external APIs; see LLM API reliability.

GPU servers from European providers can be far cheaper than hyperscalers for steady workloads; our experience with cost-efficient clusters is in Kubernetes on Hetzner.

Security and compliance

Self-hosting moves responsibility to you:

  • Authenticate every endpoint. Open inference ports are routinely found and abused.
  • Network isolation. Inference servers should not need outbound internet in production.
  • Supply chain. Download weights from official sources, verify checksums, prefer safetensors over pickle-based formats, and scan containers; see container supply chain security.
  • Licences. Open-weight licences differ: some restrict commercial use above a user threshold or for certain purposes. Record the licence of every model you deploy.
  • Logging. Prompts may contain personal data; apply the same retention and access controls as for any other system holding it; see LLM observability.
  • EU AI Act. Deploying an open model inside your product can carry obligations depending on use; see the EU AI Act guide.

Total cost: an honest comparison

Compare total cost of ownership, not GPU hourly price against token price:

Cost item API Self-hosted
Inference Per token GPU servers 24/7 (or reserved), regardless of usage
Engineering Integration only Setup, upgrades, monitoring, on-call
Quality Frontier models Open models; may need more prompt or fine-tuning work
Scaling Instant Capacity planning, procurement
Compliance Vendor DPA, data transfer assessment Your own controls, often simpler legally

A rough break-even analysis: estimate tokens per month, multiply by API prices for a model of comparable quality, and compare with GPU cost plus 0.25–0.5 FTE of engineering time. For many SMB workloads, APIs win on cost; self-hosting wins on data control and very high volume. Re-run the numbers every six months — both API prices and open model quality change quickly.

A rollout path

  1. Prototype with an API to prove the use case and build evals.
  2. Shortlist open models that fit your hardware and licence constraints; see choosing an LLM.
  3. Run them locally with Ollama against the eval set.
  4. Deploy vLLM on one GPU node for a pilot group; measure latency, throughput and quality.
  5. Harden: gateway, authentication, monitoring, autoscaling, backups of configurations.
  6. Optimise: quantization, prefix caching, speculative decoding — a small draft model proposes tokens that the large model verifies in parallel, cutting latency without changing outputs (Leviathan et al., 2022) — and fine-tuning a smaller model if volume justifies it.

FAQ

Can we run a useful model without a GPU? Small quantized models (up to ~8B) run on modern CPUs and Apple Silicon with llama.cpp or Ollama — fine for low-volume internal tools, too slow for many concurrent users.

Are open models good enough for Ukrainian? Larger multilingual open models handle Ukrainian reasonably well; small models are noticeably weaker. Test on your own tasks.

Is self-hosting automatically GDPR-compliant? No. It removes third-party transfers, but you still need a legal basis, retention limits, access control and security for the data you process.

How do we keep up with new models? Keep the application behind an OpenAI-compatible interface, keep your eval set current, and evaluate new releases quarterly.

Sources

  1. Kwon et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention; vLLM documentation.
  2. Zheng et al. (2023). SGLang: Efficient Execution of Structured Language Model Programs.
  3. Ollama and llama.cpp on GitHub.
  4. Frantar et al. (2022). GPTQ; Lin et al. (2023). AWQ.
  5. Leviathan et al. (2022). Fast Inference from Transformers via Speculative Decoding.
  6. Hugging Face. safetensors.