The short answer
Pick your GPU for on‑prem LLMs by sizing VRAM first, then performance, power and software support. For a single-seat 7B model chat or RAG with short prompts, a 12 GB card (for example, an NVIDIA RTX 4070) is often sufficient with 4‑bit quantization. For a smoother 13B experience or multiple concurrent users, plan for 24 GB VRAM (for example, RTX 3090 or 4090). Running 70B-class models or doing non‑trivial fine‑tuning typically needs 48 GB+ VRAM or multiple GPUs with a fast interconnect; many teams use a server or cloud for that tier. Always validate with the specific model’s documentation and your context length and concurrency targets before buying.
Define your workload in 5 minutes
Your GPU choice follows your workload, not the other way around. Write down:
- Task profile: chat-only LLM, RAG over your documents, embeddings, code completion, vision-language, or fine‑tuning/LoRA.
- Model family and size you prefer (for example, 7B, 13B, 34B, 70B) and whether you will quantize to 4‑bit/8‑bit.
- Context length you need today (for example, 4k, 8k, 16k, 32k tokens). Longer context increases memory for the KV cache.
- Concurrency and latency goals (for example, 2 users at <1 second first token; 10 users batch generation).
- Availability constraints: must run fully offline, or can temporarily burst to cloud if you exceed capacity?
These five notes will surface the real constraints: VRAM, memory bandwidth, and whether you need a single quiet tower or a multi‑GPU server. They also prevent over‑buying a data‑center GPU when a well‑cooled workstation will do.
VRAM first: rough sizing for common models
Model weights plus runtime memory (KV cache, activations) must fit in VRAM if you want low latency. With quantization and smart runtimes, you can do more than many expect, but the margins matter. Treat the figures below as practical starting points and confirm with your chosen model and runtime.
- 7B models: 4‑bit quantization usually runs in roughly 4–6 GB for weights. With overhead and a modest context, plan 8–10 GB VRAM minimum; 12 GB gives headroom.
- 13B models: 4‑bit weights often land near 8–12 GB. With overhead and longer contexts or small batches, 16–24 GB is typical; 24 GB is comfortable for concurrent users.
- 34B–70B models: Inference generally exceeds a single consumer card’s VRAM even with 4‑bit. Expect to shard across multiple high‑VRAM GPUs and to benefit from fast interconnect (for example, NVLink).
Two easy-to-miss drivers of VRAM: context length and concurrent requests. Doubling context doubles KV cache memory. Serving two users at once doubles working set too. If you are unsure, prototype the exact model with your planned max context and one more concurrent user than you expect.
Throughput, bandwidth and multi‑GPU choices
Beyond fitting the model, you need acceptable tokens per second. More SMs/Tensor Cores and higher memory bandwidth help, which is why a 24 GB flagship card can feel much faster than a 24 GB older-generation card. For small models at 4‑bit, memory bandwidth and kernel optimizations in the runtime (for example, vLLM, TensorRT‑LLM, llama.cpp) are often the bottleneck.
Multi‑GPU is useful when a single card cannot fit your target model or when you want more throughput. For large models, tensor or pipeline parallelism splits weights across GPUs. This works best with high‑speed interconnects such as NVLink; plain PCIe works but can limit performance. If you cannot justify multi‑GPU hardware, consider model distillation, smaller models, or a hybrid design that runs basic queries locally and bursts complex ones to the cloud.
Power, cooling and form factor in a Polish office
High‑end GPUs draw significant power and produce heat. Verify your power supply unit (PSU) capacity and connectors (for example, adequate 12VHPWR or PCIe 8‑pin), case airflow, and where the machine will sit. A quiet workstation under a desk may be fine for a single 12–24 GB card. Two or more high‑TDP GPUs are usually better in a well‑ventilated tower or short‑depth rack with dedicated cooling.
Most Polish offices use 230 V circuits and common 16 A breakers. A single workstation is usually fine on a standard outlet, but if you plan multi‑GPU or continuous heavy loads, ask a licensed electrician to assess the circuit, peak draw, and UPS needs. A UPS rated with enough VA for your PSU and monitors protects you from brief outages and brownouts. Also plan for ambient noise; blower‑style cards and small racks can be loud in open offices.
Software compatibility: CUDA, ROCm, Windows or Linux
Ecosystem support still favors NVIDIA for LLM inference and training. CUDA, TensorRT‑LLM and mature kernels in common runtimes often deliver the least friction. AMD ROCm support has improved, but always check the exact ROCm and driver versions required by your chosen frameworks. Intel GPU support is evolving; verify with your software stack before committing.
Linux tends to be the safest choice for on‑prem LLM servers due to driver stability and container tooling. Many teams use Docker with prebuilt images for vLLM, llama.cpp, or text‑generation‑webui. Windows with WSL 2 can work for development, but for 24/7 services, a minimal Linux host with NVIDIA drivers and container runtime is simpler to maintain. Whatever you choose, freeze versions once stable and document your install steps so you can reproduce them on a replacement machine.
Example builds and what they can realistically do
Quiet desktop, 12 GB GPU: A modern 12 GB card can serve a 7B instruction‑tuned model at 4‑bit, RAG over a few hundred PDFs, and 1–2 concurrent users with short prompts. Expect modest tokens per second but usable latency. Great for pilots and departmental assistants.
Workstation, 24 GB GPU: A 24 GB flagship card provides a noticeably better experience. You can host a 13B model at 4‑bit with longer context, handle a handful of simultaneous chats, and generate embeddings quickly. With careful tooling, you may run small LoRA fine‑tunes on 7B models. This is a common sweet spot for SME rollouts.
Small server, 2 x high‑VRAM GPUs with NVLink: If you truly need large context or 70B‑class inference on‑prem, consider two 48 GB+ GPUs in a chassis that supports NVLink, ample cooling and a strong PSU. This tier moves you closer to data‑center territory in cost, noise and power. Many organizations instead keep an on‑prem 13B service and route occasional heavy jobs to a cloud GPU.
Poland‑specific purchasing notes: New parts are available via major retailers and integrators; used options appear on local marketplaces. If buying used, verify that cards were not modified for mining, check thermals under load, and ask for a valid warranty or at least a test window. For business purchases, a proper VAT invoice (faktura VAT) simplifies accounting. Some data‑center GPUs require compatible motherboards, chassis spacing and auxiliary power—confirm before you click buy.
Next steps
- Write your workload sheet: task, model(s), context, concurrency, latency targets, and tolerance for cloud bursting.
- Prototype in the cloud or on a borrowed GPU with the exact model and runtime to observe VRAM use and tokens/second at your target context.
- Pick 1–2 candidate GPUs and verify case dimensions, PSU headroom, connectors, and driver support for your OS. Plan for UPS and cooling.
- Run a one‑week pilot in your office and measure real load: peak concurrent sessions, errors, and thermal behavior. Adjust before a full purchase.
- If you want an end‑to‑end plan or a build you can support long‑term, talk to Auranik’s AI & Automation team in Poland (/poland/technology/ai). We can help you choose the right model, hardware and runtime, and set up monitoring so you know when to scale.
Community content reflects individual experiences and should not be treated as legal, immigration, financial or government advice.
Know someone who may find this guide useful?