Self-Hosted LLM Guide: GPU Selection, VRAM and Deployment

Choose a GPU for self-hosted LLM inference using model-size guidance, VRAM calculations, and vLLM and Ollama benchmark data. Compare deployment options for 7B–70B models before selecting a server.

By GPU Mart Technical Team · Updated September 17, 2026

Self-hosted LLMs run language models on infrastructure you control instead of sending every request to a third-party model API. A self hosted LLM environment can provide greater control over deployment, data handling, model versions, and available GPU resources.

This guide explains how to choose an LLM server by comparing LLM GPU requirements, VRAM sizing, deployment frameworks, and benchmark results. Always validate the model, quantization, context length, expected concurrency, and software stack before selecting a server.

Choose a Self-Hosted LLM Deployment Path

Start with your self hosted LLM workload, then compare model size, VRAM requirements, concurrency, and framework compatibility before choosing an LLM server.

Choose GPU Hardware

Compare typical starting points for development, production inference, RAG, and large quantized models.

View GPU Recommendations →

Build a Production API

Use vLLM for continuous batching, OpenAI-compatible endpoints, and higher-concurrency serving.

Explore vLLM Hosting →

Develop with Ollama

Use Ollama for model evaluation, private assistants, and simpler single-user or small-team deployments.

Explore Ollama Hosting →

Quick GPU Recommendations by Workload

WorkloadTypical VRAMStarting GPU OptionsSelection Note
7B–14B development and testing16–24 GBRTX A4000 / RTX Pro 2000 / RTX Pro 4000Prioritize affordability and enough headroom for context and KV cache.
14B–35B production inference24–48 GBRTX 5090 / RTX A6000 / RTX Pro 5000Balance model fit, memory bandwidth, concurrency, and persistent operation.
RAG, agents, and multiple models48 GB+RTX Pro 5000 / RTX A6000 / A40Allow additional VRAM for embeddings, rerankers, longer context, and concurrent services.
70B-class quantized models48–96 GB+RTX Pro 6000 / A100 / H100 / Multi-GPURequirements vary substantially by quantization, context length, and concurrency.
High-concurrency production API80 GB+ or Multi-GPUH100 / A100 / Multi-GPUValidate throughput using your model and production request profile.

These are starting points, not guaranteed capacity limits. Check current configurations and pricing before ordering.

Why Teams Switch Away from Cloud AI APIs

Three pain points that surface once your self-hosted LLM text generation workloads move from prototype to production at scale.

Unpredictable Costs

AWS / GCP / OpenAI charge per token — the more you use, the higher the bill. Usage-based API costs increase with traffic, and provider limits may affect some high-volume workloads.

Unpredictable Latency

Shared or queued services can introduce latency variability. Dedicated resources can provide more predictable capacity for sustained workloads.

Data Privacy Risk

Third-party APIs process prompts outside your own server environment. Teams handling sensitive data should evaluate provider policies, architecture, and compliance requirements before deployment.

Why Self-Host? The Four Deployment Paths Compared

Understanding structural limits matters more than comparing spec sheets — especially for production LLM inference at scale.

Deployment PathExamplesKey AdvantageKey LimitationBest For
Cloud APIOpenAI / Anthropic / GeminiFast setup, no ops overheadData leaves your env; token costs scale; rate limits at peakPrototyping, very low volume
Model MarketplaceTogether AI / FireworksWide model selectionShared resources, unstable perf, limited customizationMid-volume testing
On-Prem InfrastructureOwned workstation or private rackMaximum hardware controlCapital expense, capacity planning, and ongoing operationsTeams with internal infrastructure expertise
Dedicated GPU HostingGPU MartDedicated resources, predictable monthly billing, and environment controlRequires workload sizing and basic system administrationPersistent inference and private AI environments

GPU Mart Dedicated GPU Hosting: Key Advantages

Exclusive GPU Resources

Dedicated GPU allocation avoids GPU compute contention from other tenants. CPU, memory, storage, and network specifications still vary by plan.

Greater Data Control

Run models in your own hosted server environment and control application-level data flows. Compliance depends on your configuration, policies, and legal requirements.

Predictable Dedicated Capacity

Persistent dedicated GPU capacity can reduce cold-start and shared-resource variability. Actual latency depends on the model, framework, context length, and concurrency.

Predictable Monthly Cost

Flat-rate monthly hosting can make sustained, high-utilization workloads easier to budget than per-token services. Compare total cost using your actual traffic and operations requirements.

DedicatedGPU allocation
Flat-rateMonthly billing options
99.9%Uptime SLA
<5 minSupport response, 24/7

What Hardware Determines LLM Performance?

Understanding LLM GPU requirements, including VRAM, memory bandwidth, and precision support, is essential before selecting hardware for an LLM inference server.

1 — VRAM: LLM GPU Requirements by Model Size

VRAM is usually the first constraint when sizing a self-hosted LLM. The GPU must hold the model weights, KV cache, runtime overhead, and sufficient headroom for the expected context length and concurrency. The examples below use a 14B-parameter model as a reference; larger models require proportionally more memory.

Note on Q4_K_M (Ollama default): to avoid the significant quality degradation of pure 4-bit quantization, Ollama uses a "mixed precision" approach — core weights at 6-bit, less critical weights at 4-bit. This keeps quality loss minimal while keeping a 35B model at approximately 22–23 GB VRAM.

Quantization FormatBytes / Param14B Model VRAMNotes
FP16 / BF16 (full precision)2 bytes~28 GBHighest quality, no precision loss
FP81 byte~14 GBNear-FP16 quality; requires native GPU support
INT81 byte~14 GBSlight quality loss; broad compatibility
Q4_K_M (Ollama GGUF default)~0.55 bytes~7–8 GBMixed precision (6-bit core + 4-bit other); 35B fits ~22–23 GB
INT4 / AWQ / GPTQ0.5 bytes~7 GBHeavy compression; good for constrained setups

KV Cache VRAM Requirements: Context Length & Concurrency

Beyond model weights, inference requires KV Cache memory for every active request. For Qwen2.5-14B at FP16 (24 layers, 8 KV heads, 64 dims per head):

KV Cache / Token = 2 × L × H_kv × D_head × B = 2 × 24 × 8 × 64 × 2 bytes ≈ 48 KB / Token
ConcurrencyContext Length (Tokens)KV Cache VRAMTypical Use Case
11,024~48 MBSingle-user short chat
81,024~384 MBSmall team short chat
321,024~1.5 GBMedium concurrency short chat
132,768 (32K)~1.5 GBSingle-user long document
1131,072 (128K)~6 GBSingle-user extended context

VRAM requirements formula: Total VRAM = Model Weights + KV Cache + 20–30% headroom. Always size for peak concurrency, not just model weights. This is the most common cause of OOM errors in self-hosted LLM deployments.

2 — Memory Bandwidth: How Fast Tokens Generate

Generating each token requires loading the full model weight set from VRAM into Tensor Cores. Memory bandwidth — a key LLM GPU requirement — is the hard ceiling on tok/s, not TFLOPS. A100-40G at 1,555 GB/s on a 20B FP16 model: theoretical max ≈ 38 tok/s for a single request.

Workload ProfilePrimary VRAM UsageBottleneckRecommended Strategy
Low concurrency / short contextWeights dominantMemory bandwidthHigh-bandwidth GPU: H100, RTX 5090
High concurrency / long contextKV Cache dominantCompute (Tensor Core queue)Large-VRAM GPU + batching
Offline batch processingWeights + large-batch KV CacheBandwidth & compute bothH100 / A100 + vLLM continuous batching

3 — Tensor Core Precision: Where Blackwell Pulls Ahead

On GPUs with native FP4/FP8 support: FP4 compute is 2× FP8, and FP8 is 2× FP16. FP8-quantized model weights also use half the VRAM of FP16, and FP4 uses a quarter — freeing more VRAM for context and concurrency. Only Blackwell GPUs can run LLMs at FP4 precision, achieving maximum possible throughput.

PrecisionRTX Pro 6000 (Blackwell)H100-80GA100-80GNotes
TF32234 TFLOPS—312 TFLOPSTraining common precision
FP16 / BF161,000 TFLOPS989 TFLOPS312 TFLOPSMain inference precision
INT8 / FP82,000 TFLOPS~1,979 TFLOPSNo native FP8FP8: 2× throughput, half VRAM vs FP16
FP4 / INT44,000 TFLOPSNot supportedNot supportedBlackwell-exclusive: 4× vs FP16, quarter VRAM

Best GPUs for LLM Workloads

Finding the best GPU for LLM workloads depends on model size, concurrency, and budget. GPU specifications should be checked against the exact installed model and form factor. Current pricing: gpu-mart.com/pricing.

Tensor Cores Dense: FP16 dense compute (TFLOPS).  AI TOPS: max compute at lowest supported precision (sparse).  Precision: natively supported precisions.

GPU Model VRAM Mem BW Tensor Cores
Dense
(FP16)
AI TOPS
(max
prec.)
Precision Max Model Typical Use Case Current Price Order
Data Center — Volta
V100-SXM2-16G16 GB900 GB/s21.2 TFLOPS1,248 TOPSFP16
INT8
~7B FP16Legacy production, FP16 inferenceCheck current priceOrder Now
Data Center — Ampere
A40-48G48 GB696 GB/s149.7 TFLOPS—FP16/BF16
INT8/INT4
~27B FP8Offline batch, doc analysisCheck current priceOrder Now
A100 - SXM4 - 40GB40 GB1,555 GB/s312 TFLOPS2,496 TOPSFP16/BF16
INT8/TF32
~14B FP16Mid-size inference/trainingCheck current priceOrder Now
A100 - 80GB - PCIe80 GB1,935 GB/s312 TFLOPS2,496 TOPSFP16/BF16
INT8/TF32
~40B FP16Large model productionCheck current priceOrder Now
Data Center — Hopper
H100 - 80G - PCIe80 GB2,000 GB/s756 TFLOPS2,026 TOPSFP16/BF16
FP8/INT8
~80B FP8High-concurrency production APICheck current priceOrder Now
Professional Workstation — Ampere
RTX A4000-16G16 GB448 GB/s76.7 TFLOPS153.4 TOPSFP16
INT8
~7B FP16 / 14B INT4Dev/test, single-userCheck current priceOrder Now
RTX A5000-24G24 GB768 GB/s111.1 TFLOPS222.2 TOPSFP16
INT8
~13B FP16 / 27B INT4Mid-model dev/testCheck current priceOrder Now
RTX A6000-48G48 GB768 GB/s154.9 TFLOPS309.7 TOPSFP16
INT8
~24B FP16 / 48B INT4Mid-large private deployCheck current priceOrder Now
Consumer — Ada Lovelace
RTX 4090-24G24 GB1,008 GB/s165 TFLOPS1,321 TOPSFP16/BF16
FP8/INT4/TF32
~13B FP16 / 27B INT4High-speed small modelCheck current priceOrder Now
Consumer — Blackwell
RTX 5090-32G32 GB1,792 GB/s419 TFLOPS3,352 TOPSFP16/BF16
FP8/FP4/INT4
~35B INT4Fast inference, small teamCheck current priceOrder Now
Professional — Blackwell (Recommended)
ⓘ  GPU Mart GPU VPS uses KVM PCI GPU Passthrough — exclusive GPU, no shared resources
RTX Pro 2000-16G16 GB288 GB/s136 TFLOPS545 TOPSFP16/BF16
FP8/FP4/INT4
~7B FP16Lightweight, single userCheck current priceOrder Now
RTX Pro 4000-24G24 GB672 GB/s147 TFLOPS1,178 TOPSFP16/BF16
FP8/FP4
~27B INT4Small-team mid modelCheck current priceOrder Now
RTX Pro 5000-48G48 GB1,344 GB/s268 TFLOPS—FP16/BF16
FP8/FP4
~35B INT4 (128K)Agent, RAG, multi-modelCheck current priceOrder Now
RTX Pro 6000-96G96 GB1,597 GB/s500 TFLOPS4,000 TOPSFP16/BF16
FP8/FP4
INT4
~122B INT4Large model single-cardCheck current priceOrder Now

LLM Inference Server Benchmarks

vLLM framework · Input 1,024 tokens + Output 512 tokens · GPU Mart test hardware

Benchmark scope: These results can help estimate the starting configuration for an inference server, but they are specific to the listed model, precision, request length, concurrency, framework version, driver, CUDA version, and server configuration. Validate performance with your own model and production request pattern before ordering.

Mean TTFT = Time to First Token (lower is better)  ·  P50 TTFT = median first-token latency  ·  Mean E2EL = end-to-end latency
Per-user Output Tokens/s: Average output token generation speed per request under the specified concurrency level. Reflects single-stream generation performance in a multi-user serving environment.
Aggregate Output Tokens/s: Total output token generation rate across all concurrent requests. Measures overall serving capacity excluding input tokens.

Most common enterprise deployment model · All concurrency levels shown

GPUConcurrencyMean TTFT (s)P50 TTFT (s)Per-user Output Tokens/sAggregate Output Tokens/sMean E2EL (s)
A40-48G11.7221.7775.165.1699.15
A40-48G86.6016.1864.8839.0695.03
A40-48G3212.75810.9043.0597.52158.43
A100-80G10.6300.64720.5120.5124.96
A100-80G81.0600.52620.67165.3623.61
A100-80G322.6121.06011.01352.4643.34
A6000-48G10.2710.28823.1523.1522.12
A6000-48G80.9470.91919.35154.7825.54
A6000-48G322.2271.68712.70406.4337.41
H100-80G10.1990.23640.0740.0712.78
H100-80G80.9540.71034.81278.4814.70
H100-80G321.0860.37024.26776.4619.36
RTX 5090-32G10.1640.18340.1040.1012.77
RTX 5090-32G80.5710.54931.98255.8415.45
RTX 5090-32G321.0440.39422.20710.5319.25
RTX Pro 5000-48G10.1640.18340.5540.1012.77
RTX Pro 5000-48G80.5710.54934.34255.8415.45
RTX Pro 5000-48G321.0440.39428.07710.5319.25
RTX Pro 6000-96G11.1030.14141.741.712.272
RTX Pro 6000-96G80.3950.23940.1320.612.302
RTX Pro 6000-96G321.0990.37727.2871.017.435

For 14B FP16 models, the best GPUs for LLM inference are H100, RTX 5090, and RTX Pro 5000 — all achieving ~40 tok/s single-user. A40 bottlenecks at bandwidth (~5 tok/s). A6000-48G delivers best price-to-throughput for production LLM hosting: $409/mo for 406 tok/s at 32-concurrency.

Blackwell FP8 native acceleration vs Ampere · All concurrency levels shown

GPUConcurrencyMean TTFT (s)P50 TTFT (s)Per-user Output Tokens/sAggregate Output Tokens/sMean E2EL (s)
A40-48G10.9421.08325.0625.0620.43
A40-48G81.7441.23215.74125.9331.52
A40-48G325.1271.4826.98223.2070.00
A6000-48G10.2250.20869.4269.427.37
A6000-48G80.4200.25454.70437.629.03
A6000-48G320.7680.42431.19997.9615.45
H100-80G10.1200.103122.22122.224.19
H100-80G80.2140.15087.04696.334.19
H100-80G320.5540.34845.961,470.834.19
RTX 5090-32G10.0950.080144.41144.413.55
RTX 5090-32G80.1640.121126.161,009.263.91
RTX 5090-32G320.4040.34190.302,889.665.21
RTX Pro 5000-48G10.1060.087118.62116.014.41
RTX Pro 5000-48G80.1750.111108.11804.734.90
RTX Pro 5000-48G320.3970.27981.632,269.426.66
RTX Pro 6000-96G10.0340.033133.8133.83.826
RTX Pro 6000-96G80.0610.050120.1960.54.107
RTX Pro 6000-96G320.7770.81978.22,502.76.021

For 8B-FP8 models, the best GPU for LLM inference is the RTX 5090 at 144 tok/s — nearly matching H100 (122 tok/s) at less than 1/5 the cost. RTX Pro 5000 achieves 118 tok/s with 48 GB VRAM. A6000 manages only 69 tok/s — this is Blackwell FP8 native advantage in action.

A100 vs H100 on mid-large models · FP8 quantization

GPUConcurrencyMean TTFT (s)P50 TTFT (s)Per-user Output Tokens/sAggregate Output Tokens/sMean E2EL (s)
A100-80G11.3661.32315.7515.7532.50
A100-80G84.2813.68713.22105.7637.54
A100-80G327.4807.6877.10227.1169.36
H100-80G10.3470.30837.7937.7913.55
H100-80G81.4381.52032.16257.2615.27
H100-80G322.9142.89215.61499.3930.55
RTX Pro 6000-96G10.2660.16945.845.811.183
RTX Pro 6000-96G81.9112.26935.3282.713.962
RTX Pro 6000-96G324.2554.34521.6692.622.094

H100 is 2.4× faster than A100 on 27B-FP8 (37.79 vs 15.75 tok/s). A100 hits 7.5-second TTFT at 32 concurrency — unacceptable for real-time API. For 27B+ FP8 models in production, H100 is the only correct single-GPU choice.

H100-80G only · High-concurrency limits test

GPUConcurrencyMean TTFT (s)P50 TTFT (s)Per-user Output Tokens/sAggregate Output Tokens/sMean TPOT (ms)Mean E2EL (s)
H100-80G11.0771.25815.1815.1863.9033.73
H100-80G84.9454.82211.9195.2471.2341.35
H100-80G3270.82378.2524.10131.2586.11114.83
RTX Pro 6000-96G10.3490.33121.221.2—24.163
RTX Pro 6000-96G82.8013.42317.1137.2—28.816
RTX Pro 6000-96G3217.8806.9218.7276.8—56.087

At concurrency 32, TTFT spikes to 70 seconds — 31B model on a single H100 hits severe queue buildup above 8 concurrent requests. For production, cap concurrency at 4–8 per card, or use multi-GPU deployment.

Ready to run these benchmarks on your own workload?

Compare current configurations, VRAM, server type, and availability before ordering
View AI Server Plans

Ollama Single-User Benchmarks

llama.cpp backend · Q4_K_M quantization · Suitable for self-hosted LLM development on a GPU VPS · No KV Cache pre-allocation · Input 1,024 tokens + Output 512 tokens · Single user · Average of 10 requests

VRAM Usage & Maximum Supported Models

GPU (VRAM)ModelVRAM UsedContextNotes
RTX Pro 6000 (96 GB) — 120B-class models
RTX Pro 6000qwen3.5:122b95 GB262,144 (256K)Highest VRAM usage
RTX Pro 6000gpt-oss:120b70 GB131,072 (128K)
RTX Pro 6000qwen3-coder-next:latest61 GB262,144 (256K)
RTX Pro 5000 (48 GB), A6000, A40 — 35B primary workloads
RTX Pro 5000glm-4.7-flash:latest40 GB202,752 (~200K)
RTX Pro 5000qwen3.5:35b34 GB262,144 (256K)35B — full 256K context
RTX 5090 (32 GB) — 35B with context reduction
RTX 5090gemma3:27b30 GB131,072 (128K)
RTX 5090qwen3.5:35b30 GB131,072 (128K)35B — 128K context
RTX 5090qwen3.5:35b27 GB32,768 (32K)35B — 32K context
RTX 5090gemma4:31b27 GB32,768 (32K)
RTX 5090qwen3.5:35b26 GB16,384 / 8,19235B — 16K / 8K context
RTX Pro 4000 (24 GB), A5000, RTX 4090
RTX Pro 4000qwen3.6:27b24 GB32,768 (32K)
RTX Pro 4000gemma4:26b20 GB32,768 (32K)
RTX Pro 4000deepseek-v2:16b19 GB32,768 (32K)
RTX Pro 4000qwen3.5:4b17 GB262,144 (256K)Small param, ultra-long context
RTX Pro 2000 (16 GB), A4000, P100, V100
RTX Pro 2000gpt-oss:20b14 GB32,768 (32K)
RTX Pro 2000qwen3.5:9b9.9 GB32,768 (32K)Lowest VRAM usage

Context reduction tip: set num_ctx to reduce VRAM and run 35B models on 32 GB cards:

curl http://localhost:11434/api/generate -d '{"model":"qwen3.5:35b","prompt":"hello","options":{"num_ctx":32768}}'

Generation Speed by GPU

gemma4:26b — 20 GB, 32K context
GPUAvg TTFT (s) ↓P50 TTFT (s)Avg Gen Speed tok/s ↑Avg E2E Time (s) ↓
RTX 50904.9334.917149.774.93
RTX Pro 60004.9654.897140.414.97
RTX Pro 50005.0455.006136.795.05
RTX Pro 40006.1146.079107.566.11
A60006.4826.438102.306.48
A50006.8306.80093.916.83
A407.0487.00892.247.05
gpt-oss:20b — 14 GB, 32K context
GPUAvg TTFT (s) ↓P50 TTFT (s)Avg Gen Speed tok/s ↑Avg E2E Time (s) ↓
RTX 50900.6530.619214.903.67
RTX Pro 60000.5560.558202.253.62
RTX Pro 50000.6130.597178.843.98
A60000.6420.638124.665.28
RTX Pro 40000.5530.555117.605.37
A50000.6640.620109.005.85
A400.6460.64596.456.60
RTX Pro 20000.5410.53261.699.24
qwen3.5:9b — 9.9 GB, 32K context
GPUAvg TTFT (s) ↓P50 TTFT (s)Avg Gen Speed tok/s ↑Avg E2E Time (s) ↓
RTX 50900.6180.597140.454.97
RTX Pro 60000.5970.589130.045.16
RTX Pro 50000.5760.579123.135.31
A60000.7470.73280.957.73
A400.7470.73280.957.73
RTX Pro 40000.6000.60378.597.75
A50000.7570.73870.508.60
RTX Pro 20000.7460.74942.1313.45

RTX Pro 5000 (48G Blackwell) achieves 178 tok/s on gpt-oss:20b and 123 tok/s on qwen3.5:9b — approaching RTX 5090 while offering 48 GB vs 32 GB VRAM. Best overall value for single-user LLM hosting when both speed and model capacity matter.

Inference Framework Selection

Framework choice affects throughput and latency as much as GPU selection — pick the right one for your LLM inference use case.

DimensionOllamavLLMSGLangTensorRT-LLM
Design goalLocal single-userProduction high-concurrency APIHigh-throughput structured inferenceMax throughput (NVIDIA only)
Deployment complexitySimplestMediumMediumVery high (requires compilation)
Cold start timeSeconds~62 sec~58 sec~28 min
Single-user TTFT~65 ms~10.7 ms~11–12 ms~10.5 ms
High-concurrency throughputLow (~484 tok/s)HighHigher (+17–29% vs vLLM)Highest
VRAM usageLow (INT4, no pre-alloc)High (pre-allocated)High (pre-allocated)High
FP8/FP4 supportPartialFullFullFull
OpenAI-compatible APIYesYesYesYes
GPU supportNVIDIA + Apple MNVIDIA + AMD + TPUNVIDIA + AMDNVIDIA only

Ollama

Dev / Test

One-command setup, model auto-download. Best for dev, single-user text generation, rapid evaluation.

vLLM

Production First Choice

General production API, widest model compatibility (incl. AMD/TPU). Continuous batching default on.

SGLang

RAG / Agent

RadixAttention prefix caching delivers +29% text generation throughput vs vLLM. Ideal for RAG, multi-turn, DeepSeek-class models.

TensorRT-LLM

Advanced Only

Only if: max throughput on pure NVIDIA stack AND you accept 28-min compilation per model version.

Deployment Optimization Tips

Reduce VRAM (Fix OOM)

  • Enable vLLM continuous batching (default on): dynamically merges requests, lowers peak VRAM
  • Quantize to INT4/INT8 via AWQ/GPTQ: 2–4× VRAM reduction, minimal quality loss
  • Reduce max_tokens and batch_size: cuts peak KV Cache usage
  • In Ollama: set num_ctx explicitly to allocate only what you need
  • Last resort: CPU offloading (latency penalty; only for extreme VRAM shortage)

Improve Throughput & Speed

  • Choose high-bandwidth GPUs: bandwidth caps tok/s. H100 > RTX 5090 > A100 > A6000
  • Use FP8 models + FP8-capable GPU: doubles throughput at same VRAM — also reduces weight memory footprint so more context and concurrency fits
  • Enable Speculative Decoding: small draft model assists large model, reduces TPOT
  • Multi-GPU tensor parallelism: vLLM/SGLang support --tensor-parallel-size

Auxiliary Models for <16 GB GPUs

Model TypeExample ModelTypical VRAMUse Case
EmbeddingQwen3-Embedding-8B~10 GBRAG vector encoding
Rerankerbge-reranker-large~1.7 GBRetrieval result reranking
ASRWhisper / Wav2Vec2–6 GBSpeech-to-text transcription
VLM (Vision, small)MedGemma-4B~4 GBMultimodal perception

Real Customer Deployments

Representative deployment examples. Identifying details are omitted; performance varies by model, software, and workload.

Case 1 — AI Application Company: Private Agent Platform

AI software company running Agent systems for code generation, document processing, and automated tasks 24/7. Previously on cloud API token billing — costs unpredictable, rate limits triggered at peak. The team moved to a persistent dedicated environment to control its model stack and operating schedule.

Config: RTX A6000 (48 GB) dedicated · Dual E5-2697v4 CPUs · 256 GB DDR4 · vLLM · Qwen3.5-27B GPTQ-Int4

First-token <500 ms 40–50 tok/s 168h continuous — zero failures Persistent dedicated deployment

Case 2 — Enterprise: Knowledge Base RAG

Large enterprise with millions of documents. Internal Q&A, customer service assist. Data cannot leave the internal network.

Stack: vLLM + gpt-oss-20b (29.6 GB) + Qwen3-Embedding-8B (10.3 GB) + bge-reranker-large (1.7 GB) + Weaviate / FAISS + DIFY workflow.

Avg response 1.5–2.5 sec P99 <3.5 sec 5–8 concurrent threads Cost outcome depends on utilization

Case 3 — Dev Team: Local Coding Assistant

Enterprise AI R&D team, high-frequency code generation for multiple developers. Previously on Claude/GPT APIs — code data leaving company.

Config: RTX Pro 6000 (96 GB) dedicated · 32-core CPU · 84 GB RAM · 1 Gbps unmetered · vLLM · GLM-4.7-Flash

Avg response 1–2 sec P99 <3 sec All code processed locally Cost outcome depends on utilization

Additional Industry Deployment Records

Customer TypeCore NeedGPUModel ArchitectureNotes
Medical AI vendorMultimodal clinical note generationRTX A6000Whisper + vision model + LLMMedical-grade privacy
AI medical teamImage + text joint reasoningRTX A6000MedGemma-4B-it multimodalMultimodal medical scene
AI application companyChat + memory + image + image gen AgentRTX A6000LLM + Embedding + VLM + ComfyUIMulti-model collaborative
Financial firmTime-series trading + RL + risk controlRTX A6000Transformer + RL model + FinBERTReal-time, low-latency
Creative / AI teamImage generation workflowRTX 5090ComfyUI + Stable Diffusion multi-modelBlackwell bandwidth advantage
Law firmContract OCR + semantic searchRTX 5090LLM + OCR + EmbeddingDocument privacy
Voice AI teamStable ASR serviceRTX Pro 5000Whisper / Wav2VecLow power, long-running
Enterprise AI teamKnowledge base Q&ARTX Pro 5000Embedding + LLM + RAGKnowledge stays on-prem
Sports data (BeSoccer)Multi-model parallel content genRTX Pro 5000Qwen3-8B-Q4 + Gemma-3-12B Q4Multiple models simultaneously
AI content platformText gen + image gen multi-taskRTX Pro 5000LLM + ComfyUI; Docker multi-containerIsolated deployment
AI application companyDialogue + TTS voice interactionRTX Pro 5000Qwen3.5-35B-AWQ-4bit + CosyVoice TRTGunicorn + Uvicorn
AI R&D teamText + image + voice multimodal hubRTX Pro 5000Qwen3.6-35B + ComfyUI + WhisperDocker multi-container
AI application teamVoice + vision + text multimodalRTX Pro 5000Qwen3.5-27B-VLM-INT4 (vLLM)Voice + vision input

Need help mapping your model and concurrency target to a GPU configuration?

Compare model fit, VRAM, framework requirements, and current server availability
Choose an AI GPU Server

Who This Is (and Isn't) For

Choosing the right GPU for self-hosted LLM workloads on dedicated infrastructure is not right for every team. Here is the honest breakdown.

Good Fit

  • Your application has sustained inference demand and you want to compare the predictable monthly cost of an LLM server with usage-based model API pricing
  • Workloads involve patient data, legal contracts, financial reports, or proprietary code — data privacy is non-negotiable
  • Need 24/7 always-on inference without cold-start latency or random resource preemption
  • Building RAG pipelines, AI Agents, or multi-turn applications
  • Have basic Linux ops: SSH access, able to deploy vLLM or Ollama

Not a Good Fit

  • Only need a few hours of GPU time for occasional experiments — an hourly or serverless service may be more economical
  • Need hyperscale distributed training across a large accelerator cluster — evaluate specialist cluster providers and network requirements

Frequently Asked Questions

What type of LLM server should I choose?
Choose a GPU VPS for smaller models, development, and moderate inference workloads. Choose a dedicated GPU server when you need stable resources, more VRAM, continuous operation, or higher concurrency. Consider a multi-GPU server when the model and KV cache exceed the capacity of one GPU.
What is the difference between a GPU VPS and a dedicated GPU server?
A GPU VPS can provide dedicated GPU access within a virtualized server, while a dedicated GPU server provides the entire physical machine. CPU, RAM, storage, network, operating system, and deployment time vary by plan, so check the selected configuration before ordering.
How do I calculate VRAM requirements for a self-hosted LLM?
Start with model weights, then add KV cache and operating headroom. A practical estimate is model weights plus KV cache plus 20–30% headroom. Context length, concurrency, framework, precision, and quantization can materially change the result.
Can I self-host an LLM on a GPU VPS?
Yes, when the model and runtime fit the available GPU memory and server resources. Smaller quantized models may fit on 16–24 GB GPUs, while larger models, long contexts, and higher concurrency may require 48–96 GB or multi-GPU configurations.
Should I use Ollama or vLLM?
Ollama is convenient for development, evaluation, and simpler single-user deployments. vLLM is generally better suited to production APIs that need continuous batching and higher concurrency. Validate compatibility with the exact model and version you plan to run.
Should I use one large-VRAM GPU or multiple GPUs?
Prefer one GPU when the model and required KV cache fit comfortably in memory. Multi-GPU tensor parallelism is useful when the workload exceeds single-card capacity, but it adds communication overhead and deployment complexity.
How do I deploy Hugging Face or private models?
Install the required NVIDIA driver, CUDA runtime, framework, and model dependencies, then download from Hugging Face or upload your private model through your approved transfer process. Deployment steps depend on the operating system, model license, and serving framework.
Is self-hosting cheaper than a model API?
It can be for sustained, high-utilization workloads, but not in every case. Compare monthly infrastructure cost, expected tokens, utilization, engineering time, storage, backups, monitoring, and support against current API pricing.
Where can I find current GPU server prices?
Prices, promotions, configurations, and availability change. Review the current GPU hosting pricing page before ordering.

Not sure which GPU fits your model? Talk to an expert.

Compare model fit, VRAM requirements, server type, and current availability
Ask About Your LLM Configuration

Choose a GPU for Your Self-Hosted LLM

Use your model size, quantization, context length, and concurrency target to narrow the configuration before ordering.

Dedicated GPU options LLM server and bare metal options vLLM and Ollama deployment paths