RTX 3050 - Order Now
Home / Blog / Cost & Pricing / Cost Per 1M Tokens for DeepSeek Self-Hosted: V2 16B Across Every GPU
Cost & Pricing

Cost Per 1M Tokens for DeepSeek Self-Hosted: V2 16B Across Every GPU

Detailed cost-per-million-tokens for DeepSeek-V2 16B Lite on every dedicated GPU we host, plus the multi-GPU configurations needed for V2 236B.

DeepSeek-V2 16B Lite is a Mixture-of-Experts model with 16B total params and 2.4B active per token. The MoE architecture means it runs faster than a dense 16B model would suggest — but it occupies the full VRAM of the larger weight footprint. This page lists exact cost-per-token numbers across every GPU we rent.

TL;DR

RTX 5090 + FP8 at £0.24/1M tokens is the cost leader for DeepSeek-V2 16B Lite. RTX 4090 at AWQ-INT4 is the cheapest path on Ada hardware. For DeepSeek-V2 236B, only multi-GPU configs work; expect £1.50–2.50/1M tokens self-hosted.

Methodology

Standard methodology — 60% utilisation, vLLM 0.6.3 with continuous batching, 50-thread Locust driver, Ubuntu 22.04. Models from deepseek-ai/DeepSeek-V2-Lite-Chat (BF16) and the AWQ-INT4 community port.

DeepSeek-V2 16B Lite by GPU

GPUMonthlyPrecisiontok/sCost per 1M
RTX 5060 Ti 16 GB£119AWQ-INT4320£0.34
RTX 5080£189AWQ-INT4520£0.28
RTX 3090£159AWQ-INT4410£0.28
RTX 4090£289AWQ-INT4590£0.30
RTX 5090£399AWQ-INT4780£0.30
RTX 5090£399FP8950£0.24
RTX 6000 Pro£899FP161,420£0.50

RTX 5090 with FP8 is the cost leader. RTX 3090 is the cheapest dedicated server with sufficient VRAM for AWQ-INT4.

DeepSeek-V2 236B (multi-GPU only)

V2 236B has 21B active params per token but the full 236B has to be VRAM-resident. ~470 GB FP16 weights makes it multi-GPU territory exclusively.

ConfigMonthlyPrecisiontok/sCost per 1M
4× RTX 5090£1,800AWQ-INT4 (TP=4)180£1.93
2× RTX 6000 Pro£2,200FP8 (TP=2)210£2.02
4× A100 80 GBPOAFP16 (TP=4)270POA
8× H100POA (£15K+)FP8 (TP=8)480~£3.50

For V2 236B the official DeepSeek API at £0.18/1M is hard to beat unless you need data residency. We size custom builds on request.

DeepSeek-Coder variants

VariantParamsVRAM (INT4)tok/s on 5090Cost per 1M
DeepSeek-Coder-V2-Lite16B (2.4B active)10 GB780£0.30
DeepSeek-Coder 6.7B6.7B4 GB~1,800£0.13
DeepSeek-Coder 33B33B18 GB~310£0.74
DeepSeek-Coder-V2 236B236BMulti-GPUcluster£1.50+

Cost leaders

  • V2 16B Lite cost leader: RTX 5090 + FP8 at £0.27/1M.
  • V2 16B Lite budget option: RTX 3090 / 4090 at AWQ-INT4 — slightly higher cost-per-token but cheaper monthly.
  • DeepSeek-Coder 6.7B: RTX 5090 at FP16 — £0.14/1M tokens, fastest cost leader of the lot.
  • V2 236B: use the official DeepSeek API unless you have a residency requirement.

Bottom line

For typical DeepSeek deployments, the RTX 5090 at FP8 is the right host. Drop to a 4090 or 3090 if FP8 is not available. For V2 236B and V3 671B, hosted APIs win on cost — see cost: run DeepSeek vs API for the break-even math.

Need a Dedicated GPU Server?

Deploy from RTX 3050 to RTX 5090. Full root access, NVMe storage, 1Gbps — UK datacenter.

Browse GPU Servers

gigagpu

We benchmark, deploy, and optimise GPU infrastructure for AI workloads. All data in our guides comes from real-world testing on our UK-based dedicated GPU servers.

Ready to deploy your AI workload?

Dedicated GPU servers from our UK datacenter. NVMe storage, 1Gbps networking, full root access.

Browse GPU Servers Contact Sales

Have a question? Need help?