RTX 3050 - Order Now
Home / Blog / Benchmarks / Mistral 7B Tokens/sec by GPU (Full Benchmark)
Benchmarks

Mistral 7B Tokens/sec by GPU (Full Benchmark)

Full tokens/sec benchmarks for Mistral 7B across 6 GPUs at multiple batch sizes, precisions, and inference engines. Compare throughput, latency, and cost per million tokens on dedicated servers.

Test Methodology

We benchmarked Mistral 7B v0.3 across six GPUs on GigaGPU dedicated servers. Tests used vLLM with continuous batching. Each run generates 512 output tokens from a 256-token prompt. Mistral 7B uses grouped-query attention (GQA) and sliding window attention, which gives it a throughput advantage over similarly-sized models. For interactive data, visit our tokens/sec benchmark tool.

Single-User Throughput (Batch Size 1)

GPUVRAMFP16 tok/sTTFT (ms)Latency (512 tok)Server $/hr
RTX 509032 GB148283.5 sec$1.80
RTX 508016 GB92425.6 sec$0.85
RTX 309024 GB68507.5 sec$0.45
RTX 4060 Ti16 GB52659.8 sec$0.35
RTX 40608 GB388813.5 sec$0.20
RTX 30508 GB2017025.6 sec$0.10

Mistral 7B is the fastest model in the 7-8B parameter class due to its 7B parameter count (vs 8B for LLaMA 3) and efficient GQA attention. It is 10% faster than LLaMA 3 8B and 15% faster than DeepSeek-R1 8B on the same hardware. For comparison, see our LLaMA 3 8B benchmark and DeepSeek benchmark.

Concurrent User Throughput

GPU1 User (tok/s)4 Users (total tok/s)8 Users (total tok/s)16 Users (total tok/s)
RTX 5090148450730985
RTX 508092278445560
RTX 309068205315395
RTX 4060 Ti52155238292
RTX 406038105150172
RTX 305020466068

The RTX 3090 at 8 concurrent users delivers 315 total tok/s for Mistral 7B, roughly 40 tok/s per user on average. This is fast enough for interactive chatbot deployments. For engine options, see vLLM vs TGI vs Ollama.

Quantisation Impact on Throughput

Mistral 7B on the RTX 3090 at various precisions:

PrecisionVRAM Usedtok/s (bs=1)tok/s (bs=8)Quality Impact
FP16~14 GB68315Baseline
AWQ 4-bit~5 GB86425Minimal
GPTQ 4-bit~5 GB82408Minimal
FP8~8 GB78375Negligible

Mistral 7B’s smaller parameter count means even FP16 fits on 16 GB GPUs. AWQ 4-bit quantisation makes it viable on 8 GB cards with excellent performance. For quantisation details, see our GPTQ vs AWQ vs GGUF guide.

Cost per Million Tokens

GPUFP16 Cost/1M TokensAWQ 4-bit Cost/1M TokensMistral API Equiv.
RTX 5090$3.38$2.68$2.00
RTX 5080$2.57$2.02$2.00
RTX 3090$1.84$1.45$2.00
RTX 4060 Ti$1.87$1.37$2.00
RTX 4060$1.46$0.98$2.00
RTX 3050$1.39$0.90$2.00

Self-hosted Mistral 7B beats API pricing on all GPUs at FP16 or better. The RTX 3090 at FP16 costs $1.84 per million tokens, 8% below API pricing. With AWQ quantisation, even budget GPUs deliver significant savings. See our GPU vs API cost breakdown and LLM cost calculator.

Mistral 7B vs LLaMA 3 8B vs DeepSeek-R1 8B

Side-by-side speed comparison on the RTX 3090 (FP16, bs=1):

Modeltok/sVRAMCost/1M TokensStrength
Mistral 7B68~14 GB$1.84Fastest, efficient attention
LLaMA 3 8B62~16 GB$2.02Broadest ecosystem
DeepSeek-R1 8B59~16 GB$2.12Best reasoning/coding

Mistral 7B is the fastest and cheapest to serve. LLaMA 3 8B has the broadest ecosystem support. DeepSeek-R1 8B excels at reasoning tasks. For RAG applications where speed matters, Mistral 7B is an excellent choice. See best GPU for RAG pipelines and best GPU for LangChain.

GPU Recommendations for Mistral 7B

Best overall: RTX 3090. At 68 tok/s (FP16) and $1.84 per million tokens, the RTX 3090 is the sweet spot for Mistral 7B hosting. The 14 GB model fits comfortably in 24 GB VRAM with headroom for KV cache and concurrent users.

Best for production: RTX 5090. At 148 tok/s and 32 GB VRAM, the 5090 handles high-concurrency Mistral serving with sub-4-second responses. Best for production chatbots and API endpoints.

Best budget: RTX 4060. Mistral 7B fits on 8 GB with AWQ 4-bit quantisation at 86 tok/s on the 3090 tier or 38 tok/s at FP16 on the 4060. At $0.98 per million tokens (AWQ), it is cheaper than any API.

Best mid-range: RTX 5080. At 92 tok/s the 5080 delivers fast single-user responses. The 16 GB VRAM fits FP16 Mistral 7B with room for an embedding model alongside for RAG.

Related benchmarks: LLaMA 3 8B benchmark, DeepSeek benchmark, Stable Diffusion benchmark, and best GPU for LLM inference.

Run Mistral 7B on Dedicated GPU Servers

GigaGPU provides bare-metal GPU servers optimised for Mistral inference with vLLM and Ollama. No per-token fees, no rate limits, full data privacy.

Browse GPU Servers

Need a Dedicated GPU Server?

Deploy from RTX 3050 to RTX 5090. Full root access, NVMe storage, 1Gbps — UK datacenter.

Browse GPU Servers

gigagpu

We benchmark, deploy, and optimise GPU infrastructure for AI workloads. All data in our guides comes from real-world testing on our UK-based dedicated GPU servers.

Ready to deploy your AI workload?

Dedicated GPU servers from our UK datacenter. NVMe storage, 1Gbps networking, full root access.

Browse GPU Servers Contact Sales

Have a question? Need help?