Best GPU for FLUX.1 Hosting
Black Forest Labs’ FLUX.1 is the strongest open-weight image model. It’s also memory-hungry: 24 GB at FP16 for the 12B variant. The right GPU choice depends on whether you can use the FP8 path.
The short answer: the RTX 5090 is the best GPU for self-hosting FLUX.1 dev on a dedicated server. It has the right VRAM (32 GB) for the model, modern tensor cores, and the best cost-per-token in our catalogue for this workload.
Ranking — Best to Worst for This Workload
From best to worst for this specific workload, with the reason in plain English.
RTX 5090 Top Pick
32 GB plus Blackwell native FP8 = ~6s per 1024×1024. Best cost-per-image.
32 GB · Blackwell · from £359/mo
RTX 6000 Pro 96 GB Best Quality
96 GB lets you keep multiple LoRAs hot, batch 4+, run FLUX.1 plus SDXL on one card.
96 GB · Blackwell · from £1099/mo
RTX 3090 Budget Pick
24 GB just fits FLUX.1 dev FP16. ~14s per image — workable for low-volume.
24 GB · Ampere · from £179/mo
RTX 4090 Ada Pick
24 GB Ada — similar to 3090 but ~30% faster. Premium pricing.
24 GB · Ada Lovelace · from £279/mo
RTX 5080 FP8 Tight
16 GB needs FP8 quantisation. Fine for FLUX.1 schnell, tight for dev.
16 GB · Blackwell · from £189/mo
Background & Sizing
FLUX.1 from Black Forest Labs has overtaken Stable Diffusion as the de-facto open image model. The dev variant (12B params) is the production target; schnell is the fast variant for iteration; pro is closed and not available for self-hosting.
FP16 vs FP8 vs INT4 for FLUX.1
FLUX.1 is one of the few image models where the precision ladder makes a meaningful difference:
- FP16 (~24 GB): full quality, slowest. The reference deployment.
- FP8 (~12 GB): <1% quality drop on standard prompts, 2× faster on Blackwell. The default for new deployments.
- INT4 (~7 GB): noticeable quality loss on prompt adherence; useful for speed-sensitive iteration.
For a complete deployment guide see our FLUX.1 deployment guide and the broader image generator hosting page.
Frequently Asked Questions
The questions buyers actually ask before committing to a GPU server.
FLUX.1 dev or schnell for production?
Dev for quality, schnell for iteration. Most teams ship both — schnell as a draft mode, dev for final renders.
Does FLUX.1 work with ComfyUI?
Yes — first-class support. Pre-installed on our image-generation server template.
LoRA training on FLUX.1?
Possible on 24 GB+ but tight. The RTX 6000 Pro 96 GB is comfortable for FLUX LoRA training.
How does it compare to SDXL?
FLUX.1 has dramatically better prompt adherence and text rendering. SDXL is faster on older hardware. Most new deployments pick FLUX.
Related Pages
Pages our visitors typically read next.
Ready to deploy?
Same-day deployment on in-stock GPUs. Talk to a specialist who actually understands your workload.