Can the RTX 5090 Run FLUX.1?
Yes — and the 5090 is currently the best single GPU we host for FLUX.1. Native Blackwell FP8 support drops a 1024×1024 to ~6 seconds while leaving room for ControlNets and LoRAs.
YesThe RTX 5090 (32 GB) runs FLUX.1 dev at FP16 with 6 GB of VRAM headroom for KV cache and concurrent batching.
Detailed Breakdown
FLUX.1 dev (12B parameters) at FP16 needs ~24 GB of VRAM. The RTX 5090 has 32 GB and Blackwell’s hardware FP8 path. That combination gives you:
- FLUX.1 dev FP16 — fits with 8 GB headroom. Great for ControlNet stacks.
- FLUX.1 dev FP8 (Blackwell native) — ~12 GB. Generates 1024×1024 in ~6 seconds.
- FLUX.1 schnell — same 12B but distilled for 4-step generation. Renders in ~2 seconds.
- Multiple LoRAs hot-loaded — 8-10 LoRAs can stay resident in VRAM for fast switching.
- Batch 4 simultaneous — 4× 1024×1024 in ~10 seconds at FP8.
Comparison: 5090 vs 4090 vs 3090 for FLUX.1
- RTX 3090 — 24 GB, no FP8 hardware. Renders 1024×1024 in ~14 s. Workable but slow.
- RTX 4090 — 24 GB, software FP8. Renders in ~9 s.
- RTX 5090 — 32 GB, hardware FP8. Renders in ~6 s.
For batch image-generation farms, the 5090 plus ComfyUI workflows is the best price-per-image of any card we rent.
Frequently Asked Questions
The questions buyers actually ask before committing to a GPU server.
Does FLUX.1 dev need FP16?
No — FP8 is essentially free quality-wise on the 5090’s native FP8 path. Most production deployments run FP8.
Can I train a FLUX.1 LoRA on a 5090?
Yes but tight. 32 GB is enough for LoRA training at rank 16-32. For comfortable LoRA training use a 6000 Pro.
FLUX.1 pro — can I host it?
No — pro is closed-source / API only. Self-hosting is dev or schnell.
Related Pages
Pages our visitors typically read next.
Ready to deploy?
Same-day deployment on in-stock GPUs. Talk to a specialist who actually understands your workload.