Table of Contents
April 2026 Image Generation Benchmarks
Image generation performance has improved since our last benchmark round thanks to optimised inference pipelines and new model architectures. This April 2026 update covers the latest generation speeds, VRAM requirements, and batch throughput for the most popular models on dedicated GPU servers.
All tests used the default recommended sampling schedules with FP16 precision. ComfyUI was used as the generation frontend with PyTorch 2.x backends.
Single Image Generation Speed
Time to generate a single 1024×1024 image:
| Model | Steps | RTX 3090 | RTX 5090 | RTX 5090 | RTX 6000 Pro |
|---|---|---|---|---|---|
| FLUX.1 [dev] | 28 | 14.5 s | 8.2 s | 5.8 s | 9.5 s |
| FLUX.1 [schnell] | 4 | 2.8 s | 1.4 s | 0.9 s | 1.8 s |
| SD 3.5 Large | 28 | 11.2 s | 6.5 s | 4.5 s | 7.8 s |
| SDXL (1-step turbo) | 1 | 0.7 s | 0.4 s | 0.25 s | 0.5 s |
The RTX 5090 delivers a 30-35% speed improvement over the RTX 5090 across all models, driven by higher memory bandwidth and improved tensor cores.
Batch Throughput by GPU
Images per hour at 1024×1024 with continuous batch processing:
| Model | RTX 3090 | RTX 5090 | RTX 5090 | RTX 6000 Pro (batch=4) |
|---|---|---|---|---|
| FLUX.1 [dev] | 248/hr | 440/hr | 620/hr | 580/hr |
| FLUX.1 [schnell] | 1,285/hr | 2,570/hr | 4,000/hr | 3,200/hr |
| SD 3.5 Large | 321/hr | 554/hr | 800/hr | 690/hr |
The RTX 6000 Pro’s 48 GB VRAM enables batch sizes of 4 for FLUX.1 models, boosting hourly throughput close to the RTX 5090 despite lower per-image speed. For volume image generation, VRAM enables batching which drives throughput. See the cost per 1000 images analysis for the economic breakdown.
VRAM Usage at Different Resolutions
| Model | 512×512 | 1024×1024 | 2048×2048 |
|---|---|---|---|
| FLUX.1 [dev] | 16 GB | 22 GB | OOM on 24 GB |
| SD 3.5 Large | 12 GB | 18 GB | 38 GB |
| SDXL | 6 GB | 8 GB | 18 GB |
FLUX.1 requires nearly the full 24 GB of an RTX 5090 at 1024×1024, leaving no room for batching. For high-resolution or batch generation with FLUX.1, the RTX 6000 Pro or RTX 5090 (32 GB) provide necessary headroom.
Cost per Image by GPU
Based on GigaGPU dedicated hosting rates, running 24/7:
| GPU | Monthly Cost | FLUX.1 [dev] per image | FLUX.1 [schnell] per image |
|---|---|---|---|
| RTX 3090 | ~$175 | $0.0010 | $0.00019 |
| RTX 5090 | ~$250 | $0.00079 | $0.00013 |
| RTX 5090 | ~$425 | $0.00095 | $0.00015 |
The RTX 5090 delivers the lowest cost per image overall. Compare these rates to commercial APIs using the GPU vs API cost comparison tool.
Generate Images at Scale on Dedicated Hardware
No per-image fees. Run FLUX.1, Stable Diffusion, or any model with full control over your generation pipeline.
Browse GPU ServersHardware Recommendations
For FLUX.1 [dev] quality at production scale, the RTX 5090 is the price-performance winner. For maximum throughput, the RTX 6000 Pro with batch generation matches higher-end cards. For real-time or interactive generation, FLUX.1 [schnell] on an RTX 5090 produces images in under 1.5 seconds. Visit the best image generation models guide for model selection and the GPU comparisons for broader hardware analysis.
For multi-GPU batch generation, two RTX 5090s double throughput linearly. The benchmarks section tracks these numbers as new models and optimisations are released.