RTX 3050 - Order Now
Home / Blog / Benchmarks / Image Generation Benchmark Update: April 2026
Benchmarks

Image Generation Benchmark Update: April 2026

Updated April 2026 benchmarks for AI image generation models across GPUs. Covers FLUX.1, Stable Diffusion 3.5, and SDXL generation speed, VRAM usage, and batch throughput.

April 2026 Image Generation Benchmarks

Image generation performance has improved since our last benchmark round thanks to optimised inference pipelines and new model architectures. This April 2026 update covers the latest generation speeds, VRAM requirements, and batch throughput for the most popular models on dedicated GPU servers.

All tests used the default recommended sampling schedules with FP16 precision. ComfyUI was used as the generation frontend with PyTorch 2.x backends.

Single Image Generation Speed

Time to generate a single 1024×1024 image:

Model Steps RTX 3090 RTX 5090 RTX 5090 RTX 6000 Pro
FLUX.1 [dev] 28 14.5 s 8.2 s 5.8 s 9.5 s
FLUX.1 [schnell] 4 2.8 s 1.4 s 0.9 s 1.8 s
SD 3.5 Large 28 11.2 s 6.5 s 4.5 s 7.8 s
SDXL (1-step turbo) 1 0.7 s 0.4 s 0.25 s 0.5 s

The RTX 5090 delivers a 30-35% speed improvement over the RTX 5090 across all models, driven by higher memory bandwidth and improved tensor cores.

Batch Throughput by GPU

Images per hour at 1024×1024 with continuous batch processing:

Model RTX 3090 RTX 5090 RTX 5090 RTX 6000 Pro (batch=4)
FLUX.1 [dev] 248/hr 440/hr 620/hr 580/hr
FLUX.1 [schnell] 1,285/hr 2,570/hr 4,000/hr 3,200/hr
SD 3.5 Large 321/hr 554/hr 800/hr 690/hr

The RTX 6000 Pro’s 48 GB VRAM enables batch sizes of 4 for FLUX.1 models, boosting hourly throughput close to the RTX 5090 despite lower per-image speed. For volume image generation, VRAM enables batching which drives throughput. See the cost per 1000 images analysis for the economic breakdown.

VRAM Usage at Different Resolutions

Model 512×512 1024×1024 2048×2048
FLUX.1 [dev] 16 GB 22 GB OOM on 24 GB
SD 3.5 Large 12 GB 18 GB 38 GB
SDXL 6 GB 8 GB 18 GB

FLUX.1 requires nearly the full 24 GB of an RTX 5090 at 1024×1024, leaving no room for batching. For high-resolution or batch generation with FLUX.1, the RTX 6000 Pro or RTX 5090 (32 GB) provide necessary headroom.

Cost per Image by GPU

Based on GigaGPU dedicated hosting rates, running 24/7:

GPU Monthly Cost FLUX.1 [dev] per image FLUX.1 [schnell] per image
RTX 3090 ~$175 $0.0010 $0.00019
RTX 5090 ~$250 $0.00079 $0.00013
RTX 5090 ~$425 $0.00095 $0.00015

The RTX 5090 delivers the lowest cost per image overall. Compare these rates to commercial APIs using the GPU vs API cost comparison tool.

Generate Images at Scale on Dedicated Hardware

No per-image fees. Run FLUX.1, Stable Diffusion, or any model with full control over your generation pipeline.

Browse GPU Servers

Hardware Recommendations

For FLUX.1 [dev] quality at production scale, the RTX 5090 is the price-performance winner. For maximum throughput, the RTX 6000 Pro with batch generation matches higher-end cards. For real-time or interactive generation, FLUX.1 [schnell] on an RTX 5090 produces images in under 1.5 seconds. Visit the best image generation models guide for model selection and the GPU comparisons for broader hardware analysis.

For multi-GPU batch generation, two RTX 5090s double throughput linearly. The benchmarks section tracks these numbers as new models and optimisations are released.

Need a Dedicated GPU Server?

Deploy from RTX 3050 to RTX 5090. Full root access, NVMe storage, 1Gbps — UK datacenter.

Browse GPU Servers

gigagpu

We benchmark, deploy, and optimise GPU infrastructure for AI workloads. All data in our guides comes from real-world testing on our UK-based dedicated GPU servers.

Ready to deploy your AI workload?

Dedicated GPU servers from our UK datacenter. NVMe storage, 1Gbps networking, full root access.

Browse GPU Servers Contact Sales

Have a question? Need help?