Table of Contents
Test Methodology
We benchmarked three Stable Diffusion model generations across six GPUs on GigaGPU dedicated servers. Tests used ComfyUI with FP16 precision and xformers attention. All images are batch size 1 (sequential generation). For interactive exploration, visit our GPU comparisons tool.
SD 1.5 Benchmarks (512×512)
SD 1.5 at 512×512 resolution, 30 steps, Euler a sampler. This remains the most popular configuration for fine-tuned models and ControlNet workflows.
| GPU | VRAM | Sec/Image | Images/hr | it/s | Server $/hr |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 0.8 | 4,500 | 37.5 | $1.80 |
| RTX 5080 | 16 GB | 1.5 | 2,400 | 20.0 | $0.85 |
| RTX 3090 | 24 GB | 1.7 | 2,118 | 17.6 | $0.45 |
| RTX 4060 Ti | 16 GB | 2.4 | 1,500 | 12.5 | $0.35 |
| RTX 4060 | 8 GB | 3.5 | 1,029 | 8.6 | $0.20 |
| RTX 3050 | 8 GB | 7.2 | 500 | 4.2 | $0.10 |
The RTX 5090 generates SD 1.5 images in under a second. The RTX 3090 produces over 2,000 images per hour, making it a cost-effective workhorse for batch generation. For a broader GPU comparison, see our best GPU for Stable Diffusion guide.
SDXL Benchmarks (1024×1024)
SDXL at 1024×1024, 30 steps, Euler a. SDXL requires significantly more VRAM and compute than SD 1.5.
| GPU | VRAM | Sec/Image | Images/hr | it/s | Server $/hr |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 3.2 | 1,125 | 9.4 | $1.80 |
| RTX 5080 | 16 GB | 6.4 | 563 | 4.7 | $0.85 |
| RTX 3090 | 24 GB | 7.1 | 507 | 4.2 | $0.45 |
| RTX 4060 Ti | 16 GB | 10.2 | 353 | 2.9 | $0.35 |
| RTX 4060 | 8 GB | OOM | — | — | $0.20 |
| RTX 3050 | 8 GB | OOM | — | — | $0.10 |
SDXL requires 10+ GB VRAM, excluding 8 GB cards. The RTX 3090’s 24 GB VRAM handles SDXL with room for ControlNet and other extensions. The RTX 4060 Ti fits SDXL but at slower speeds.
Flux.1 dev Benchmarks
Flux.1 dev at 1024×1024, 28 steps, Euler sampler. Flux uses a transformer-based architecture that is more compute-intensive than SDXL.
| GPU | VRAM | Sec/Image | Images/hr | it/s | Server $/hr |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 8.5 | 424 | 3.3 | $1.80 |
| RTX 5080 | 16 GB | 18.2 | 198 | 1.5 | $0.85 |
| RTX 3090 | 24 GB | 19.8 | 182 | 1.4 | $0.45 |
| RTX 4060 Ti | 16 GB | OOM* | — | — | $0.35 |
| RTX 4060 | 8 GB | OOM | — | — | $0.20 |
| RTX 3050 | 8 GB | OOM | — | — | $0.10 |
*Flux.1 dev can fit on 16 GB with aggressive offloading but runs 3-4x slower. Not practical for production.
Flux requires 24+ GB VRAM for production-quality generation. The RTX 5090’s 32 GB is ideal for Flux with headroom for extensions. For Flux deployment, see our Flux.1 hosting guide.
Cost per Image by GPU and Model
| GPU | SD 1.5 ($/image) | SDXL ($/image) | Flux.1 ($/image) |
|---|---|---|---|
| RTX 5090 | $0.0004 | $0.0016 | $0.0043 |
| RTX 5080 | $0.0004 | $0.0015 | $0.0043 |
| RTX 3090 | $0.0002 | $0.0009 | $0.0025 |
| RTX 4060 Ti | $0.0002 | $0.0010 | OOM |
| RTX 4060 | $0.0002 | OOM | OOM |
| RTX 3050 | $0.0002 | OOM | OOM |
Self-hosted Stable Diffusion costs a fraction of a penny per image on all GPUs. Even Flux.1, the most expensive model, costs only $0.0025 per image on the RTX 3090. Compare with cloud image generation APIs charging $0.02-$0.10 per image. See our cost analysis and cheapest GPU for AI inference for broader cost comparisons.
ComfyUI vs A1111 Speed Impact
The UI choice affects generation speed. ComfyUI’s optimised execution is consistently 33-41% faster than A1111. See our full ComfyUI vs A1111 comparison for detailed analysis.
| GPU | SD 1.5 ComfyUI (sec) | SD 1.5 A1111 (sec) | SDXL ComfyUI (sec) | SDXL A1111 (sec) |
|---|---|---|---|---|
| RTX 5090 | 0.8 | 1.1 | 3.2 | 4.5 |
| RTX 3090 | 1.7 | 2.3 | 7.1 | 9.6 |
| RTX 5080 | 1.5 | 2.1 | 6.4 | 8.7 |
| RTX 4060 Ti | 2.4 | 3.2 | 10.2 | 13.8 |
All benchmarks in this article use ComfyUI. If you use A1111, expect roughly 35% slower generation times. For video generation benchmarks, see our AI video generation GPU guide.
GPU Recommendations for Stable Diffusion
Best overall: RTX 3090. Handles SD 1.5, SDXL, and Flux.1 on a single 24 GB card. At $0.0009 per SDXL image and 507 images/hr, it is the best value for production image generation. The go-to GPU for Stable Diffusion hosting.
Best for speed: RTX 5090. Generates 4,500 SD 1.5 images per hour and 1,125 SDXL images per hour. The 32 GB VRAM handles Flux with complex workflows. Best for high-volume or low-latency applications.
Best for SDXL on a budget: RTX 4060 Ti. The cheapest GPU that runs SDXL at $0.35/hr. Slower than the 3090 but produces 353 SDXL images per hour, which is sufficient for moderate workloads.
Best for SD 1.5 budget: RTX 4060. At $0.20/hr and 1,029 images per hour, the RTX 4060 is the cheapest path to SD 1.5 generation. Cannot run SDXL or Flux.
For related benchmarks, see our LLaMA 3 8B benchmark, Coqui TTS latency benchmark, deep learning training GPUs, and running multiple AI models. For the complete image hosting stack, see image generator hosting.
Generate Images on Dedicated GPU Servers
GigaGPU provides bare-metal GPU servers with ComfyUI and Stable Diffusion pre-installed. No per-image fees, no shared resources, just raw GPU power for image generation.
Browse GPU Servers