RTX 3050 - Order Now
Home / Blog / Benchmarks / XTTS-v2 Latency by GPU
Benchmarks

XTTS-v2 Latency by GPU

Benchmark data for Coqui XTTS-v2 text-to-speech latency across six GPUs with voice cloning performance and cost analysis for dedicated GPU hosting.

XTTS-v2 Benchmark Overview

Coqui XTTS-v2 is a multilingual text-to-speech model with voice cloning capabilities, able to replicate a speaker’s voice from a short audio reference. This makes it uniquely valuable for personalised TTS applications. Running it on a dedicated GPU server is essential for consistent latency in production.

We measured end-to-end latency on GigaGPU servers using a 15-word English sentence with a 6-second voice reference clip. XTTS-v2 requires approximately 2.5 GB of VRAM. For comparisons with other TTS models, see our TTS latency benchmarks page.

Latency Results by GPU

GPUVRAMXTTS-v2 Latency (ms)Notes
RTX 30506 GB2,400msFunctional but slow
RTX 40608 GB1,450msAcceptable for non-real-time
RTX 4060 Ti16 GB1,050msJust above 1 second
RTX 309024 GB720msGood for interactive use
RTX 508016 GB480msSub-500ms generation
RTX 509032 GB310msBest latency tested

XTTS-v2 sits between Kokoro (fast, no voice cloning) and Bark (slow, very expressive) in terms of latency. The RTX 5090 at 310ms provides responsive voice-cloned speech suitable for interactive applications.

Streaming vs Full Generation

XTTS-v2 supports streaming mode, which begins audio output before the full sentence is generated. Below we compare time-to-first-audio (streaming) vs full generation latency.

GPUFull Generation (ms)Time to First Audio (ms)
RTX 3090720280
RTX 5080480185
RTX 5090310120

Streaming mode reduces perceived latency significantly. The RTX 5090 delivers first audio in 120ms, which feels nearly instantaneous to users.

Cost Efficiency Analysis

GPULatency (ms)Approx. Monthly CostGen/s per Pound
RTX 30502,400~£450.0093
RTX 40601,450~£600.0115
RTX 4060 Ti1,050~£750.0127
RTX 3090720~£1100.0126
RTX 5080480~£1600.0130
RTX 5090310~£2500.0129

The RTX 5080 leads on cost efficiency by a narrow margin. For the best GPU for TTS with voice cloning, it is the recommended choice.

GPU Recommendations

  • Budget: RTX 4060 Ti — 1.05 seconds per sentence is workable for audiobook and content creation.
  • Best value: RTX 5080 — sub-500ms full generation with streaming at 185ms to first audio.
  • Lowest latency: RTX 5090 — 310ms full, 120ms streaming for real-time voice cloning.
  • Development: RTX 4060 — affordable for prototyping voice cloning applications.

For faster TTS without voice cloning, see the Kokoro TTS benchmark. For maximum expressiveness, check the Bark TTS results. Browse all benchmarks in the Benchmarks category.

Conclusion

XTTS-v2 is the go-to model for voice cloning TTS, offering personalised speech synthesis in multiple languages. With streaming mode on the RTX 5080 or RTX 5090, it achieves latency low enough for interactive voice applications while maintaining speaker similarity from just a short reference clip.

Voice Cloning TTS on Dedicated GPU Servers

Deploy XTTS-v2 with low latency on bare-metal GPU hardware. UK hosting with full root access.

Browse GPU Servers

Need a Dedicated GPU Server?

Deploy from RTX 3050 to RTX 5090. Full root access, NVMe storage, 1Gbps — UK datacenter.

Browse GPU Servers

gigagpu

We benchmark, deploy, and optimise GPU infrastructure for AI workloads. All data in our guides comes from real-world testing on our UK-based dedicated GPU servers.

Ready to deploy your AI workload?

Dedicated GPU servers from our UK datacenter. NVMe storage, 1Gbps networking, full root access.

Browse GPU Servers Contact Sales

Have a question? Need help?