Table of Contents
XTTS-v2 Benchmark Overview
Coqui XTTS-v2 is a multilingual text-to-speech model with voice cloning capabilities, able to replicate a speaker’s voice from a short audio reference. This makes it uniquely valuable for personalised TTS applications. Running it on a dedicated GPU server is essential for consistent latency in production.
We measured end-to-end latency on GigaGPU servers using a 15-word English sentence with a 6-second voice reference clip. XTTS-v2 requires approximately 2.5 GB of VRAM. For comparisons with other TTS models, see our TTS latency benchmarks page.
Latency Results by GPU
| GPU | VRAM | XTTS-v2 Latency (ms) | Notes |
|---|---|---|---|
| RTX 3050 | 6 GB | 2,400ms | Functional but slow |
| RTX 4060 | 8 GB | 1,450ms | Acceptable for non-real-time |
| RTX 4060 Ti | 16 GB | 1,050ms | Just above 1 second |
| RTX 3090 | 24 GB | 720ms | Good for interactive use |
| RTX 5080 | 16 GB | 480ms | Sub-500ms generation |
| RTX 5090 | 32 GB | 310ms | Best latency tested |
XTTS-v2 sits between Kokoro (fast, no voice cloning) and Bark (slow, very expressive) in terms of latency. The RTX 5090 at 310ms provides responsive voice-cloned speech suitable for interactive applications.
Streaming vs Full Generation
XTTS-v2 supports streaming mode, which begins audio output before the full sentence is generated. Below we compare time-to-first-audio (streaming) vs full generation latency.
| GPU | Full Generation (ms) | Time to First Audio (ms) |
|---|---|---|
| RTX 3090 | 720 | 280 |
| RTX 5080 | 480 | 185 |
| RTX 5090 | 310 | 120 |
Streaming mode reduces perceived latency significantly. The RTX 5090 delivers first audio in 120ms, which feels nearly instantaneous to users.
Cost Efficiency Analysis
| GPU | Latency (ms) | Approx. Monthly Cost | Gen/s per Pound |
|---|---|---|---|
| RTX 3050 | 2,400 | ~£45 | 0.0093 |
| RTX 4060 | 1,450 | ~£60 | 0.0115 |
| RTX 4060 Ti | 1,050 | ~£75 | 0.0127 |
| RTX 3090 | 720 | ~£110 | 0.0126 |
| RTX 5080 | 480 | ~£160 | 0.0130 |
| RTX 5090 | 310 | ~£250 | 0.0129 |
The RTX 5080 leads on cost efficiency by a narrow margin. For the best GPU for TTS with voice cloning, it is the recommended choice.
GPU Recommendations
- Budget: RTX 4060 Ti — 1.05 seconds per sentence is workable for audiobook and content creation.
- Best value: RTX 5080 — sub-500ms full generation with streaming at 185ms to first audio.
- Lowest latency: RTX 5090 — 310ms full, 120ms streaming for real-time voice cloning.
- Development: RTX 4060 — affordable for prototyping voice cloning applications.
For faster TTS without voice cloning, see the Kokoro TTS benchmark. For maximum expressiveness, check the Bark TTS results. Browse all benchmarks in the Benchmarks category.
Conclusion
XTTS-v2 is the go-to model for voice cloning TTS, offering personalised speech synthesis in multiple languages. With streaming mode on the RTX 5080 or RTX 5090, it achieves latency low enough for interactive voice applications while maintaining speaker similarity from just a short reference clip.
Voice Cloning TTS on Dedicated GPU Servers
Deploy XTTS-v2 with low latency on bare-metal GPU hardware. UK hosting with full root access.
Browse GPU Servers