RTX 3050 - Order Now
Whisper Large-v3 · faster-whisper · Real-Time

Whisper Hosting — Self-Hosted Speech-to-Text

Deploy OpenAI Whisper Large-v3 (and faster-whisper, distil-whisper, Whisper-Streaming) on a dedicated GPU server. Transcribe at 3–10× real-time, keep all audio on your hardware, and replace the OpenAI Audio API at a fraction of the cost.

Whisper Large-v3 included 3–10× real-time Audio never leaves your server From £69/mo
Large-v3
Highest quality
3–10×
Real-time factor
£69
/mo entry
0
Per-minute fees

What’s Pre-Installed

Everything you need to run production AI workloads on dedicated hardware in the UK.

Whisper Large-v3

OpenAI’s flagship speech-to-text model — 99 languages, robust to noise, punctuation and casing built in. Pre-downloaded on every server.

faster-whisper (CTranslate2)

4× faster than the reference implementation, half the VRAM, no quality loss. The default backend on our Whisper image.

distil-whisper

Distilled small/medium variants for English-only workloads — 6× faster than Large-v3 with <1% WER regression.

Streaming endpoints

Whisper-Streaming for live captioning. Sub-second time-to-first-word latency on RTX 3090+.

OpenAI-compatible API

Drop-in replacement for /v1/audio/transcriptions. Existing SDKs (openai-python, Pipecat, LiveKit Agents) work unchanged.

VAD + diarisation

WhisperX, Silero VAD and pyannote diarisation pre-installed for multi-speaker workflows.

Whisper Performance by GPU

Real-time factor (RTF) measures how many seconds of audio the model transcribes per second of wall time. Higher is better; 1× = real-time.

ModelParamsFP16 VRAMINT4 VRAMRecommended
RTX 3050 6 GBWhisper Large-v3 FP16~3× real-timeFits comfortablyfrom £79/mo
RTX 3060 12 GBWhisper Large-v3 FP16~4× real-timeHeadroom for parallel streamsfrom £99/mo
RTX 4060 8 GBWhisper Large-v3 FP16~5× real-timeBest entry-tier RTFfrom £109/mo
RTX 3090 24 GBWhisper Large-v3 + LLM~6× real-timeRun Whisper + Llama 3 on same cardfrom £179/mo
RTX 5080 16 GBWhisper Large-v3 + 7B LLM~7× real-timeLowest single-stream latencyfrom £189/mo
RTX 5090 32 GBMulti-stream Whisper~9× real-time / stream, 16 streamsVoice agent farmfrom £399/mo

Common Whisper Deployments

Real customer workloads we run on this hardware every day.

Voice agent backend

Whisper for STT, an LLM for reasoning, a TTS for output — all on the same server. Sub-second time-to-first-audio for live conversations.

PipecatLiveKit AgentsTwilio Media Streams

Batch transcription

Podcast archives, call recordings, video courseware. Run Whisper at 8× real-time and clear a 1,000-hour backlog overnight.

PodcastsCoursewareCompliance archives

Meeting / call analytics

Live captioning + speaker diarisation + LLM summary. Pipecat or a custom pipeline running on a 5090 handles 16 concurrent calls.

Sales call analyticsMeeting recapCaptioning

Regulated transcription

Healthcare consultations, legal depositions, government interviews. Audio stays on your hardware — see private AI hosting.

HealthcareLegalPublic sector

Deep Dive

Why self-host Whisper instead of using OpenAI’s API

OpenAI charges £0.004/minute for Whisper. That’s cheap until you process a few thousand hours of audio — a podcast network, a legal-discovery pipeline, a call-centre QA team. At 1,000 hours/month you’re already at £240/mo on the OpenAI bill, and you’ve sent every minute of customer audio to a US-based provider.

A dedicated RTX 3050 hosts Whisper Large-v3 for £69/mo. It will transcribe ~720 hours/day at 3× RTF. The economics flip the moment you have any meaningful volume — and the privacy story is dramatically better.

Real-time vs batch — different deployment shapes

Batch. Throw a backlog of audio at a faster-whisper worker, let it churn at 5–8× real-time. Single 3090 or 5080. No streaming, no VAD, no fuss.

Streaming. Whisper-Streaming or WhisperX with a sliding window, processing 200ms chunks. Lower throughput per stream, but sub-second time-to-first-word. Multi-stream serving needs a 5080 or 5090.

Voice agent. STT + LLM + TTS pipeline, latency budget under 1 second total. Best on RTX 5090 (32 GB) so all three models stay hot in VRAM. See our voice agent hosting page.

Frequently Asked Questions

The questions buyers actually ask before committing to a GPU server.

Which Whisper model is included?

All of them — Large-v3, Large-v3-Turbo, distil-whisper-large-v3, faster-whisper-large-v3 — pre-downloaded. You pick which to serve.

Does it support 99 languages?

Yes. Whisper Large-v3 is multilingual by default. Distil variants are English-only.

Can I run it alongside an LLM on the same GPU?

Yes — Whisper uses ~6 GB of VRAM at FP16. Pair with a 7B LLM on a 24 GB+ card and you have a complete voice agent stack on one server.

How does it handle multi-speaker audio?

WhisperX + pyannote-audio diarisation gives speaker-aware transcripts. Pre-installed.

OpenAI-compatible API?

Yes — exposes /v1/audio/transcriptions with the same JSON shape. Existing OpenAI SDK calls work unchanged.

What about real-time / live captioning?

Whisper-Streaming is included. Time-to-first-word is ~600ms on a 5080, sub-300ms on a 5090.

Self-host Whisper today.

Pre-installed Whisper Large-v3 on your own GPU server. Real-time transcription, OpenAI-compatible API, fixed monthly bill. From £69/mo.

Have a question? Need help?