Layer 4 · self-hosting reality check
What it actually takes to self-host faster-whisper
The docs say 2 GB. In practice you want 5 GB VRAM for large-v3 in float16; 2 GB with int8. Here is the honest version — real requirements, real monthly cost, what you will be maintaining, and the one thing that catches people out.
Usually reached from Deepgram alternatives, where faster-whisper is one of the picks.
| RAM — documented minimum | 2 GB |
|---|---|
| RAM — what it really needs | 5 GB VRAM for large-v3 in float16; 2 GB with int8 |
| CPU | 4 vCPU if running on CPU — usable but slow |
| Disk | 3 GB for the large model |
| Monthly cost | $0 on hardware you own. A one-off archive transcription is a few dollars of rented GPU against a four-figure API invoice. |
| Setup time | 30 minutes |
| How you install it | pip install faster-whisper; it uses CTranslate2 under the hood |
| Ongoing maintenance | Very low. |
| Where it stops scaling | Hours of audio per hour of GPU time. Batch throughput is excellent; concurrency needs multiple workers. |
The thing that catches people out
Whisper hallucinates on silence — long pauses, music or dead air produce confident invented sentences, often repeated phrases from training data. Deepgram does not do this. Enable the built-in VAD filter and set `no_speech_threshold`, or you will ship transcripts containing text nobody said, which is worse than a gap.
When not to self-host faster-whisper
You need real-time streaming with sub-second interim results. Whisper is architecturally batch-oriented and streaming wrappers trade latency against accuracy.
Every guide here carries this section. A site that only ever tells you to self-host is selling something — the useful answer is sometimes no.
Other Layer 4 self-hosting guides
- Self-hosting Ollama8 GB VRAM for a 7B model at usable speed; 24 GB for 30B-class
- Self-hosting vLLM24 GB VRAM minimum for useful production serving
- Self-hosting QdrantVectors × dimensions × 4 bytes, in RAM, plus overhead — 1M × 768d is roughly 3 GB
- Self-hosting pgvector8 GB — the HNSW index wants to be resident
- Self-hosting LlamaIndex4 GB for the app; your vector store is the real cost
- Self-hosting Langfuse4 GB
Common questions
- How much RAM does faster-whisper actually need?
- 5 GB VRAM for large-v3 in float16; 2 GB with int8 in practice. The documented minimum is 2 GB, which is the figure at which the process starts rather than the figure at which it works under real use. 4 vCPU if running on CPU — usable but slow alongside it.
- What does self-hosting faster-whisper cost per month?
- $0 on hardware you own. A one-off archive transcription is a few dollars of rented GPU against a four-figure API invoice. This is commodity VPS pricing and excludes your time, which is the larger cost for most people — budget for very low.
- How long does it take to set up faster-whisper?
- 30 minutes, via pip install faster-whisper; it uses CTranslate2 under the hood.
- When should I NOT self-host faster-whisper?
- You need real-time streaming with sub-second interim results. Whisper is architecturally batch-oriented and streaming wrappers trade latency against accuracy.
- What is the most common mistake when self-hosting faster-whisper?
- Whisper hallucinates on silence — long pauses, music or dead air produce confident invented sentences, often repeated phrases from training data. Deepgram does not do this. Enable the built-in VAD filter and set `no_speech_threshold`, or you will ship transcripts containing text nobody said, which is worse than a gap.