macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCR

About

How we rank & score
Head-to-head · Speech Recognition & Transcription

Vosk vs Whisper

Both are alternatives to Deepgram. Here's how they stack up — verified facts, no spin.

Also searched as Whisper vs Vosk — same comparison, one verdict.

95

Vosk

Real-time transcription on a CPU. Runs on a Raspberry Pi.

OPEN SOURCEApache-2.0SELF-HOSTLOCAL-FIRST

Vosk is a lightweight offline speech-recognition toolkit built for constrained environments — it runs in real time on a CPU, works on Android, iOS and single-board computers, and its models are tens of megabytes rather than gigabytes. It will not match Whisper on accuracy for difficult audio. It is the right answer when you need speech recognition on a device, offline, with no GPU and no network, which is a requirement no managed API can satisfy at any price. Apache-2.0.

94

Whisper

The model that changed the category. MIT, and free to run forever.

OPEN SOURCEMITSELF-HOSTLOCAL-FIRST

Whisper is OpenAI's speech-recognition model, released openly under MIT — weights included. It handles around a hundred languages, is robust to accents and background noise, and does translation as well as transcription. Releasing it is what collapsed the economics of this category: a model competitive with the commercial APIs became something anyone could download. The reference implementation is slower than the optimised runtimes, but it is the simplest thing that works and the baseline everything else is measured against.

Side by side

 VoskWhisper
Sovereignty Score9594
Open sourceYesYes
Self-hostableYesYes
Local-firstYesYes
LicenseApache-2.0MIT
PricingFree, Apache-2.0.Free — model weights and code both MIT.
The verdict

Vosk edges it on the Sovereignty Score, but the right pick depends on the trade-offs below.

Vosk

Strengths

  • +Real-time on CPU — no GPU required at all
  • +Tiny models; runs on phones and single-board computers
  • +Fully offline, which some products require by law
  • +Around twenty languages supported

Trade-offs

  • Noticeably less accurate than Whisper on hard audio
  • Smaller language coverage
  • No built-in diarization

Whisper

Strengths

  • +Genuinely open weights under MIT, not a restricted community licence
  • +About a hundred languages, robust to noise and accent
  • +Reference implementation — simplest possible starting point
  • +Translation to English included

Trade-offs

  • Slower and hungrier than faster-whisper for identical output
  • Weak word-level timestamps without help
  • No speaker diarization
See all 5 Deepgram alternatives →

Related alternative guides

Facts verified 2026-08-11. Licenses and pricing change — spotted something out of date? That's a correction we want.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.