macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCR

About

How we rank & score
Best of · 5 tools ranked

The best Speech Recognition & Transcription

Every option ranked — open-source, self-hostable, and commercial — by our transparent Sovereignty Score, with honest trade-offs so you choose what fits you, not us.

Turning audio into text. This is the category where open weights most decisively caught the paid APIs: Whisper-class models run on a laptop and cost nothing per hour, while the incumbents still bill by the minute.

L4Models & tooling — layer 4 of the AI stack
  1. 1

    faster-whisper

    Top pickOpen source

    Whisper, several times faster, on less memory. The practical default.

    Free. Hardware you already own; a laptop handles the smaller models. · in our Deepgram comparison →

    95
    sovereignty
    What is faster-whisper? →
  2. 2

    Vosk

    Open source

    Real-time transcription on a CPU. Runs on a Raspberry Pi.

    Free, Apache-2.0. · in our Deepgram comparison →

    95
    sovereignty
    What is Vosk? →
  3. 3

    Whisper

    Open source

    The model that changed the category. MIT, and free to run forever.

    Free — model weights and code both MIT. · in our Deepgram comparison →

    94
    sovereignty
    What is Whisper? →
  4. 4

    WhisperX

    Open source

    Accurate word timestamps and speaker labels — Whisper's two weak spots, fixed.

    Free. Diarization models may require accepting separate terms. · in our Deepgram comparison →

    90
    sovereignty
    What is WhisperX? →
  5. 5

    NVIDIA NeMo ASR

    Open source

    The strongest open streaming and diarization story.

    Free and Apache-2.0; assumes NVIDIA GPU hardware. · in our Deepgram comparison →

    89
    sovereignty
    What is NVIDIA NeMo ASR? →

Replacing a specific tool?

Head-to-head comparisons for each popular speech recognition & transcription product.

Straight head-to-heads

Two speech recognition & transcription tools, side by side — verified facts and a plain verdict.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.