macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCR

About

How we rank & score
Best of · 5 tools ranked

The best LLM Evaluation & Testing

Every option ranked — open-source, self-hostable, and commercial — by our transparent Sovereignty Score, with honest trade-offs so you choose what fits you, not us.

How you find out whether a prompt change made things better or worse. Without it, every model upgrade is a guess. The open tools run in CI next to your unit tests; the platforms add a hosted UI and per-seat pricing.

L4Models & tooling — layer 4 of the AI stack
  1. 1

    Promptfoo

    Top pickOpen source

    Prompt regressions fail the build. Declarative evals that live in CI.

    Free and MIT, unlimited seats. An enterprise tier exists for larger organisations. · in our Braintrust comparison →

    94
    sovereignty
    What is Promptfoo? →
  2. 2

    Inspect AI

    Open source

    The rigorous one, from the UK AI Safety Institute.

    Free, MIT, publicly funded. · in our Braintrust comparison →

    93
    sovereignty
    What is Inspect AI? →
  3. 3

    DeepEval

    Open source

    Evals as pytest tests, with the research metrics already implemented.

    Free and Apache-2.0; optional paid cloud dashboard. · in our Braintrust comparison →

    92
    sovereignty
    What is DeepEval? →
  4. 4

    Ragas

    Open source

    Purpose-built RAG evaluation — measures retrieval and generation separately.

    Free, Apache-2.0. · in our Braintrust comparison →

    92
    sovereignty
    What is Ragas? →
  5. 5

    OpenAI Evals

    Open source

    The original benchmark harness. Simple, standard, still useful.

    Free, MIT. You pay for whatever model API the evals call. · in our Braintrust comparison →

    85
    sovereignty
    What is OpenAI Evals? →

Replacing a specific tool?

Head-to-head comparisons for each popular llm evaluation & testing product.

Straight head-to-heads

Two llm evaluation & testing tools, side by side — verified facts and a plain verdict.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.