macrostack
Browse

The AI stack

Categories

Local & Sovereign AINotes & KnowledgeObservability & MonitoringPassword ManagersWeb AnalyticsTeam ChatSmart HomeNetworking & RoutersVideo ConferencingCloud Storage & SyncPhotos & MediaAPI DevelopmentImage EditingWorkflow Automation & iPaaSDeveloper Tools & ContainersOffice & Productivity SuitesNo-Code DatabasesCode Hosting & Git ForgesProject ManagementEmail Marketing & NewslettersScheduling & BookingError Tracking & Exception MonitoringLog Management & SIEMVPN & PrivacyEmail & Secure MailVector Databases & AI SearchLLM & Agent FrameworksDomains & Web HostingData Removal & PrivacyAuthentication & IdentityHelp Desk & Customer SupportCloud & VPSKubernetes & Container PlatformsEmbedding ModelsPDF & DocumentsAI Coding AssistantsAI Voice & SpeechLLM Observability & EvaluationLLM Gateways & RoutingCloud GPU & AI ComputeCI/CD & build automationData & pipeline orchestrationModel serving & inferenceAI agent frameworksBackend as a serviceSecrets managementFeature flags & experimentationProduct analyticsSearch infrastructureUptime & status monitoringAffiliate & partner platformsVisitor identification & personalisationWikis & internal docsIdentity & access managementData warehouses & analytics enginesCustomer data platformsCRMObject storageBI & dashboardsE-signatureWhiteboards & diagrammingIn-memory data stores & cachingPlatform as a serviceTransactional & bulk emailHeadless CMSDesign & prototypingE-commerce platformsInternal tools & admin panelsManaged databasesForms & surveysFine-Tuning & Model TrainingRAG & Retrieval PlatformsLLM Evaluation & TestingAI Guardrails & Content SafetySpeech Recognition & TranscriptionExperiment Tracking & ML OpsDocument AI & OCRCompliance automation & security posture

About

How we rank & score
Best of · 4 tools ranked

The best LLM Observability & Evaluation

Every option ranked — open-source, self-hostable, and commercial — by our transparent Sovereignty Score, with honest trade-offs so you choose what fits you, not us.

Tracing, evaluation and cost tracking for LLM applications — hosted platforms compared against tools you can run yourself. What a request actually cost, why an answer was wrong, and whether a prompt change made things better.

L4Models & tooling — layer 4 of the AI stack
  1. 1

    Langfuse

    Top pickOpen source

    The self-hosted default for LLM tracing — MIT core, framework-agnostic.

    Free to self-host — Docker Compose or Kubernetes, costs are your own infrastructure. A managed cloud with a free tier and paid plans exists if you would rather not run it. · in our LangSmith comparison →

    90
    sovereignty
    What is Langfuse? →
  2. 2

    OpenLLMetry

    Open source

    Apache-2.0 OpenTelemetry instrumentation — send traces anywhere you already send them.

    Free — Apache-2.0 instrumentation libraries. You pay only whatever backend you already send telemetry to; a hosted Traceloop backend is available separately. · in our LangSmith comparison →

    92
    sovereignty
    What is OpenLLMetry? →
  3. 3

    Helicone

    Open source

    Gateway and observability in one — one integration instead of two.

    Free and self-hostable under Apache-2.0. Hosted service has a free tier of roughly 10,000 requests per month, with paid plans starting around $79/month. · in our LangSmith comparison →

    88
    sovereignty
    What is Helicone? →
  4. 4

    Phoenix (Arize)

    Commercial

    Strong evaluation tooling — but read the licence before you build on it.

    Free to self-host under ELv2 for your own internal use. Arize's commercial platform is priced separately. · in our LangSmith comparison →

    62
    sovereignty
    What is Phoenix (Arize)? →

Is it free?

The free tier, its real limits and where you start paying — read from each vendor's own pricing page and dated.

Run these yourself

What each one actually needs — real RAM, honest running cost, and the setup time nobody quotes.

Replacing a specific tool?

Head-to-head comparisons for each popular llm observability & evaluation product.

Straight head-to-heads

Two llm observability & evaluation tools, side by side — verified facts and a plain verdict.

The Macrostack brief

New swaps, worth your inbox.

A short, occasional email when we add a high-intent alternative or ship a new head-to-head. No spam, no selling your address — unsubscribe in one click.