LLM Observability & Evals
See what your prompts, models and agents actually did in production, and test changes against real examples before you ship them.
The real choice
Open-source platform you can self-host, or a hosted eval product. Open-source tools keep traces on your own servers and cost nothing to run yourself, but you maintain them. Hosted eval platforms give you polished experiment and scoring workflows for a monthly bill, with your data in their cloud.
Pick by situation
Tap the ones that are you — the tools that fit light up below.- You want tracing, prompts and evals in one tool you can self-host for free
Langfuse - You already use OpenTelemetry or want vendor-neutral instrumentation
Arize Phoenix - You want request logs and costs today by changing one base URL
Helicone - Your app is built on LangChain or LangGraph
LangSmith - Your main question is whether a prompt or model change made outputs better or worse
Braintrust
The contenders
Grouped by the side of the choice they answer, not ranked. Open a row for when to use it, the trade-off and what makers say.
Langfuseplatform · open source · free tierTracing, prompt management and evals in one tool you can run on your own server for free.

33+43 open source$29Core1.8M/wk+81% vs npm
Use it whenYou want to own your trace data or keep costs flat as volume grows.
Trade-offSelf-hosting means running Postgres, ClickHouse, Redis and object storage alongside it; otherwise it's a usage-based cloud plan.
llms.txtCLIUsed by Agentplace.io, AutonomyAI, cognee, Conduit, Finyuus and 28 more · in 43 open-source projects.
Marc gave me an in-person onboarding in SF - I found an issue in our LLM provider config just 30 minutes after the onboarding thanks to Langfuse. 10/10 recommendation
Arize Phoenixservice · self-hostable · free tierOpenTelemetry-based tracing and evals you can start locally in a notebook, then self-host or move to its cloud.19 open source—80.4k/wk+180% vs npm
Use it whenYou already use OpenTelemetry or want vendor-neutral instrumentation.
Trade-offSource-available under the Elastic License rather than a permissive open-source license.
llms.txtNo maker's product we track shows it for this yet · in 19 open-source projects.
Heliconeplatform · open source · free tierGetting request logs, costs and latency by changing one base URL, with no SDK instrumentation.


13$116Pro3.8k/wk
Use it whenYou want visibility today and your app makes direct model calls.
Trade-offProxy logging sees individual requests well but agent steps and eval workflows less deeply; acquired by Mintlify in March 2026, so check its roadmap before building on it.
llms.txtUsed by Alai, Codebuff, Fume, Image Ally, Persana and 8 more.
Helicone AI offers open-source observability tools tailored for developers working with LLMs. It simplifies debugging and optimization, providing valuable insights into AI model performance.
LangSmithplatform · free tierApps built on LangChain or LangGraph, where tracing works with almost no setup.

11+49 open source$475Developer6.1M/wk−4% vs npm
Use it whenYou already use the LangChain stack.
Trade-offPaid per seat beyond the free tier; self-hosting is enterprise-only.
llms.txtUsed by Astrid, DryMerge, GitLaw, Intryc, lmChatGPTtfy and 6 more · in 49 open-source projects.
LangSmith’s real-time analytics and versioning keep our AI agents rock-solid -- so everything just works better.
Braintrustplatform · free tierEval-driven work — scoring outputs and comparing prompts and models side by side in experiments.

5+10 open source—1.4M/wk+117% vs npm
Cost: the cheapest plan that fits traces 100k traces, from list prices. Try your own numbers →
Cost as you grow
Each contender's cheapest usable plan as usage rises.
Who switches to what
Public pull requests on GitHub since Oct 2024 whose title says "X to Y" — real code changes, by developers in general rather than makers only. Pick a flow to see its pull requests.
- Migrate observability from LangSmith to Langfuse and trim logged fieldsHOSH19/HarnessLab · 2026-08-29
- feat: migrate observability from LangSmith to Langfuse Cloud (v4 SDK)icekarim/momo-assistant · 2026-08-05
- feat(trace): migrate LangSmith to Langfuse v4 and fix env key priority1935494577/Xiaoxin-Agentic-RAG-System · 2026-08-01
- [codex] Migrate from LangSmith to Langfuse, remove Python legacy, fix review issuesBunnyRabbit8mile/codex-tee · 2026-07-14
- Reapply "feat: migrate observability from LangSmith to Langfuse"sinuarlowbaby/RAG-PDF-Chatbot · 2026-07-14
- refactor: migrate observability from LangSmith to Langfusewei-yiting/fin-lab-x · 2026-03-18
- Migrate observability from LangSmith to LangfuseAneeshPulukkul/hybrid-rag-solution · 2026-03-14
- feat:migrate LLM observabilty from LangSmith to LangFuse100-hours-a-week/17-JinyUs-Q-Feed-AI · 2026-02-26
- Refactor: Migrate from Langsmith to Langfusedylangamachefl/fantasy-football-chatbot-v2 · 2025-12-01
Before you choose
Start by logging every request with its prompt, output, latency and cost — a day of real traces teaches more than any benchmark. Save the bad outputs you find as a small dataset, and rerun it every time you change a prompt or model. Add automated scoring only after you've read enough traces to know what "good" means for your app.
- Logging full prompts and outputs that contain users' personal data to a third-party service without masking it or listing the vendor as a subprocessor.
- Changing a prompt in production with no saved examples to rerun, so a fix for one case silently breaks five others.
- Trusting an LLM-as-judge score before checking that it agrees with your own judgment on a few dozen examples.
Other options
Real choices most makers here won't need to weigh.
Decided alongside
What the 58 makers' products here chose for their other decisions.