LangSmith vs MLflow

Two sides of the LLM observability & evals decision: hosted eval platform and open-source platform. When each fits, what it costs, who moves from one to the other, and what makers who chose it say.

Ask your AI about this, with this page as the source:ChatGPT ↗Claude ↗Perplexity ↗

Which fits you

Choose LangSmith if
  • Your app is built on LangChain or LangGraph

Use it whenYou already use the LangChain stack.

Trade-offPaid per seat beyond the free tier; self-hosting is enterprise-only.

Choose MLflow if
  • Tracing, evaluating and versioning prompts for LLM apps in the same open-source platform many teams already use for ML experiments.

Use it whenYou already use MLflow or Databricks for models, or want OpenTelemetry-based tracing and evals you can self-host.

Trade-offGrew out of classic ML tooling, so the UI and concepts are broader than an LLM-only tool; you run the tracking server yourself unless you use a managed version.

At a glance

LangSmithMLflow
Used by11 makers' products · 49 open-source projectsNo maker's product yet · 15 open-source projects
Cost at default usagetraces 100k traces$475/mo Developer—
Downloads6.1M/wk−4% vs npm7.9k/wk
PricingFree tier with a monthly trace limit; paid per-seat plan with usage overage; custom Enterprise, including self-hosted. · paid from $39/seat/moFree (open source); managed versions are offered by Databricks and cloud providers.
Free tierYesYes
Open sourceNoYes · self-hostable
Incidents, 90 daysfrom its status page0no public status feed

Cost as you grow

Both cost $0 up to 1k traces; from 10k traces MLflow costs less ($0 vs $25); and still does at 10M traces ($0 vs $49,975). They're different kinds of tool — hosted eval platform and open-source platform — so the prices don't buy the same thing.

$0$1,000$5,000$20,0001k10k50k100k500k1M5M10M
MLflowLangSmithx: traces per month (trace) · cheapest usable plan at each point, list prices · try your own numbers
The numbers, plan by plan
Traces per monthLangSmithMLflow
1,000$0 Developer$0 Self-hosted (open source)
10,000$25 Developer$0 Self-hosted (open source)
50,000$225 Developer$0 Self-hosted (open source)
100,000$475 Developer$0 Self-hosted (open source)
500,000$2,475 Developer$0 Self-hosted (open source)
1,000,000$4,975 Developer$0 Self-hosted (open source)
5,000,000$24,975 Developer$0 Self-hosted (open source)
10,000,000$49,975 Developer$0 Self-hosted (open source)

From each vendor's pricing page: LangSmith, MLflow.

What makers say

Makers on using it for LLM observability & evals, from Product Hunt and Starter Story interviews, each linked to the source. Products with a page of their own and fuller notes first.

On LangSmith
LangSmith’s real-time analytics and versioning keep our AI agents rock-solid -- so everything just works better.
Watchman AI, the makerSep 2026 ↗
You can't build AI agents without monitoring. Metadata filtering is strong.
DryMerge, the makerSep 2026 ↗
I deployed the manage Pig agent on LangGraph and it's been smooth sailing!
Pig, the makerSep 2026 ↗
5 more on the LangSmith page →
On MLflow

No maker quote about MLflow for LLM observability & evals yet.

Loved and watch-outs

Themes that recur in makers' words and Hacker News comments, each linked to what it summarises.

LangSmith
Most loved
  • Tracing through the whole prompt path replaces guesswork when diagnosing agent issues. PH
  • Traces, datasets, annotation queues and evals live in one platform widely used by enterprises. PHHN
  • Pairs tightly with LangChain and LangGraph for debugging stateful agent workflows. PHHNHN 2
Watch-outs
  • It keeps pulling users toward the LangChain platform, and LangChain docs push LangSmith, which feels like lock-in. HNHN 2HN 3HN 4
  • Traces show which agent failed but not why, so root-cause analysis stays manual. HNHN 2HN 3HN 4
  • Viewing your own traces requires a cloud account, with no local-first option. HNHN 2
On Product Hunt: 4.8★, 19 reviews · mentioned most: monitoring AI model performance, chain sequence debugging, evals
MLflowNothing that recurs in what we collected yet.

Who uses each

MLflow0 makers' products
None tracked yet; 15 open-source projects declare it.

What makers pair each with

With LangSmith
Hosting
LangfuseOpen-source platformTracing, prompt management and evals in one tool you can run on your own server for free.vs LangSmith →vs MLflow →
Arize PhoenixOpen-source platformOpenTelemetry-based tracing and evals you can start locally in a notebook, then self-host or move to its cloud.vs LangSmith →
HeliconeProxy loggingGetting request logs, costs and latency by changing one base URL, with no SDK instrumentation.vs LangSmith →
BraintrustHosted eval platformEval-driven work — scoring outputs and comparing prompts and models side by side in experiments.vs LangSmith →