Braintrust
Braintrust is an AI evaluation platform that helps engineering teams test, score, and monitor large language model applications.
About this data
Updated August 3, 2026
Overall Pulse Score
+3 over this period
A 0-100 index summarizing the tone of 133 relevant public mentions gathered from public online communities across 14 weeks in the selected period. It measures online sentiment, not a rating of the product's quality.
Weekly Sentiment Trend
Pulse Score by week over the selected period. Each point is one complete week of mentions.
This week in public discussion
Sentiment around Braintrust over the recent period skewed negative, with bugs and reliability drawing the most attention among commenters. Discussion focused heavily on specific issues including a Zod schema mismatch in the list-experiments MCP tool, gevent incompatibility with the sync API, and a tracing leak tied to Microsoft Agent Framework integration. Several mentions praised specific features and noted favorable competitor comparisons, but complaint themes outnumbered praise themes, keeping the pulse score in a middling range.
Read the deeper analysisAI-generated summary of public online discussion during this period. It reflects the tone of that discussion, not facts about the product or our views.
Sentiment mix by week
How the tone of public discussion splits each week.
Ringed points mark weeks with unusually high discussion volume, more than double this product's typical week.
Most-discussed praise
Most-discussed complaints
Themes across the selected period, with mention counts.
How Braintrust compares
Pulse Score over the selected period versus the top tracked competitors in Coding.
Where the mentions come from
Share of the 133 relevant public mentions in the selected period, by source.
Sample public mentions
Showing 5 of 133 analyzed public mentions in this period, with links to the original source. We do not reproduce full threads.
“Fix experiment-dataset linking when running evals with a dataset. We recently had this bug reported in Ruby, and my analysis shows this is also present in Java. https://github.com/braintrustdata/braintrust-sdk-ruby/pull/103”
“[Bug] UseBraintrustTracing leaks Microsoft Agent Framework local history sentinel with RequirePerServiceCallChatHistoryPersistence. When using Braintrust.Sdk.AgentFramework with Microsoft Agent Framework and RequirePerServiceCallChatHistoryPersistence = true, UseBraintrustTracing...”
“Evals: automatically compare runs against prior or selected baseline. ## Context AgentV should make eval comparisons first-class. During WTG Braintrust/Phoenix UX benchmarking, Braintrust automatically compared a new experiment against the prior compatible experiment, and Phoenix...”
“DatasetCase.metadata() always empty when fetching dataset rows from Braintrust. When executing a Playground in braintrust, remote eval task is always passed an empty metadata map for each dataset, even if those datasets have metadata. Reproduction steps: 1. Create a dataset with ...”
“list-experiments MCP tool throws: repo_info.dirty is required in Zod schema but absent in API responses. ## Description The list-experiments MCP tool fails entirely when any returned experiment has a repo_info object that lacks the dirty field. This happens whenever evals are run...”
241+ more analyzed mentions, full history, and theme breakdowns are part of Pro.
Get ProDeeper analysis
- Bug reports and reliability complaints dominated discussion volume over the four-week window.
- Sentiment moved in sharp swings rather than a clear trend, with no sustained improvement taking hold across the period.
- Commenters were split on integration quality, with some praising it and others citing specific SDK and framework failures.
- Favorable competitor comparisons gave Braintrust a positive presence in evaluation-tooling conversations despite the criticism.
| Praise theme | Mentions |
|---|---|
| Strong features | 27 |
| Good integrations | 11 |
| Compared to rivals | 10 |
| AI quality | 5 |
| Feature requests | 4 |
| Complaint theme | Mentions |
|---|---|
| Bugs | 39 |
| Missing features | 18 |
| Reliability | 16 |
| Lacking integrations | 14 |
| Compared to rivals | 9 |
Discussion of Braintrust over the past four weeks has been shaped primarily by two competing forces: genuine enthusiasm for the product's role in AI evaluation workflows and a persistent undercurrent of frustration tied to bugs and reliability. Bug reports were the single most cited theme across all mentions, with commenters describing specific failure modes including schema validation errors in MCP tooling, SDK incompatibilities with gevent monkey patching, and tracing leaks when integrated with third-party agent frameworks. Reliability concerns ran closely behind, suggesting that for users attempting to embed Braintrust deeper into CI pipelines and production observability stacks, confidence in consistent behavior remains a sticking point.
On the positive side, feature praise was the second most prominent theme, and several mentions framed Braintrust favorably in direct competitor comparisons against tools like DeepEval, promptfoo, and LangSmith. Discussion suggested that commenters see the product as a credible later-stage observability layer rather than a first-gate evaluation runner, which points to a reasonably well-defined perceived niche even among skeptics.
The score trajectory across the window tells a volatile story. Sentiment dipped sharply in mid-May before recovering strongly in late May on a surge of mentions, then oscillated through June, softening again toward the final weeks of measured activity. The overall direction from the start of the window to the most recent high-volume periods is largely flat with notable swings, suggesting that spikes in discussion tend to surface concentrated bug reports rather than sustained positive momentum.
Opinion was most divided around integration quality. Some commenters praised integrations as a strength while others surfaced concrete failures tied to SDK and framework compatibility. Pricing drew a small but notable share of complaints, hinting at value-perception tension among at least some users.
AI-generated summary of public online discussion during this period. It reflects the tone of that discussion, not facts about the product or our views.
Member perspectives
Individual opinions from Pro members, posted over time. These are personal member views, not aggregated sentiment data.
Overall Pulse Score
+3 over this period
A 0-100 index summarizing the tone of 133 relevant public mentions gathered from public online communities across 14 weeks in the selected period. It measures online sentiment, not a rating of the product's quality.
Data summary
Compare with another tool
Braintrust
45
Trainual
88
Score-level preview from live weekly tracking.
Are you Braintrust?
Get a private enterprise dashboard for your product - full history, every source, theme deep-dives, and weekly alerts. You can also respond to the data shown here.
Explore the enterprise dashboardAffiliate disclosure
Some links on this site may be affiliate links. If you click one and make a purchase, we may earn a commission at no extra cost to you. Learn more.
Is Braintrust your product?
See everything behind this page - full history, every source, theme deep-dives, and weekly alerts - in a private enterprise dashboard.
Request early accessCompare with similar tools
Safari MCP Server
A Model Context Protocol server that gives AI assistants programmatic access to Safari browser automation for developers.
Free
View DetailsRevenueCat
A platform that manages in-app purchases, subscriptions, and revenue analytics for iOS and Android app developers.
Free tier; paid plans available
View Details