Back to Coding
Braintrust logo

Braintrust

Braintrust is an AI evaluation platform that helps engineering teams test, score, and monitor large language model applications.

Primary category: Coding
About this data
This page reflects public online discussion, collected and scored by automated systems and summarized using AI. It is not a statement of fact, not an audit, and not our own opinion of the product. Automated analysis can be incomplete or wrong, and scores carry the limitations described in our methodology. Companies can respond with their own perspective. See how this is calculated.

Updated August 3, 2026

Overall Pulse Score

45
Pulse Score

+3 over this period

A 0-100 index summarizing the tone of 133 relevant public mentions gathered from public online communities across 14 weeks in the selected period. It measures online sentiment, not a rating of the product's quality.

Weekly Sentiment Trend

Pulse Score by week over the selected period. Each point is one complete week of mentions.

Download chart

This week in public discussion

Sentiment around Braintrust over the recent period skewed negative, with bugs and reliability drawing the most attention among commenters. Discussion focused heavily on specific issues including a Zod schema mismatch in the list-experiments MCP tool, gevent incompatibility with the sync API, and a tracing leak tied to Microsoft Agent Framework integration. Several mentions praised specific features and noted favorable competitor comparisons, but complaint themes outnumbered praise themes, keeping the pulse score in a middling range.

Read the deeper analysis

AI-generated summary of public online discussion during this period. It reflects the tone of that discussion, not facts about the product or our views.

Sentiment mix by week

How the tone of public discussion splits each week.

Ringed points mark weeks with unusually high discussion volume, more than double this product's typical week.

Most-discussed praise

Strong features27
Good integrations11
Compared to rivals10
AI quality5
Feature requests4

Most-discussed complaints

Bugs39
Missing features18
Reliability16
Lacking integrations14
Compared to rivals9

Themes across the selected period, with mention counts.

How Braintrust compares

Pulse Score over the selected period versus the top tracked competitors in Coding.

Where the mentions come from

Share of the 133 relevant public mentions in the selected period, by source.

GitHub96% (128)
Hacker News4% (5)

Sample public mentions

Showing 5 of 133 analyzed public mentions in this period, with links to the original source. We do not reproduce full threads.

Fix experiment-dataset linking when running evals with a dataset. We recently had this bug reported in Ruby, and my analysis shows this is also present in Java. https://github.com/braintrustdata/braintrust-sdk-ruby/pull/103

GitHubFeb 17, 2026

[Bug] UseBraintrustTracing leaks Microsoft Agent Framework local history sentinel with RequirePerServiceCallChatHistoryPersistence. When using Braintrust.Sdk.AgentFramework with Microsoft Agent Framework and RequirePerServiceCallChatHistoryPersistence = true, UseBraintrustTracing...

GitHubJun 21, 2026

Evals: automatically compare runs against prior or selected baseline. ## Context AgentV should make eval comparisons first-class. During WTG Braintrust/Phoenix UX benchmarking, Braintrust automatically compared a new experiment against the prior compatible experiment, and Phoenix...

GitHubJun 15, 2026

DatasetCase.metadata() always empty when fetching dataset rows from Braintrust. When executing a Playground in braintrust, remote eval task is always passed an empty metadata map for each dataset, even if those datasets have metadata. Reproduction steps: 1. Create a dataset with ...

GitHubJun 17, 2026

list-experiments MCP tool throws: repo_info.dirty is required in Zod schema but absent in API responses. ## Description The list-experiments MCP tool fails entirely when any returned experiment has a repo_info object that lacks the dirty field. This happens whenever evals are run...

GitHubJun 24, 2026

241+ more analyzed mentions, full history, and theme breakdowns are part of Pro.

Get Pro

Deeper analysis

  • Bug reports and reliability complaints dominated discussion volume over the four-week window.
  • Sentiment moved in sharp swings rather than a clear trend, with no sustained improvement taking hold across the period.
  • Commenters were split on integration quality, with some praising it and others citing specific SDK and framework failures.
  • Favorable competitor comparisons gave Braintrust a positive presence in evaluation-tooling conversations despite the criticism.
Praise themeMentions
Strong features27
Good integrations11
Compared to rivals10
AI quality5
Feature requests4
Complaint themeMentions
Bugs39
Missing features18
Reliability16
Lacking integrations14
Compared to rivals9

Discussion of Braintrust over the past four weeks has been shaped primarily by two competing forces: genuine enthusiasm for the product's role in AI evaluation workflows and a persistent undercurrent of frustration tied to bugs and reliability. Bug reports were the single most cited theme across all mentions, with commenters describing specific failure modes including schema validation errors in MCP tooling, SDK incompatibilities with gevent monkey patching, and tracing leaks when integrated with third-party agent frameworks. Reliability concerns ran closely behind, suggesting that for users attempting to embed Braintrust deeper into CI pipelines and production observability stacks, confidence in consistent behavior remains a sticking point.

On the positive side, feature praise was the second most prominent theme, and several mentions framed Braintrust favorably in direct competitor comparisons against tools like DeepEval, promptfoo, and LangSmith. Discussion suggested that commenters see the product as a credible later-stage observability layer rather than a first-gate evaluation runner, which points to a reasonably well-defined perceived niche even among skeptics.

The score trajectory across the window tells a volatile story. Sentiment dipped sharply in mid-May before recovering strongly in late May on a surge of mentions, then oscillated through June, softening again toward the final weeks of measured activity. The overall direction from the start of the window to the most recent high-volume periods is largely flat with notable swings, suggesting that spikes in discussion tend to surface concentrated bug reports rather than sustained positive momentum.

Opinion was most divided around integration quality. Some commenters praised integrations as a strength while others surfaced concrete failures tied to SDK and framework compatibility. Pricing drew a small but notable share of complaints, hinting at value-perception tension among at least some users.

AI-generated summary of public online discussion during this period. It reflects the tone of that discussion, not facts about the product or our views.

Member perspectives

Individual opinions from Pro members, posted over time. These are personal member views, not aggregated sentiment data.

Data summary

Total mentions analyzed (all time)
246
Mentions in selected period
133
Weeks in range
14
vs Coding average (47)
Below by 2
Pricing
Free tier; paid plans available
Sources
GitHub (128), Hacker News (5)

Compare with another tool

Braintrust

45

Trainual

88

Full comparison

Score-level preview from live weekly tracking.

Are you Braintrust?

Get a private enterprise dashboard for your product - full history, every source, theme deep-dives, and weekly alerts. You can also respond to the data shown here.

Explore the enterprise dashboard

Try Braintrust

Visit the official website to get started

Visit site

Affiliate disclosure

Some links on this site may be affiliate links. If you click one and make a purchase, we may earn a commission at no extra cost to you. Learn more.

Is Braintrust your product?

See everything behind this page - full history, every source, theme deep-dives, and weekly alerts - in a private enterprise dashboard.

Request early access

Compare with similar tools

Safari MCP Server logo

Safari MCP Server

84

A Model Context Protocol server that gives AI assistants programmatic access to Safari browser automation for developers.

Strong features
Limited data

Free

View Details
RevenueCat logo

RevenueCat

75

A platform that manages in-app purchases, subscriptions, and revenue analytics for iOS and Android app developers.

Good integrations
Limited data

Free tier; paid plans available

View Details