Compare

    Syntropylabs vs Braintrust

    Their strengths as we found them in their documentation; ours as the flow book states them. Nine rows, three states.

    Braintrust is the most complete evaluation product in this comparison: immutable experiments with git metadata, named dataset snapshots, a compare view that grades a change as Improvement, Regression, Tradeoff or Tie, and online scoring at trace, span or session scope with backfill. Its tracing stops at the LLM span: there is no services or endpoints layer, and several UI ceilings (exports capped at 1,000 rows, dataset search capped at 1,000 records) show in daily use. Syntropylabs starts from the other end, backend tracing with the model calls inside it, and adds coding-agent traces, scenario simulation and red-teaming that Braintrust does not have.

    Feature table

    Cells about Syntropylabs cite the flow book that specifies the product (the F- and C- ids are its flow numbers); a partial is stated with its limit. Cells about the other product come only from our documented teardown of its public documentation, dated 2026-09-06, with the source linked below. Where that teardown is silent, the cell says “not documented here”, never “no”.

    Feature comparison: Syntropylabs versus Braintrust
    FeatureSyntropylabsBraintrust
    Tracing depth
    yes

    Trace waterfall, sessions, users, error groups, logs, and services → endpoint → request in one tool (F-TR-02, F-TR-04 to F-TR-07). Span cost is computed from the price sheet in the browser and shown as approximate (F-TR-02).

    yes

    Logs with Spans tree, Thread, Timeline and Debugger layouts; ANY_SPAN() SQL predicates across the span tree; Topics and Patterns clustered across traces. No APM layer (services, endpoints, service map); exports are capped at 1,000 rows.

    Coding-agent tracing
    partial

    Claude Code, Codex, Gemini CLI, Cursor, Windsurf and OpenCode in one command; one trace per turn with subagents nested (C-INS-01). Cost arrives on a separate record merged onto the model call, so a lost record leaves a turn priced by the catalogue estimate; tool output for non-shell tools needs the CLI hook (F-TR-17).

    not documented here

    The teardown records “Copy pattern as prompt” for a coding agent, not tracing of coding-agent sessions.

    Evaluation runs
    yes

    Runs from a dataset with existing or generated outputs, per-row tri-state scores with judge reasons, compare against a baseline, gates, and SDK-created runs that open like any other (F-EV-01 to F-EV-07).

    yes

    Immutable Experiments with a baseline, git metadata and a parameters version; compare with an Improvement / Regression / Tradeoff / Tie grade, a regressions filter and diff modes; trials per row; eval-action posts a PR comment.

    Online rules
    yes

    Per-environment rules with evaluators, judge, interval, sampling, filters and a daily budget that pauses the rule (F-EV-08, F-EV-09). Deterministic assertion rules attached to an online rule still run through the judge path (F-EVL-04).

    yes

    Online scoring scoped to Trace, Span or Group (session), a SQL filter, a sampling rate, Test automation, and Re-process to backfill.

    Datasets and snapshots
    yes

    CSV and JSONL import, rows from traces, snapshots that runs pin to, and runs over time on the dataset (F-DS-01 to F-DS-03, C-DS-01). The snapshot diff shows counts and row references, not before/after values, and rows cannot be restored from a snapshot (F-DS-14).

    yes

    Named dataset snapshots that experiments pin to; a Runs tab with score over time; Add to dataset with copy vs trace reference and an origin link. Add-to-dataset captures the root span only.

    Simulation
    yes

    No-code, code and connector agents; generated or hand-written scenarios; live runs; LLM-judge scoring; run comparison (F-SIM-01 to F-SIM-09). Deterministic scores carried on SDK-run spans are not rendered yet (F-SIM-10).

    no

    Nothing generates scenarios or drives an agent in a sandbox.

    Red-team
    yes

    Vulnerabilities × attack strategies on the sandbox runner, a risk report, and every materialized attack kept as a trace that can become a regression row (C-SIM-01).

    not documented here
    Self-hosting

    Hosted at syntropylabs.ai. On-premise or VPC deployment is an Enterprise-plan line; there is no community self-hosted edition.

    BYOC is listed on Enterprise. No community self-hosted edition is documented here.

    Pricing model

    Free: tracing for one project, 7-day retention, 10,000 spans a month soft cap, coding-agent traces included. Pro $20 a month: evaluation, datasets, simulation. Enterprise: custom. Upgrades are requested in-app; self-serve billing is not built (F-ORG-06).

    Starter $0: $10 model credits, 1 GB data, 10k scores, 14-day retention, one human-review score type per project. Pro $249/mo: 30-day retention, custom dashboards, environments, dataset snapshots, permission groups. Enterprise: object-level ACLs, audit logs, SAML, BYOC.

    fig. 1 · nine rows, three states · reviewed 2026-09-15

    Choose Braintrust if…

    • You want experiments with git metadata, a Tradeoff / Tie grade and trials per row today; our compare view does not grade a comparison and per-row trials are partial (C-EV-01).
    • You want to score online at span or session-group scope with SQL filters and backfill; our online rules are per environment and trace-scoped.
    • You want Topics and Patterns clustered across traces; ours is per-trace Optimize, and cross-trace clustering is not built (C-TR-01).
    • You want an AI assistant inside logs and playgrounds; we have none.

    Choose Syntropylabs if…

    • You need services, endpoints, logs and error groups in the same tool as the LLM spans.
    • You want Claude Code, Codex, Gemini CLI, Cursor, Windsurf or OpenCode sessions as traces in one command, subagents nested where they ran.
    • You want to generate scenarios, run them against your real agent in a sandbox, score the run with a judge, and red-team the same agent.
    • You want tracing free on every plan; Braintrust’s dataset snapshots and environments start at Pro ($249/mo).

    Next steps

    Try the tracing side first

    Tracing is free for one project on every plan, coding-agent traces included. Evaluations, datasets and simulation are Pro and up.