Docs

    Get started

    Introduction

    What SyntropyLabs is, what the EvalKit SDK records, and how the app is organized.

    SyntropyLabs is a platform for teams that ship LLM applications and agents. The EvalKit SDK (syntropylabs-evalkit on PyPI and npm) runs inside your application. It records every model call, tool call, HTTP request, database query and log line as a trace.

    The platform stores those traces. It can also:

    • drive an agent through generated scenarios in a sandbox
    • score traces with evaluators: LLM judges, code and assertions
    • run the same evaluators over a dataset, so a change can be compared against a baseline
    • put traces, sessions and dataset rows in review queues for people to score

    Coding agents such as Claude Code, Codex, Gemini CLI, Cursor, Windsurf and OpenCode can be traced too, with one CLI command and no application code.

    How the app is organized

    An Organization owns Projects, and a Project owns Environments. An Environment is a development, staging or production deployment with its own environment key (tk_live_…). A trace belongs to the Environment whose key sent it.

    Datasets, runs, Models, prompts, agents and review queues belong to the Project. Evaluators, collections and Provider connections belong to the Organization, and every project in it shares them. Concepts defines each term.

    Sidebar itemWhat you do thereDocs
    HomeThe Getting started checklist, then the selected Environment's health over a time range you pickYour first trace
    TracesRead and filter traces, sessions, users, errors, logs, services, topics and agent versions, with live tail and exportTraces
    EvaluationsRuns over a dataset, comparisons, gates, online rules for an Environment, and calibrationEvaluations
    DatasetsCSV and JSONL import, rows built from traces, generated outputs, snapshots and golden setsDatasets
    EvaluatorsThe managed library, your own LLM judges, code evaluators, assertions and collectionsEvaluators
    SimulationsAgents, scenario sets, sandbox runs, scoring and red-team runsSimulations
    PlaygroundChat, pointwise and pairwise scoring, grids over a dataset, and model comparePlayground
    PromptsVersioned prompts with labels, fetched by name from codePrompts
    AnnotateScore types, review queues, labels, and RL and preference datasetsAnnotate
    AlertsThreshold alerts, monitors and the notification bellAlerts
    CostModel spend and judge spend for an Environment, and the online-rule budgetCost
    SettingsProject name, Environments and keys, and Models. For the organization: members, Providers, and plan and usageSettings

    What needs a plan feature

    Home, Traces, Cost and Settings are available to every organization. Evaluations, Datasets, Evaluators, Simulations, Playground, Prompts, Annotate and Alerts are plan features.

    When your organization's plan does not include one, its sidebar item stays in place with a lock and opens a panel that explains the feature. Docs pages about a gated feature show the same lock in the docs sidebar. Plan, usage and locks explains the fifteen feature keys.

    Where to start

    1. 1

      Create an Environment and copy its key

      Signing up already created a Default Project with a development Environment. To add staging and production, see Create an environment and get its key.

    2. 2

      Install the SDK

      Call init() once with the key, in Python or TypeScript. See Install the SDK.

    3. 3

      Read your first trace

      Find the model call in the waterfall and read its prompt, completion, tokens and estimated cost. See Your first trace.

    4. 4

      Score something

      Pick evaluators from the managed library and judge your recent traces. See Your first evaluation.

    To trace a coding agent instead of an application, start with Trace a coding agent.

    EvalKit is built by SyntropyLabs. Published on PyPI and npm.