Get started
Introduction
What SyntropyLabs is, what the EvalKit SDK records, and how the app is organized.
SyntropyLabs is a platform for teams that ship LLM applications and agents. The EvalKit SDK (syntropylabs-evalkit on PyPI and npm) runs inside your application. It records every model call, tool call, HTTP request, database query and log line as a trace.
The platform stores those traces. It can also:
- drive an agent through generated scenarios in a sandbox
- score traces with evaluators: LLM judges, code and assertions
- run the same evaluators over a dataset, so a change can be compared against a baseline
- put traces, sessions and dataset rows in review queues for people to score
Coding agents such as Claude Code, Codex, Gemini CLI, Cursor, Windsurf and OpenCode can be traced too, with one CLI command and no application code.
How the app is organized
An Organization owns Projects, and a Project owns Environments. An Environment is a development, staging or production deployment with its own environment key (tk_live_…). A trace belongs to the Environment whose key sent it.
Datasets, runs, Models, prompts, agents and review queues belong to the Project. Evaluators, collections and Provider connections belong to the Organization, and every project in it shares them. Concepts defines each term.
| Sidebar item | What you do there | Docs |
|---|---|---|
| Home | The Getting started checklist, then the selected Environment's health over a time range you pick | Your first trace |
| Traces | Read and filter traces, sessions, users, errors, logs, services, topics and agent versions, with live tail and export | Traces |
| Evaluations | Runs over a dataset, comparisons, gates, online rules for an Environment, and calibration | Evaluations |
| Datasets | CSV and JSONL import, rows built from traces, generated outputs, snapshots and golden sets | Datasets |
| Evaluators | The managed library, your own LLM judges, code evaluators, assertions and collections | Evaluators |
| Simulations | Agents, scenario sets, sandbox runs, scoring and red-team runs | Simulations |
| Playground | Chat, pointwise and pairwise scoring, grids over a dataset, and model compare | Playground |
| Prompts | Versioned prompts with labels, fetched by name from code | Prompts |
| Annotate | Score types, review queues, labels, and RL and preference datasets | Annotate |
| Alerts | Threshold alerts, monitors and the notification bell | Alerts |
| Cost | Model spend and judge spend for an Environment, and the online-rule budget | Cost |
| Settings | Project name, Environments and keys, and Models. For the organization: members, Providers, and plan and usage | Settings |
What needs a plan feature
Home, Traces, Cost and Settings are available to every organization. Evaluations, Datasets, Evaluators, Simulations, Playground, Prompts, Annotate and Alerts are plan features.
When your organization's plan does not include one, its sidebar item stays in place with a lock and opens a panel that explains the feature. Docs pages about a gated feature show the same lock in the docs sidebar. Plan, usage and locks explains the fifteen feature keys.
Where to start
- 1
Create an Environment and copy its key
Signing up already created a Default Project with a development Environment. To add staging and production, see Create an environment and get its key.
- 2
Install the SDK
Call
init()once with the key, in Python or TypeScript. See Install the SDK. - 3
Read your first trace
Find the model call in the waterfall and read its prompt, completion, tokens and estimated cost. See Your first trace.
- 4
Score something
Pick evaluators from the managed library and judge your recent traces. See Your first evaluation.
To trace a coding agent instead of an application, start with Trace a coding agent.