Docs

    Get started

    Concepts

    Organization, Project, Environment, trace, span, run, evaluator, dataset, online rule and the other objects the app uses, with their API and former names.

    Each term below is the word the app uses. Where the API or an older version of the product used another name, the last column gives it, so old links and scripts still make sense.

    Organization, project and keys

    TermWhat it isAPI or former name
    OrganizationThe account. Owns members and their roles, Provider connections, evaluators, collections and a plan (free, pro or enterprise). Created automatically at signup.orgs
    ProjectWhat you are building. Owns Environments, datasets, runs, Models, prompts, agents, simulation runs, alerts, monitors, review queues and topics.projects
    EnvironmentOne deployment of the project (development, staging or production) with its own environment key and its own traces. Traces, Online rules, Alerts and Cost read one Environment at a time, picked with the switcher in the top bar.trace project (/trace-projects), addressed by tenantId in trace read URLs
    Environment keyThe secret the SDK sends traces with: tk_live_…. One per Environment. It stays copyable in Settings and can be rotated.subscription key (subscription_key, X-Subscription-Key)
    ProviderAn LLM vendor credential, stored once on the Organization. It is validated with a real call and never returned to the browser. Models are picked from it.provider-connections
    ModelA project-level pointer to a provider model (provider, modelId, parameters, modalities). Every model dropdown in the app lists Models, including the judge, agent model, Playground and generation pickers.model configuration (/models, modelConfigId)

    Tracing

    TermWhat it isAPI or former name
    TraceOne request or one agent turn: a tree of spans under one trace id. Stored in the trace store under its Environment.traceId
    SpanOne timed operation in a trace. Types: llm_call, tool_call, http_call, function_call, db_query, log, eval_result, and agent for a coding-agent turn or subagent.spanType
    SessionAll traces that share a sessionId, which is one conversation. Your code sets it through the SDK. The Sessions view lists sessions with their turns, errors, tokens and last seen time.session_id, sessionId
    UserAll traces that share a userId, which is one end user. The Users view rolls them up. A deviceId groups traffic before sign-in.user_id, userId, device_id
    Agent versionA fingerprint of an llm_call's system prompt, tool names, model and sampling parameters, computed at ingest. The same configuration gets the same id. A changed prompt, tool set or model is a new version.agentVersion
    Content captureWhether the SDK sent prompts, completions, arguments and results. With capture off (capture_content=False), spans keep tokens, latency and cost and are marked evalkit.content_captured=false. The trace page then says how to turn capture on.capture_content, captureContent

    Evaluation

    TermWhat it isAPI or former name
    EvaluatorOne check, owned by the Organization. Its kind is custom_prompt or statistical (LLM judges), assertion (deterministic) or code (Python run in a sandbox). Managed evaluators are maintained by the platform. Duplicate one to edit its prompt. Evaluators covers scopes, thresholds, reference modes and output types.evaluation rule (/evaluation-rules, ruleId)
    CollectionA named set of evaluators, each with a weight, a threshold and a hard flag, plus an aggregation (weighted, mean or min). Runs, online rules, simulation scoring and the Playground grid use collections.evaluator-collections
    RunOne scoring of a set of rows: the dataset, the Model that produced the outputs, the evaluators and judge, a gate and a baseline. source: ui when started in the app, source: sdk when posted by Eval().evaluation job (/evaluation-jobs, jobId)
    ComparisonTwo completed runs on the same dataset version, row by row: improved, regressed, unchanged or indeterminate. Computed only when the dataset hash or snapshot matches.GET /evaluation-jobs/:id/compare/:baselineId
    GateA CI verdict on a run: minimum pass rate, minimum average score, per-evaluator minimums and maximum regressions against the baseline. It passes, fails, or is indeterminate when a check could not be decided.gate, gateResult
    Online ruleContinuous evaluation of an Environment's live traces: evaluators, judge, interval, sampling rate, filters, scope and a daily judge budget. One per Environment.online eval config (/playground/online-eval/:traceProjectId)
    Trace evaluationThe scores an evaluation left on a trace, manual or from an online rule, with a reason per evaluator and the judge cost. Stored in the control plane, not in the trace store, so the trace list joins them in the browser.TraceEvaluation
    Tri-state scoreEvery score cell is a number, no verdict or judge error. No verdict means the judge could not decide: a reference was missing, the reply could not be parsed, or the evaluator does not support the input. Judge error means the judge call failed. Neither is ever drawn as zero.status: ok, unparseable, judge_error, skipped, unsupported
    CalibrationHow often a judge agrees with humans: Cohen's κ, Pearson, precision, recall and a confusion matrix, for each evaluator and evaluator version. Not computed below five pairs./judge-calibration

    Datasets, review and simulation

    TermWhat it isAPI or former name
    DatasetProject-scoped rows (input, target, output, system instructions, extra columns) of type text, conversation or voice. Its content hash over input and target decides whether two runs are comparable./datasets
    Row statusdraft → finalized → golden, or archived. Golden means a person verified the expected output. A golden row needs a non-empty target.dataset_rows.status
    SnapshotAn immutable, named version of a dataset that a run can pin to. A run pinned to a snapshot is compared by snapshot, not by hash./datasets/:id/snapshots
    Golden datasetA dataset whose targets a person verified. It is imported or exported as JSONL in the Gemini contents or OpenAI messages shape.POST /datasets/import, GET /datasets/:id/export
    ScoreOne typed judgment about a trace, span, session, run or row: numeric, categorical, boolean or text. Its source is human, judge, deterministic, sdk or end_user. End-user thumbs and ratings sent from code are scores./scores, /v1/feedback
    Score typeA project's definition of what reviewers score: kind, range or categories, and whether a text answer is written to the row's expected output.score-configs
    Review queueA list of traces, sessions or rows to review. It has an assignment mode (unassigned, single, round robin or random) and can have rules that add live traces automatically./review-queues
    LabelThe rubric a reviewer leaves on a trace: four sliders from 1 to 5 (correctness, helpfulness, efficiency, safety) and a note. It derives a reward from 0 to 1 and feeds calibration and RL datasets.RL label (/rl-labels)
    Agent (simulation)What a simulation drives: nocode (instructions, tools, context), code (a Python entrypoint(ctx)) or connector (vendor connectors, answered by a stateful world simulator that knows the real vendor tool schemas)./agents
    Scenario setA versioned, immutable set of generated or hand-written multi-turn test conversations for an agent.scenario_sets
    Simulation runOne sandbox launch of an agent over a scenario set in an Environment. Every scenario becomes a trace tagged as a simulation./simulation-runs
    PromptA named prompt with immutable numbered versions and movable labels (production, staging). Code fetches it by name, and it is stamped on the spans it drives./prompts

    EvalKit is built by SyntropyLabs. Published on PyPI and npm.