Docs

    Help

    FAQ

    Short answers to the questions people ask most often about EvalKit.

    What is EvalKit?

    EvalKit is the SyntropyLabs SDK for tracing, simulating and evaluating AI agents and LLM applications. It is available for Python and TypeScript (Node.js).

    How do I install the EvalKit SDK?

    For Python, run "pip install syntropylabs-evalkit" and import evalkit. For TypeScript (Node.js), run "npm install syntropylabs-evalkit". Then call init() once, as early as possible, with your environment key.

    What gets traced automatically?

    In Python, evalkit.init() patches the LLM SDKs you have installed, among them OpenAI, Anthropic, Bedrock, Google, Cohere, Mistral and Groq. It also traces HTTP clients and database drivers, and records the tool calls the model makes. By default it traces your own code too: every function in the modules under the folder of the file that calls init(). Pass function_tracing=False to turn that off. The TypeScript SDK patches the same kinds of libraries, except Mistral, and traces your own functions only where you wrap them, for example with traceFunction.

    Can I bring my own LLM API key (BYOK)?

    Yes. In the SDK, scenario generation and simulation scoring take provider, model and api_key in Python, or apiKey in TypeScript. Generation falls back to a hosted default model when you pass no key, and scoring always needs one. Any provider in the catalog works. Runs you start in the app need no key in code, because they use the Provider connected to your organization.

    Where do provider API keys live in the dashboard?

    On the organization. Connect a Provider once in Settings → Organization → Providers, and every project in the organization can use its models. A project can still point a Model at a different Provider when it needs to bill a separate account. Keys are encrypted at rest and are never returned to the browser.

    Which providers are supported, and how current is the model list?

    The catalog lists 49 providers. The model labs include OpenAI, Anthropic, Google, xAI, Mistral, Cohere, DeepSeek, Moonshot (Kimi), Z.ai (GLM), MiniMax, Qwen and Perplexity. The clouds include AWS Bedrock, Vertex AI, Azure OpenAI, Databricks, watsonx and SageMaker. Aggregators and fast-inference hosts include OpenRouter, Together, Fireworks, Groq, Cerebras, SambaNova, DeepInfra and Nebius. For self-hosted models there are Ollama, vLLM, LM Studio and any OpenAI-compatible base URL. Voice and media vendors include ElevenLabs, Deepgram, AssemblyAI, Cartesia, Replicate and fal.ai. Bedrock takes four kinds of credentials: IAM access keys, a Bedrock API key, an assumed role or instance profile, and a custom endpoint. Model lists, context windows and prices re-sync from the LiteLLM price sheet every 12 hours, so a model becomes selectable within 12 hours of LiteLLM listing it.

    What's the difference between offline and online evaluation?

    Offline evaluation runs when you start it, on traces you select in the app or by calling evalkit.evaluate() in code. Online evaluation is an online rule on an Environment. Every 1, 5, 15, 30 or 60 minutes it scores the Environment's new traces, either all of them or a sampled share.

    How does scenario simulation work?

    EvalKit generates multi-turn user scenarios from your agent's system prompt and tools. Then simulate_user in Python, or simulateUser in TypeScript, plays each scenario against your real agent one turn at a time, and the results are scored.

    Where do I view traces and simulation results?

    In the app. Traces lists every trace, and a trace's Waterfall tab shows all of its spans. Simulations shows each run with its scenario scores, per-turn traces, tool trajectory and LLM-judge ratings.

    Why does my trace show several scores instead of one?

    The trace held more than one conversation. Either its spans carried several session ids, or a simulation ran several scenarios under one root. Each conversation is scored on its own, as a part. The trace itself gets no score, because an average of separate conversations describes none of them. The trace list shows parts passed out of parts scored, such as 2/3, and each session's page shows its own part on that turn.

    Why did every agent get a new version after upgrading?

    A version is a fingerprint of the model, system prompt, tools, temperature and top-p. The Python SDK from 0.3.1 and the TypeScript SDK from 0.3.2 record the tools and sampling parameters on each model call, and the trace service also reads them from the request body. Calls that were fingerprinted without them now include them. That starts one new version entry for each agent configuration, once. The earlier entry stays listed for the time it was seen, and nothing in your agent changed.

    Do I need to self-host anything to use EvalKit?

    No. By default the SDK sends traces to the hosted ingest endpoint and calls the hosted API, both at https://api.syntropylabs.ai. To self-host, set base_url and api_url in Python, or baseUrl and apiUrl in TypeScript, when you call init(). The Configuration reference describes both.

    What languages does the SDK support?

    Python and TypeScript (Node.js). Both are published as syntropylabs-evalkit, on PyPI and npm, and in Python the import name is evalkit. Both cover tracing, evaluation and scenario simulation. The install guide in the app also has Go and Java tabs with snippets that create spans by hand.

    EvalKit is built by SyntropyLabs. Published on PyPI and npm.