Docs

    SDK reference

    TypeScript SDK

    syntropylabs-evalkit for Node.js: init, what is traced, framework middleware, tracing your own code, manual spans, identity, offline evaluation, Eval() runs, feedback, prompts, scenario simulation, OpenTelemetry coexistence.

    EvalKit for Node.js is the syntropylabs-evalkit package on npm. Most of this reference is still being written; sessions and agents are documented below.

    Sessions and agents

    A session is one conversation. Every span started inside withSession(...) belongs to that session. An agent is the name you give your agent in code. After setAgent(...), every trace carries that name, and Traces → Agents shows the agent's sessions, versions and verdicts. You never declare a version: the trace service detects it from the model, system prompt, tools, temperature and top-p.

    TypeScript
    import { setAgent, startTrace, withSession } from "syntropylabs-evalkit";
    
    setAgent("support-bot"); // every trace after this belongs to support-bot
    setAgent(null);          // clears
    
    await withSession({ sessionId: "conv_55", userId: "u_8123" }, async () => {
      await agent.reply(message);
    });
    
    const turn = startTrace("chat-turn", { agentName: "support-bot" }); // this trace only
    • When withSession returns, the previous session comes back. On Node the scope follows the async context, so concurrent requests never see each other's session.
    • Spans carry the agent as evalkit.agent_name and gen_ai.agent.name. A span started with { "gen_ai.agent.name": "researcher" } belongs to that sub-agent.
    • A span that already carries session.id, conversation.id, gen_ai.conversation.id or evalkit.session_id keeps its own session.
    • OpenAI and Anthropic model-call spans record temperature, top_p, max tokens and the request body (gen_ai.request.body). From 0.3.2, model-call spans also list the tools the request offered in gen_ai.request.tools, for OpenAI, Anthropic, Bedrock, Google, Vertex AI, Cohere, Ollama and LangChain. The trace service reads the tools from that attribute and falls back to the request body, so a tool change gives a new version.
    • From 0.3.1, a tool call the model only requested is a gen_ai.tool.call event on its model-call span. The tool your code runs through traceTool is the tool_call span, so each call is counted once. When the provider runs a tool itself and returns its result, that tool still gets its own tool_call span.

    EvalKit is built by SyntropyLabs. Published on PyPI and npm.