Docs

    SDK reference

    Python SDK

    EvalKit for Python: init, what is traced, tracing your own code, manual spans, identity, offline evaluation, Eval() runs, feedback and rewards, prompts and scenario simulation.

    EvalKit for Python is the syntropylabs-evalkit package, imported as evalkit. Most of this reference is still being written; sessions and agents are documented below.

    Sessions and agents

    A session is one conversation. Every span opened inside evalkit.session(...) belongs to that session. An agent is the name you give your agent in code. After evalkit.set_agent(...), every trace carries that name, and Traces → Agents shows the agent's sessions, versions and verdicts. You never declare a version: the trace service detects it from the model, system prompt, tools, temperature and top-p.

    Python
    import evalkit
    
    evalkit.set_agent("support-bot")                 # every trace after this belongs to support-bot
    
    with evalkit.session("conv_55", user_id="u_8123"):
        answer = agent.run(message)
    
    @evalkit.session("conv_55")                      # also a decorator, sync or async
    async def handle_turn(message): ...
    
    trace_id, end, ctx = evalkit.start_trace("turn", agent="planner")                   # one trace only
    end_step, _ = evalkit.start_span("research", {"gen_ai.agent.name": "researcher"})  # a sub-agent
    • On exit, the session and user that were in force before the block come back. Concurrent requests and tasks never see each other's session. user_id is optional.
    • set_agent is scoped like set_user, per thread and per asyncio task. It returns a token that evalkit.clear_agent(token) uses to restore the previous name. evalkit.current_agent() reads the current name.
    • Spans carry the agent as evalkit.agent_name and gen_ai.agent.name. A span that already carries session.id, conversation.id, gen_ai.conversation.id or evalkit.session_id keeps its own session.
    • From 0.3.1, OpenAI chat, OpenAI Responses and Anthropic calls record the names of the tools they were given in gen_ai.request.tools, built-in tools such as web_search included. They also record temperature, top_p and the max-tokens setting, so a change of tools or sampling shows as a new version. An Anthropic system prompt given as a list of blocks is recorded as its text. Right after the upgrade, each agent configuration starts one new version, because these parameters are now part of it.

    EvalKit is built by SyntropyLabs. Published on PyPI and npm.