Syntropylabs. Your agent evaluation copilot.
From one line of code to full-stack evals.
pip install syntropylabs-evalkitOne line, fully instrumented.
stl.init() patches every LLM client, HTTP layer, and DB adapter your agent touches. No manual spans.
Every call becomes a trace.
Tool calls, LLM calls, DB reads — stitched into one waterfall per run, not fourteen disconnected logs.
Every output gets scored.
LLM-as-judge evaluators run automatically on every trace — online in production, or offline in batch.
Regressions get caught before ship.
Compare pass rates across deploys. Know a prompt change hurt quality before your users do.
One connected pipeline, not five tools
The SDK feeds the harness. The harness feeds the dashboard. Nothing to stitch together yourself.
SDK
One init() call instruments your entire agent stack — LLM clients, HTTP, SQL, logging. Python and TypeScript, zero deps.
- Auto-patches OpenAI, Anthropic, Gemini, Bedrock
- W3C traceparent propagation
- Session + device context
Agent Harness evaluation
The evaluation core. Run batch jobs, score traces online, compare models pairwise — text, voice, image, and video.
- Online + batch evaluation
- Voice agent eval via LiveKit
- Image / video quality scoring
Dashboard
Trace waterfall, session explorer, eval results, cost breakdown, and model catalog — everything in one place.
- Trace waterfall per agent run
- Per-session cost analytics
- Regression comparison across deploys
Why it's different
Deep evaluations
Grounded tool-use checks catch fabricated actions and hallucinated calls an LLM judge alone would miss — not just pass/fail.
Built for real agents
Full parent/child span tree with self-time aware latency profiling — built for multi-step, multi-turn, tool-calling agents.
Any LLM, any framework
OpenAI, Anthropic, Gemini, Bedrock, and more — auto-patched. Zero vendor lock-in, zero manual instrumentation.
Private by design
Your traces stay in your project. On-premise deployment and SSO/SAML available for teams that need it.
The complete evaluation stack
Trace every span, score every output, catch every regression — from prototype to production.
Agent Tracing
Every LLM call, tool invocation, HTTP hop, and DB query lands as a typed span — one waterfall per run.
Zero-Config SDK
One init() call. All providers patched automatically — plus a central model catalog to manage access across all of them.
Online Evaluation
Every production trace scored automatically as it arrives. No cron jobs.
Voice Agent Eval
Generate audio, score naturalness + instruction-following, replay rows. Connect live in-browser.
Batch Evaluation
Run a dataset through your agent, score per row, regenerate failures individually.
Image & Video Eval
Generate, then LLM-as-judge score across visual quality dimensions.
Evaluator Collections
Group LLM-as-judge prompts into reusable collections. Apply to any project or set globally. Three rule types: model-graded, custom prompt, statistical.
One call.
Everything Tracked.
init() patches every LLM client, HTTP layer, and database adapter your agent touches. Traces start flowing in seconds — no manual spans, no config files.
- pip install syntropylabs-evalkit / npm install syntropylabs-evalkit
- Auto-instruments OpenAI, Anthropic, Gemini, Bedrock, HTTP, SQL, Redis, Mongoose
- Online eval: attach evaluators to a project, every trace gets scored automatically
- W3C traceparent — frontend and backend stitched into one trace, no extra work
Simple, transparent pricing
Start free. Scale up when you need it.
Ship AI agents
you can trust.
Instrument in minutes. Score every output. Catch failures before your users do.
