Model Selection

    Run pairwise comparisons across model configurations and build a decision log of which model won, and why.

    1

    Add candidate models

    Configure each model you’re considering from the dashboard.

    2

    Run pairwise comparisons

    Use Playground → Pairwise to run the same prompt against two models with an LLM-as-judge picking a winner and giving reasoning.

    3

    Or batch-compare on a dataset

    Run a batch evaluation with multiple models against the same dataset and rules for a broader signal than a handful of prompts.

    4

    Check the leaderboard

    Cross-reference with the dashboard’s built-in benchmark leaderboard for latency and throughput, not just quality.

    EvalKit is built by Syntropylabs. Published on PyPI and npm.