[

EVALS GUIDE

]

Learn LLM evals: the free guide.

17 chapters on finding out whether your AI feature works and whether a change made it better. From the first error analysis to evals in CI, sourced throughout and without a vendor's tool ranking.

  • Border Shape
  • Border Shape

[

17 CHAPTERS

]

What's in the guide.

What's in the guide.

Written for engineers running LLM features in production. Free to read, no sign-up. Seven of the 17 chapters:

  • Border Shape
  • Border Shape

[

THE OUTCOME

]

You end up with a table you can trust.

You end up with a table you can trust.

That is where the guide takes you: every model against the same frozen set, quality and cost side by side, and a gate that decides what ships. Among the models that pass, the cheapest one wins.

Classifying insurance claims · 412 cases

Gate: quality ≥ 92%

Model

Quality

p50 latency

€ per 1,000 tasks

€ / 1,000

Gate

Frontier API (today)

Baseline

95.1%

2.1 s

€8.40

PASS

Mistral Large

93.4%

1.5 s

€2.20

PASS

Qwen3 32B

Recommended

92.6%

0.9 s

€0.70

PASS

Llama 3.3 70B

91.2%

1.1 s

€0.80

FAIL

Gemma 3 27B

88.0%

0.8 s

€0.40

FAIL

Illustrative example. Your report shows your task, your set and your models.

  • Border Shape
  • Border Shape

[

WITH KONTINENT

]

Rather not do it alone? We build the eval set with you.

The eval set comes from your operations.

Real cases, checked by the people who decide today what counts as correct. The set is frozen so every number stays comparable later.

Graders you can trust.

Where an answer can be checked in code, code checks it. Where an LLM judge is needed, it is calibrated against human judgments before it decides anything about a model.

The cheapest model that clears the bar.

The gateway puts 170+ models behind one OpenAI-compatible API. Each of them can be measured against your set, so you choose with numbers instead of gut feeling.

  • Border Shape
  • Border Shape

Questions about LLM evals

What are LLM evals?

Systematic tests that measure whether an AI application does its job. Unlike public benchmarks, they test your application on your cases: your prompts, your documents, your tolerance for errors. Without them you cannot tell whether a change made anything better.

Why aren't public benchmarks enough?

What LLM evaluation tools are there?

How many test cases does an eval set need?

Can an LLM act as a judge?

Is the Evals Guide free?

What do evals have to do with cost?

What are LLM evals?

Systematic tests that measure whether an AI application does its job. Unlike public benchmarks, they test your application on your cases: your prompts, your documents, your tolerance for errors. Without them you cannot tell whether a change made anything better.

Why aren't public benchmarks enough?

What LLM evaluation tools are there?

How many test cases does an eval set need?

Can an LLM act as a judge?

Is the Evals Guide free?

What do evals have to do with cost?

  • Border Shape
  • Border Shape

[

GET STARTED

]

Start with chapter one.

Or talk to us if you would rather not build your eval set alone. We build it with your domain experts and use it to find the model that solves your task at the lowest cost.

  • Border Shape
  • Border Shape