[

Training

]

Your data. Your evals. Your model.

Post-training on your real cases, measured by evals defined by your professionals. Trained and hosted in the EU, at a fraction of the cost of a frontier API, and the weights are yours.

  • Border Shape
  • Border Shape

[

KONTINENT TRAINING

]

Every run visible, every checkpoint measurable.

Datasets, training runs, evals, and checkpoints in one place. See at any time what the model has learned, how much it costs, and whether it beats the baseline.

Runs

Training

Evals

Project

Datasets

Models

Checkpoints

Deployments

Endpoints

Optimization

Governance

Audit log

Export

Documentation

Runs / Training / claims-triage-08

EU · Frankfurt

Run 08

claims-triage-08

qwen3-32b · 8 Traces

Running · Step 15 of 24

Grader: caseworker gold, 4 tasks per step, eval every 5 steps against the same frozen set

Pause run

Grader Score

0.502

+0.196 since step 0

GPU hours

128

H200, Frankfurt

Tokens consumed

8.4M

in 15 steps

Grader score per step

Train vs. Eval, same frozen set

Run 08 · live

Step

Grader

Tokens in

Tokens out

eval@0

0.306

659,945

3,773

train@0

0.362

4,093,720

15,197

train@5

0.460

253,060

9,497

eval@9

0.445

131,746

6,096

eval@14

0.495

123,951

5,915

train@15

0.502

198,135

11,351

Records leaving the EU: 0

Last checkpoint: v15 saved

View checkpoints

  • Border Shape
  • Border Shape

[

EVAL-GATE

]

Before the training, measurements are taken.

We measure before we train.

The evaluation set is available before the first training session.

It is created from your real cases and then frozen. After that, any progress is measurable and every claim is verifiable.

Training takes place in your harness.

The learning environment is the environment in which the model will operate later: your tools, your prompts, your formats. What matters in training also matters in production.

Only what beats the baseline is rolled out.

Every checkpoint runs against the same set. New versions go behind the same endpoint, allowing rollbacks at any time without affecting your integration.

  • Border Shape
  • Border Shape
  • Border Shape

[

THE PROCESS

]

From the first measurement to your own model.

No proof of concept that leads nowhere. Four steps with an exit criterion after the second.

01

Evaluation Set

We gather real cases with your specialists, define what is considered correct, and freeze the set.

Week 1 to 2

01

Evaluation Set

We gather real cases with your specialists, define what is considered correct, and freeze the set.

Week 1 to 2

02

Baseline

Your current setup and a strong open model are running against the same set. Only after that will we talk about training.

Week 2

02

Baseline

Your current setup and a strong open model are running against the same set. Only after that will we talk about training.

Week 2

03

Training

Post-training on your data, in your harness, on dedicated GPUs in the EU. Every checkpoint is verified.

Week 3 to 6

03

Training

Post-training on your data, in your harness, on dedicated GPUs in the EU. Every checkpoint is verified.

Week 3 to 6

04

Deployment

The first checkpoint that beats the baseline goes to an endpoint. After that, the improvement cycle begins.

from week 6

04

Deployment

The first checkpoint that beats the baseline goes to an endpoint. After that, the improvement cycle begins.

from week 6

  • Border Shape
  • Border Shape

Questions about training

What data do you need for training?

The most important are real cases with a decision: processed transactions, verified documents, corrected answers. A few hundred well-annotated cases are often sufficient for the evaluation set, and a few thousand for training. What exactly is needed will be clarified by the baseline audit before any data is moved.

Fine-tuning, distillation or RL: which do you use?

Which base model are you training on?

What happens if the model does not beat the baseline?

Can personal data be used for training?

How long until the first model?

Who owns the weights and eval set?

What data do you need for training?

The most important are real cases with a decision: processed transactions, verified documents, corrected answers. A few hundred well-annotated cases are often sufficient for the evaluation set, and a few thousand for training. What exactly is needed will be clarified by the baseline audit before any data is moved.

Fine-tuning, distillation or RL: which do you use?

Which base model are you training on?

What happens if the model does not beat the baseline?

Can personal data be used for training?

How long until the first model?

Who owns the weights and eval set?

  • Border Shape
  • Border Shape

[

LET'S GO

]

Start with the eval set.

In a consultation, we will clarify whether your task can be trained, what data is required for it, and what it costs to run your own model.

  • Border Shape
  • Border Shape