[
Training
]
Your data. Your evals. Your model.
Post-training on your real cases, measured by evals defined by your professionals. Trained and hosted in the EU, at a fraction of the cost of a frontier API, and the weights are yours.
[
KONTINENT TRAINING
]
Every run visible, every checkpoint measurable.
Datasets, training runs, evals, and checkpoints in one place. See at any time what the model has learned, how much it costs, and whether it beats the baseline.
Runs / Training / claims-triage-08
EU · Frankfurt
Run 08
claims-triage-08
qwen3-32b · 8 Traces
Running · Step 15 of 24
Grader: caseworker gold, 4 tasks per step, eval every 5 steps against the same frozen set
Pause run
Grader Score
0.502
+0.196 since step 0
GPU hours
128
H200, Frankfurt
Tokens consumed
8.4M
in 15 steps
Grader score per step
Step
Grader
Tokens out
eval@0
0.306
3,773
train@0
0.362
15,197
train@5
0.460
9,497
eval@9
0.445
6,096
eval@14
0.495
5,915
train@15
0.502
11,351
Records leaving the EU: 0
Last checkpoint: v15 saved
View checkpoints
[
EVAL-GATE
]
The evaluation set is available before the first training session.
It is created from your real cases and then frozen. After that, any progress is measurable and every claim is verifiable.
Training takes place in your harness.
The learning environment is the environment in which the model will operate later: your tools, your prompts, your formats. What matters in training also matters in production.
Only what beats the baseline is rolled out.
Every checkpoint runs against the same set. New versions go behind the same endpoint, allowing rollbacks at any time without affecting your integration.
[
THE PROCESS
]
From the first measurement to your own model.
No proof of concept that leads nowhere. Four steps with an exit criterion after the second.
Questions about training
[
LET'S GO
]
Start with the eval set.
In a consultation, we will clarify whether your task can be trained, what data is required for it, and what it costs to run your own model.