[

Dedicated Inference

]

Dedicated GPUs in the EU,
for any open model.

Qwen, Llama, Mistral or your own model on reserved GPUs in the EU, behind an OpenAI-compatible endpoint. No shared traffic, stable latency, billed by capacity rather than per token. You don't need to train anything first.

  • Border Shape
  • Border Shape

[

DEDICATED INFERENCE

]

From an open model to your own endpoint.

You pick a model, we start it on reserved GPUs in the EU and configure it for your workload. A newer version or a model of your own goes behind the same endpoint later.

Deployments

Endpoints

Checkpoints

Model

Overview

Evals

Training

Traffic

Traces

Retraining

Governance

Audit log

Policies

Documentation

Deployments / Endpoints / qwen3-32b-prod

EU · Frankfurt

Dedicated

qwen3-32b-prod

Qwen3 32B

Live

Reserved GPUs in Frankfurt, OpenAI-compatible endpoint, rollback anytime.

Switch model

p50 latency

240 ms

p99 610 ms in interactive deployment

Throughput

1,400 Tok/s

across all replicas

Availability

99.95%

last 30 days

Deployments

Same model, three configurations

Frankfurt Region

Deployment

Purpose

Hardware

Utilization

Replicas

Prod

Interactive

4× H200

62%

3

Batch

Nightly batch

2× H200

88%

2

Canary

A/B on 5%

1× H200

11%

1

New model versions go behind the same endpoint. If a metric drops in the canary, the deployment rolls back automatically.

Traces captured: 1.2 million per week

Zero Data Retention: active

Export traces

  • Border Shape
  • Border Shape

[

DEDICATED INFERENCE

]

Designed for your workload, not the average.

Open model or your own: it runs on reserved GPUs in the EU, and every deployment is configured for a latency and throughput target rather than a default.

Serving configuration per workload.

Batch size, context length, and concurrency are tuned to meet your latency and throughput goals. Interactive applications and batch jobs receive separate deployments of the same model.

Autoscaling along your usage curve.

Capacity follows demand, with reserved GPUs as a baseline and defined upper and lower limits. You do not pay for peak loads that never occur.

Region and isolation per deployment.

You determine the region during the design phase. Single-tenant and regional binding for data residency requirements, with zero data retention available upon request.

Availability targets and monitoring.

SLAs tailored to your product requirements, with alerting on latency, error rate, and saturation before your users notice.

  • Border Shape

[

NOT ENOUGH TRAFFIC YET?

]

You can run open models in the EU without your own GPUs.

Until reserved capacity pays off, use Llama, Mistral or Qwen through the Kontinent gateway: European providers, one API key, billed per token. When your volume grows, switch to dedicated GPUs through the same OpenAI-compatible API.

  • Border Shape
  • Border Shape

[

OPTIONAL: TRAINING

]

Your traffic can become your own model later.

Not required, but an option: running an open model on dedicated GPUs already gives you the data for a specialised model. Only what you approve is captured.

01

Deliver.

Every request runs through your endpoint. If you choose, prompts, responses and tool outputs are captured, in the EU and under your control.

01

Deliver.

Every request runs through your endpoint. If you choose, prompts, responses and tool outputs are captured, in the EU and under your control.

02

Improve.

Real traffic generates training data: self-distillation against error cases, preference optimization on corrections, reinforcement learning for new skills.

02

Improve.

Real traffic generates training data: self-distillation against error cases, preference optimization on corrections, reinforcement learning for new skills.

03

Redeploy.

The next checkpoint goes to the same endpoint minutes after training. API, routing, access rights, and monitoring remain unchanged.

03

Redeploy.

The next checkpoint goes to the same endpoint minutes after training. API, routing, access rights, and monitoring remain unchanged.

  • Border Shape
  • Border Shape

Questions about dedicated inference

Which models can I run on dedicated GPUs?

Open models like Qwen, Llama, Mistral or DeepSeek, and equally a model you have trained yourself. We start it on reserved GPUs in the EU, and you call it through an OpenAI-compatible API.

Do I need to train a model to use dedicated inference?

How is a deployment configured for my workload?

Do I pay per token or per GPU?

When is dedicated worth it over the gateway?

What happens when I switch to a new model version?

Can I move to my own model later?

  • Border Shape
  • Border Shape

[

GET STARTED

]

Speak with an engineer.

We look at your workload, pick the right model with you and size the deployment.

  • Border Shape
  • Border Shape