[
Dedicated Inference
]
Dedicated GPUs in the EU,
for any open model.
Qwen, Llama, Mistral or your own model on reserved GPUs in the EU, behind an OpenAI-compatible endpoint. No shared traffic, stable latency, billed by capacity rather than per token. You don't need to train anything first.
[
DEDICATED INFERENCE
]
From an open model to your own endpoint.
You pick a model, we start it on reserved GPUs in the EU and configure it for your workload. A newer version or a model of your own goes behind the same endpoint later.
Deployments / Endpoints / qwen3-32b-prod
EU · Frankfurt
Dedicated
qwen3-32b-prod
Qwen3 32B
Live
Reserved GPUs in Frankfurt, OpenAI-compatible endpoint, rollback anytime.
Switch model
p50 latency
240 ms
p99 610 ms in interactive deployment
Throughput
1,400 Tok/s
across all replicas
Availability
99.95%
last 30 days
Deployments
Deployment
Purpose
Utilization
Prod
Interactive
62%
Batch
Nightly batch
88%
Canary
A/B on 5%
11%
New model versions go behind the same endpoint. If a metric drops in the canary, the deployment rolls back automatically.
Traces captured: 1.2 million per week
Zero Data Retention: active
Export traces
[
DEDICATED INFERENCE
]
Designed for your workload, not the average.
Open model or your own: it runs on reserved GPUs in the EU, and every deployment is configured for a latency and throughput target rather than a default.
Serving configuration per workload.
Batch size, context length, and concurrency are tuned to meet your latency and throughput goals. Interactive applications and batch jobs receive separate deployments of the same model.
Autoscaling along your usage curve.
Capacity follows demand, with reserved GPUs as a baseline and defined upper and lower limits. You do not pay for peak loads that never occur.
Region and isolation per deployment.
You determine the region during the design phase. Single-tenant and regional binding for data residency requirements, with zero data retention available upon request.
Availability targets and monitoring.
SLAs tailored to your product requirements, with alerting on latency, error rate, and saturation before your users notice.
[
NOT ENOUGH TRAFFIC YET?
]
You can run open models in the EU without your own GPUs.
Until reserved capacity pays off, use Llama, Mistral or Qwen through the Kontinent gateway: European providers, one API key, billed per token. When your volume grows, switch to dedicated GPUs through the same OpenAI-compatible API.
[
OPTIONAL: TRAINING
]
Your traffic can become your own model later.
Not required, but an option: running an open model on dedicated GPUs already gives you the data for a specialised model. Only what you approve is captured.
Questions about dedicated inference
Which models can I run on dedicated GPUs?
Open models like Qwen, Llama, Mistral or DeepSeek, and equally a model you have trained yourself. We start it on reserved GPUs in the EU, and you call it through an OpenAI-compatible API.
Do I need to train a model to use dedicated inference?
How is a deployment configured for my workload?
Do I pay per token or per GPU?
When is dedicated worth it over the gateway?
What happens when I switch to a new model version?
Can I move to my own model later?
[
GET STARTED
]
Speak with an engineer.
We look at your workload, pick the right model with you and size the deployment.