[

LLM Gateway

]

One API key. 14 providers. 170+ models.

The Kontinent Gateway speaks the OpenAI API and sends every request to the cheapest model that can still solve it. All providers in the EU, with DPA, failover, and full cost transparency.

  • Border Shape
  • Border Shape

[

THE LLM GATEWAY

]

Every request finds the cheapest model that solves it.

You define a sequence of models. The gateway starts at the cheapest level and only escalates if a check fails. Your own trained model can be the first level.

Gateway

Routes

API keys

Models

Catalog

Your models

Provider

Deployments

Endpoints

Autoscaling

Governance

Audit log

Policies

Documentation

Gateway / Routes / prod-eu

EU · Frankfurt

Zero Data Retention

prod-eu

OpenAI-compatible

Cascade active

One route, 14 providers, one key. Fallback and failover are part of the route.

Edit route

Cost per 1M tokens

−74%

86% of requests resolve at Level 1

Latency p50

380 ms

across all levels

Availability

99.95%

Failover via 14 providers

Routing cascade

Order of models per request

Last 24 hours

Level

Model

Provider · Region

Share

€ / 1M

Level 1

claims-triage · Your model

Kontinent · Frankfurt

86%

€0.40

Level 2

Mistral Large

Mistral · Paris

11%

€2.00

Level 3

Llama 3.3 70B

IONOS · Berlin

3%

€0.70

Failover: in case of timeout or error, the next provider of the same level takes over, without retry in the client.

Records leaving the EU: 0

Zero Data Retention: aktiv

Export route

  • Border Shape
  • Border Shape

[

WHAT THE GATEWAY DOES

]

Four things you would otherwise have to build yourself.

Routing, failover, billing, and governance across 14 providers. Once in the gateway instead of 14 times in your code.

API

An OpenAI-compatible interface.

Swap the base URL, set the key, and you're done. Your SDKs, agent frameworks, and tools continue to run unchanged, including streaming, tool calls, and structured outputs.

API

An OpenAI-compatible interface.

Swap the base URL, set the key, and you're done. Your SDKs, agent frameworks, and tools continue to run unchanged, including streaming, tool calls, and structured outputs.

ROUTING

Cascade routing instead of one model for everything.

Every request starts at the most cost-effective tier that can reliably resolve it. Only if a check fails does the gateway escalate to the next tier.

ROUTING

Cascade routing instead of one model for everything.

Every request starts at the most cost-effective tier that can reliably resolve it. Only if a check fails does the gateway escalate to the next tier.

FAILOVER

Eleven providers behind one route.

If a provider fails or rate limits are reached, the next one in the same region takes over. Your client won't notice a thing.

FAILOVER

Eleven providers behind one route.

If a provider fails or rate limits are reached, the next one in the same region takes over. Your client won't notice a thing.

CONTROL

Costs, latency, and tokens per route.

You can see what is being spent by team, route, and model. You set budgets, rate limits, and policies in the gateway, not in ten different provider consoles.

CONTROL

Costs, latency, and tokens per route.

You can see what is being spent by team, route, and model. You set budgets, rate limits, and policies in the gateway, not in ten different provider consoles.

  • Border Shape
  • Border Shape

[

THE PROVIDERS

]

Eleven European providers behind one route.

Ten European providers behind one route.

All under EU law, all with data processing agreements (DPA), all accessible via the same key. If a new provider is added, nothing changes in your integration.

Inceptron
Inceptron

Sweden

Infercom

Luxembourg

IONOS

Germany

Mistral AI

France

Nebius
Nebius

Netherlands

Nebula
Nebula

Netherlands

OVHcloud
OVHcloud

France

Scaleway
Scaleway

France

Regolo AI

Italy

TensorX
TensorX

Ireland

  • Border Shape
  • Border Shape

[

SAVINGS

]

Calculate how much open models can save you.

Compare your current API spend with the list prices of open models on European infrastructure.

Which model are you using today?
Which open model would you switch to?
Monthly spend
€/ month
What does your workload look like?
Tokens per request
Input
Output
that’s 3.862.247 Requests per month with this spend
Recommended model
GLM 5.2
High-throughput open model for general-purpose production workloads.

Cost reduction

82 %

Today

150.000 € / mo.

Savings per month

123.449 €

Savings over three years: 7 Mio. €

With Kontinent

26.551 € / mo.

List pricesper 1 million tokens
ModelCached InputInputOutput
GPT 5.6 Sol · Today0,50 €5,00 €30,00 €
GLM 5.2 · Kontinent0,14 €1,40 €4,40 €
How savings grow
Year 1Year 2Year 3
Current trajectory1,8 Mio. €2,7 Mio. €4,1 Mio. €
With Kontinent319 Tsd. €478 Tsd. €717 Tsd. €
Savings1,5 Mio. €2,2 Mio. €3,3 Mio. €

The results are a rough estimate for internal discussions. They are based on public list prices, which are subject to change at any time. Kontinent does not guarantee any specific level of cost savings or other financial benefit.

  • Border Shape

[

HIGH VOLUME?

]

At a certain volume, your own GPUs pay off.

If your traffic stays high, run the same open model on reserved GPUs in the EU: stable latency, no shared traffic, billed by capacity. The OpenAI-compatible API stays the same.

  • Border Shape
  • Border Shape

Questions about the LLM Gateway

How do I integrate the gateway?

Via the OpenAI-compatible API: replace the Base URL and Key, and your existing SDKs and agent frameworks will continue to run without changes. Streaming, tool calls, and structured outputs are supported.

How exactly does Cascade Routing work?

What happens if a provider fails?

Does my data really stay in the EU?

Do I need my own model to use the gateway?

How is billing handled?

When is dedicated serving worth it compared to the gateway?

  • Border Shape
  • Border Shape

[

GET STARTED

]

Speak with an engineer.

We will analyze your traffic, calculate the exact savings, and build your first route with you.

  • Border Shape
  • Border Shape