News KrakenD 3.0 Is Here: AI Router, Semantic Cache, and On-the-Fly Stream Manipulation

Document updated on Sep 30, 2026

AI Router for LLM Provider Selection

The AI Router selects which LLM provider from your provider inventory handles each request. It matches requests with reusable gates (header or policy conditions), organizes them into an ordered decision tree, and can hand the final choice to a classifier that picks the best provider from a candidate list. A failover guarantees a provider even when nothing matches. Because routes reference providers by name, the client sends its request to a single endpoint and never selects a model or vendor; the router decides.

Requires the provider inventory
Routes and classifiers reference providers by name, so define your providers first in the LLM provider inventory (ai/providers). The router resolves the vendor, model, endpoint, and credentials from that inventory.

How the AI Router Works

You configure routing in two places. Gates and classifiers are shared, reusable definitions declared once at the service level under the ai/classifiers namespace. The routing rules that use them live in a backend’s extra_config under the ai/router namespace, in an ordered routes list.

For each request, KrakenD evaluates the routes from top to bottom and uses the first one that matches. A route matches when all of its gates pass; a route with no gates always matches, so you can use it as an inline fallback. When the matching route names a provider, the router forwards the request to that provider from the inventory. When it names a classifier and a list of providers, the classifier chooses among them. Routes can also nest, forming a decision tree: a gated route holds its own routes, evaluated only when its gates pass.

Route on signals the client cannot forge
Header gates match on request headers, which clients control. Route on values your gateway sets, for example a plan or tier header injected after validating an API key or JWT, so a client cannot select a more expensive model by sending a header itself.

Defining Gates

A gate is a named matching condition declared under ai/classifiers.gates. Routes reference gates by name, and you can reuse the same gate across many routes. Each gate has a name, a rule that selects the matching mechanism (header or policy), and the configuration for that rule.

Header Gates

A header gate matches when a request header equals a value. Set header.name and header.value:

{
  "name": "plan_enterprise",
  "rule": "header",
  "header": {
    "name": "X-Plan",
    "value": "enterprise"
  }
}

The gate above matches requests that carry X-Plan: enterprise.

Policy Gates

A policy gate matches when a CEL expression evaluates to true. Set policy.value to the expression, using helper functions such as hasHeader and getHeader (see the policy built-in functions):

{
  "name": "enterprise_policy",
  "rule": "policy",
  "policy": {
    "value": "getHeader('X-Plan') == 'enterprise'"
  }
}

Policy gates handle conditions a header match cannot express, such as comparing values or combining several checks.

Defining Classifiers

A classifier delegates the provider choice to an external engine that scores the request against a set of candidate providers and returns the best fit. Declare classifiers under ai/classifiers.classifiers. Each has a name, an engine, and the engine’s configuration. The supported engine is notdiamond, which calls the Not Diamond model selection endpoint to pick the best provider:

{
  "name": "smart",
  "engine": "notdiamond",
  "notdiamond": {
    "credentials": "xxx",
    "req_content": "req_body.messages"
  }
}

The credentials authenticate against the Not Diamond service; set them through an environment variable. The req_content field selects the request content sent to the classifier, and it works exactly like the same field in the Semantic Cache. It defaults to the full req_body, and you can point it at any request input source:

  • req_body: the request body. Use dot-notation to reach nested fields, for example req_body.messages.
  • req_params: a URL path parameter. The key is capitalized, so {id} is exposed as req_params.Id.
  • req_query_string: a query string value, for example req_query_string.user.
  • req_headers: a request header, for example req_headers.x-user.

The tradeoff option tells Not Diamond whether to favor cost (the default) or latency, while cost_quality_tradeoff balances cost against quality with an integer. They are mutually exclusive: set one or the other, not both.

The following fields are available under the service-level ai/classifiers namespace:

Fields of AI Router Gates and Classifiers
* required fields

classifiers array of objects
The list of classifiers available to the routes. A classifier delegates the provider choice to an external engine that scores the request against the candidate providers of a route and returns the best fit.
Each item of classifiers accepts the following properties:
engine *
The external engine that chooses the provider. The notdiamond engine calls the Not Diamond model selection endpoint, and requires you to add its configuration under the notdiamond key.
Possible values are: "notdiamond"
name * string
A unique name for this classifier. Routes reference the classifier through this value in their classifier field.
Example: "smart"
notdiamond object
The configuration of the notdiamond engine, which calls the Not Diamond model selection endpoint to pick the best provider from the candidates of the route. Set either tradeoff or cost_quality_tradeoff, but not both.
base_url string
The base URL of the Not Diamond API. Override it only when you target a custom or proxied deployment of the service.
Example: "https://api.example.com"
cost_quality_tradeoff integer
An integer that balances cost against quality in the model selection. You cannot use it together with tradeoff.
Defaults to 0
credentials * string
The API key that authenticates KrakenD against the Not Diamond service. Set it through an environment variable instead of writing it in the configuration file.
disable_hash_content boolean
KrakenD asks the Not Diamond API to hash the request content it receives, for extra security. Set this flag to true to turn that option off, and Not Diamond uses its own default of not hashing the content.
Defaults to false
metric
The metric the engine optimizes for when it chooses the provider.
Possible values are: "accuracy"
Defaults to "accuracy"
req_content string
The part of the request that KrakenD sends to the classifier. It works like the same field in the semantic cache: use req_body with dot-notation to reach nested fields, or req_params, req_query_string, and req_headers. Path parameters are capitalized, so {id} becomes req_params.Id. When you don’t set it, KrakenD sends the entire request body.
Examples: "req_body.messages" , "req_params.Id" , "req_headers.x-user" , "req_query_string.user"
tradeoff
The optimization preference of the model selection. Use cost to favor cheaper providers or latency to favor faster ones. You cannot use it together with cost_quality_tradeoff.
Possible values are: "cost" , "latency"
Defaults to "cost"
gates array of objects
The list of gates available to the routes. A gate is a named matching condition, and a route matches only when all the gates it lists pass. Header gates match on values the client controls, so route on headers your gateway sets after validating an API key or a JWT.
Each item of gates accepts the following properties:
header object
The configuration of a header rule. The gate passes when the request header name equals value. Add the header to the input_headers of the endpoint so the router can see it.
name * string
The name of the request header you want to inspect.
Example: "X-Plan"
value * string
The value the header must have for the gate to pass.
Example: "enterprise"
name * string
A unique name for this gate. Routes reference the gate through this value in their gates list.
Example: "plan_enterprise"
policy object
The configuration of a policy rule. The gate passes when the CEL expression in value evaluates to true.
value * string
The CEL expression to evaluate. You can use the built-in functions of the security policies, such as hasHeader and getHeader, to compare values or combine several checks that a header rule cannot express.
Example: "getHeader('X-Plan') == 'enterprise'"
rule *
The mechanism the gate uses to match the request. Use header to compare a request header with a value, or policy to evaluate a CEL expression. Each rule requires its configuration under the key with the same name.
Possible values are: "header" , "policy"

Routing Requests

Add the ai/router namespace to a backend’s extra_config. It holds an ordered routes list and an optional failover provider.

Matching with Gates

Each route lists the gates it requires, and matches only when all of them pass. A route with no gates always matches, so place it last as an inline fallback. A matching route forwards the request to its provider:

{
  "ai/router": {
    "routes": [
      {
        "provider": "coding-basic",
        "gates": ["plan_pro"]
      },
      {
        "provider": "cheap-quick"
      }
    ]
  }
}

Pro-plan requests go to coding-basic; every other request falls through to cheap-quick.

Nested Routes and Decision Trees

The routes field is recursive. A gated route can hold its own routes, evaluated only when its gates pass, and those can nest further to any depth, forming a decision tree. A route that only groups other routes needs no provider; the matching leaf route selects it:

{
  "ai/router": {
    "routes": [
      {
        "gates": ["plan_enterprise"],
        "routes": [
          {
            "provider": "coding-team",
            "gates": ["usecase_code"]
          },
          {
            "provider": "gemini-flash"
          }
        ]
      },
      {
        "provider": "cheap-quick"
      }
    ]
  }
}

Enterprise-plan requests enter the branch and route to coding-team for code tasks (gate usecase_code), or to the branch fallback gemini-flash. Any other plan skips the branch and hits the top-level fallback cheap-quick.

Classifier Routes

Instead of a single provider, a route can name a classifier and a list of candidate providers. The classifier chooses the best provider for each request from that list. A route cannot set both provider and classifier:

{
  "ai/router": {
    "failover": "gpt-default",
    "routes": [
      {
        "gates": ["enterprise_policy"],
        "classifier": "smart",
        "providers": [
          "cheap-quick",
          "gemini-flash",
          "coding-basic"
        ],
        "failover": "coding-basic"
      }
    ]
  }
}

Enterprise requests are handed to the smart classifier, which picks among cheap-quick, gemini-flash, and coding-basic. If the classifier cannot return a choice, the route’s own failover sends the request to coding-basic.

Failover

The router supports a failover at two levels:

  • Route failover: a failover inside a classifier route names the provider used when that route’s classifier cannot return a choice. It requires a classifier and takes precedence over the router-level failover.
  • Router failover: the failover next to routes names the provider used when no route matches, or when a classifier fails and its route sets no failover of its own. It is the safety net for the whole router, and an alternative to a trailing route without gates.

The following fields are available under the backend-level ai/router namespace:

Fields of AI Router
* required fields

failover string
The name of the provider in the ai/providers inventory that handles the request when no route matches, or when a classifier cannot return a choice and its route sets no failover of its own. It is the safety net of the whole router, and an alternative to a trailing route without gates.
Example: "gpt-default"
routes * array
The ordered list of routes. KrakenD evaluates them from top to bottom and uses the first one that matches. A route without gates always matches, so place it last as an inline fallback.

Each object in the routes list accepts the following fields:

FieldDescription
gatesThe names of the gates declared in ai/classifiers that must all pass for the route to match. Omit it to create a route that always matches.
providerThe name of the provider in the inventory that handles the request when the route matches. You cannot use it together with classifier.
classifierThe name of a classifier declared in ai/classifiers that chooses the best provider from the providers list. You cannot use it together with provider.
providersThe names of the providers the classifier chooses from. It requires a classifier.
failoverThe provider that handles the request when the classifier of this route cannot return a choice. It requires a classifier and takes precedence over the router-level failover.
routesThe ordered list of nested routes that KrakenD evaluates only when the gates of this route pass. Nested routes can go to any depth, and a route that only groups other routes needs no provider.

Configuration Example

The following configuration defines the providers, gates, and a classifier once at the service level, then uses them across two backends. The first backend routes by plan and use case with a decision tree; the second delegates the choice to the classifier and sets a failover:

{
  "$schema": "https://www.krakend.io/schema/krakend.json",
  "version": 4,
  "extra_config": {
    "ai/providers": {
      "providers": [
        {
          "name": "cheap-quick",
          "provider": "openai",
          "model": "gpt-4o-mini",
          "credentials": "xxx"
        },
        {
          "name": "gpt-default",
          "provider": "openai",
          "credentials": "xxx"
        },
        {
          "name": "gemini-flash",
          "provider": "gemini",
          "model": "gemini-2.5-flash",
          "credentials": "xxx"
        },
        {
          "name": "coding-basic",
          "provider": "anthropic",
          "model": "claude-haiku-4-5",
          "credentials": "xxx"
        },
        {
          "name": "coding-team",
          "provider": "anthropic",
          "model": "claude-sonnet-4-6",
          "credentials": "xxx"
        }
      ]
    },
    "ai/classifiers": {
      "gates": [
        {
          "name": "plan_enterprise",
          "rule": "header",
          "header": {
            "name": "X-Plan",
            "value": "enterprise"
          }
        },
        {
          "name": "plan_pro",
          "rule": "header",
          "header": {
            "name": "X-Plan",
            "value": "pro"
          }
        },
        {
          "name": "usecase_code",
          "rule": "header",
          "header": {
            "name": "X-Use-Case",
            "value": "code"
          }
        },
        {
          "name": "enterprise_policy",
          "rule": "policy",
          "policy": {
            "value": "getHeader('X-Plan') == 'enterprise'"
          }
        }
      ],
      "classifiers": [
        {
          "name": "smart",
          "engine": "notdiamond",
          "notdiamond": {
            "credentials": "xxx",
            "req_content": "req_body.messages"
          }
        }
      ]
    }
  },
  "endpoints": [
    {
      "endpoint": "/v1/chat",
      "method": "POST",
      "input_headers": [
        "X-Plan",
        "X-Use-Case"
      ],
      "backend": [
        {
          "host": [
            "https://api.example.com"
          ],
          "url_pattern": "/v1/chat",
          "@comment": "host and url_pattern are placeholders; the selected provider determines the real upstream.",
          "extra_config": {
            "ai/router": {
              "routes": [
                {
                  "gates": ["plan_enterprise"],
                  "routes": [
                    {
                      "provider": "coding-team",
                      "gates": ["usecase_code"]
                    },
                    {
                      "provider": "gemini-flash"
                    }
                  ]
                },
                {
                  "provider": "coding-basic",
                  "gates": ["plan_pro"]
                },
                {
                  "provider": "cheap-quick"
                }
              ]
            }
          }
        }
      ]
    },
    {
      "endpoint": "/v1/route",
      "method": "POST",
      "input_headers": [
        "*"
      ],
      "backend": [
        {
          "host": [
            "https://api.example.com"
          ],
          "url_pattern": "/v1/route",
          "extra_config": {
            "ai/router": {
              "failover": "gpt-default",
              "routes": [
                {
                  "gates": ["enterprise_policy"],
                  "classifier": "smart",
                  "providers": [
                    "cheap-quick",
                    "gemini-flash",
                    "coding-basic"
                  ]
                }
              ]
            }
          }
        }
      ]
    }
  ]
}

On the /v1/chat endpoint, enterprise-plan requests enter the nested branch and route to coding-team for code tasks or to gemini-flash otherwise, pro-plan requests go to coding-basic, and any other request falls through to cheap-quick. On the /v1/route endpoint, enterprise requests are classified by smart across three candidate providers, and gpt-default serves as the failover when no route matches or the classifier cannot decide.

The AI Router is the inventory-based way to choose a provider per request. For the other routing strategies, such as conditional or path-based routing, see LLM Routing.

Unresolved issues?

The documentation is only a piece of the help you can get! Whether you are looking for Open Source or Enterprise support, see more support channels that can help you.

See all support channels