News KrakenD 3.0 Is Here: AI Router, Semantic Cache, and On-the-Fly Stream Manipulation

Product Updates

10 min read

KrakenD 3.0 Is Here: AI Router, Semantic Cache, and On-the-Fly Stream Manipulation

by Toni Pinel

post image

Today we release KrakenD 3.0, for both the Community Edition and the Enterprise Edition. It’s the first version of the next generation of KrakenD, which we are calling Series 3, and the biggest jump in the gateway’s capabilities since we launched. Most of the features below are part of the Enterprise Edition. If you run the open source version, jump to KrakenD Community Edition 3.0 or read the KrakenD CE v3.0 release notes. This post walks through what ships in 3.0 and why it matters, plus a preview of what’s coming next in the Series 3 line.

Streaming With On-the-Fly Manipulation

Streaming becomes a first-class part of the gateway core in 3.0. A streamed response integrates with KrakenD’s native features the same way a regular response does: your quota rules can read from it, and you can manipulate it as it flows through the gateway.

This is a key difference from the major AI gateways on the market. They treat a stream as a pipe: once it starts, the response either skips the response policies or has to be buffered in full before anything can change it. KrakenD can modify the stream on the fly, message by message, while it’s still reaching the client.

With LLM streaming, token-by-token responses flow through without buffering the whole completion first, which is critical for anything chat-shaped. You get that responsiveness without giving up the controls you apply to every other response.

AI Router: Dynamic Model Selection and Failover

The headline AI feature is the AI Router. Instead of statically declaring which LLM handles which type of request, 3.0 lets you route dynamically in two ways:

  • The expression-based router picks the model with gates that match request headers or CEL expressions, the same language you already use in KrakenD’s security policies.
  • The prompt classifier, an integration with the Not Diamond router, reads the prompt itself and decides which model fits best.

Send simple queries to a cheap, fast model and complex ones to a frontier model, automatically, based on the content of the request.

The router also has a failover: when no route matches, or the classifier cannot decide, the request goes to the provider you choose, so every request gets an answer. Combined with the fallback proxy strategy and quota fallbacks, you also decide what happens when a provider goes down or a user spends all their quota, like falling back to a free local model.

This is the piece that turns KrakenD from “a gateway that can reach LLMs” into “a gateway that manages your LLM spend, latency, and reliability for you.” See the AI Router documentation to set it up.

Semantic Cache: Stop Paying Twice for the Same Answer

The Semantic Cache stores LLM responses and serves them again when a new prompt means the same thing as a previous one, even when the wording is different. “How do I reset my password?” and “I forgot my password, what do I do?” get the same cached answer, and the second one never reaches your model provider.

Every cache hit is a call you don’t pay for and a response that doesn’t wait for the model to generate it. It pays off most where people ask the same questions in different words: support assistants, internal knowledge bots, and FAQ-style agents. See the Semantic Cache documentation to set it up.

Prompt Guard at the Gateway

Prompt Guard catches prompt injection and unsafe inputs before they reach your model provider. You combine three kinds of guards: regular expressions for known patterns, CEL policies for rules based on the request, and an external classifier that scores each prompt. Every guard has a severity level, and you decide whether a request goes through or gets blocked when the external classifier is unavailable.

Because it runs in KrakenD, the same protection applies to every application and team sending traffic through the gateway, instead of each one implementing its own filters. See the Prompt Guard documentation to set it up.

One Interface for Every LLM Provider

  • OpenAI and Anthropic-compatible responses, so a client built against OpenAI’s or Anthropic’s API format can talk to a different provider underneath without knowing it. Point a tool like Claude Code at KrakenD instead of Anthropic directly, and route it to Amazon Bedrock behind the scenes. The client never has to change.
  • Native Alibaba Cloud / Qwen support, joining OpenAI, Anthropic, and the rest of the provider list.
  • LLM Provider Inventory: define each LLM provider once under ai/providers, credentials included, and reference it from any endpoint instead of repeating the same configuration everywhere.

Put together, these are the pieces teams have been building themselves in application code. 3.0 moves them into the gateway, where they apply to every service behind it at once.

MCP Server: Authentication and Tool Filtering

KrakenD’s MCP Server gets the two features teams asked for most before exposing their APIs to agents:

  • Authentication, following the MCP authorization specification. KrakenD publishes the protected resource metadata on its well-known endpoint, a discovery method every compliant MCP client must support, so clients find your authorization server on their own and come back with a token KrakenD validates.
  • Tool filtering, so each client only sees the tools meant for it. You assign audiences to your tools, and KrakenD reads the audience from a request header. Servers exposing a large number of tools no longer burn context, or budget, on tools a client will never call.

MCP tools also gain dynamic routing, access to the incoming Host header, and support for binary responses.

Other Improvements

  • The fallback proxy strategy: when an endpoint has several backends, KrakenD tries them in order and returns the first successful response.
  • Gzip compression profiles: choose speed, balanced, or compression with the new gzip_profile option to favor throughput or smaller responses.
  • The krakend report command packages your configuration, with sensitive fields redacted, into a single file you can share with our support team.
  • OpenAPI generation supports reusable parameters through components_parameters and several response types for the same status code.
  • LLM metrics now include the model that served each request.

KrakenD Community Edition 3.0

KrakenD Community Edition 3.0 ships today too, as a leaner and cleaner gateway:

  • Wildcards and array indexes in JWT claims: roles, scopes, and propagated claims accept paths like resource_access.*.roles, so you can read claims whose position in the token isn’t fixed.
  • The HTTP method in endpoint logs, so GET /foo and POST /foo are no longer indistinguishable.
  • A smaller dependency tree: removing deprecated components takes dozens of dependencies out of the binary, which means a smaller attack surface and fewer security advisories to track.

The KrakenD CE v3.0 release notes list every change, including the deprecated components this version removes.

What’s Next in the Series 3 Line

3.0 is the first release in Series 3. Here’s what’s landing right behind it:

  • A rebuilt routing engine, with more flexible matching rules and priority-based failover routing, so you can define overlapping or conditional routes without the workarounds teams have needed until now.
  • A new internal event system: KrakenD fires events at relevant points across the platform, and you’ll be able to subscribe to them to build your own workflows, think notifying on quota usage or monitoring circuit breaker state changes.
  • Native TOON payload compression, cutting the wire cost of verbose LLM payloads. Today’s TOON requires a little extra work in the configuration.
  • More native AI Router criteria, including a built-in classifier based on embedders, plus broader compatibility with external classifiers.
  • Full MCP Server protocol coverage, with resources, prompts, Tasks, and notifications, keeping pace with the latest MCP specification.
  • An OpenAPI spec upgrade to the latest version.
  • External secrets management, so secrets can live in a vault instead of plain configuration.

Upgrading to KrakenD 3.0

As a major version, 3.0 changes the configuration syntax version: every configuration file must now declare "version": 4 instead of the "version": 3 used across the 2.x series, or KrakenD refuses to start. It also removes components that were deprecated long ago. Most setups upgrade with a few configuration changes, not a re-architecture. The KrakenD CE v3.0 release notes and the changelog list every removal and its replacement, and the upgrade guide walks you through the changes.

The one change that needs planning is for Community Edition users with plugins, which move to Enterprise Edition in 3.0. We explained why in Dropping plugin support in KrakenD Open Source and Lura. If your setup depends on plugins, talk to us before you upgrade. We have a migration path, including a free architecture review, so this doesn’t catch you mid-deploy.

If you’re new to 3.0, the AI Router and Semantic Cache are the two features worth trying first.

Running Community Edition and using plugins? Check your migration path →

Curious what the AI Router can do for your model costs? Talk to us about AI Gateway →

KrakenD 3.0, the first release in Series 3, is available today. Check the full changelog and the upgrade guide to get started.

🚀 Summary of changes for EEv3.0

The first release of Series 3 brings the AI Router, Semantic Cache, Prompt Guard, on-the-fly stream manipulation, MCP authentication and tool filtering, and more!

  • New AI Router (ai/router and ai/classifiers) to select the LLM provider per request using header or CEL gates, the Not Diamond prompt classifier, and a failover provider.
  • New LLM Provider Inventory (ai/providers) to define each LLM provider once at the service level and reference it from ai/llm with the provider option.
  • New Semantic Cache (ai/semantic-cache) to return stored responses for requests with the same meaning, using a local ONNX embedding model and Redis as vector storage.
  • New Prompt Guard (ai/prompt-guard) to block prompt injection and unsafe inputs using regular expressions, CEL policies, or an external classifier.
  • Added Alibaba Cloud / Qwen support to the AI Gateway ai/llm component.
  • Prompt Guard reports the krakend.promptguard.check OpenTelemetry metric, counting each check with the guard that blocked the request, and adds its result to the current trace.
  • LLM streaming through ai/llm for all supported vendors, with the format option to normalize the stream to the OpenAI or Anthropic format.
  • New sse, ndjson, and eventstream encodings that parse each message of a stream, so modifier/jmespath, modifier/response-body, and governance/quota can manipulate streamed responses on the fly.
  • MCP Server authentication following the MCP authorization specification, advertising the authorization server through the well-known protected resource metadata endpoint.
  • MCP Server tool filtering by audience, using tool_audience_source in the server and audiences in each tool.
  • MCP tools support dynamic routing, receive the Host header and input_headers.* parameters, and can return binary responses.
  • New fallback proxy strategy for the endpoint’s proxy strategy, which tries the backends in order and returns the first successful response.
  • New gzip_profile option to choose between balanced (default), speed, and compression.
  • New krakend report command that packages the configuration, with sensitive fields redacted, to share it with the support team.
  • OpenAPI generation supports reusable parameters with components_parameters and multiple response types for the same status code.
  • JWT validation accepts wildcards and array indexes in the paths of roles, scopes, and propagated claims.
  • Endpoint logs include the HTTP method of the request.
  • LLM metrics include the model that served the request.
  • The license info subcommand prints the organization of the licensee.
  • The audit command takes the HTTP circuit breaker into account in rule 3.1.3.
  • OpenAPI export skips parameters that are only a * wildcard.
  • AI Gateway components and AWS Lambda backends no longer need a host in the backend definition.
  • Fixed the Logstash formatter removing the prefix of log lines coming from external components.
  • All previous configurations used version: 3. Migrate your configuration and use version: 4 instead.
  • Removed the deprecated built-in plugins basic auth, virtual host, redirect, HTTP proxy, live, content replacer, minimum response, and Redis rate limit. Use their native replacements, such as Basic Authentication, Virtual Hosts, and content replacement.
  • Removed OpenCensus (telemetry/opencensus) and all its exporters. Use OpenTelemetry instead.
  • Removed Telemetry Port (telemetry/metrics). KrakenD could listen to an alternative port for metrics, but this is no longer supported.
  • Removed the InfluxDB component (telemetry/influx). Use OpenTelemetry instead.
  • Legacy repository-style namespaces (e.g., github_com/devopsfaith/krakend-gologging) are no longer accepted. Use their current names (e.g., telemetry/logging).
  • ai/llm no longer uses Martian to authenticate against providers. Declare the credentials in ai/llm or ai/providers.
  • Removed the noop and override license validators.

Upgrading to the latest version is always advised.

Categories: Product Updates

Stay up to date with KrakenD releases and important updates