Document updated on Sep 30, 2026
Prompt Guard for LLM and Agent Endpoints
Prompt Guard inspects the body payload of incoming requests and blocks content that matches known attack patterns, such as prompt injection, jailbreaks, or data exfiltration attempts. It applies one or more guards to the request and rejects the call when a guard flags the payload as malicious. Because it operates on the raw request body, you can add it to any endpoint that accepts a body, not only to those connected to an LLM backend. KrakenD ships with a set of default guards that protect your endpoints out of the box, and you can extend or restrict them through configuration.
policy guard can additionally reference request metadata such as headers, query strings, and the URL. Review which guards you enable per endpoint and keep the default guards active unless you have a specific reason to disable them.How Prompt Guard Works
KrakenD maintains a global shared registry that holds both the default guards and the guards you define in your configuration. When a request reaches a protected endpoint, the gateway selects the guards that apply, combines them into a single guard, and evaluates the request body against it.
Three kinds of guards are available:
- Regex: matches the body against a list of regular expressions.
- Classifier: sends the body to an external service that returns a score.
- Policy: evaluates CEL expressions against the body and request metadata.
The default guards are a large set of regular expressions covering common attack categories. KrakenD merges all regex guards (defaults and your own) into a single set so it can test every pattern in one pass, which keeps evaluation fast even when the number of patterns is high.
Enabling Prompt Guard
Prompt Guard uses the ai/prompt-guard namespace in the extra_config. The default guards are active on any endpoint or backend where you add the namespace, even with an empty object. To take full control, declare your own guards at the service level and then choose which guards apply on each endpoint or backend.
A configuration has two parts:
- The guard definitions, placed in the root
extra_config, where you describe each custom guard. - The guard selection, placed in the
extra_configof each endpoint or backend, where you list the guards to apply by name.
Defining Guards
You define each custom guard at the service level inside the ai/prompt-guard namespace, under the guards list. Every guard declaration has a name to reference it, a kind (regex, classifier, or policy), a severity from critical (highest) to low (lowest), and a config object whose fields depend on the kind.
Severity levels matter when selecting or disabling guards: you can disable all default guards of a given severity (see Disabling Default Guards).
The following fields are available at the service level:
Fields of AI Prompt Guard
guardsarray of objects- The list of custom guards available to endpoints and backends. Each guard has a unique
name, akindthat defines how it evaluates the request, aseverity, and aconfigwhose fields depend on thekind.Each item of guards accepts the following properties:config* object- The settings of the guard. Its fields depend on the
kind:patternsfor aregexguard,urlandscore_thresholdfor aclassifierguard, andexpressionsfor apolicyguard. kind*- How the guard evaluates the request. A
regexguard matches the body against a list of regular expressions, aclassifierguard sends the body to an external service that returns a score, and apolicyguard evaluates CEL expressions against the body and the request metadata.Possible values are:"regex","classifier","policy" name* string- A unique name for this guard. Endpoints and backends reference the guard through this value in their
guardslist.Examples:"block_jailbreak_terms","injection_classifier" severity*- The severity level assigned to the guard, from
critical(highest) tolow(lowest). Endpoints and backends can disable all the default guards of a given severity by adding!critical,!high,!medium, or!lowto theirguardslist.Possible values are:"critical","high","medium","low"
Regex Guards
A regex guard matches the request body against a list of regular expressions. If any pattern matches, the guard blocks the request. Define the patterns under config.patterns:
{
"name": "block_jailbreak_terms",
"kind": "regex",
"severity": "high",
"config": {
"patterns": [
"(?i)ignore (all )?previous instructions",
"(?i)disregard .*(system )?prompt"
]
}
}
The guard above blocks any request whose body matches either pattern. KrakenD combines these patterns with the patterns of all other regex guards into a single set evaluated in one pass.
Classifier Guards
A classifier guard delegates the decision to an external service. KrakenD sends the body to the configured url, and the service responds with a JSON payload that contains a score between 0.0 and 1.0, both included. A score of 0.0 means “do not block”, and 1.0 is the maximum confidence that the payload should be blocked:
{ "score": 0.23 }
You configure a classifier guard with the endpoint to call and the threshold that decides the outcome:
{
"name": "injection_classifier",
"kind": "classifier",
"severity": "high",
"config": {
"score_threshold": 0.6,
"on_failure_block": false,
"url": "https://example.com/classify"
}
}
The guard blocks the request when the returned score is equal to or greater than score_threshold. The on_failure_block flag controls the behavior when the external service cannot be reached or returns an error: set it to true to block the request on failure, or false to let it through.
Policy Guards
A policy guard evaluates a list of CEL expressions against the request. The body is exposed as a single body string, not as parsed JSON. A meta object exposes the request context with the same data available to other Security Policies:
headersparamsmethodpathqueryurl
{
"name": "payload_policy",
"kind": "policy",
"severity": "medium",
"config": {
"expressions": [
"meta.method == 'POST'",
"size(body) <= 8192"
]
}
}
Each expression returns a boolean and decides whether the request is allowed to continue. The policy above only lets through POST requests whose body is 8 KB or smaller. Use body to inspect the raw payload and meta to reason about the request headers, parameters, method, path, query string, and URL.
Selecting Guards per Endpoint or Backend
After defining guards at the service level, add the ai/prompt-guard namespace to the endpoint or backend you want to protect, and list the guards to apply by name. You can protect either the endpoint or an individual backend.
The following endpoint applies a regex guard and a classifier guard defined earlier:
{
"endpoint": "/guard/on/endpoint",
"method": "POST",
"input_headers": [
"*"
],
"backend": [
{
"url_pattern": "/v1/chat/completions",
"method": "POST",
"encoding": "json"
}
],
"extra_config": {
"ai/prompt-guard": {
"guards": [
"block_jailbreak_terms",
"injection_classifier"
]
}
}
}
To protect a specific backend instead of the whole endpoint, move the namespace into the backend’s extra_config:
{
"endpoint": "/guard/on/backend",
"method": "POST",
"input_headers": [
"*"
],
"backend": [
{
"url_pattern": "/v1/chat/completions",
"method": "POST",
"extra_config": {
"ai/prompt-guard": {
"guards": [
"block_jailbreak_terms",
"injection_classifier"
]
}
}
}
]
}
Both examples keep the default guards active in addition to the named guards you selected. The following fields are available at the endpoint and backend levels:
Fields of AI Prompt Guard Selection
guardsarray of strings- The guards to apply, in addition to the default ones. Add the
nameof a guard declared at the service level to apply it, or prefix an entry with!to disable default guards:!all_defaultsremoves all of them,!critical,!high,!medium, and!lowremove those of that severity, and!followed by a group name, such as!unicode_steganography, removes that group.
Disabling Default Guards
The guards list also accepts negations to remove default guards for an endpoint or backend. Prefix the name with an exclamation mark ! to disable it:
!all_defaults: Removes all default guards. Only the guards you explicitly select remain. If this is the only entry, KrakenD applies an internal no-op guard, effectively disabling Prompt Guard for that endpoint.!low,!medium,!high,!critical: Disable all default guards of that severity level.!group_name: Disable a specific default guard group, for example!unicode_steganography.
For instance, the following selection applies two named guards while disabling every default guard with medium severity and the unicode_steganography group:
{
"ai/prompt-guard": {
"guards": [
"block_jailbreak_terms",
"injection_classifier",
"!medium",
"!unicode_steganography"
]
}
}
Default Guard Groups
The default guards below are grouped by the kind of attack they detect, to make the list easier to scan. Disable an individual guard group with the ! notation (for example, !unicode_steganography); the category headings are not configuration values.
Prompt injection and instruction manipulation
cascade_amplificationcontext_hijackingfew_shot_hijackindirect_injectioninstruction_overrideinstruction_piggybackingmulti_turnoutput_manipulationprompt_extractionrepetition_bypassrole_manipulation
Jailbreak and safety bypass
bypass_coachingcognitive_manipulationemotional_manipulationjailbreaklanguage_switch_evasionsafety_bypassscenario_jailbreakurgency_manipulation
Social engineering and impersonation
authority_impersonationdm_social_engineeringphishingsocial_engineeringsystem_impersonationsystem_mimicry
Agent, tool, and memory abuse
action_gate_bypassagent_payment_hijackagent_sovereigntyapproval_expansionauto_approve_exploitcognitive_rootkithooks_hijackingmcp_abusememory_manipulationmemory_poisoningrecursive_delegationsemantic_wormsubagent_exploit
Code execution and system attacks
code_injectionfork_bombgitignore_bypassremote_code_executionreverse_shellsql_injectionssh_key_injectionsupply_chain_injectionsystem_destructionsystem_file_accessxss
Data exfiltration
covert_exfiltrationdata_exfiltration
Obfuscation and hidden content
hidden_textobfuscated_payloadtoken_smugglingunicode_steganographyunicode_tag_injection
Prompt Guard Metrics and Traces
When you enable OpenTelemetry, Prompt Guard reports the krakend.promptguard.check counter, which increases every time a prompt guard runs. It carries the following attributes:
blocker: the category or name of the guard that blocked the request. It is empty when the request passes.http_route: the endpoint that received the request.http_method: the HTTP method of the request.krakend_backend: theurl_patternof the backend, when you set Prompt Guard at the backend level.
Prompt Guard also adds its result as an attribute of the current trace.
Full Configuration Example
The following configuration defines four guards at the service level and applies a subset of them, together with some negations, on both an endpoint and a backend:
{
"$schema": "https://www.krakend.io/schema/krakend.json",
"version": 4,
"name": "prompt-guard-example",
"timeout": "5s",
"cache_ttl": "300s",
"host": [
"http://llm_backend:9876"
],
"extra_config": {
"ai/prompt-guard": {
"guards": [
{
"name": "block_jailbreak_terms",
"kind": "regex",
"severity": "high",
"config": {
"patterns": [
"(?i)ignore (all )?previous instructions",
"(?i)disregard .*(system )?prompt"
]
}
},
{
"name": "block_sensitive_terms",
"kind": "regex",
"severity": "medium",
"config": {
"patterns": [
"(?i)reveal .*(system )?prompt",
"(?i)you are now .*(DAN|jailbroken)"
]
}
},
{
"name": "injection_classifier",
"kind": "classifier",
"severity": "high",
"config": {
"score_threshold": 0.6,
"on_failure_block": false,
"url": "https://example.com/classify"
}
},
{
"name": "payload_policy",
"kind": "policy",
"severity": "medium",
"config": {
"expressions": [
"meta.method == 'POST'",
"size(body) <= 8192"
]
}
}
]
}
},
"endpoints": [
{
"endpoint": "/guard/on/endpoint",
"method": "POST",
"input_headers": [
"*"
],
"backend": [
{
"url_pattern": "/v1/chat/completions",
"method": "POST",
"encoding": "json"
}
],
"extra_config": {
"ai/prompt-guard": {
"guards": [
"block_jailbreak_terms",
"injection_classifier",
"!medium",
"!unicode_steganography"
]
}
}
},
{
"endpoint": "/guard/on/backend",
"method": "POST",
"input_headers": [
"*"
],
"backend": [
{
"url_pattern": "/v1/chat/completions",
"method": "POST",
"extra_config": {
"ai/prompt-guard": {
"guards": [
"block_jailbreak_terms",
"injection_classifier",
"!medium",
"!unicode_steganography"
]
}
}
}
]
}
]
}
The root extra_config defines the guards block_jailbreak_terms, block_sensitive_terms, injection_classifier, and payload_policy. Each endpoint selects block_jailbreak_terms and injection_classifier by name, keeps the remaining default guards, and disables both the medium-severity defaults and the unicode_steganography group. The first endpoint guards the whole endpoint, while the second guards a single backend. This feature complements AI Security and AI Governance to harden the requests reaching your LLMs and agents.
