An Azure OpenAI 400 response with error code content_filter and inner error ResponsibleAIPolicyViolation means the guardrail (content filter) attached to your deployment classified the prompt at a category and severity it is configured to block, or the prompt matched a blocking blocklist. To fix it, read innererror.content_filter_result to see which category has "filtered": true, then either change the input, raise that category's threshold in a custom content filter, adjust Prompt Shields, or remove the blocklist term that matched. Turning filtering off for completions requires approval for modified content filtering, but threshold changes and prompt-side changes don't.
Who this is for and what you will have
This guide is for developers and AI platform owners whose legitimate business prompts are being rejected: security teams analysing attack descriptions, clinical or safety workloads that discuss injury, HR tools that quote harassment complaints, or RAG pipelines that pass retrieved documents into the prompt. At the end you will have:
- A reliable way to identify the filter and category behind each rejection.
- A custom content filter created in the Foundry portal or through the REST API and assigned to a deployment.
- Blocklists you can add, test and remove without guesswork.
- Application code that handles filtered prompts and filtered completions separately.
How content filtering works
The content filtering system runs prompts and completions through an ensemble of classification models powered by Azure AI Content Safety. It doesn't apply to embedding models or audio models such as Whisper.
| Filter | Type | Applies to | Default |
|---|---|---|---|
| Hate and fairness, sexual, violence, self-harm | Severity: safe, low, medium, high | Prompts and completions | Filter at medium |
| Prompt Shields for direct attacks (jailbreak) | Binary | User prompt | On |
| Prompt Shields for indirect attacks | Binary | User prompt (documents) | Off |
| Protected material: text | Binary | Completion | On |
| Protected material: code | Binary | Completion | On |
| Groundedness | Preview | Completion | Off |
| Personally identifiable information | Preview | Completion | Off |
Content at the safe level is labelled in annotations but never filtered. With the default medium threshold, content detected at medium or high is filtered and low or safe content passes.
The behaviour your application sees depends on where the hit happens:
- Prompt filtered: the API call fails with HTTP 400.
- Completion filtered (non-streaming): HTTP 200, no content (rarely a partial result), and
finish_reasonset tocontent_filter. - Completion filtered (streaming): chunks stream until filtered content is detected; the last chunk for that choice has
finish_reasonset tocontent_filter. - Filter unavailable: the request completes unfiltered, and
content_filter_resultscontains an error object.
Prerequisites
- An Azure OpenAI or Foundry resource with a model deployment.
- Cognitive Services Contributor or Owner on the resource to manage blocklists and content filters; without it the management API returns 403.
- Azure CLI 2.50 or later for
az account get-access-tokenandaz rest. - A handful of real prompts that are being rejected, with any personal or sensitive data removed.
Step 1: Read the error and find the category
A blocked prompt returns a body like this documented example, where a custom blocklist caused the rejection:
{
"error": {
"message": "The response was filtered due to the prompt triggering Azure OpenAI's content management policy. Please modify your prompt and retry. To learn more about our content filtering policies please read our documentation: https://go.microsoft.com/fwlink/?linkid=2198766",
"type": null,
"param": "prompt",
"code": "content_filter",
"status": 400,
"innererror": {
"code": "ResponsibleAIPolicyViolation",
"content_filter_result": {
"custom_blocklists": {
"details": [{ "filtered": true, "id": "pizza" }],
"filtered": true
}
}
}
}
}The content_filter_result object uses the same keys as the annotations: hate, sexual, violence and self_harm carry filtered and severity, while binary models such as jailbreak, profanity and custom_blocklists carry filtered (and detected where applicable). Log the whole object for every rejection, then pull out the entries that blocked the call:
jq '.error.innererror.content_filter_result
| to_entries[]
| select(.value.filtered == true)
| {category: .key, detail: .value}' rejected-response.jsonIn Python with the OpenAI library, a 400 is raised as openai.BadRequestError; keep the response body in your logs so the category isn't lost:
import json
import openai
try:
response = client.chat.completions.create(model="gpt-4o-prod", messages=messages)
except openai.BadRequestError as e:
body = e.body or {}
if body.get("code") == "content_filter":
result = body.get("innererror", {}).get("content_filter_result", {})
blocked = {k: v for k, v in result.items() if isinstance(v, dict) and v.get("filtered")}
print(json.dumps(blocked, indent=2))
raiseMap the category to the action:
| Category with filtered true | Typical cause | Go to |
|---|---|---|
hate, sexual, violence, self_harm | Prompt text reached the deployment's threshold | Step 3 or Step 4 |
jailbreak | Prompt Shields read the input as an attempt to override system rules | Step 5 |
custom_blocklists or profanity | A blocklist term matched | Step 6 |
Step 2: Check completions and annotations too
A deployment can accept the prompt and still filter the output. Check finish_reason on every choice and inspect content_filter_results on the choice and prompt_filter_results at the top level. Annotations show the category and severity even for content that wasn't blocked, which tells you how close a request came to a threshold. Microsoft documents annotations as returned with preview API versions from 2023-06-01-preview onward and with the GA version 2024-02-01.
"prompt_filter_results": [
{
"prompt_index": 0,
"content_filter_results": {
"hate": { "filtered": false, "severity": "safe" },
"jailbreak": { "detected": false, "filtered": false },
"self_harm": { "filtered": false, "severity": "safe" },
"sexual": { "filtered": false, "severity": "safe" },
"violence": { "filtered": false, "severity": "safe" }
}
}
]If you see content_filter_results containing an error with code content_filter_error and message "The contents are not filtered", the filter didn't run for that call. Treat it as unmoderated output.
Step 3: Reproduce the severity in isolation
Before changing configuration, confirm which category and severity the text lands in. In the Foundry portal, the Content Safety Try it out page (under Guardrails + controls) lets you paste text into Moderate text content and see each category with a severity of 0 (safe), 2 (low), 4 (medium) or 6 (high).
For repeatable tests, call the Azure AI Content Safety text analysis API from a Content Safety resource. The caller needs Cognitive Services User or higher on that resource:
curl --location --request POST "$CONTENT_SAFETY_ENDPOINT/contentsafety/text:analyze?api-version=2024-09-01" \
--header "Ocp-Apim-Subscription-Key: $CONTENT_SAFETY_KEY" \
--header "Content-Type: application/json" \
--data-raw '{
"text": "Describe how the ransomware disabled backups before encrypting the file server.",
"categories": ["Hate", "Sexual", "SelfHarm", "Violence"],
"outputType": "FourSeverityLevels"
}'The response lists categoriesAnalysis entries with a category and a severity value. Use it to compare rephrasings and to decide whether a threshold change is needed at all. Often a shorter, more neutral system prompt or removing irrelevant retrieved passages brings the text under the threshold without any configuration change.
Step 4: Create a custom content filter
Content filters are configured on the resource and then assigned to one or more deployments. All customers can move each harm category between three thresholds, separately for prompts and completions:
| Threshold setting | What is filtered |
|---|---|
| Low, medium, high | Strictest: low, medium and high are filtered |
| Medium, high | Default: low passes, medium and high are filtered |
| High | Only high is filtered |
| No filters | Nothing blocked or annotated; completions require approval |
| Annotate only | Not blocked, annotations returned; completions require approval |
To create one in the portal (Foundry classic experience):
- Open your project and select Guardrails + controls, then the Content filters tab.
- Select + Create content filter, enter a name and select the connection.
- On Input filters, move only the slider for the category that is blocking legitimate prompts. Leave the others at their defaults.
- On Output filters, keep completion thresholds unless you have evidence that outputs are the problem.
- On Connection, associate the filter with the deployment, then review and select Create filter.
To do the same through Azure Resource Manager, create a RAI policy. This example relaxes only the prompt-side violence threshold to high and keeps everything else at the default levels:
TOKEN=$(az account get-access-token --query accessToken -o tsv)
ACCOUNT_ID="/subscriptions/$SUBSCRIPTION_ID/resourceGroups/rg-ai-prod/providers/Microsoft.CognitiveServices/accounts/contoso-aoai"
curl -X PUT "https://management.azure.com$ACCOUNT_ID/raiPolicies/secops-analysis?api-version=2024-10-01" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{
"properties": {
"basePolicyName": "Microsoft.Default",
"contentFilters": [
{ "name": "Violence", "blocking": true, "enabled": true, "severityThreshold": "High", "source": "Prompt" },
{ "name": "Violence", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
{ "name": "Hate", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Prompt" },
{ "name": "Hate", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
{ "name": "Sexual", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Prompt" },
{ "name": "Sexual", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
{ "name": "Selfharm", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Prompt" },
{ "name": "Selfharm", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
{ "name": "Jailbreak", "blocking": true, "enabled": true, "source": "Prompt" },
{ "name": "Protected Material Text", "blocking": true, "enabled": true, "source": "Completion" },
{ "name": "Protected Material Code", "blocking": true, "enabled": true, "source": "Completion" }
]
}
}'Assign it by setting raiPolicyName on the deployment. Include the existing SKU and model so nothing else changes:
curl -X PUT "https://management.azure.com$ACCOUNT_ID/deployments/gpt-4o-secops?api-version=2024-10-01" \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"sku": {"name": "GlobalStandard", "capacity": 50},
"properties": {"model": {"format": "OpenAI", "name": "gpt-4o", "version": "2024-11-20"},
"raiPolicyName": "secops-analysis"}}'A dedicated deployment for the workload that needs the relaxed filter keeps the default policy on everything else. Alternatively, send the x-policy-id header with the filter name on specific requests; the request-level configuration overrides the deployment-level one for that call. Request-time selection isn't available for image input scenarios, where the default filter is used.
Step 5: Deal with Prompt Shields false positives
A jailbreak hit means Prompt Shields read the user prompt as an attempt to change system rules, replace the model's persona, or request encoded output. Legitimate prompts can trigger it when they contain instructions aimed at the model inside user content, for example pasted role-play scripts or security test cases.
- Move instructions into the system message and keep the user turn to the actual request.
- For RAG and email or document processing, use document embedding and formatting so retrieved content is marked as data. Microsoft requires this for indirect attack detection, which is off by default.
- Every customer can switch Prompt Shields and the protected material models on or off, or choose Annotate only so detections are reported without blocking. Annotate-only keeps the signal in your logs while you decide.
The protected material code model may be required for Customer Copyright Commitment coverage, so check that before turning it off.
Step 6: Manage blocklists
Blocklists filter terms specific to your organisation, such as internal project names. When custom_blocklists is the category that fired, the fix is in the list, not the model.
Create a list and add an item through the management API:
curl --request PUT "https://management.azure.com$ACCOUNT_ID/raiBlocklists/internal-codenames?api-version=2024-10-01" \
--header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" \
--data-raw '{"properties": {"description": "Unreleased project names"}}'
curl --request PUT "https://management.azure.com$ACCOUNT_ID/raiBlocklists/internal-codenames/raiBlocklistItems/item-001?api-version=2024-10-01" \
--header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" \
--data-raw '{"properties": {"pattern": "Project Fabrikam", "isRegex": false}}'Attach the list to a content filter by adding a customBlocklists entry with blocklistName, blocking and source (Prompt or Completion) to the RAI policy, or by enabling Blocklist on the input or output filter page in the portal. To stop a false positive, delete the item:
curl --request DELETE "https://management.azure.com$ACCOUNT_ID/raiBlocklists/internal-codenames/raiBlocklistItems/item-001?api-version=2024-10-01" \
--header "Authorization: Bearer $TOKEN"Limits and behaviour to remember: a list holds at most 10,000 terms, an item can be up to 1,000 characters, new terms take around 5 minutes to apply, and regex patterns are case-sensitive by default. Anchors such as ^ and $ might not behave as expected in streaming scenarios.
Verify the change
- Wait about 5 minutes after blocklist changes, then rerun the prompts that failed.
- Confirm the response is 200 and that
prompt_filter_resultsshows the category you relaxed as"filtered": falsewith its severity still reported. - Send a prompt that should still be blocked, for example one at high severity in the same category, and confirm it returns 400.
- In the playground, use Filters Feedback to report remaining false positives, without including private or sensitive information.
Troubleshooting
| Error or symptom | Cause | Fix |
|---|---|---|
400 content_filter, ResponsibleAIPolicyViolation | Prompt filtered by category, Prompt Shields or blocklist | Read content_filter_result and follow Steps 3 to 6 |
InvalidContentFilterPolicy: "Your request contains invalid content filter policy" | x-policy-id names a filter that doesn't exist | Use the exact name of an existing content filter |
403 from raiBlocklists or raiPolicies | Missing management role | Assign Cognitive Services Contributor or Owner |
| Blocklist change not visible | Propagation delay | Wait up to 5 minutes and retest |
finish_reason is content_filter with partial text | Completion filtered mid-generation | Handle as incomplete output; don't show it as a full answer |
content_filter_error: "The contents are not filtered" | Filtering system unavailable for that call | Treat output as unmoderated and apply your own checks |
| Option to disable completion filtering is missing | Requires approved modified content filtering, available only to managed customers | Set the completion threshold to high instead |
Checklist
- Log the full
content_filter_resultfor every 400 andcontent_filter_resultsfor every filtered choice. - Identify the single category that blocks legitimate traffic before changing anything.
- Rephrase, trim retrieved context or move instructions to the system message first.
- Create a custom filter that changes only that category, and assign it to a dedicated deployment.
- Use annotate-only for Prompt Shields while you measure false positives.
- Keep blocklists small, documented and tested after each edit.
- Retest both allowed and still-blocked examples after every change.
Content filter results are also a useful quality signal in evaluation pipelines; see production LLMOps for enterprise RAG for where they fit alongside groundedness and relevance metrics.
References
- Content filtering for Microsoft Foundry Models
- Configure content filters
- Guardrail annotations
- How to use blocklists in Microsoft Foundry models
- Default safety policies in Microsoft Foundry
- Rai Policies - Create Or Update REST API
- Deployments - Create Or Update REST API
- Quickstart: Analyze text content with Azure AI Content Safety