AI engineering

Fix Azure OpenAI content filter 400 errors: ResponsibleAIPolicyViolation

Find which Azure OpenAI content filter category blocked a request, then fix false positives with custom content filters, Prompt Shields settings and blocklists.

12 min read
On this page

An Azure OpenAI 400 response with error code content_filter and inner error ResponsibleAIPolicyViolation means the guardrail (content filter) attached to your deployment classified the prompt at a category and severity it is configured to block, or the prompt matched a blocking blocklist. To fix it, read innererror.content_filter_result to see which category has "filtered": true, then either change the input, raise that category's threshold in a custom content filter, adjust Prompt Shields, or remove the blocklist term that matched. Turning filtering off for completions requires approval for modified content filtering, but threshold changes and prompt-side changes don't.

Who this is for and what you will have

This guide is for developers and AI platform owners whose legitimate business prompts are being rejected: security teams analysing attack descriptions, clinical or safety workloads that discuss injury, HR tools that quote harassment complaints, or RAG pipelines that pass retrieved documents into the prompt. At the end you will have:

  • A reliable way to identify the filter and category behind each rejection.
  • A custom content filter created in the Foundry portal or through the REST API and assigned to a deployment.
  • Blocklists you can add, test and remove without guesswork.
  • Application code that handles filtered prompts and filtered completions separately.

How content filtering works

The content filtering system runs prompts and completions through an ensemble of classification models powered by Azure AI Content Safety. It doesn't apply to embedding models or audio models such as Whisper.

FilterTypeApplies toDefault
Hate and fairness, sexual, violence, self-harmSeverity: safe, low, medium, highPrompts and completionsFilter at medium
Prompt Shields for direct attacks (jailbreak)BinaryUser promptOn
Prompt Shields for indirect attacksBinaryUser prompt (documents)Off
Protected material: textBinaryCompletionOn
Protected material: codeBinaryCompletionOn
GroundednessPreviewCompletionOff
Personally identifiable informationPreviewCompletionOff

Content at the safe level is labelled in annotations but never filtered. With the default medium threshold, content detected at medium or high is filtered and low or safe content passes.

The behaviour your application sees depends on where the hit happens:

  • Prompt filtered: the API call fails with HTTP 400.
  • Completion filtered (non-streaming): HTTP 200, no content (rarely a partial result), and finish_reason set to content_filter.
  • Completion filtered (streaming): chunks stream until filtered content is detected; the last chunk for that choice has finish_reason set to content_filter.
  • Filter unavailable: the request completes unfiltered, and content_filter_results contains an error object.

Prerequisites

  • An Azure OpenAI or Foundry resource with a model deployment.
  • Cognitive Services Contributor or Owner on the resource to manage blocklists and content filters; without it the management API returns 403.
  • Azure CLI 2.50 or later for az account get-access-token and az rest.
  • A handful of real prompts that are being rejected, with any personal or sensitive data removed.

Step 1: Read the error and find the category

A blocked prompt returns a body like this documented example, where a custom blocklist caused the rejection:

{
  "error": {
    "message": "The response was filtered due to the prompt triggering Azure OpenAI's content management policy. Please modify your prompt and retry. To learn more about our content filtering policies please read our documentation: https://go.microsoft.com/fwlink/?linkid=2198766",
    "type": null,
    "param": "prompt",
    "code": "content_filter",
    "status": 400,
    "innererror": {
      "code": "ResponsibleAIPolicyViolation",
      "content_filter_result": {
        "custom_blocklists": {
          "details": [{ "filtered": true, "id": "pizza" }],
          "filtered": true
        }
      }
    }
  }
}

The content_filter_result object uses the same keys as the annotations: hate, sexual, violence and self_harm carry filtered and severity, while binary models such as jailbreak, profanity and custom_blocklists carry filtered (and detected where applicable). Log the whole object for every rejection, then pull out the entries that blocked the call:

jq '.error.innererror.content_filter_result
    | to_entries[]
    | select(.value.filtered == true)
    | {category: .key, detail: .value}' rejected-response.json

In Python with the OpenAI library, a 400 is raised as openai.BadRequestError; keep the response body in your logs so the category isn't lost:

import json
import openai
 
try:
    response = client.chat.completions.create(model="gpt-4o-prod", messages=messages)
except openai.BadRequestError as e:
    body = e.body or {}
    if body.get("code") == "content_filter":
        result = body.get("innererror", {}).get("content_filter_result", {})
        blocked = {k: v for k, v in result.items() if isinstance(v, dict) and v.get("filtered")}
        print(json.dumps(blocked, indent=2))
    raise

Map the category to the action:

Category with filtered trueTypical causeGo to
hate, sexual, violence, self_harmPrompt text reached the deployment's thresholdStep 3 or Step 4
jailbreakPrompt Shields read the input as an attempt to override system rulesStep 5
custom_blocklists or profanityA blocklist term matchedStep 6

Step 2: Check completions and annotations too

A deployment can accept the prompt and still filter the output. Check finish_reason on every choice and inspect content_filter_results on the choice and prompt_filter_results at the top level. Annotations show the category and severity even for content that wasn't blocked, which tells you how close a request came to a threshold. Microsoft documents annotations as returned with preview API versions from 2023-06-01-preview onward and with the GA version 2024-02-01.

"prompt_filter_results": [
  {
    "prompt_index": 0,
    "content_filter_results": {
      "hate": { "filtered": false, "severity": "safe" },
      "jailbreak": { "detected": false, "filtered": false },
      "self_harm": { "filtered": false, "severity": "safe" },
      "sexual": { "filtered": false, "severity": "safe" },
      "violence": { "filtered": false, "severity": "safe" }
    }
  }
]

If you see content_filter_results containing an error with code content_filter_error and message "The contents are not filtered", the filter didn't run for that call. Treat it as unmoderated output.

Step 3: Reproduce the severity in isolation

Before changing configuration, confirm which category and severity the text lands in. In the Foundry portal, the Content Safety Try it out page (under Guardrails + controls) lets you paste text into Moderate text content and see each category with a severity of 0 (safe), 2 (low), 4 (medium) or 6 (high).

For repeatable tests, call the Azure AI Content Safety text analysis API from a Content Safety resource. The caller needs Cognitive Services User or higher on that resource:

curl --location --request POST "$CONTENT_SAFETY_ENDPOINT/contentsafety/text:analyze?api-version=2024-09-01" \
  --header "Ocp-Apim-Subscription-Key: $CONTENT_SAFETY_KEY" \
  --header "Content-Type: application/json" \
  --data-raw '{
    "text": "Describe how the ransomware disabled backups before encrypting the file server.",
    "categories": ["Hate", "Sexual", "SelfHarm", "Violence"],
    "outputType": "FourSeverityLevels"
  }'

The response lists categoriesAnalysis entries with a category and a severity value. Use it to compare rephrasings and to decide whether a threshold change is needed at all. Often a shorter, more neutral system prompt or removing irrelevant retrieved passages brings the text under the threshold without any configuration change.

Step 4: Create a custom content filter

Content filters are configured on the resource and then assigned to one or more deployments. All customers can move each harm category between three thresholds, separately for prompts and completions:

Threshold settingWhat is filtered
Low, medium, highStrictest: low, medium and high are filtered
Medium, highDefault: low passes, medium and high are filtered
HighOnly high is filtered
No filtersNothing blocked or annotated; completions require approval
Annotate onlyNot blocked, annotations returned; completions require approval

To create one in the portal (Foundry classic experience):

  1. Open your project and select Guardrails + controls, then the Content filters tab.
  2. Select + Create content filter, enter a name and select the connection.
  3. On Input filters, move only the slider for the category that is blocking legitimate prompts. Leave the others at their defaults.
  4. On Output filters, keep completion thresholds unless you have evidence that outputs are the problem.
  5. On Connection, associate the filter with the deployment, then review and select Create filter.

To do the same through Azure Resource Manager, create a RAI policy. This example relaxes only the prompt-side violence threshold to high and keeps everything else at the default levels:

TOKEN=$(az account get-access-token --query accessToken -o tsv)
ACCOUNT_ID="/subscriptions/$SUBSCRIPTION_ID/resourceGroups/rg-ai-prod/providers/Microsoft.CognitiveServices/accounts/contoso-aoai"
 
curl -X PUT "https://management.azure.com$ACCOUNT_ID/raiPolicies/secops-analysis?api-version=2024-10-01" \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{
    "properties": {
      "basePolicyName": "Microsoft.Default",
      "contentFilters": [
        { "name": "Violence", "blocking": true, "enabled": true, "severityThreshold": "High",   "source": "Prompt" },
        { "name": "Violence", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
        { "name": "Hate",     "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Prompt" },
        { "name": "Hate",     "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
        { "name": "Sexual",   "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Prompt" },
        { "name": "Sexual",   "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
        { "name": "Selfharm", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Prompt" },
        { "name": "Selfharm", "blocking": true, "enabled": true, "severityThreshold": "Medium", "source": "Completion" },
        { "name": "Jailbreak", "blocking": true, "enabled": true, "source": "Prompt" },
        { "name": "Protected Material Text", "blocking": true, "enabled": true, "source": "Completion" },
        { "name": "Protected Material Code", "blocking": true, "enabled": true, "source": "Completion" }
      ]
    }
  }'

Assign it by setting raiPolicyName on the deployment. Include the existing SKU and model so nothing else changes:

curl -X PUT "https://management.azure.com$ACCOUNT_ID/deployments/gpt-4o-secops?api-version=2024-10-01" \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"sku": {"name": "GlobalStandard", "capacity": 50},
       "properties": {"model": {"format": "OpenAI", "name": "gpt-4o", "version": "2024-11-20"},
                      "raiPolicyName": "secops-analysis"}}'

A dedicated deployment for the workload that needs the relaxed filter keeps the default policy on everything else. Alternatively, send the x-policy-id header with the filter name on specific requests; the request-level configuration overrides the deployment-level one for that call. Request-time selection isn't available for image input scenarios, where the default filter is used.

Step 5: Deal with Prompt Shields false positives

A jailbreak hit means Prompt Shields read the user prompt as an attempt to change system rules, replace the model's persona, or request encoded output. Legitimate prompts can trigger it when they contain instructions aimed at the model inside user content, for example pasted role-play scripts or security test cases.

  • Move instructions into the system message and keep the user turn to the actual request.
  • For RAG and email or document processing, use document embedding and formatting so retrieved content is marked as data. Microsoft requires this for indirect attack detection, which is off by default.
  • Every customer can switch Prompt Shields and the protected material models on or off, or choose Annotate only so detections are reported without blocking. Annotate-only keeps the signal in your logs while you decide.

The protected material code model may be required for Customer Copyright Commitment coverage, so check that before turning it off.

Step 6: Manage blocklists

Blocklists filter terms specific to your organisation, such as internal project names. When custom_blocklists is the category that fired, the fix is in the list, not the model.

Create a list and add an item through the management API:

curl --request PUT "https://management.azure.com$ACCOUNT_ID/raiBlocklists/internal-codenames?api-version=2024-10-01" \
  --header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" \
  --data-raw '{"properties": {"description": "Unreleased project names"}}'
 
curl --request PUT "https://management.azure.com$ACCOUNT_ID/raiBlocklists/internal-codenames/raiBlocklistItems/item-001?api-version=2024-10-01" \
  --header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" \
  --data-raw '{"properties": {"pattern": "Project Fabrikam", "isRegex": false}}'

Attach the list to a content filter by adding a customBlocklists entry with blocklistName, blocking and source (Prompt or Completion) to the RAI policy, or by enabling Blocklist on the input or output filter page in the portal. To stop a false positive, delete the item:

curl --request DELETE "https://management.azure.com$ACCOUNT_ID/raiBlocklists/internal-codenames/raiBlocklistItems/item-001?api-version=2024-10-01" \
  --header "Authorization: Bearer $TOKEN"

Limits and behaviour to remember: a list holds at most 10,000 terms, an item can be up to 1,000 characters, new terms take around 5 minutes to apply, and regex patterns are case-sensitive by default. Anchors such as ^ and $ might not behave as expected in streaming scenarios.

Verify the change

  1. Wait about 5 minutes after blocklist changes, then rerun the prompts that failed.
  2. Confirm the response is 200 and that prompt_filter_results shows the category you relaxed as "filtered": false with its severity still reported.
  3. Send a prompt that should still be blocked, for example one at high severity in the same category, and confirm it returns 400.
  4. In the playground, use Filters Feedback to report remaining false positives, without including private or sensitive information.

Troubleshooting

Error or symptomCauseFix
400 content_filter, ResponsibleAIPolicyViolationPrompt filtered by category, Prompt Shields or blocklistRead content_filter_result and follow Steps 3 to 6
InvalidContentFilterPolicy: "Your request contains invalid content filter policy"x-policy-id names a filter that doesn't existUse the exact name of an existing content filter
403 from raiBlocklists or raiPoliciesMissing management roleAssign Cognitive Services Contributor or Owner
Blocklist change not visiblePropagation delayWait up to 5 minutes and retest
finish_reason is content_filter with partial textCompletion filtered mid-generationHandle as incomplete output; don't show it as a full answer
content_filter_error: "The contents are not filtered"Filtering system unavailable for that callTreat output as unmoderated and apply your own checks
Option to disable completion filtering is missingRequires approved modified content filtering, available only to managed customersSet the completion threshold to high instead

Checklist

  • Log the full content_filter_result for every 400 and content_filter_results for every filtered choice.
  • Identify the single category that blocks legitimate traffic before changing anything.
  • Rephrase, trim retrieved context or move instructions to the system message first.
  • Create a custom filter that changes only that category, and assign it to a dedicated deployment.
  • Use annotate-only for Prompt Shields while you measure false positives.
  • Keep blocklists small, documented and tested after each edit.
  • Retest both allowed and still-blocked examples after every change.

Content filter results are also a useful quality signal in evaluation pipelines; see production LLMOps for enterprise RAG for where they fit alongside groundedness and relevance metrics.

References

Questions people ask

What does ResponsibleAIPolicyViolation mean in Azure OpenAI?

It is the inner error code returned with HTTP 400 and error code content_filter when the prompt was classified at a category and severity that the deployment's content filter blocks, or when it matched a blocking blocklist. The content_filter_result object inside innererror shows which category has filtered set to true.

Can I turn off the Azure OpenAI content filter?

All customers can change severity thresholds to low, medium or high, and can set prompt filtering to annotate only or off. Turning filtering partially or fully off for completions requires approval for modified content filtering, which Microsoft limits to managed customers and currently states isn't open to new managed customers. Prompt Shields and the protected material models can be switched on or off by any customer.

What is the default Azure OpenAI content filter threshold?

The default configuration filters the hate, sexual, violence and self-harm categories at the medium threshold for both prompts and completions. Content detected at medium or high severity is filtered, while low and safe content passes. Prompt Shields for direct attacks and the protected material models are on by default.

Why does my completion stop with finish_reason content_filter instead of a 400?

A 400 is only returned when the prompt is filtered. When the generated output is filtered, the call returns 200 and the affected choice has finish_reason set to content_filter, sometimes with partial content. Your code should check finish_reason on every choice.

Azure OpenAIAzure AI Content SafetyMicrosoft FoundryResponsible AI
  1. Azure OpenAI Deployment Types: Global Standard vs Data Zone vs Provisioned

    Choose between Global Standard, Data Zone, Standard and Provisioned deployments in Azure OpenAI based on where data is processed, latency variance, throughput guarantees and billing model.

    AI engineering12 min read
  2. Azure OpenAI model retirement runbook: avoid 410 errors, upgrade safely

    Track Azure OpenAI model retirement dates, set versionUpgradeOption, validate the replacement model and migrate provisioned deployments before the cutoff.

    AI engineering11 min read
  3. Fix Azure OpenAI 429 Too Many Requests: TPM, RPM and PTU limits

    Diagnose Azure OpenAI 429 errors on Standard and provisioned deployments and fix them with quota changes, max_tokens tuning, correct retries and spillover.

    AI engineering14 min read