
8/26/2025 · Radwa Radwan, Mathias Deschamps
What this post added
Introduces unsafe content moderation integrated into Cloudflare Firewall for AI, leveraging Llama Guard 3 for real-time detection and blocking of harmful prompts targeting LLM endpoints. This feature provides a model-agnostic, edge-native policy layer for AI security, enabling unified detection, analytics, and topic enforcement without application code modification. It utilizes a new asynchronous architecture with parallel, non-blocking requests to detection modules (PII, unsafe topics) deployed on Workers AI with GPUs, ensuring scalability and minimal latency. Rules can be enforced via custom rules in the WAF to log or block based on detected unsafe topics.