LLM Firewall
The protective perimeter that keeps your AI agent safe, on-brand, and cost-controlled.
Overview
The LLM Firewall is the layer of protection that sits between your customers and your AI. It screens risky incoming messages before they reach the AI, catches unsafe or off-brand replies before they send, and caps runaway conversations before they burn credits. It's made up of several settings you'll find grouped together on the Agent Configuration → General page, and it's on by default for every workspace — you can tune each part or leave the safe defaults in place.
Features
Filters risky incoming messages
Message Filtering and Topic Auto-Responses block spam, abuse, prompt-injection attempts, one-time codes, and off-topic noise before they ever reach the AI — and answer common questions instantly with your own templates. You can add your own rules for what to ignore.
Catches unsafe replies before they send
Built-in guardrails check every reply for leaked system instructions, over-long messages, and promises the AI can't keep. If the AI ever fails mid-reply, a safety net sends a polite fallback in the customer's language instead of going silent — and you're not charged for the failed turn.
Caps runaway conversations
The Session Limit & Cooldown puts a ceiling on how many times the AI replies in one conversation per day. If a chat goes in circles — or another bot messages your account — it pauses automatically and resumes later, so a single conversation can't quietly drain your credits.
Hands off to your team when needed
The AI pauses the moment a teammate replies and resumes on its own, or hands off on demand with an instant alert — so people and AI never talk over each other, and the tricky cases always reach a human.
Tips
Frequently asked questions
Was this page helpful?