Skip to Main Content

AgentShield

Built at AI Engineer World's Fair Hackathon 2026 · Jun 27, 2026 · San Francisco, CA

AgentShield — Demo video

AgentShield Arena is a self-improving guardrails system that solves the problem of securing customer-facing AI agents against adversarial attacks. # The problem: Companies deploying AI agents (support bots, sales assistants, workflow agents) need protection against misuse. While generic prompt injection and jailbreak defenses exist, the real danger lies in business-rule violations — attacks that exploit the specific logic, permissions, and workflows of your system. A refund bot has different vulnerabilities than a medical-records assistant. Every AI agent has a unique attack surface defined by its tools, permissions, and business rules. Static, one-size-fits-all guardrails cannot cover this. # Our solution: An automated adversarial arena that creates a red-team/blue-team co-evolution loop, tailored to each specific AI system: - An Attack Agent probes the target with increasingly sophisticated exploits — not just generic jailbreaks, but system-specific business-rule bypasses (e.g., splitting a $500 refund into multiple sub-$100 requests to circumvent approval thresholds). - A Defender learns generalized patterns from every successful attack and applies them as runtime guardrails. - Each round, attacks get smarter and defenses get stronger — the system adapts itself. # What makes us special: We focus on business-logic violations — the exploits that generic guardrails miss entirely. Every client's AI agent is different: different tools, different rules, different users, different risks. AgentShield doesn't ship a static ruleset — it self-improves by adapting to each client's unique system, discovering and defending against the specific threats that matter for their agent. The adversarial loop ensures guardrails are always evolving, never decaying. # The hardened Defender then ships as the production guardrails layer — arriving in production already battle-tested against the exact attack surface it will face, with protection that keeps getting better after deployment.

Team