# AgentShield

- **Event:** [AI Engineer World's Fair Hackathon 2026](https://cerebralvalley.ai/e/aiewf-hackathon-2026)
- **When:** Jun 27 at 9:00 AM – Jun 28 at 5:00 PM (PDT)
- **Where:** San Francisco, CA
- **Team:** [Kamil Zych](https://cerebralvalley.ai/u/vonHousen)
- **GitHub:** https://github.com/vonHousen/agent-shield-arena
- **Demo video:** https://youtu.be/KvL2YOh7yb4
- **Gallery:** https://cerebralvalley.ai/e/aiewf-hackathon-2026/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/aiewf-hackathon-2026/hackathon/gallery/30

AgentShield Arena is a self-improving guardrails system that solves the problem of securing customer-facing AI agents against adversarial attacks.

# The problem: Companies deploying AI agents (support bots, sales assistants, workflow agents) need protection against misuse. While generic prompt injection and jailbreak defenses exist, the real danger lies in business-rule violations — attacks that exploit the specific logic, permissions, and workflows of your system. A refund bot has different vulnerabilities than a medical-records assistant. Every AI agent has a unique attack surface defined by its tools, permissions, and business rules. Static, one-size-fits-all guardrails cannot cover this.

# Our solution: An automated adversarial arena that creates a red-team/blue-team co-evolution loop, tailored to each specific AI system:

- An Attack Agent probes the target with increasingly sophisticated exploits — not just generic jailbreaks, but system-specific business-rule bypasses (e.g., splitting a $500 refund into multiple sub-$100 requests to circumvent approval thresholds).
- A Defender learns generalized patterns from every successful attack and applies them as runtime guardrails.
- Each round, attacks get smarter and defenses get stronger — the system adapts itself.

# What makes us special: We focus on business-logic violations — the exploits that generic guardrails miss entirely. Every client's AI agent is different: different tools, different rules, different users, different risks. AgentShield doesn't ship a static ruleset — it self-improves by adapting to each client's unique system, discovering and defending against the specific threats that matter for their agent. The adversarial loop ensures guardrails are always evolving, never decaying.

# The hardened Defender then ships as the production guardrails layer — arriving in production already battle-tested against the exact attack surface it will face, with protection that keeps getting better after deployment.

---

Markdown version of https://cerebralvalley.ai/e/aiewf-hackathon-2026/hackathon/gallery/30. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
