Badlands
Built at National Security Hackathon (by Army xTech) · May 2, 2026 · San Francisco, CA
Badlands is a cyber self-play environment for long-horizon, co-evolving attacker and defender evaluation. It helps mission owners see how model behavior, capability, cost, and mission risk evolve over weeks or months instead of relying on one-off benchmarks. Badlands simulates a mission system with users, identity, applications, files, tickets, services, telemetry, deadlines, and operational disruption. Attacker, defender, and mission-user agents interact inside that stateful world over repeated episodes. The core idea is co-evolution: attackers and defenders adapt through role-visible feedback and role-isolated memory while green/user activity keeps mission pressure alive. This lets teams study how agent behavior changes as context, memory, tokens, latency, test-time compute, and system state accumulate. The problem we solve is continuous cyber capability measurement. As AI systems become more capable and inference time compute becomes more valuable than model weights, mission owners need to know whether cyber risk is rising, whether defenders are improving, and where automation creates new operational harm. Badlands makes that measurable and affordable by running locally on OpenAI-compatible model endpoints, compacting long trajectories when needed, and replaying every score from canonical JSONL evidence.