# RL-Recruiters

- **Event:** [OpenEnv Hackathon SF](https://cerebralvalley.ai/e/openenv-hackathon-sf)
- **When:** Mar 7 at 9:00 AM – Mar 8 at 12:00 AM (PST)
- **Where:** Shack15, San Francisco, CA
- **Team:** [Karthik Raja Anandan](https://cerebralvalley.ai/u/kitrakrev), [Shubham Gaur](https://cerebralvalley.ai/u/sgaur2)
- **GitHub:** https://github.com/kitrakrev/rl-recruits/tree/fix/oom-and-prompt-truncation
- **Website:** https://github.com/kitrakrev/rl-recruits/blob/fix/oom-and-prompt-truncation/training/train_grpo.py
- **Demo video:** https://youtube.com/watch?v=1OhIqftmkT8&si=zF1U4ZqiZUZWpdlJ
- **Hugging Face:** https://huggingface.co/spaces/sgaur2/staffing-agency
- **Gallery:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/87

We built an OpenEnv RL environment that trains an LLM to act as a Staffing Agency CEO over a 52-week simulated year. Inspired by VendingBench, it features brutal economics: hired candidates sitting on the "bench" bleed weekly salaries, while massive profits only unlock when multi-role corporate projects are fully staffed before strict deadlines. The agent starts blind. It must spend actions to "interview" candidates, triggering a background LLM Judge to score hidden skills and flag culture risks.

The Problem It Solves:
1. Long-Horizon Planning (Statement 2): We push LLMs beyond shallow, single-turn reasoning. The agent must survive sparse, delayed rewards and manage a complex P&L over 52 steps, avoiding the trap of blindly hiring everyone.
2. HR Workflow (Scale AI): Traditional staffing agencies are bloated and slow. We decentralize human capital routing so every solo freelancer can operate with the capacity of a million-dollar agency.

---

Markdown version of https://cerebralvalley.ai/e/openenv-hackathon-sf/hackathon/gallery/87. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
