# JITAgent

- **Event:** [The Harness Engineering & Model Wrangling Hackathon](https://cerebralvalley.ai/e/mongodb-nyc-hackathon)
- **When:** Sat, Sep 26 at 9:00 AM – 10:00 PM (EDT)
- **Where:** New York, NY
- **Team:** [Susan Poudel](https://cerebralvalley.ai/u/Susanpdl), [Shaurabh Ghimire](https://cerebralvalley.ai/u/ghimireshaurabh)
- **GitHub:** https://github.com/ShaurabhGhimire/AgentJIT
- **Demo video:** https://drive.google.com/file/d/15-0KIulGUoXbnnNpClKsVPUNmGTul9mT/view?usp=sharing
- **Gallery:** https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery
- **Page:** https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/82

# AgentJIT: a JIT compiler for LLM agents

## The problem
Production agents pay full LLM cost and latency for every task, even the thousandth one that follows the same path as the first. Existing shortcuts (caching, plan replay, skill libraries) have no validity checks, so they fail silently when the world changes. And when a shortcut breaks mid-task, restarting in the LLM is unsafe: if a refund was already issued or an email already sent, restarting repeats it.

## Our solution
AgentJIT treats an agent like code in a JIT compiler.

- **Interpret, then compile.** The LLM handles tasks at first while every tool call is traced. When a task path becomes frequent and stable, the traces are generalized into a template, an LLM drafts code for it, and the code is accepted only if it exactly reproduces every recorded trace.
- **Learned guards.** Validity checks (e.g., "currency is USD," "order has one shipment") are inferred statistically from real traffic, never written by the LLM, and placed before any irreversible action wherever possible.
- **Safe deoptimization.** When a guard fails mid-task, the LLM takes over from that exact point with a record of what already happened. Every side effect goes through a journal, so a refund or email is never repeated.
- **Measured silent errors.** Some changes trip no guard. A sample of compiled runs is re-checked against the LLM, and each routine carries a statistical bound on its error rate. When the bound slips, the routine is demoted and recompiled.
- **Knows when not to compile.** Task types too varied to compile safely stay with the LLM.

## Why it's technically new
Skill libraries and caches reuse behavior but have no guards, no fallback semantics, and no error bounds. JIT compilers have all three, but never had to deoptimize after irreversible real-world actions. AgentJIT's core contribution is safe fallback across committed side effects, combined with a measured bound on silent errors.

## MongoDB Atlas
- **Vector Search** routes each task to its compiled routine.
- **A unique index on the side-effect journal** guarantees each action is recorded once, which makes mid-task handoff safe.
- **Change streams** hot-swap new routines into running workers and trigger recompiles.
- **Aggregation pipelines** find frequent, stable task paths.
- **Transactions** promote and roll back routine versions atomically.
- **Time series collections** track cost, failures, and error bounds live.

## More from The Harness Engineering & Model Wrangling Hackathon

- [Team Darwin](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/79)
- [Raaya](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/80)
- [Tomok](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/81)
- [Tokeneyezed](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/83)
- [DDN](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/84)
- [satyam](https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/85)

---

Markdown version of https://cerebralvalley.ai/e/mongodb-nyc-hackathon/hackathon/gallery/82. Site index for agents: https://cerebralvalley.ai/llms.txt · full text: https://cerebralvalley.ai/llms-full.txt
