AI Engineer World's Fair Hackathon 2026
Jun 27, 2026 · San Francisco, CA
SplatForge is a robot manipulation agent that trains itself inside the worlds it reconstructs. Point a phone at a tabletop and it turns the scene into a…

https://www.podman.live/ PodMan is the pair programmer for your AI-native dev. team. Writing code got cheap; knowing who's writing what didn't. With…

Rote is a self-improving computer-use agent for real desktop work. When a task is new, Gemini Computer Use drives the browser or Mac app live; Rote records…

Electrical design engineers often have to place and route boards by hand — agonizing over where the buck switching node goes, how far the crystal sits from…

EvoLoRA is an auditable, bounded self-improvement loop for LoRA fine-tuning. You provide a plain-English goal; a MiniMax agent plans the evaluation set, the…

FocalPoint is a chat interface that watches how you read and adapts in real time — no feedback buttons, no ratings, no extra effort from the user. Using your…

This is a test submission.
Agent Arena is an evolutionary tournament where AI agents compete on the same real-world web task, learn from the winner, and get smarter every round. Three…
A Recursive Self-Improving AI that fine-tunes its own weights from failures — meta agent, target agent, LoRA, repeat.
Earthquake simulator for historical DeepShake analysis, shake prediction and damage estimation, follow-up house repair and related architecture enhancement…

An RL agent (GRPO + LoRA) that learns the person-specific privacy attacks universal safety filters miss — eliciting "allowed-but-compromising" leaks judged…

DeFi/Crypto Macro-Regime RSI Engine—an AI system that reads the economy like a team of macro strategists and explains what it means for crypto. Eight…

Introducing Omnicon RSI (Recursive Self-Improvement) - taking your golden summaries from Omnicon - and using them to improve and benchmark conversation…
[Do not include in judging - incomplete result as of deadline, just adding thought process] RecursiveLearner — training-free recursive self-improvement via…
THIS PROJECT IS APPLYING FOR- BEST USE OF GEMINI Self-Taught Operator Self-Taught Operator is an AI agent, built on Gemini 3.5 Flash's computer-use…

Every domain expert who owns "what good looks like" — a compliance lead, a security reviewer — is drowning, hand-checking AI answers one at a time, and the…

Meet Veritas: a silent AI agent that fact-checks live meetings/debates in real time. It plugs directly into LiveKit, handles over 70 languages, and matches…

SIA × GraphWorldModels A self-improving AI loop where every improvement is measured against reality. The problem A deployed AI agent performs at some level…

A self improving system to QA our changes with evidence, and add reviews in github. It uses Google Gemini Managed Agents and Computer use to test de app and…

A satellite classifier that gets smarter from its mistakes — without ever retraining the model. SubStrata labels satellite imagery (Sentinel-2, Google…

Engram is a continuous-learning memory system that lets an LLM keep serving users while deciding what to store in RAG, what durable beliefs to internalize…

OpsGym is a self-improving policy engine for operations teams
The unit economics regarding customer support bots using frontier models from OpenAI and OpenAI won't ever make sense. Research labs reinvest gains into…

Personalized feature creation and discovery Pain point 1. Don't like how an app is organized, either too complex or too simple Pain point 2. Global updates…

Recallroom is a copilot for learners. Learning in this new age can be made so much easier. Watching lectures for students, or anything academic doesn't need…
LLM driven world and reward design and improvement based on policy feedback.
An autonomous visual-agentic data factory that finds a small model's UI-coding blind spots, turns them into a training set, and LoRA-fine-tunes the model on…

We built a recursively self-improving infrastructure for solving complex supply chain optimizations that monolithic LLM's overlook due to the multi-objective…

Self-improving SRE System https://whale-app-pn7gj.ondigitalocean.app/
Tagline Drift detection and automated self-improvement for AI agents: it detects when an agent's accuracy silently degrades and makes the agent learn from…

Superprediction is a continual learning system over prediction markets. We perform rollouts at inference time with Gemma E4B, which has search capabilities by…

Solux is an AI-powered geospatial screening platform for solar and solar-plus-storage development in India and the United States. Solar developers often…
AgentShield Arena is a self-improving guardrails system that solves the problem of securing customer-facing AI agents against adversarial attacks. # The…

GPUSitter is an on call engineer that responds to failures in GPU datacenter clusters. \ We built tools for the agent to inspect NVIDIA sensor data, and…

The Self-Healing Proxy — autonomous runtime debugging, patching, and hot-swapping for production AI systems.
Four ghost agents learn to catch a Pac-Man on a grid. After each round, each ghost reflects on what happened and rewrites its strategy — in plain English…

I built Autodrone, an LLM-driven drone agent that takes a natural-language mission and flies a simulated drone through an Unreal/AirSim environment. The core…

VECHE makes computer-use reliable and cheap on legacy software with no API, no DOM, no MCP. It is made possible by continual learning: a swarm of vision…
GravityMedShield - Human-in-the-Loop Recursive Medical De-identification Medical AI is experiencing an unprecedented boom, but it faces a massive regulatory…

A real paralegal does one task at a time. ABLE runs four Gemini Computer Use agents in parallel — simultaneously navigating a court portal, researching case…

Recursive Shield protects browser-use AI agents from audio prompt-injection attacks during sensitive workflows. We built a fake paper-brokerage environment…
HarnessDreams a macOS health app for AI coding harnesses: while your tools rest, it runs a "Sleep Cycle" that scores how well you and your agent collaborated…
Autonomous SRE agents fail in production for two reasons: a single-shot LLM can't reliably diagnose and fix a real incident, and there's no honest way to know…

TEMPER evaluates the harness wrapped around a model — the system prompt, skill files, and tool definitions — not the model itself. It identifies where that…

Rune is Stack Overflow for the agentic era. Coding agents like Claude Code, Codex, Cursor, and other AI coding harnesses are solving bugs inside private…

AI tutors are everywhere. But are they actually working? Today's AI tutors are patient, available 24/7, and far less intimidating than a human teacher. Yet…

YOLO WallStreet is a modular AI trading research system built to forecast short-term stock moves by combining market data, time-series modeling, news…

DARWIN — an AI that researches quant trading alphas and measurably gets better at it. Each generation it proposes new trading signals, kills the weak, breeds…
Self-learning Judge and Jury for Question Answering Tasks. To improve the llm-as-judge using recursive self-improvement methods for QA use cases.

Skill-Loop for Gemini Computer Use A training-free loop on Gemini 3.5 Flash Computer Use. A kept skill must have passed a verified retry. We measure two…

SeaForge is a Gemini-powered self-learning loop made for Autonomous Surface Vessel manufacturers to test their ships against real Navy mission briefs. Gemini…

Agents can already modify prompts, write tools, store memories, and run experiments. But without rigorous infrastructure, self-improvement is hard to trust…

Most AI tutors reset every session. They don't remember that you struggled with derivatives, or that a human tutor's graph-first explanation actually worked…

Recursive Autonomous Customer Experience Moves the needle from answering a question to solving the underlying ticket.
An agent that draws. You become the artist. ShaderMind is a drawing tool — but the hand holding the pen is an agent. It generates live GLSL sketches; you…

The trust layer for production web agents. A support agent issues real refunds in a live browser; Gemini 3.5 Flash predicts each screen, a self-trained Gemma…
An AI discovers an unknown planet's physics and learns how to discover the physics laws by updating its weights and then records discovered facts into the…

Cadenza is a composition feedback tool that learns each composer's individual voice over time, becoming more useful the more it's used. Most music feedback…

An agentic coding manager that continuously improves to deliver high quality PR with minimum human input. It tracks software engineer requests into coding…

Gambit — a self-improving auto-negotiator. Selling secondhand is a haggling tax: most people anchor wrong, fold early, and leave money on the table, while…
Kun is a mission-control cockpit and runtime for autonomous ML experiment loops — and the open standard those trajectories are logged in. Run-centric tools…

Existing Multiagent systems are typically built around predefined agent roles and static workflow. They have limited ability to create new specialist agents…
LaunchLens helps teams safely test and launch AI agents. Today, agent behavior is often defined across messy chats, docs, shared demos, and last-minute…

Yunaki : shared memory and self-evolving skills for agent teams. Your code, your model, your memory — shared across your team
governance and context for autonomous enterprise agents

Companies are starting to manage teams of AI agents like teams of people — and can't answer the question that follows: is all that AI work actually moving the…
## StoryForge: AI-Powered Children's Storybook Generator **From imagination to illustrated eBook in minutes.** StoryForge is an end-to-end AI pipeline that…

RAX is a real-time, low-latency robotic agent framework powered by Gemini 3.1 Flash Audio and LiveKit Agents. By leveraging a persistent, bidirectional…

Voice Bridge is a real-time, multimodal lip-reading assistive-communication tool for people with ALS. As ALS progressively destroys speech (80–85% of patients…

Nidra (Sanskrit for "sleep") is a self-hosted personal assistant that learns who you are from what you actually do online, then sleeps on it. A privacy-first…
