Skip to Main Content

AI Engineer World's Fair Hackathon 2026

Jun 27, 2026 · San Francisco, CA

SplatForge is a robot manipulation agent that trains itself inside the worlds it reconstructs. Point a phone at a tabletop and it turns the scene into a…

SplatForge project preview

https://www.podman.live/ PodMan is the pair programmer for your AI-native dev. team. Writing code got cheap; knowing who's writing what didn't. With…

PodMan: Pair Programmer ❤️ project preview

Rote is a self-improving computer-use agent for real desktop work. When a task is new, Gemini Computer Use drives the browser or Mac app live; Rote records…

rote project preview

Electrical design engineers often have to place and route boards by hand — agonizing over where the buck switching node goes, how far the crystal sits from…

field-ratchet: a recursively self-improving AI for PCB design. project preview

EvoLoRA is an auditable, bounded self-improvement loop for LoRA fine-tuning. You provide a plain-English goal; a MiniMax agent plans the evaluation set, the…

EvoLoRA project preview

FocalPoint is a chat interface that watches how you read and adapts in real time — no feedback buttons, no ratings, no extra effort from the user. Using your…

Focal point project preview

This is a test submission.

youtube.com/…

Agent Arena is an evolutionary tournament where AI agents compete on the same real-world web task, learn from the winner, and get smarter every round. Three…

www.loom.com/…

A Recursive Self-Improving AI that fine-tunes its own weights from failures — meta agent, target agent, LoRA, repeat.

github.com/…

Earthquake simulator for historical DeepShake analysis, shake prediction and damage estimation, follow-up house repair and related architecture enhancement…

shake project preview

An RL agent (GRPO + LoRA) that learns the person-specific privacy attacks universal safety filters miss — eliciting "allowed-but-compromising" leaks judged…

Atlantis project preview

DeFi/Crypto Macro-Regime RSI Engine—an AI system that reads the economy like a team of macro strategists and explains what it means for crypto. Eight…

agentic_ontology project preview

Introducing Omnicon RSI (Recursive Self-Improvement) - taking your golden summaries from Omnicon - and using them to improve and benchmark conversation…

www.loom.com/…

[Do not include in judging - incomplete result as of deadline, just adding thought process] RecursiveLearner — training-free recursive self-improvement via…

youtube.com/…

THIS PROJECT IS APPLYING FOR- BEST USE OF GEMINI Self-Taught Operator Self-Taught Operator is an AI agent, built on Gemini 3.5 Flash's computer-use…

Kevin Durant project preview

Every domain expert who owns "what good looks like" — a compliance lead, a security reviewer — is drowning, hand-checking AI answers one at a time, and the…

Bonsai project preview

Meet Veritas: a silent AI agent that fact-checks live meetings/debates in real time. It plugs directly into LiveKit, handles over 70 languages, and matches…

Veritas project preview

SIA × GraphWorldModels A self-improving AI loop where every improvement is measured against reality. The problem A deployed AI agent performs at some level…

Two Levers project preview

A self improving system to QA our changes with evidence, and add reviews in github. It uses Google Gemini Managed Agents and Computer use to test de app and…

Auto QA project preview

A satellite classifier that gets smarter from its mistakes — without ever retraining the model. SubStrata labels satellite imagery (Sentinel-2, Google…

SubStrata project preview

Engram is a continuous-learning memory system that lets an LLM keep serving users while deciding what to store in RAG, what durable beliefs to internalize…

PiPlan project preview

OpsGym is a self-improving policy engine for operations teams

OpsGym project preview

The unit economics regarding customer support bots using frontier models from OpenAI and OpenAI won't ever make sense. Research labs reinvest gains into…

AcmeBox project preview

Personalized feature creation and discovery Pain point 1. Don't like how an app is organized, either too complex or too simple Pain point 2. Global updates…

autodiscover project preview

Recallroom is a copilot for learners. Learning in this new age can be made so much easier. Watching lectures for students, or anything academic doesn't need…

www.loom.com/…

LLM driven world and reward design and improvement based on policy feedback.

RSIW project preview

An autonomous visual-agentic data factory that finds a small model's UI-coding blind spots, turns them into a training set, and LoRA-fine-tunes the model on…

BrickByBrick project preview

We built a recursively self-improving infrastructure for solving complex supply chain optimizations that monolithic LLM's overlook due to the multi-objective…

Bron project preview

Self-improving SRE System https://whale-app-pn7gj.ondigitalocean.app/

Darwin SRE project preview

Tagline Drift detection and automated self-improvement for AI agents: it detects when an agent's accuracy silently degrades and makes the agent learn from…

SQLWriters project preview

Superprediction is a continual learning system over prediction markets. We perform rollouts at inference time with Gemma E4B, which has search capabilities by…

Ignore All Previous Instructions project preview

Solux is an AI-powered geospatial screening platform for solar and solar-plus-storage development in India and the United States. Solar developers often…

Solux project preview

AgentShield Arena is a self-improving guardrails system that solves the problem of securing customer-facing AI agents against adversarial attacks. # The…

AgentShield project preview

GPUSitter is an on call engineer that responds to failures in GPU datacenter clusters. \ We built tools for the agent to inspect NVIDIA sensor data, and…

GPU Sitter project preview

The Self-Healing Proxy — autonomous runtime debugging, patching, and hot-swapping for production AI systems.

drive.google.com/…

Four ghost agents learn to catch a Pac-Man on a grid. After each round, each ghost reflects on what happened and rewrites its strategy — in plain English…

OmegaPacman: Interpretable Multi-Agent Learning project preview

I built Autodrone, an LLM-driven drone agent that takes a natural-language mission and flies a simulated drone through an Unreal/AirSim environment. The core…

Autodrone project preview

VECHE makes computer-use reliable and cheap on legacy software with no API, no DOM, no MCP. It is made possible by continual learning: a swarm of vision…

gyazo.com/…

GravityMedShield - Human-in-the-Loop Recursive Medical De-identification Medical AI is experiencing an unprecedented boom, but it faces a massive regulatory…

GravityMedShield project preview

A real paralegal does one task at a time. ABLE runs four Gemini Computer Use agents in parallel — simultaneously navigating a court portal, researching case…

Aperacium project preview

Recursive Shield protects browser-use AI agents from audio prompt-injection attacks during sensitive workflows. We built a fake paper-brokerage environment…

www.loom.com/…

HarnessDreams a macOS health app for AI coding harnesses: while your tools rest, it runs a "Sleep Cycle" that scores how well you and your agent collaborated…

Dream Harness project preview

Autonomous SRE agents fail in production for two reasons: a single-shot LLM can't reliably diagnose and fix a real incident, and there's no honest way to know…

A.S.F project preview

TEMPER evaluates the harness wrapped around a model — the system prompt, skill files, and tool definitions — not the model itself. It identifies where that…

Temper project preview

Rune is Stack Overflow for the agentic era. Coding agents like Claude Code, Codex, Cursor, and other AI coding harnesses are solving bugs inside private…

Rune project preview

AI tutors are everywhere. But are they actually working? Today's AI tutors are patient, available 24/7, and far less intimidating than a human teacher. Yet…

tutorAI project preview

YOLO WallStreet is a modular AI trading research system built to forecast short-term stock moves by combining market data, time-series modeling, news…

Yolo WallStreet Model project preview

DARWIN — an AI that researches quant trading alphas and measurably gets better at it. Each generation it proposes new trading signals, kills the weak, breeds…

www.loom.com/…

Self-learning Judge and Jury for Question Answering Tasks. To improve the llm-as-judge using recursive self-improvement methods for QA use cases.

Judy project preview

Skill-Loop for Gemini Computer Use A training-free loop on Gemini 3.5 Flash Computer Use. A kept skill must have passed a verified retry. We measure two…

Gemini CUA Skill Loop project preview

SeaForge is a Gemini-powered self-learning loop made for Autonomous Surface Vessel manufacturers to test their ships against real Navy mission briefs. Gemini…

SeaForge project preview

Agents can already modify prompts, write tools, store memories, and run experiments. But without rigorous infrastructure, self-improvement is hard to trust…

Agent University project preview

Most AI tutors reset every session. They don't remember that you struggled with derivatives, or that a human tutor's graph-first explanation actually worked…

Tutor Loop project preview

Recursive Autonomous Customer Experience Moves the needle from answering a question to solving the underlying ticket.

www.loom.com/…

An agent that draws. You become the artist. ShaderMind is a drawing tool — but the hand holding the pen is an agent. It generates live GLSL sketches; you…

Shadermind project preview

The trust layer for production web agents. A support agent issues real refunds in a live browser; Gemini 3.5 Flash predicts each screen, a self-trained Gemma…

drive.google.com/…

An AI discovers an unknown planet's physics and learns how to discover the physics laws by updating its weights and then records discovered facts into the…

IntoTheUnknown project preview

Cadenza is a composition feedback tool that learns each composer's individual voice over time, becoming more useful the more it's used. Most music feedback…

Cadenza project preview

An agentic coding manager that continuously improves to deliver high quality PR with minimum human input. It tracks software engineer requests into coding…

Marine Gravity project preview

Gambit — a self-improving auto-negotiator. Selling secondhand is a haggling tax: most people anchor wrong, fold early, and leave money on the table, while…

gambit.nudepineapple.com/…

Kun is a mission-control cockpit and runtime for autonomous ML experiment loops — and the open standard those trajectories are logged in. Run-centric tools…

9320 project preview

Existing Multiagent systems are typically built around predefined agent roles and static workflow. They have limited ability to create new specialist agents…

github.com/…

LaunchLens helps teams safely test and launch AI agents. Today, agent behavior is often defined across messy chats, docs, shared demos, and last-minute…

LaunchLens project preview

Yunaki : shared memory and self-evolving skills for agent teams. Your code, your model, your memory — shared across your team

www.loom.com/…

governance and context for autonomous enterprise agents

RevMem project preview

Companies are starting to manage teams of AI agents like teams of people — and can't answer the question that follows: is all that AI work actually moving the…

liminal project preview

## StoryForge: AI-Powered Children's Storybook Generator **From imagination to illustrated eBook in minutes.** StoryForge is an end-to-end AI pipeline that…

Muse the Agent project preview

RAX is a real-time, low-latency robotic agent framework powered by Gemini 3.1 Flash Audio and LiveKit Agents. By leveraging a persistent, bidirectional…

RAX project preview

Voice Bridge is a real-time, multimodal lip-reading assistive-communication tool for people with ALS. As ALS progressively destroys speech (80–85% of patients…

VOICE-BRIDGE project preview

Nidra (Sanskrit for "sleep") is a self-hosted personal assistant that learns who you are from what you actually do online, then sleeps on it. A privacy-first…

Nidra project preview