WeaveHacks 3: Self-Improving Agents Hackathon with Weights & Biases
Jan 31, 2026 · San Francisco, CA
Cloud platforms like AWS and GCP are powerful, but their interfaces can be complex and unintuitive - especially for new users. Our project makes navigating…

Turn image into GLSL fragment shader code through self-improving agentic loops. https://shader-38ceag7cg-jessies-projects-a62a485b.vercel.app

Chameleon of the Dungeons is an intelligent login security suite designed primarily for ethical hackers and penetration testers to simulate, analyze, and test…

Frontier LLMs are generalists trying to squeeze every benchmark. We showcase a deployed application of "Learning to Discover at Test Time" (Yuksekgonul et al…
n live sports and real-time media, AI agents often fail because they are stateless, relying on static prompts that cannot be adjusted mid-broadcast when…

- Self-learning AI email agent that evolves with you over time - Discord-connected, Gmail-powered workflow - Builds a living relationship map from your…

** github and basic model calling was created before hackday. all core functionarities are created during the weekend ** Archon is a self-improving AI…

This project demonstrates how an AI agent predicts collisions from dashcam footage via a file upload and warns the user of the risk of injury or fatality…

Darwin is an autonomous feedback-to-fix pipeline that gets better at writing code the more humans review it. It scrapes user feedback from Reddit, forums…

The Agent Performance Profiler is a self-improving meta-agent that monitors, analyzes, and optimizes other AI agents. It captures execution traces, detects…

Agent Smith is an open-source operating layer for AI-native teams, letting anyone deploy approved AI agents on demand while engineering guardrails enforce…
Mirror, Mirror is an AI-powered fashion app that acts as a personal stylist, tackling the daily "what should I wear?" dilemma. It learns your style…
"good night" gives your agents REM sleep it looks at your past interactions, finds patterns of failure, and fixes them writes new skills, updates memory /…
Mafia ACE is a self-improving agent system where AI agents play Mafia and autonomously evolve their strategies through the ACE (Agentic Context Engineering)…
Lobster Pot is a safety-first agent guardrail system for OpenClaw that forces all actions into a sandboxed environment "lobster trap" before execution, logs…

Psi is a playground for building and testing software directly from natural language. With a single prompt, Psi scaffolds full projects from scratch, executes…

Self Improving Graphs for LGC-MARL

NVD has been the “single source of truth” for vulnerability intelligence—but budget cuts and downsizing have slowed coverage, leaving security teams with…

Cache: Self-Improving Agent Memory for Claude Code AI coding agents lose all context between sessions—successful patterns vanish, mistakes get repeated…

Most AI music tools are static—they ship a model and that's it. The prompt you write is the prompt you get. Loopism takes a different approach: agents that…
We introduce Stake Chess, a novel chess variant that adds hidden information through a biased staking mechanism, reshaping strategic play under uncertainty…
Project Description: We built a memory test for AI agents. Hundreds of bots on Moltbook share memory strategies with each other, but nobody can verify if…

SmartChat is an AI agent built on the CoALA framework. It uses a three-tier memory system to give context-aware responses and improve from feedback. It has…
QAgent is a self‑healing QA agent that continuously tests web apps, diagnoses failures, generates minimal code fixes, and verifies those fixes automatically…

Zypharmini is a World model powered agent that learns how to design chips. A working chips. it starts with generating random rtl and then use the tools from…
QuantSF is an agent that acts like a quant and creates investment thesis for the stock market. You get a slack message only when the agent detects a big swing…

Jimmy 2.0 is a system that utilizes the YOLO vision model and dataset to identify objects it does NOT know, and then through a pipeline of Claude/GPT models…

Vapor - a platform where you can fine tune models for a specific usecase for browserbase automations, which will save cost over time

Zero Me is a voice-controlled personal assistant with a cute floating blob UI that can manage your todos, calendar, documents, and emails via Notion and…
Full browser Catan game built with React, TypeScript, Tailwind, and Supabase. Includes hex board, roads, settlements, cities, dev cards, trading, robber…

Protect My Weaver is an AI agent observatory and cost management platform built on Weights & Biases Weave. While Weave provides distributed tracing, we…

Since I have some AI evals experience but never tried evaluating agents before, I took this hackathon time to get started with the entire agent evals…

JORDAN is a self-learning autonomous technical sales agent that eliminates deal-killing delays by providing real-time technical validation during live calls…

"Your web, ranked for you." AI-powered Chrome extension that personalizes any webpage by highlighting the content most relevant to you.
WebScout is a self-improving web scraping agent that learns from every success and failure. Traditional scrapers break silently when sites change and repeat…

Problem We lose hours and millions of tokens getting stuck on conceptually similar issues in Claude Code. It gets even worse in teams where knowledge is…

We built a self-improving AI memory backend that prevents “fresh start” chats by continuously capturing a user’s goals, preferences, progress, and browsing…

OneShot is an AI-powered development platform that takes a user's idea (like "build me a todo app") and automatically handles the entire software…

Fractal is an infinite curiosity engine that fundamentally inverts the search paradigm: instead of finding answers, it helps you discover better questions…

coding/completing tasks anywhere

Mothbot is a selfevolving agent which maintains a space ship. It can call tools in chain instea dof running diagnostic on hull, then oxygen system and then…

LoopLess is a self-improving browser automation agent. It uses BrowserBase for cloud browser sessions with live streaming and permanent recordings. Gemini…

OpenEnv for thr VLM to play Duck Hunt game, ready for GRPO training and eval Part of research that called: Hardware-Aware Horizon Minimization for VLM Game…

We don't want to automate decision making. Since AI can write code, everyone knows planning becomes the most important step. However, no one wants to spend…

Microcode is an RLM-powered terminal agent built with DSPy & Modaic. It features a self-reflection loop where one instance monitors another's failures during…

Voice assistant for ADHD/Autism support using Gemini Live API with Weave observability, custom evaluation scorers, and self-improving memory…
Philoagent is a reasoning agent that answers philosophical questions. Chain of Thought + File System as Context + Swarm Browserbase MCP + Weave Inference and…

Clarity is a self-improving, constraint-aware agent system that transforms raw photos of real spaces into realistic renovation visualizations without…

SIMPSEN is an autonomous software engine that builds and improves products without human developers. Give it a mission, roadmap, and a CEO personality. The…

This project acts as a memory extension for real life. It listens to conversations, extracts important details about the people you meet, and remembers them…

**DML (Deterministic Memory Layer)** is an MCP server and Claude skill that gives AI agents structured, auditable, event-driven memory. Instead of appending…
ClinXplain is an advanced medical platform designed to automate clinical documentation, streamline patient management, and accelerate medical research using…

FULLSEND An autonomous GTM agent that ships ideas continuously, builds its own tools, and gets smarter over time.

OpenClaw Trace is a recursive self-improvement pipeline for OpenClaw / Clawdbot session traces. We mine real session traces for grounded signals - errors…
A browser-based terminal coding agent multiplexer; to run agents where you want and control them from whatever device your prefer (mobile, desktop…
Self-improving agents via accumulated reflection. Each iteration: diagnose failure (µf) → minimal skill edit → repeat. The key: Recursive LMs operate over…

An AI-powered research and mapping app that turns natural-language questions into live map updates. You ask about places, routes, or current events in plain…
My project is called Haggler and it is a negotiation/refund agent which autonomously self-improves such that you can get refunds without having the pressure…

A B2B internal platform for building, observing, and improving autonomous agents. It provides trace-level visibility into agent runs, prompt and playbook…
a self-reliant multi-agent system that fuses satellite signals and real-time phone calls to continuously evolve optimal wildfire prediction and response…

An AI-powered mental health documentation and exploration platform for therapists and potential therapy clients. For Therapists (/dap): • Verbally…
