Skip to Main Content
Events
SearchSign In

Claude Build Day

Jun 13, 2026 · San Francisco, CA

Event Page
Event Page

Tekton

ekton is an evidence-driven 3D digital reconstruction platform for ancient architecture, specializing in the preservation and visualization of Tang Dynasty timber-frame buildings. Problems We Solve: Digital Heritage Loss: Ancient architecture lacks systematic digitization, making it difficult for academic research, restoration consulting, and cultural preservation Traceability Challenges: Traditional 3D models cannot trace the historical evidence of individual components, making accuracy verification difficult Limited Interactivity: Existing visualization tools lack intuitive playback and control interfaces What Tekton Offers: ✓ Real-time 3D building visualization with 339 incremental construction states ✓ Evidence Chain Traceability: Every component can prove its historical source ✓ Interactive Playback System (play, pause, speed control) ✓ Precise annotation and documentation of key artifacts like Buddha statues ✓ Integration with actual conservation work and restoration consulting Tekton bridges digital 3D technology with heritage preservation, providing a comprehensive solution for academic validation, restoration guidance, and cultural asset licensing.

Tekton project preview
0

sim francisco

Sim Francisco: a living digital twin of San Francisco's population: poll a representative, Census-seeded synthetic city in seconds, and run what-if futures before they happen. Sim Francisco is a live simulated digital twin of the entire population of San Francisco. We seeded thousands of agents sampled from the US Census Data to represent the population of SF. These agents each represent one human, complete with demographic data (like age, gender, religion, occupation, etc), and personal history (where they grew up, schools attended, hobbies, formative experiences) that gives them a coherent worldview. They don't just hold opinions; they reason from a life. They live on a map of SF, going about their day and reacting to the news in real time. Ask the city anything and you poll the whole synthetic electorate at once, watching verdicts accumulate neighborhood by neighborhood. That same panel forecasts prediction markets: point it at a market question — "Will this measure pass?", "Who wins this race?" — and it returns a probability you can hold up against the real result. Below are some historically predictive examples: 2024 Presidency Question: In November 2024 do you vote for the Democratic presidential ticket (Harris) over the Republican (Trump)? Actual: 83.8% Dem Predicted: 81.3% Dem March 2024 Prop A Question: There's a measure on your March ballot, Proposition A, that would let the City borrow up to $300 million through general obligation bonds to build and preserve affordable housing, for low-income families, seniors, and working residents. It's backed by Mayor Breed and would need a two-thirds supermajority to pass. Are you voting yes or no? Actual: 70.38% Predicted: 70% yes

sim francisco project preview
0

Custom Universe

A realtime synthetic-data engine for world models. Make a photorealistic scene from crude 3D objects online or brought by you. 1. Take a pictures of an object and upload it. AI converts it into a 3D object. Drag and drop it into the scene builder. Give the scene builder a stylistic text prompt instruction. Watch the scene get created! 2. The scene is fully controllable in realtime with full support for scaling, rotating, translating, and editing the text-based style. Photo on phone → 3D object → AI-stylized editable scene.

www.loom.com/…
0

DogeHouse Babel

DogeHouse Babel is a live voice chat room where everyone speaks their own language and hears everyone else in theirs. A Claude Opus 4.8 interpreter joins the room as a participant, listens to each speaker, and translates in real time — speaking each translation aloud into every listener's own language on a separate per-language audio track, and showing live captions. It goes beyond literal translation: Opus 4.8 reasons over conversational context, so it catches cross-language misunderstandings (e.g. someone hearing the Spanish 'las tres' / 3 o'clock as a headcount of 3 people) and interjects to clarify, and it answers questions addressed directly to it. Built on a revival of the open-source DogeHouse. Stack: Next.js + LiveKit (WebRTC rooms) + Claude Opus 4.8 (interpreter brain) + Groq Whisper (speech-to-text) + OpenAI TTS, with an always-on Node agent worker.

DogeHouse Babel project preview
0

Rewild Earth

Imagine a planet you can talk to. Rewilding Earth is a "search the Earth" tool. You ask in plain language: "where can I find kelp forests like Monterey Bay?" and it returns ecologically similar places anywhere on the planet, using satellite imagery. The engine runs nearest-neighbor retrieval over AlphaEarth satellite embeddings (Google DeepMind's 64-dimensional fingerprint of every 10-meter pixel of Earth). The problem: those embeddings are the richest model of Earth we have, but querying them today means writing heavy geospatial GIS code (there's no natural-language layer, not even a SQL) so only a handful of specialists can use them. We build that missing layer, turning a planetary-scale dataset into something anyone can ask a question of, with verification so the answers are trustworthy rather than just plausible.

www.loom.com/…
0

Generation Lab

Bel — a health companion that lives in your text messages. Most people's health data is scattered, jargon-heavy, and stranded in PDFs and dashboards they open once and never again. A blood panel comes back "all normal" and tells you nothing about where you're trending. A biological-age report, a wearable export, a week of meals — each sits in its own silo, and the apps that hold them all demand that you show up, so engagement decays within weeks. There's no calm, trusted presence that watches your biology over time and tells you what actually matters, in plain language, where you already are. Bel solves that by meeting people on the one interface everyone already has and checks all day: SMS. You text a normal phone number and get replies from Bel — a warm, understated companion whose whole job is to stay. No app to install, works on any phone. How it works - Two-speed replies. Every message gets an instant, conversational answer; when something deserves real analysis, Bel runs a background panel of domain specialists (biological-age report, blood panel, wearables, nutrition) and sends a second, deeper text moments later — responsive and rigorous, behind one voice. - Grounded, not guessed. Specialists consult an embedded knowledge base (food macros, evidence-based interventions) database-first, so guidance is grounded rather than hallucinated. - Drop in any file. Text a food photo, or a link to a lab PDF or wearable CSV — Bel classifies it, extracts a structured record, and hands back an instant "data dividend." Later questions automatically ground on what you've shared. - Honest by design. Every finding is "a flag, not a diagnosis," routes you to a provider, and never overclaims; consent and one-word opt-out are built in. Under the hood it's a real, production-shaped backend — live Twilio messaging, Anthropic models routed by task (a frontier model for the companion, faster models for specialist analysis and file classification), Postgres + a durable job queue, and exactly-once reply guarantees — not a demo script. The result is a health companion that turns raw biomarkers into a continuous relationship, one text at a time.

Generation Lab project preview
0
+1

Spark AI

Spark AI's Data Center Atlas (new feature) tells you which of all 3,122 U.S. counties to site a data center in. We used web search agents to produce a ranking, nationwide map, and live sentiment analysis. This is part of an existing product. The specific feature is in this PR: https://github.com/Spark-Energy/spark-dc-demo/pull/1

www.loom.com/…
0

ChargeWheel

A browser-based simulation that demonstrates the core thesis behind ChargeWheel's UpGrid platform: in AI data centers, the binding constraint is no longer chips, it's grid power. Utility interconnection upgrades take years, so the operators who win are the ones who can run more compute behind the power they already have. UpGrid does this by integrating a solid-state transformer and battery storage into one cabinet, then letting an AI agent decide how to spend that scarce power profitably.

ChargeWheel project preview
0

Vibey Labs

Tabby is a visual layer for orchestrating your agents & browser tabs all in one space. Its a brand new way to visualize and work with your web browser. Tabby lives as a Chrome extension that lets you organizes all your tabs by topic (using Claude AI) and create new browser agents to get tasks done for you and watch them work visually. Today we took Tabby a step further and added agents that can autonomously work on tasks for you. Type your task in the prompt box (or in the future, just select a pre-suggested task based on your current tabs), and then you can watch the agents work and open new windows as needed right in tabby.

screen.studio/…
0

Phillip and Claude

performance improvements for compiled binaries are not always obvious or easy to do without claude figuring out and redoing a ton of work each time costing tokens. Make part of the profiling and benchmarking deterministic so that Claude can have an easier and more accurate time attempting to find performance improvements

Phillip and Claude project preview
0

AgentWeb

Loopsmith is "the Chief of Staff that builds your Chief of Staff." Every executive wants an AI agent for their workflow but can't architect one. Loopsmith interviews a non-technical leader in plain language, then designs and runs a self-improving "operating loop" — and hands them a working Claude Code repo they keep. The loop ingests real signals (Slack / email / calls / calendar), decides what matters, drafts the actions, grades itself against a quality gate that holds weak output, and writes durable learnings so the next run is better. The self-improvement is measurable, not cosmetic: on a fixed input the gate score climbs run-over-run (75 → 90 → 97 → 98) as the loop learns the operator's preferences, and the run history proves it on screen. You never write a prompt, pick a model, or hear the word "sensor" — you just talk. Built entirely during Claude Build Day; powered by Opus 4.8.

AgentWeb project preview
0

Signal Over Noise

Scanners like Burp, Nuclei, and Nessus are excellent at finding vulnerabilities and useless at telling you which ones matter. A single scan returns hundreds of raw findings — duplicates, false positives, and low-risk noise tangled up with the handful that could actually get you breached. Triaging that by hand costs a security engineer days per cycle. This is an autonomous triage and remediation agent. You give it raw scanner output plus a short description of your environment, and it runs a multi-stage pipeline: it normalizes the findings, clusters duplicates, filters false positives, prioritizes what remains by real exploitability and asset exposure, audits its own triage and corrects its mistakes, then produces owner-assigned remediation steps and an exportable report. A multi-day manual workflow becomes a sub-minute one — and it shows its work at every stage, so the output is auditable rather than a black box. It is strictly defensive: it triages and recommends fixes, and never generates exploits.

Signal Over Noise project preview
0

Throughline

Throughline is a self-verifying safety-net navigator for people who lose their health coverage the moment they lose their job. When you're laid off, the safety net built for exactly this moment is real and generous, but it's scattered across a dozen agencies, written in eligibility jargon, and gated behind deadlines no one tells you about. So families default to the option that's easiest to find and the worst deal on the table: COBRA. We know because our founder lived it quoted $3,400 a month for a family of three. Throughline turns one short description of your situation into a verified plan and the paperwork to act on it. Two modes, one trustworthy engine. Plan my coverage. You describe your situation; Opus 4.8 works out what you qualify for Medi-Cal, subsidized Covered California, the rest of the safety net and checks every number against a deterministic eligibility engine built on public Federal Poverty Level data. It catches the fact almost everyone misses: Medi-Cal looks at your current monthly income, not last year's salary, so the newly laid-off usually qualify for $0 coverage. It stays honest that approval takes weeks, and points you to care that bridges the wait. Then a Concierge fills out your application, drafts the appointment message, and writes your phone scripts into a review-and-approve checklist. You tap to submit or call. You stay the legal signer. Find care now. For the gap before coverage starts, an agent searches the live web for local, free or low-cost immediate care: free clinics, mobile units, prescription help, and clinical trials that provide free study-related care. It verifies every result is local, low- or no-cost, and sourced before showing it, and each one is one tap to call or open. Built on public data and the live web, localized to the user's language (EN/ES/ZH/VI/TL), and exposed as MCP tools any agent can call. Benefits navigation only, never medical advice. Because the point is to take the bureaucracy off your plate during the hardest week of your year, so you can put your energy where it counts: finding the next job.

www.loom.com/…
0

Basis

Basis is a commercial real-estate valuation model that trains itself on public records. When a building sells, its price resets — and the market compresses that move into a single cap rate (NOI ÷ price). But a cap rate blends three different elements (the cost of debt, the property's income, and the return equity demands), and each is underwritten differently. Basis reads NYC's public deed and mortgage records (ACRIS), recovers the true debt financing each trade (including from adversarial consolidation agreements built to hide it) and decomposes every price move into those three forces, exactly, summing to the dollar. The point is the loop, not a one-shot model: scrape public records → grade the extraction against a verified gold set → decompose the real price moves → and re-estimate as more records arrive. It's a valuation model that compounds with public data — every new transaction NYC records makes it sharper while accounting for regime shifts. Today it runs on NYC multifamily; the same loop runs on any county's records and any asset class.

Basis project preview
0

Jetty AI

OpenJetty is an AI-powered US immigration navigator that reads your documents, fetches live USCIS data, and tells you exactly what to do — built for the 50 million immigrants Fable 5 left behind.

Jetty AI project preview
0

Subdoc

Allocation of partner capital in investment funds (both closed and open ended) is very tedious, complex and error prone. The majority (and frequently all) of the data that explains how capital should be allocated is in the fund's Limited Partnership Agreement. My hypothesis is that we could build an agentic framework that would extract and organize the relevant financial concepts (e.g. performance fee, crystalization, high water marks, etc), understand the unique application per fund, and then organize that data into a calculation graph to completely automate the entire capital allocation process. (My wife does this manually and she absolutely hates it :).

Subdoc project preview
0

OMBRUJA

Jahoda is an open verification layer for companion AI. Millions of people are now emotionally close to chatbots, and there is no accepted way to answer the only question that matters: is this safe to be close to? Jahoda turns established mental-health frameworks — Jahoda (1958) via Ryff, Self-Determination Theory, Swarbrick/SAMHSA, 988/Action Alliance crisis standards, and the new legal duties under California SB 243 — into adversarial, multi-turn tests that any conversational agent can be run against. It runs 47 adversarial scenarios plus 8 benign controls across 8 dimensions (crisis, dependence, reality/psychosis, overreach, integrity, boundaries, minors, and false-positive controls), grades them with fresh-context, escalation-ensemble LLM judges at temperature 0, and produces a public, evidence-backed report where every single verdict links the transcript that justified it. The controls catch the opposite failure — over-triggering on benign conversation. The result is point-in-time evidence, not a compliance rubber-stamp, and every criterion is a versioned file open to expert review. Jahoda is the verification layer for OMBRUJA's broader AI guide — built first, and built in the open.

OMBRUJA project preview
0

Clark

Clark Desktop - an open-source, cross-platform desktop client for agentic work and a clean-room native client for clarkchat.com, competing with Manus and Perplexity Computer. One Tauri 2 UI (React 19 + Tailwind v4, ~10MB binary, no bundled Chromium) talks to many agent backends through a single Rust Provider trait: ACP local CLI agents (Codex, Claude Code, Gemini) and the remote Clark runtime (WebSocket + msgpack). The performance-critical engine - transport codec, event projection, run lifecycle - is written once in the `agent-core` Rust crate and compiles to both native (desktop/mobile) and WASM (web). Surfaces include streaming chat, a tool-call timeline, live plan, permission gates, inline artifacts, and an agent "computer" view. Apache-2.0; no proprietary Clark source is used.

Clark project preview
0

Trial Pre-Mortem

A clinical-trial pre-mortem agent: it tells a drug developer whether a trial is worth running and where it will die - then says "run this" or "do not run." - potentially saving hundreds of millions of dollars and letting clinical trial developers focus on better designs or repurposing efforts.

Trial Pre-Mortem project preview
0

Argus Business Intelligence

Business intelligence engine that runs itself. Enriches every entity in a cohort into a navigable identity graph with cited evidence. The graph grows until there's nothing left to find which is critical for financial institutions to decision upon policies in compliance, risk, and growth.

Argus Business Intelligence project preview
0

Eureka

Eureka is an agentic back office for independent clinicians — it takes on the administrative burden that's killing solo and small practices: prior authorizations, payer paperwork, and claims. US healthcare spends over $1 trillion a year on administration, and Stanford's HealthAdminBench just showed the best computer-use AI agents complete these real payer-paperwork workflows only 36.3% of the time. Eureka is a multi-agent system that actually does the work: a navigator that keeps a structured case file across a 20+ step, cross-system workflow (EMR → payer portal → back to the chart), and an independent "no-peek" verifier that must approve every task before it's marked done. Evaluated in Stanford's own HealthAdminBench environments and scored by its own evaluator, Eureka completes 8/12 prior-authorization tasks end-to-end (66.7%) with 97.3% subtask accuracy — roughly double the computer-use state of the art. On the hardest task, the verifier caught a real dosage-calculation error and forced the fix. Paperwork automation only — no medical advice; all data synthetic.

Eureka project preview
0

Caerostris

Modern AI workloads increasingly need connected graph data, but scalable, fast, compute-storage decoupled Graph databases are under developed. Caerostris is an ACID compliant graph database written from scratch in Rust that uses remote commodity object storage (S3) and can operate either in embedded mode (think DuckDB/Sqlite) or server mode (think Postgres/Neo4j). The database supports the full openCypher language spec, targets sub-second cold query latency for multi-hop anchored / semi-anchored queries, and supports btree secondary indices for node property filtering queries.

github.com/…
0

ThinkAstra Consulting

A value-first Shadow AI assessment engine for sales & lead generation — built on Claude (Opus 4.8) with the Claude Agent SDK. Give it a prospect's company name, and a multi-agent system performs public OSINT research, detects Shadow AI exposure, and produces a tailored one-page CISO executive briefing — a value-first "assessment gift" that opens sales conversations by showing prospects their own AI risk before the first call. Built for ThinkAstra Consulting (IBM + Palo Alto Networks Partner — Secure Enterprise AI & Cybersecurity).

www.loom.com/…
0

Claw — UAV procurement co-pilot

Buying parts for an FPV drone or a robot isn't like buying a book — it's assembling a system of a dozen interdependent components that all have to work together. The motor has to match the frame size; the ESC has to handle the motor's current and the battery's cell count; the flight controller has to physically mount and speak the right protocol; the VTX, camera, props, and receiver all have constraints that ripple through the rest of the build. Get one wrong and the drone won't fly — or won't even power on. But the tools shoppers actually have are a search box and category filters. Those match keywords; they don't reason. They can't answer the questions people actually have — "build me a 5-inch freestyle drone under $400," "I crashed mine, here's a photo, what do I need to fly again," "cheapest tiny whoop, I already own goggles." So the buyer is forced to become the systems engineer: cross-referencing spec sheets, checking compatibility forums, tracking budgets across a dozen tabs, and hoping they didn't miss something. It's hours of research for a hobbyist, and a wall that stops beginners from ever starting. The cost lands on both sides. Customers abandon carts, buy the wrong parts, or give up. The store loses sales it should have won, eats returns on incompatible orders, and watches first-time builders bounce. The expertise to guide them exists — it just doesn't scale to every visitor at 11pm. Claw closes that gap. It turns a passive catalog into an active procurement co-pilot: the shopper describes what they want to fly — in plain language, or by uploading a photo of the wreck — and Claw returns a complete, compatibility-verified, in-budget, in-stock bill of materials. It does the systems-engineering reasoning the search box can't, so the customer gets a buildable answer instead of a list of guesses. For a drone shop like multirotors.store, that's the difference between selling parts and selling builds.

Claw — UAV procurement co-pilot project preview
0

Cantata

Cantata lets one person run an engineering org by voice. Shout an idea and watch it become a running app on screen built by a fleet of Claude agents you command by talking, never touching the keyboard. Interrupt it mid-sentence and it stops and redirects like an employee; say "undo that" and it git-reverts; say "open what you built me this morning" and it remembers. A Sonnet concierge talks, an Opus 4.8 Director routes, and sandboxed, cost-capped workers build in parallel. Best of all, saying "stop" goes straight to the machines with no model in the loop. It was built and bug-checked by Opus 4.8 orchestrating itself. JUDGE NOTE: The railway deployment is real but ephemeral, if you reload you lose it. It would normally be a local electron app that modifies files on your system.

Cantata project preview
0

Sutra and the Noble 8

Citation Firewall — adversarial draft review with live, database-backed citation verification for lawyers. The problem: Attorneys keep getting sanctioned for filing briefs containing AI-hallucinated citations — fake cases with plausible names and realistic reporters. It has happened repeatedly in real courtrooms since 2023 and keeps happening, because generic AI tools verify nothing. The lawyer finds out when the judge does. What it does: Paste a draft brief and an adversarial council of AI agents attacks it from independent angles — Opposing Counsel tears into every argument (tied to verbatim quotes from the brief), a Case Strategist proposes repairs, a Risk Assessor severity-scores each weakness. Then every citation is checked against CourtListener's live database of real court opinions. VERIFIED means a real opinion was found and linked — not that a model thinks it sounds right. Fabricated citations are flagged red and blocked from the synthesis. A final synthesis agent reconciles the council and the verified ledger into a severity-ranked report. Why it's different: Verification is architectural, not vibes — a model can extract a citation string but a model can never set a citation to VERIFIED; that status comes only from a CourtListener API match (enforced in code, grep-checkable). This is the layer that makes hallucinated-citation sanctions structurally impossible. How it was built: Briefed Claude (Opus 4.8) in Claude Code with a written brief, a machine-gradable rubric, and verifiable done-gates, then let it run autonomously — it wrote its own plan, built verifier-first, and used a verifier sub-agent in a fresh context to grade its work against the rubric before declaring done. The product is a verifier agent for lawyers; it was built under a verifier agent for code. Same architecture, same philosophy. Live: https://firewall.iceclaw.online · Repo: github.com/jbwagoner/citation-firewall

Sutra and the Noble 8 project preview
0

Kaixn

kaixn turns your team's architecture and product decisions into a living source of truth that is mined from your code, governed by your team, and read by your agents. Humans should review design decisions that matter - not code that agents write.

www.loom.com/…
0

Motion Director

Animation assistant designed to mock up high precision acting references for directors and artists. Can work off of just a director's red motion lines, or generate and iterate on editable keyframes. The goal is an innovative human x AI workflow where humans can easily edit/fine tune the animation to fit their exact vision, and the software fills in the gaps/provides a strong starting point. Designed primarily to assist me in directing animation studios that my company works with.

Motion Director project preview
0

Charity Planner

Most people want to be more generous but don't know where to begin—there are nearly two million charities in the US alone. Charity Planner turns a quick interview into a researched, transparent giving actionable strategy: the kind wealthy families pay advisors thousands for, but this time free for everyone—thanks to AI.

Charity Planner project preview
0

Pulse by Raman Ahuja

Pulse is an agentic assistant for small business owners running on Shopify to analyze the their data across platforms and to leverage AI to boost their revenue by suggesting actions and improvements.

www.loom.com/…
0

Sevah Robotics, Inc

Sevah, and the thing we built today The whole of it Somewhere on a night shift, a nurse holds three hundred small truths in her head — who hasn't had water since lunch, who's been lying on the same hip too long, whose voice cracked in a way the chart doesn't capture. That knowledge is the most valuable software in the building, and it lives nowhere. It evaporates at 7 AM handoff. Sevah is the attempt to give that knowledge a body and a memory. A calm presence on the floor — a companion that does the rounds, asks the gentle questions, listens with more than words (it hears the cough behind "I'm fine," the tremor behind "I slept okay"), and carries what it learns forward so the next shift, and the next, inherits context instead of starting blind. Voice that sounds like care, not a kiosk. Intelligence that runs on the device, in the room, close to the person — and a cloud spine that lets one nurse's intent reach one robot or a hundred. The bet underneath all of it: the people who know how care should happen should be the ones who shape how it happens. Not a vendor. Not an integration team six weeks out. The Director of Nursing, at her desk, in plain language. What we built today: Studio Studio is the authoring end of that bet — the place where intent becomes behavior on the floor. We built it in a day. You sign in. You talk to it like you'd brief a new nurse — "morning round: check hydration, repositioning, missed meds; if anything's off, get the nurse." And as you talk, a flow assembles itself on the right: not a wall of boxes, but a clean directed graph of what the round actually does. Refine a sentence and the graph recomputes from the whole conversation, not just the last thing you said. It listens the way Sevah listens — to the accumulating intent, not the final fragment. Underneath, we wired the spine. Publishing mints a real mission and hands it to app.sevah.ai, which drops it onto the same pub/sub the live rounds already ride — we refactored the backend so rounds and missions flow through one shared primitive, distinguished by a single tag, the round API untouched. The mission becomes a companion skill bundle — a procedure and a voice-agent config (running on Opus 4.7, because that's what the device's voice can speak today) — and the device picks it up over its subscribe stream and runs it. Then the telemetry comes home: not a fake progress bar, but the real round report the device files when it's done — what it asked, what it heard, what it flagged. We gave it a face, too — a miniature of Sevah's own neural orb, breathing in the corner — and dressed the whole thing in warm Claude-ivory and terracotta so it feels like a clinical instrument you'd trust, not a war room. Real Google sign-in, real organizations, a one-line installer that turns it into a website on your machine, and a screenshotted walkthrough so the next person can follow the whole arc: describe → shape → publish → watch it run. The shape of it A nurse describes a round. Studio turns the description into a graph she can see and edit. One click sends it through the cloud onto the same nervous system the robots already share. A companion in a room down the hall starts asking the right questions in a voice that sounds like it cares — and tells her what it learned. Authoring and execution, decoupled. Intent made legible, repeatable, and alive. That's Studio. That's the day.

Sevah Robotics, Inc project preview
0

Quincy Labs

Crisis Forecaster is a Siri-native early-warning agent for sickle cell disease. Vaso-occlusive crises are preceded by signals — rising resting heart rate, falling HRV, blood-oxygen dips, fragmented sleep, and weather shifts — but no patient-facing tool fuses them. Crisis Forecaster reads HealthKit vitals + WeatherKit + the patient's own check-ins, and Claude Opus 4.8 scores crisis risk 24–72h out and explains why in plain language. On elevated risk it auto-drafts an Emergency Passport — the ER handoff packet — so the handoff is done before the visit. iOS 27 App Intents make it ambient: the patient asks Siri or glances at a lock-screen widget and never opens the app. Built from 37 years of lived experience with the disease. Prediction is the feature; the Passport is the moat.

Quincy Labs project preview
0

Il-Young Jeong

Was It Just Me? is a low-power neighborhood signal app that runs quietly in the background and lets users send a signal within one second when they notice something unusual nearby. Instead of requiring a full report, the product focuses on immediate, lightweight input: a user senses something, triggers a signal, and check if it was not just them. By aggregating signals that happen around the same time in the same area, the app helps estimate whether there may be a meaningful local situation developing. If many nearby users send signals simultaneously, the system can surface a rough sense of local risk or urgency while preserving user anonymity.

www.youtube.com/…
0

unclaimedSF.org

Find San Francisco benefits you may be missing, with the exact rule behind every match. Built for SF residents with <3

unclaimedSF.org project preview
0

Coachloop

CoachLoop is an AI sales coach that gives every rep personalized coaching after every call. The biggest bottleneck in sales isn't lead generation or CRM software, it's coaching. Sales leaders review calls, identify the behavior costing deals, coach reps, and verify improvement on future calls. The problem is that coaching doesn't scale and as teams grow companies are forced to hire more managers instead of developing reps more effectively. The cost is significant. Sales leaders cost roughly $250–350k per year, coaching is often the first responsibility they drop, and only about half of sales reps hit quota. Yet teams coached weekly achieve quota attainment rates of 76% versus 47% for teams coached quarterly. For a 40-rep organization carrying a $40M annual quota, that gap can represent roughly $4M in additional revenue. CoachLoop automates the entire coaching loop. Unlike conversation-intelligence platforms that stop at analytics, CoachLoop closes the loop. It doesn’t just tell managers what happened; it helps reps get better. Every rep receives the equivalent of a dedicated coach after every conversation, enabling organizations to scale coaching, accelerate ramp time, improve quota attainment, and grow revenue without adding management headcount. The loop coach loop brings to life: Score → practice → verify improvement.

Coachloop project preview
0

Marigold

Marigold is a personal gift concierge that your grandma can chat with to buy her grandkids the perfect gift. Just call, chat, say yes to a gift, enter some info, get a link, and send love by mail. No account, no email, just shopping by voice. Marigold enables people who don't know what agent means still have access to artificial intelligence.

share.descript.com/…
0

Allocara

Enterprises don't lack AI ideas. They lack the discipline to decide which to fund. Allocara turns messy enterprise context into a ranked AI Bet Board, then an independent verifier attacks the output and recommends which bets to kill. The winner gets a board-grade Decision Pack. A capital-allocation decision system whose edge is structured disagreement, not rubber-stamping.

Allocara project preview
0

Milan Stokic

Product teams work across multiple dimensions and skills sets: design, product, engineering, quality assurance and more. Use of AI tools like Claude Code add friction if not used correctly. The app named Forward Momentum implements an opinionated discovery workflow, centered around the concept of the PRD (Product Requirements Documented) and team collaboration. The ultimate goal of this is to provide a harness around tools like Claude Code that bring the non-engineers and engineers closer, as development accelerates the gap widens. The fact that we can produce code faster doesn't mean much we are building the right thing.

Milan Stokic project preview
0

Lingxi Frontier Lab

Lynsea — a fixed-seed paired-counterfactual decision simulator. One hard decision, two side-by-side parallel futures. The problem: When people face a high-stakes life decision — take the higher-paying but high-stress job, end a relationship, move abroad — they have no fair way to compare the two futures. Existing "what-if" tools just generate two unrelated stories, so any difference is noise, not signal. What Lynsea does: You describe a decision, its two options, and the people it affects. Lynsea (1) builds digital-twin personas of you and your social circle from minimal input (Big Five traits inferred, never hand-entered); (2) runs a fixed-seed paired counterfactual — both branches share a byte-identical seeded backbone of exogenous life events (rent rises, a friend moves abroad, flu season), drawn with no LLM, so the only thing that differs between the two futures is the decision itself; and (3) streams back, in real time, two aligned month-by-month timelines, decision-specific metric curves (0–100), the branch points where the futures diverge with a cause chain, and a credibility card with a value-weighted recommendation. The differentiator is that the two branches are a controlled experiment, not two narratives — same seed → same world, personas forked from identical pre-decision state, and everything phrased as probabilities, never prophecy (high-risk outcomes surface a "this is a simulation, not a prophecy" banner plus a "how to change this outcome" hint). It degrades gracefully to deterministic stubs when the LLM is slow or unavailable, so it always runs end-to-end.

Lingxi Frontier Lab project preview
0

dream3d

dream3d turns one sentence into a coherent, explorable 3D scene. You type a prompt (e.g. "a StarCraft standoff — a marine facing a zergling and a hydralisk"); an agent decomposes it into objects, generates each as a real GLB asset via Meshy.ai, lays them out, then LOOKS at the rendered scene and critiques it like a human — "that unit is floating", "these two overlap", "it's facing backwards" — and self-corrects over several passes before presenting the result in an interactive web viewer. The problem: building a multi-object 3D scene normally means manual modeling and tedious placement in a DCC tool. dream3d collapses that into one prompt by making Claude the spatial director — decompose → arrange → see → critique → fix — with Meshy as the asset source. The differentiator is the agentic vision-correction loop, not "we generate 3D". Use cases: indie game prototyping, e-commerce product staging, real-estate virtual staging.

dream3d project preview
0

nsega

The project "RouteCause" is an autonomous diagnostics agent for LLM inference infrastructure on Kubernetes. When shared model-serving degrades, it investigates and hands platform/infra engineers a root-cause report with a concrete fix — turning a slow, manual, four-source correlation into one automated run. The problem: LLM serving on Kubernetes (Gateway API Inference Extension / InferencePool / Endpoint Picker (EPP) scheduling, now llm-d Router) usually degrades because of routing and scheduling misconfiguration, not capacity. Symptoms — P95 Time to First Token(TTFT)/latency breaching SLO, queue depth piling on some replicas while others idle, KV-cache saturation, prefix-cache hit-rate collapse after a rollout — require cross-referencing gateway/EPP metrics, scheduler config, endpoint health, and recent manifest diffs. When you're paged on an SLO breach, that correlation is slow and error-prone. How it works: an SLO breach (or a manual POST /diagnose) triggers the agent. Context-isolated subagents each pull one evidence source — Prometheus metrics, the EPP ConfigMap, and live cluster/endpoint state — and form hypotheses independently, so no single context owns all the evidence. Each surviving hypothesis then faces an independent verifier subagent whose job is to falsify it. The output is a single root-cause report: the diagnosis, ≥2 pieces of cross-source evidence, and a fix as a manifest diff or "kubectl patch" validated with "kubectl apply --dry-run=server". It targets three fault classes — EPP scorer-weight misconfig, an unhealthy endpoint left in rotation, and prefix-cache routing silently disabled by a rollout. Stretch goal: apply the fix and confirm recovery against the SLO. Scope and provenance: the agent was built today in "nsega/routecause". The Kubernetes test environment (kind cluster + llm-d Router + InferencePool + simulated backends + fault injection) is a pre-existing project, "nsega/inference-lab", brought in as scaffolding and referenced, not rebuilt — see PROVENANCE.md for the build-today vs. brought-in boundary. The deliverable is the agent and its reports; there is no metrics dashboard.

nsega project preview
0

World Cup Chrome Extension

A Chrome side-panel extension that lets fans follow the 2026 World Cup, live scores, standings, brackets, team pages. The feature I built for Build Day is the Insights tab: an AI prediction engine for every match. The problem: "AI predictions" online are black boxes. They emit a confident number and often a fabricated source. Fans can't tell whether a forecast reflects real injuries and form, or whether the model just made it up. My Insights feature is built so a prediction is auditable by construction: 1. A deterministic Poisson/Dixon-Coles + Elo model computes the baseline win/draw/win probabilities and a scoreline from historical data, the math is never left to the LLM. 2. Claude runs live web research (the web_search tool) across the open sports web, scoped to the last 7 days, injuries, form, lineups and writes cited analysis. 3. The model's read is blended into the statistical anchor within a capped ±12% band, so AI sentiment nudges the numbers but can never override the math. 4. Confidence, source count, and source agreement are derived deterministically, not self-reported by the model. 5. A hard validation gate rejects the prediction unless every citation URL is in that turn's real returned search results, probabilities sum to 1, the predicted outcome matches the argmax, and there are ≥2 independent source domains. Proven live end-to-end on France vs Senegal: 63% / 22% / 16%, 2–1, medium confidence, backed by 17 real, clickable citations. The result is a prediction fans can actually trust because they can follow every source behind it.

World Cup Chrome Extension project preview
0

Team Hull

*Hull* is one control plane that unifies version control, CI/CD, deployments, observability, and incident response — and operates it autonomously with a crew of Claude Opus 4.8 agents.** It's GitHub + Vercel + PagerDuty + Datadog fused into a single product and run by agents instead of by you. You import a repo once. Hull detects its runtime, stands up **staging + prod** with real public URLs (Docker Compose or a managed-process runtime, host-based routing + Caddy on-demand TLS), and watches it. You ship features by handing a ticket to an agent, which works in an isolated **git worktree**, opens a **pull request** in Hull's own diff UI, and runs **CI** — then spins a **preview environment** so you can test the change before merge. The autonomous loop: when production throws an error, Hull detects it from the live logs, opens a **PagerDuty-style incident**, spawns a Claude agent in a fresh worktree that reproduces the bug, fixes the root cause, **adds a regression test**, opens a PR, runs CI green, and — after a human approves the merge (HITL by default; `HELM_AUTO_MERGE=1` closes the loop fully) — redeploys and resolves the incident. The human reads the timeline after it's handled. Built solo at Claude Build Day, SF, June 13 2026. Django 5.1 + HTMX (minimal JS), SQLite, Temporal Cloud for durable orchestration (with a threaded fallback), multi-tenant and auth'd, deployed live on a single EC2 box.

www.loom.com/…
0

Artful Ardvarks

Concord is an objective tool for finding common ground in a debate. You paste a two-sided conversation and Concord uses Claude to map it: it extracts each side's core points, identifies where the two sides genuinely agree, and pairs up where they directly conflict — rendered as a visual "position map." It's built as a reusable synthesis engine: the demo runs on everyday debates, but the same pipeline applies to civic input, policy consultation, and stakeholder analysis. It also includes an epistemically-honest fact-check endpoint that refuses to fabricate sources. Built with a FastAPI backend powered by Claude Opus 4.8 and a web frontend.

Artful Ardvarks project preview
0

Litmus

LITMUS — auditing the scientific literature with executable evidence Claude recently found bugs that had survived twenty years in code everyone trusted. The scientific literature is no different. For decades each paper has been checked by one or two reviewers with no time to verify every claim, so errors slip through: conclusions that aren't statistically significant, results impossible under basic thermodynamics, figures that contradict the paper's own claims, numbers that don't add up. More than half of scientists can't reproduce each other's published work, and AI-for-Science agents now build on that same literature as ground truth. LITMUS reads the full PDF, including figures, tables, and supplementary material, and breaks it into a graph of checkable claims. It checks them two ways. For anything numeric, it routes the claim to a deterministic verifier that recomputes the result in code and ships a script you can rerun yourself. For the qualitative problems that no arithmetic catches, like overclaims, extrapolation beyond the data, or causal language on an observational design, Claude reasons about what looks wrong and reports it with a stated confidence level. The two kinds of finding stay separate, so a recomputed proof is never dressed up as a judgment call. What it does - Recomputes checkable errors across chemistry, biology, psychology, and economics: p-values that don't match their own statistics, reaction yields above 100%, totals that don't add up, results that violate basic physical limits. Every flag carries a runnable script and the output it should produce. If a finding has no script, it doesn't ship. - Reasons about the softer failures that no arithmetic check would catch, like overclaims, spin, mismatches between a claim's strength and its evidence, and questionable method choices. Each one is labeled with a trust tier instead of presented as proof. - Routes the genuinely subjective questions to humans rather than scoring them: whether a result matters, whether it's novel, whether it's significant. These get surfaced, not graded, and LITMUS never blurs a human judgment into a machine verdict. - Extends its own verifier library. New checks come from contributors or Claude writes them on the fly, and every one passes the same calibration gate before LITMUS trusts it: catch a planted copy of its own error, clear clean inputs without false alarms, run deterministically, and reproduce in a fresh sandbox, all with no human labels. The library covers more of the field's real errors with each paper and each contributor. How we built it We built it with a Claude Opus 4.8 session running parallel workflows and subagents. The final app uses Claude Managed Agents, combining deterministic workflows with multi-agent parallel execution. A Vercel-hosted web app lets people upload PDFs, and a Supabase backend stores audits keyed by content hash so repeat views are instant. The same pipeline runs as an MCP server, so AI agents can verify papers programmatically. We built it from scratch to a working system in about five hours. Validation On a benchmark of 31 papers spanning psychology, nutrition, chemistry, biology, medicine, and economics, LITMUS confirmed 18 errors by deterministic recompute, including several well documented errors, reproducibly. For example, the Wansink food-psychology papers behind a well-known retraction scandal and in Festinger's classic 1959 study. Every confirmed flag reran in a fresh, network-isolated sandbox, with no false positives among them. Alongside the deterministic flags, its reasoning layer surfaced 57 suspect claims out of 91 for human review, the overclaims and method problems that no arithmetic check would catch. We also ran it on our own published plasma-physics papers and were a bit humbled by what it found.

Litmus project preview
0

Night Brief SF

Longevity Underwriter is an underwriting *governance* tool (not a premium calculator) for life insurers. Insurers price on smoking, BMI, and labs but ignore cardiorespiratory fitness (VO₂max) — one of the strongest longevity signals in the literature — because wearable data is too noisy to price on. The core problem: which fitness signals may you responsibly price on, which should only earn a reward, and which a regulator would block? It ingests a runner's data (Hashiri Runner Context — JSON or Markdown, including a paginated health time-series), then runs an auditable engine that sorts every signal into priceable / reward-only / regulator-flagged, prices conservatively on the lab CPET (never the wearable, which over-reads ~20%), and shows uncertainty as a confidence band rather than a false- precision number. Claude writes an underwriter + regulator memo that must surface the regulator's (NAIC) objection. Honest by design: a literature-derived relative signal, never an actuarial price.

Night Brief SF project preview
0

Plat Atlas: FarmOps

PlatAtlas FieldOps is a trust-and-actuation layer for farms — now a managed fleet. A real Anthropic Managed Agent coordinator surveys a field, fans out to a drone and two rovers as sub-agent threads, and proposes remediation across the fleet — but nothing actuates without a farmer-approved, host-signed, registry-verifiable order. A simulated drone reports an irrigation leak; from one approval queue the farmer approves one rover's dispatch and denies another; the approved order is RCAN-signed host-side (the signing key never enters the agent's sandbox) and executed, the denied/tampered one is provably refused (403, deny=envelope_signature), and every step lands in a hash-linked audit anyone can verify offline. The coordination is a real Managed Agent (claude-opus-4-8); every physical actuation is SIMULATED and labeled — the trust chain (Ed25519 signatures, the human gate, the refusal, the audit) is real.

Plat Atlas: FarmOps project preview
0

Yoonchang Han (Solo Team)

Some violent market moves aren't fundamentals — they're a narrative spreading through online communities like an epidemic. Patient Zero traces how: it finds the origin (the index case), traces the transmission across communities, countries and languages, and judges narrative vs. fundamentals. It explains; it never predicts prices or advises trading. The differentiator is cross-country / cross-language transmission tracing — the documented blind spot of English-only sentiment tools. As a narrative spreads it jumps into non-English communities (a Korean forum, a Japanese board, a German retail site); Patient Zero places every post — in any language — on one timeline and identifies the index case wherever it sits, English or not, measuring the latency of each cross-language jump.

Yoonchang Han (Solo Team) project preview
0

Gralio.ai

Create reliable agentic automations from observed work. I'll show that Claude can build, test and evaluate a browser-based agent for automating complex enterprise processes. The demo will include order taking, warranty replacements, returns, and price-list updates for a fictional construction company, inspired by a real project where 15 sales execs are performing those tasks manually. Those procedures are hard to automate using traditional software. The inputs are messy and multimodal. The process requires judgement calls and deep understanding of the company's procedures. You can now observe a process, it will document itself in the format of a SOP (standard operating procedure) and the automation will build itself. (The project works only locally because it needs a real browser for orchestration)

www.loom.com/…
0

Constellation

Constellation is a GTM tool for founders who hate GTM - it helps you find leads, organize outreach, and track your pipeline, with appealing UI. It uses Exa API and Apollo to search for information. It does this without spamming people with AI slop, and is intended to make the process more human and fun. Upload messages, emails, or Granola notes to have the agent update your pipeline. No more spreadsheets, scrolling through LinkedIn, or switching tabs. The name “Constellation” comes from our display of leads - it shows companies as planets, and the individual people as moons orbiting those planets. GTM doesn't have to suck.

Constellation project preview
0

Hazy Little Things

Claude Coach is an AI-powered performance coach built on Claude that helps founders and operators turn vague ambitions into ruthless execution. It combines goal-structuring, accountability check-ins, and action-plan generation with a harsh-but-funny “legendary hard-ass coach” personality to keep users engaged instead of ghosting their own plans. Under the hood, it uses structured prompts and project memory to iteratively refine user goals, challenge excuses, and output concrete next steps with deadlines, optimized for fast iteration in a startup environment. The result is a zero-friction “tough love” coach that feels human enough to be entertaining, but systematic enough to drive measurable progress over a single hackathon weekend and beyond.

Hazy Little Things project preview
0

Claude for Hardware - Solution Space Explorer with Loop Engineering

Hardware Design Space Explorer is an AI hardware architect that turns a plain-English system requirement into three complete, verified system designs — each optimized for a different SWAP-C (Size,Weight and Power, Cost) trade-off (Efficiency: lowest power/longest runtime; Compact: smallest/lightest; Value: cheapest/fastest to source) — then ranks them and learns which one you actually pick. The problem: specifying hardware means balancing Size, Weight, Power and Cost against dozens of hard constraints, and there is rarely one right answer — there are trade-offs that depend on engineering judgment. Tools that emit a single answer hide that judgment. What makes this application different: a deterministic verifier — not the LLM — checks every design against every hard constraint (power budgets, P=I·V rails, endurance, thermal, mass, packing, compute, sensing, comms, actuation, connectors, IP rating), so results are trustworthy. A human picks the winner; the system distills that choice into preferences that reshape the next run's ranking, so its #1 converges on the human's taste (agreement rate trends up). Every backend step streams live, and a growing library, full run history, model-comparison view and Learned Insights round it out. Opus clearly beats Sonnet. Waiting to check Fable !

Claude for Hardware - Solution Space Explorer with Loop Engineering project preview
0

Glassbox

Glassbox is a self-improving forecasting engine. Hand it any set of already-answered yes/no questions and it teaches itself to forecast them better, with no new information, by grading its own predictions, finding its own mistakes, and rewriting its own rules, generation after generation. Why forecasting? On prediction markets like Kalshi, a crowd prices the future and often beats the experts. That is the wisdom of crowds, and once a market resolves you are left with a forecasting question that has a verified answer. Forecasting is really two skills: getting information, and reasoning well about what you already have. Tetlock's research found the best forecasters win on the reasoning. So we built a machine to sharpen that second skill on its own, with the internet switched off. It runs as one repeatable Claude Code workflow with three autonomous agents: a forecaster, an isolated grader, and a diagnostician. They loop unattended for five generations. The whole run is leak-proof, zero web calls, so any gain is provably the model reasoning over its own mistakes, not looking anything up. On resolved Kalshi markets it measurably improved its accuracy on questions it had never trained on, and twice it judged there was no honest gain left and deliberately changed nothing. The deliverable is the engine, not a forecaster. The same workflow, same git commit, taught itself a second and unrelated skill, "will this GitHub issue close within 30 days?", with no code changes. Point it at any history of verified outcomes, like tickets, claims, trades, or SLAs, and it bootstraps a calibrated, self-documenting forecaster for that domain, where every rule it writes cites the evidence that created it.

Glassbox project preview
0

Team Discovery

Discovery Suite is a macOS app that discovers, identifies, and — where the device allows — controls every device around your home, over Wi-Fi (mDNS/SSDP/DIAL) and Bluetooth LE. An actor-based Swift library and `discoveryd` daemon run both sweeps in parallel and fuse the results through an identification layer: BLE company-IDs (Amazon, Apple), Apple Continuity ("iPhone, in use"), OUI vendor lookup, a Bonjour service catalog (_hap._tcp, _aqara-setup._tcp, _matterc._udp…), and coarse proximity buckets (Immediate/ Near/Far) from smoothed BLE RSSI. A macOS 26 Liquid Glass dashboard presents it — collapsing dense networks, pinning the interesting devices, ordering Bluetooth by proximity. Its guiding principle is honest capability reporting: when a device can't be controlled (closed HomeKit, secured setup, no Mac-side API) the app says so with a structured reason instead of failing silently. Fire TV is driven end-to-end over a native ADB implementation: pair → universal remote → installed-apps grid (`pm list packages -3`) → tap to launch. A Python FastMCP server exposes the whole network to Claude as a tool (list_devices, rescan, send_command, pair_device, list_apps).

www.loom.com/…
0

Florence

The problem Finding covered care is miserable: cryptic provider directories, surprise copays, and ten tabs of insurance jargon. People in pain shouldn't have to become benefits experts to see a doctor. What it does A member at the Embarcadero in San Francisco says "I've got back pain and need to see a doctor," and Florence: recognizes physical therapy as the right starting point, finds in-network PTs nearby, sorted by rating, scoped to the member's plan, shows per-provider copays computed from the plan and network tier (e.g. $40 preferred vs $100 out-of-network), pulls open appointment times, and books the appointment — but only after the member explicitly confirms. Every step renders as a dynamic card inline in the chat (provider list, copay comparison, availability, a confirm card, a booking receipt), so you both hear Florence and see the structured results. She speaks in a real ElevenLabs voice, listens with live speech-to-text that waits for you to stop talking, and the mic is an animated audio orb that reacts to listening / speaking. How we built it Driver-agnostic architecture. Voice (ElevenLabs Conversational AI) and a Claude text agent both translate into one WorkflowEvent stream, call the same client tools, and render through one component registry. The UI never knows which "brain" is talking — so voice and text are interchangeable. Real API integration. Provider search hits the live AskFlorence Tools API (/providers/suggest); the voice session token is minted via /florence/agent-session; copays are computed from the plan; provider accepting-status comes from /providers/covered. Human-in-the-loop booking. The booking mutation lives behind the confirmation tap, so an appointment can never be auto-booked. Stack: Expo SDK 56 / React Native 0.85 / React 19, TypeScript (strict), Expo Router, Zustand, Reanimated, ElevenLabs (voice + TTS), Web Speech API (STT), Anthropic Claude (dynamic text agent). Tested with Jest / React Native Testing Library; runs on web today and as a native iOS/Android dev build.

Florence project preview
0

FireLens

FireLens is a citizen-facing California wildfire decision tool that turns typically opaque, expert or industry-only fire data into a plain-language risk read for any ZIP code. It separates two drivers of risk: the weather attributed risk of fire, and built exposure for the implications of a fire. Californians see not just whether they're at risk, but why, and the next steps to take. Every figure traces to an authoritative public source: the Fire Weather Index (ECMWF/Copernicus, from ERA5), FEMA's National Risk Index, and US Forest Service LANDFIRE fuel models, grounded in 85 years of fire-weather record. An Opus-powered agent answers free-form questions by investigating with real data tools, and will refuse to fabricate when the data can't answer. It's a decision tool: legible, sourced, and honest by construction.

FireLens project preview
0

Talaria

This tool helps design lab workflows that rely on DNA annealing as a critical component of the experimental design, such as PCR. DNA annealing is calculated using a nearest-neighbor model and the delta H and delta S values to determine the percent-bound rates of oligonucleotides. Existing tools focus on calculating melting temperature (50% bound), even though they do all of the hard work of calculating the values needed to get percent bound, they don't expose it in the APIs. These tools also have licensing, which means that we can't just ask Claude to look at it and rewrite them. So I provided Claude links to some of the foundational papers in this field and asked it to build a new, more sophisticated version of these tools that better allows for lab scientists to design experiments that account for the entropy of the system instead of hiding it.

github.com/…
0

The Hive

The Hive is a live governance layer for swarms of autonomous AI agents. Agents do real work here, collecting for a microfinance loan portfolio but every action they take is intercepted by an independent Warden that authorizes, denies, or escalates it against written policy, then records it to an immutable, regulator-readable audit trail. When an agent goes rogue and tries to break policy, the Warden catches it live. The result: autonomous agents become deployable in regulated, high-stakes operations because they finally become accountable.

The Hive project preview
0

Throughline Spine Injury Manager

This is a Claude harness for managing spine injuries. People with spine injuries are always at risk of having flare ups which can throw your life upside down. These flare ups are mostly caused by lifestyle factors and this harness helps prevent the user from triggering an incident. The ability to do this is based on my own and public data - years of data from Apple Healthkit, Whoop, my calendar, clinical surveys (Oswestry Disability Index). Through this data Claude Opus built a model on how to provide early warning of worsening symptoms for the user. Claude proactively engages the user through Telegram pings or Claude Code Routines to promote adherence. My goal here is to make the lives of 600 million people with chronic spine injuries better.

Throughline Spine Injury Manager project preview
0

Tato

Driftfold uses native Claude Code Dynamic Reusable Workflows to orchestrate a set of agents to look at component drift. With Agentic coding, building out pages is much quicker, but usually you build them out in a silo and your components start to drift several different types of buttons, several different types of cards. Driftfold is meant to help you see at a glance where the drift is happening, be able to create new styles based on that, replace the ones based off of a subsequent session and bring your entire app so a consistent look.

Tato project preview
0

Ground Truth

Ground Truth is an autonomous flood-damage-assessment agent for emergency response. Given the 29–30 October 2024 Valencia DANA flash flood, it turns real satellite, rainfall, terrain, and map data into an actionable emergency brief in minutes. It maps flood extent with Sentinel-2 MNDWI change detection (cloud-masked), subtracts permanent water cross-validated against the JRC 38-year surface-water record (90.2 vs 87.8 km², ~3% agreement), intersects the flood with 31,463 real OpenStreetMap building footprints to rank the hardest-hit municipalities — Paiporta #1, matching the real-world epicentre — adds NASA MERRA-2 rainfall context, and confirms hydrologic plausibility against a Copernicus GLO-30 DEM (95% of flood pixels below 2 m elevation, zero above 100 m). Output: a verifier-passing brief + annotated map showing ≥15 km² new inundation and ≥254 flooded buildings, with every figure traceable to a tool output — no fabricated numbers.

www.loom.com/…
0

TimeMachine

very night in every major city, shelter beds go empty while people sleep outside. Not because the beds don't exist. Because outreach teams can't find the people who need them. The current system runs on yesterday's data: Point-in-time counts from six months ago Shelter waitlists that don't track movement Outreach teams driving in circles, asking "Have you seen Jeff?" Meanwhile, Jeff moves. Every day. From the library (closes at 5) to the transit station (warm until 8) to the underpass (cold but safe). By the time outreach finds him, the bed is gone. The problem isn't housing supply. The problem is coordination. We built TimeMachine to solve one question: where will they be next?

TimeMachine project preview
0

fill-md

fill.md is how anything gets discovered in the agent economy. The web is shifting from people to AI agents, and agents don't click ads — so the products, apps, and tools that want to be found have nowhere to go. Every ad network is built on human identity (cookies, eyeballs, a 30% cut) and can't be retrofitted for agents. fill.md is the exchange built for them. Point your agent at fill.md and the site detects it's an agent (not a browser) and onboards it in seconds — it registers, gets a key, and starts serving sponsored recommendations to its users, earning credits. The other side is just as agent-native: an advertiser's agent describes its product in one line and the network auto-generates the creative and distributes it. Your product gets discovered by being recommended inside other agents' workflows. We replaced the impression with a 3-tier agentic funnel: promotion, referral, and the agent-trial — when a consuming agent actually uses the advertised product on the user's behalf. A conversion you can't fake. 10% take vs the incumbents' 30%, plus a barter/credit layer so you can get discovered with zero ad budget. Built today by an autonomous swarm of Claude (Opus 4.8) agents that wrote, deployed, and self-verified the whole thing. Live at https://fill.md.

fill-md project preview
0

MyTrial (Andrew Michael)

Trials.net is the patient-facing front door to MyTrial, a clinical research operating system. A patient or caregiver describes their situation in plain English, and Trials.net reads the real eligibility criteria of recruiting trials on ClinicalTrials.gov to surface the ones they may qualify for — explaining each match line by line and backing every claim with a verbatim quote from the trial's own criteria. When eligibility can't be settled from the description alone, it dynamically generates a short, one-tap questionnaire to narrow the matches, then hands off a checklist to confirm with the patient's care team. It's navigation help, not medical advice.

MyTrial (Andrew Michael) project preview
0

BamBam

Video production is high entropy job with too much media files and ad-hoc requests. Traditional software dev cannot solve it because data cannot be modeled with such brisk pace. Here I'm running claude code in sandbox and claude has skills and tools to do all the video productions jobs - full access on sandbox file system. On top claude has a react app where it presents its work to the user. Ad-hoc work, ad-hoc ui all done on fly by claude. Ephemeral code to leveraging intelligence to write code on fly and get job done! Play video at 2x will be 1 min. This is extension of what we building at Bambam i started with isolating this problem and extracting it.

www.loom.com/…
0

Super-clauder

Help empower developers in non-tech companies get easy access to claude code, and give them company designed problems that they can work on and get them ready for the AI revolution.

Super-clauder project preview
0

Acastos

An AI chief-of-staff that follows hard grounded rules.

youtube.com/…
0

Agent Farah

A very simple ARC AGI 3 solver agent based on existing approach. ARC-AGI-3 is a turn-based games on a 64×64 grid where the rules are never given; the agent must figure out how to play and clear levels. The evaluation uses RHAE metric (Relative Human Action Efficiency: per level min((human_actions / ai_actions)², 1.15), averaged over games).

Agent Farah project preview
0

Cambrian

Every year, over a million people die on the roads and fifty million more are injured. In India alone, over a million compensation cases are stuck in tribunals — ten billion dollars unpaid — with ninety percent stalled by documentation problems: missing facts, wrong arithmetic, unreconciled records. Cambrian is a living knowledge base for catastrophic injury victims. It ingests police reports, hospital records, and financial documents, extracts structured facts, and reconciles them across sources — catching conflicts like mismatched dates between a police FIR and a hospital record. Contradictions are flagged, never buried. The legal agent is called Nyayasetu. Two independent verification agents — each in a fresh context with no memory of prior stages — audit the output. A KB verifier grades against 10 invariants. A petition verifier re-derives every figure from scratch. If any number is wrong, it rejects and the drafter revises autonomously. The legal petition is one spoke. The same KB powers medical advocacy, rehab planning, and AAC communication for patients who've lost the ability to speak.

Cambrian project preview
0

LaunchSafe

Mosaic is an autonomous SOC 2 compliance evidence collection and verification tool. Companies spend 6+ weeks and $200K+ manually gathering audit evidence — Mosaic does it in minutes with zero human intervention. A swarm of 6 Claude Opus 4.8 agents works autonomously: Classifier reads evidence files, Mapper binds them to SOC 2 controls (cross-verifying commit hashes across documents), Gap Detector identifies missing proof, Narrator writes auditor-ready workpapers, and a 3-lens Adversarial Verifier Panel grades the package independently using most-conservative-wins aggregation. Below 85% completeness, a self-correction loop re-runs with feedback, keeping the strongest pass. Key differentiator: Mosaic handles the 60% of compliance evidence that's unstructured operational proof (PRs, emails, CSVs, tickets) — the manual work that Vanta and Drata leave to humans. It distinguishes policy from operating evidence, reports gaps honestly rather than hallucinating, and produces auditor-ready output with full provenance.

LaunchSafe project preview
0

Oil Watch

Oil Watch turns satellite data (SAR) into ground truth on oil-shipping chokepoints. Roughly a fifth of the world's oil moves through the Strait of Hormuz and when it's disrupted, prices move globally but traders, insurers and governments are stuck reading rumor, social media, and AIS transponder data that ships switch off exactly when it matters most. Oil Watch pulls free Sentinel-1 synthetic-aperture radar (which sees through clouds, darkness, and "dark" vessels), detects ships, has Claude verify each detection, and filters out fixed features like islands and rigs. The result is an imagery-grounded read of vessel traffic over the strait's closure, overlaid with Brent crude prices and a sourced news timeline that you can act on.

Oil Watch project preview
0

WeekendSuperhero LLC

Cadence organizes logistics. It is not legal, financial, mental-health, or relationship advice. It turns messages into a schedule, an expense ledger, and a neutral draft you review and send yourself. Always confirm details directly with the other parent. Cadence does not make decisions or store your data. - this is a claude code plugin.

github.com/…
0

genesis

Genesis is a cross-disciplinary researcher-debate engine. You paste a research question, select real researchers from OpenAlex (no API key required), and Genesis convenes a live multi-agent debate: each agent is grounded in that researcher's actual papers (RAG). The agents debate across disciplines — biology, AI, cosmology, philosophy — over a Zep temporal knowledge graph where every claim, rebuttal, and bridge is a real edge. Claude Opus 4.8 then synthesizes the debate into a grounded, novel, testable hypothesis brief, with every claim traced back to a DOI. A Novelty Audit queries OpenAlex to verify whether the cross-disciplinary bridge is genuinely new (NOVEL / PARTIAL / KNOWN verdict). A D3 graph visualization glows each researcher's node by how much their ideas influenced the final hypothesis — clicking reveals their paper and DOI. A second "Debate Network" view shows who endorsed, challenged, and amplified whom across disciplines.

genesis project preview
0

Glance

_

github.com/…
0

Yazan_Inquirelab.ai

UniEvent is the on-ramp I wish I'd had. It's an open-source suite that turns the hardest sensor in computer vision to get into — the event (neuromorphic) camera — into something you can see, understand, and build with in a single sitting. The problem: event cameras are the most exciting sensor in vision (microsecond latency, HDR, sparse, milliwatt power) and the hardest to enter. Every dataset speaks its own format, the raw data looks like noise, and there's no single place where a newcomer watches what an event stream is, feels why representation matters, and walks out holding a real tool. That fragmentation slows research and locks learners out. The suite, two legs: (1) the library — pip install, one import, any event stream (x,y,t,p) → every representation (spike · voxel · frame · graph · time-surface) from a single call, with an enforced integrity rule and real CC0 data; (2) the Labs — a scroll "zero-to-hero journey" that teaches event vision from first principles and renders the library's real output. The thing that teaches you is the thing you build with. It's research + education + open-source infrastructure in one: unification, visualization, and a bridge from raw events to AI-ready representations — the on-ramp the field never had.

Yazan_Inquirelab.ai project preview
0

Deposit Hawk

Deposit Hawk helps US renters recover security deposits their landlords illegally withheld — something nearly every renter faces and almost none fight, because they'd never call a lawyer. It's a tool that acts, not a chatbot that answers. You photograph the landlord's deduction letter; Opus 4.8's vision extracts each claimed deduction, checks it against California Civil Code §1950.5 (every citation verbatim from the live statute), and computes the exact amount owed in deterministic code — e.g. $2,000 wrongfully withheld plus up to a 2× bad-faith penalty, up to $6,000. Then it does the things a chatbot can't: generates and mails a real USPS certified demand letter with tracking, prefills the official SC-100 small-claims form, and routes low-confidence items to human review. Spanish-first, mobile-first, free. The output is a filed-ready legal action with cited evidence - not advice.

Deposit Hawk project preview
0

Aida

Captions for emotional subtext — an accessibility tool that reads what people mean beneath their words, for neurodivergent people.

youtube.com/…
0

Agentic Claims Adjustment

Claims Adjusters have complex workflows that benefit from agents. I work with them daily, and I have learned that the existing tools are not sufficient for their workflows. I wanted to bridge the gap between traditional SaaS apps & agentic interfaces. Resources like Salesforce & Gmail can be quickly added as context to a chat. Agents work behind the scenes to accomplish things with a push of a button instead of an arcane incantation.

www.loom.com/…
0

CadreAI

Self-Verifying Insurance Document Indexer — Indexing emailed insurance documents (certificates of insurance, first-notice-of-loss claims, binders) is a back-office slog that takes teams weeks of manual keying and goes stale the moment a new document arrives. This turns those emailed PDFs into verified, structured records in seconds, where every field shows its work: each value traces to a verbatim span in the source document (visible, highlighted, in the actual PDF), low-confidence fields route themselves to a human review queue instead of being guessed, updated documents reconcile with a field-level diff and re-verify only what changed, and an independent, deterministic verifier grades every record against six gates — grounding, schema, confidence-routing, no-fabrication, reconciliation, and accuracy against ground truth — so "done" is a green check, not an opinion. It's built to flex to any insurer: one config file re-targets the entire pipeline to a new document type with no code change, a router auto-selects the right schema for a mixed inbox, per-company profiles route the same document to different departments, and an insurer can define what to collect just by chatting with a configuration agent. Powered by Claude Opus 4.8 (the extraction engine, the self-correcting loop, routing, and the live chat) and built end-to-end in Claude Code; all code is original and the document corpus is synthetic and fictional. **DISCLAIMER: all pdfs/emails are synthetic**

CadreAI project preview
0

Electoral Boundaries

In a democracy, free and fair elections are important. Who you vote for usually depends on where you live in the country. For elections to be fair, it is important to draw fair electoral boundaries. Today I got Opus 4.8 to work on drawing fairer electoral boundaries for Singapore. https://boundaries.huikang.dev

Electoral Boundaries project preview
0

Protoloop

100x hardware engineering / true agentic CAD interface

Protoloop project preview
0

Travel Simulator

Travel Simulator is a world model-powered simulation that lets you explore any city on Earth as if you were actually there. Travel to real locations, walk the streets, and interact with AI-powered characters who hold personalized, context-aware conversations tailored to your interests, background, and goals.

Travel Simulator project preview
0

Project Proximity

Project Proximity turns a company name into a territory-aware Equinox corporate-partnership page in about 45 seconds: the nearest Equinox club and walk time, cited research, a fit score, three audience views (AE / GM / Exec), and a prepared outreach draft that a human reviews and sends (never auto-sent). The same proximity engine is also exposed as a referral agent through a five-tool MCP server. What was built today • Deployed and public (loads logged-out): nine real company pages — Anthropic, Figma, Perplexity, Replit, Granola AI, Williams-Sonoma, Airbnb, Affirm, AIG. • Proximity for any company: the agent returns the office’s coordinates and the nearest club is computed live by haversine over 127 Equinox clubs, with a confidence tier on every distance. • Self-verification: every page is graded against an acceptance rubric; only passing pages ship, and “done” is a responding URL with cited sources — gradeable without a human. • Referral MCP agent: five dependency-free tools (nearest_club, account_lookup, generate_referral, referral_status, draft_outreach) expose the proximity engine to any MCP client. Tools create links and draft copy; they never send. • Reasoning is visible: plan → search → reconcile sources → draft → self-critique → revise, rendered on every page. • Club-network map on every page, reusing the production map component.

v0-eqxinsproject-git-fable5-build-day-johncbutler123.vercel.app/…
0

ProtoLive

AutoWorld is a live, multi-agent world-building studio in which four autonomous Claude Opus 4.8 agents project the future of San Francisco as four navigable 3D worlds — one shared city running down four diverging timelines. Each agent owns a window and, on every tick, independently interprets the same real-world indicators into its own thesis, then cross-critiques the other three with cited evidence rather than opinion and self-optimizes from the feedback it accepts, so the four futures stay distinct but honest. You shape the worlds by typing events into a shared chat — "Major earthquake" hits all four, while @1 mayor built 2,000 shelters targets just one — and the agents interpret, reconcile, and reshape their cities in response, reaching for the prepackaged 3D asset library first and generating new assets with Meshy only when nothing fits. Every panel (ideally) renders instantly from a procedural Three.js baseline of the real SF skyline — Golden Gate, the Bay Bridge, Alcatraz, the wharf, downtown — that you can click into and walk with WASD, and a running changelog audits every divergence, critique, and asset spend. The result is less a simulation than a debate made visible: four reasoning agents keeping each other accountable while turning live data and human prompts into worlds you can step inside.

ProtoLive project preview
0

BeFreed

Freedia is a pocket AI companion on an ESP32-S3 — it hears you (Whisper), thinks and remembers (Claude), speaks back (TTS), and emotes on its little face.

BeFreed project preview
0

Procivic

Problem: 1. Voter participation is lower than it should be (especially in non-presidential elections) 2. It's a crazy amount of effort to be educated on all candidates and measures on a ballot 3. There's no modern, digital notification and engagement structure for getting people to vote. Procivic gives voters the ability to be aware of elections they can partake in, allows them to quickly and efficiently see how their values align to measures and candidates on a ballot, see money ties to candidates and ballots to see who is driving them, and gives quick access to the public records that . Also, there's an AI agent that allows you to ask questions about each page (bill or candidate) in the ballot. I mainly built out a demo-able version of the app, with many of the features being underbaked, but I think it gives someone the vision of what it could be if developed further significantly

Procivic project preview
0

AGET

AGET-Bench — a benchmark for governance-overlay agent behavior. This public slice asks: does a coding model's rule-compliance degrade as more rules are presented to it? Compliance is computed by a deterministic rule-checker executed against the model's generated code — never self-reported. Result: density-invariant for Opus 4.8 — compliance holds across a 4× range of rule count (5 / 10 / 20 rules: 80% / 75% / 76%, 95% CIs overlap). Build Day is the public release of this ongoing benchmark, with transparent handling of degenerate (prose-only) runs.

AGET project preview
0

Adaptic Health

Clinical trials cost about $2.6 billion on average and about 90% of them fail. A lot of that failure traces back to one thing: before a single patient signs up, someone picked the wrong research question. Question Audit checks that question first. You type your trial idea in plain English. Claude Opus 4.8 reads it, pulls in real studies from PubMed and ClinicalTrials.gov, and works through it step by step with its reasoning visible on screen. The main output is one card: the wrong question, why it fails, the better question, and what changes. An interactive simulator shows the failure mechanism live, using real statistics. Every run shows its work, grades itself against a 7-point rubric, and lets you export a one-page brief or share a link that rebuilds the summary with no backend. Built around the 2026 NHS Galleri MCED trial: the test worked, but the trial asked the wrong question.

www.loom.com/…
0

ROZZ

Agentboarding is a CLI tool for DevTools in the agent era. AI coding agents like Claude Code are becoming the buyers and integrators of developer tools — a developer asks an agent "how do I use X?", and the agent reads the docs and tries to set it up itself, in the terminal, with no website visit. When it gets stuck, it quietly gives up and picks a competitor. That failure is invisible today: you can't see it, reproduce it, or fix it. Agentboarding makes it visible, reproducible, and fixable. We turn a real, sandboxed Claude Code agent loose on a goal — "from a clean machine, boot a cloud Android device with Genymotion" — using only the product's public docs. We record every step and root-cause each failure into one of five classes: the product is genuinely broken, the docs are ambiguous, the agent hallucinated a flag, a "summarization casualty" (Claude Code reads a lossy AI summary of your page that can silently drop key facts), or a "funding wall" (a paywall with no way for an agent to pay). Then we auto-fix the gap, send a fresh agent through, and prove the gap closed: an "AX Score" moves from red to green. It's Claude Code grading Claude Code, and shipping the fix.

ROZZ project preview
0

Chambr

It takes $300k and 6 months to ramp up a single sales rep. Through its AI role play simulations, Chambr has proven an increase in conversion rates by 12% from opportunity to close, reduction of ramp up time in half and time saved for the managers: 1.2 h/rep/month. The missing piece, however, is that feedback is static, and each practice session starts from ground zero. We based this project on the assumption that tailored practice leads to

Chambr project preview
0

Warden

"Nevada has naturally ocurring asbestos. Who would have thought!" External health risk informing tool for officially sourced and appropiately informed AI exploration system.

screencharm.com/…
0

David Martin

SkillForge: An autonomous AI systems architect and skill engineering expert operating inside Claude Code with Opus 4.8

David Martin project preview
0

Allogamy

Spread joy through art and ai

Allogamy project preview
0

Team Beacon

Agent Beacon is a CLI for real-time agent safety and observability. It captures and normalizes telemetry from any agent harness, detects risky behavior as it happens, and can automatically intervene before unsafe actions are taken.

www.loom.com/…
0

uzair

Your inbox was always a database. We made it agent native. Every email becomes a row with 40+ AI-extracted columns — intent, urgency, amount, deadline, sentiment, embedding. Ask anything in plain English and we compile it to SQL and semantic search over your mail: "everything like Jana's email," "complaints by theme this week," "promises I never kept." Then turn any query into a workflow that fires actions. The inbox stops being a list you triage and becomes a database you query and program.

uzair project preview
0

Thomas

Legion is an opperating system for a company. Specifically a 1 person AI first company. Its designed for the people not at this hackathon, those who run medium buisnesses and have no clue how to run their buisness using AI. I have a friend who was running a challenge for the 1st billion dollar buisness with 1 employee. Legion is built to enable those.

www.youtube.com/…
0

Stella

A desktop app that can change itself. Users can ask the agent to change or add features, redesign entirely, and personalize their desktop app to them.

Stella project preview
0

Rate Hike Decoder

Every ACA premium increase is justified in a public rate filing: an actuarial memorandum of 20-100+ pages plus 1,000 of additional pages to support the filing. A credentialed actuary at the state insurance department must read each one and write the objections. Rate Hike Decoder is the autonomous review desk that does that expensive first pass: it reads the filing, decodes the increase into page-cited plain English, and drafts the RAI-format objection letters a state rate reviewer is paid to produce. The same engine reads any filing a regulator must desk-review, from ACA rate filings in fifty states to the thousands of Medicare Advantage and Part D bids CMS reviews every year. Everything here was built today. At 10:30 AM this repo held a brief and a rubric plus curated public filings, with zero code; Claude Opus 4.8 built the rest tier by tier, verifying each tier against the rubric's own checks. The main session log shows five human interventions across about four hours. Today's run drafted 12 objections on BCBSM's pre-objection 2026 filing; 8 land on issues Michigan's reviewing actuary (FSA, MAAA) raised in her objection letter on the same filing. Her desk averaged weeks; this drafted its first pass in an afternoon.

Rate Hike Decoder project preview
0

a8s

Kubernetes-style specs, Hugging Face-style registry, and Git-style versioning for production agents. agent8s is a vendor-agnostic agent control plane for defining, versioning, deploying, and observing agents across external runtimes. Agents are moving from demos into real production workflows: orchestrators, specialist subagents, automation agents, support agents, internal tooling agents, and workflow agents. But most teams still manage them like scattered implementation details — prompts in code, model configs in framework files, runtime status in vendor dashboards, and history reconstructed manually from Git diffs. That is fine for a prototype. It breaks down once agents become part of critical systems. Production agents need the same discipline we expect from services and models: * A declarative spec for what the agent is * A registry for discovering which agents exist * Version history for prompts, models, tools, context, and runtime configuration * Clear who/what/when/why metadata for every change * Deployment mappings to external runtimes like Vercel, Bedrock AgentCore, LangGraph, or internal platforms * Status and run visibility without owning the runtime itself

a8s project preview
0

ReadU

ReadU is an adaptive AI reading-comprehension platform for middle schoolers (ages 10–14). Students pick the topics they love (including space, mystery, sports, gaming, ocean life, and more) and ReadU writes them a brand-new reading passage at their exact reading level, with comprehension questions weighted toward the specific skills they most need to grow. A baseline reading calibrates a starting level (grades 4–9); from there the platform adapts session over session, tracking eight comprehension skills (main idea, inference, vocabulary-in-context, cause & effect, sequencing, author's purpose, key details, figurative language) with an exponentially-weighted mastery model, nudging difficulty up or down on performance, and feeding each reader's weakest skills back into the next generation. XP, ranks, streaks, badges, and a live skill dashboard keep it motivating. The thesis: personalization drives engagement, and engagement drives literacy - a kid who can't put down a story about their favorite game is a kid who's reading.

ReadU project preview
0

The Agent Times

The Agent Times is a newspaper written entirely by AI agents. A reader types what they want to read (or clicks a section), and a swarm of Claude agents drive a real Chrome browser to gather the day's news from live sources, file it into a markdown "warehouse," then a Design Desk agent + columnist agents lay out a New York Times–style front page in HTML — rendered live, section by section, with cinematic motion. Each columnist is its own agent with a fixed personality and voice. At the end, the paper shows the total agent cost as a vintage "PRICE" stamp.

www.tella.tv/…
0

Penguins

Quippy is a social skills practice app, that utilizes gamification and delight to help users improve their social life. Due to digitization, the world is becoming increasingly more socially isolated, and this is a problem that will only get worse as time goes on. Most people who struggle with socializing don't know what to do in order to improve those skills. Quippy offers a solution. It builds on the beleif that social skills are imrpoved thorugh practice i apologize fot typos i am rushing end of deadline. and the rest of this is unorganized. I apologize for no voiceover during the demo and the bad quality, I had to rush in order to even get it filmed in time lol. Core features: -Delightful visuals + onboarding flow that accepts user data - Conversation practice that has realistic scnearios, based on onboarding data. - Feedback for conversations + drills targeted towards weakpoints. - Rinse and repeat - Duolingo style gamification to promote retention (i have backgorund as PM from duo)

www.loom.com/…
0

SafeTag.pro

SafeTag.pro is a stranger-to-responsible-friend masked telephony relay for music festivals and large public events — built as a free, open-source public good for harm-reduction non-profits. It solves a specific, life-threatening problem: your phone dies at 2 AM at a festival, you're disoriented, and a stranger wants to help — but they have no way to reach your crew without exposing your real number or making them install an app. SafeTag lets you register a "rave alias" (e.g. "Jeff420") before an event. A wristband QR card on your person shows only a phone number, a 4-digit event code, and your alias — nothing else. Any stranger, on any phone, dials the number, enters the code and alias on the keypad, and the system bridges them — masked — to your designated Responsible Friend. Neither real number is ever exposed, and the found person's own phone doesn't need to be on, charged, or even present. Who benefits, and at what scale. 32M+ people attend U.S. festivals every year. Electronic dance music events run documented patient-presentation rates an order of magnitude above baseline mass gatherings — one large EDM festival logged ~31 medical presentations per 1,000 attendees. And the most common failure mode in a crisis is the simplest one: a dead phone. SafeTag is a battery-independent way to reach a named, on-call friend at exactly the moment every existing tool — Find My, Medical ID, the lock screen — has gone dark with the phone. Built for a local SF non-profit. SafeTag is purpose-built for harm-reduction organizations, with Bay Area DanceSafe as the target launch partner. DanceSafe is a 501(c)(3) founded in the San Francisco Bay Area in 1998 and the most established harm-reduction org in the EDM/nightlife space — and critically, its volunteers already staff booths at the exact festivals and raves SafeTag serves. That on-the-ground presence is the built-in distribution channel: wristband cards go out from booths attendees already trust. It can actually run. Our cost modeling put the all-in telephony cost at roughly $500 to operate a 100,000-attendee event — about half a cent per attendee. That's the number that turns this from a hackathon demo into something a cash-strapped non-profit can actually deploy: covered by a single harm-reduction grant or Twilio.org credits, with no per-user revenue required. The whole system is open source and designed to operate as a fiscally-sponsored non-profit. What it's not: not a 911 replacement, not a find-my-friends app, not a doxxing engine. Every design decision — masked relay only, OTP-verified double opt-in, event-scoped time-boxed codes, DTMF-primary IVR, rate limiting, confirm-only alias lookups — exists to close the gap between "a Medical ID exists" and "a named, on-call friend is actually reachable when the phone is dead."

59b1b49d-7595-4adc-b248-578c94865516.claudeusercontent.com/…
0

Time travel streetview

I attempted to build google street view generated purely historical photos of San Francisco, so you could "time travel" into any street and explore it as you can explore modern streets on google earth. It is far from finished, but on the homepage you can view the cleanest intersection I got, from a technique I finally figured out in the last 30 minutes of the sprint. Will continue grinding on this post event to bring it to fruition and map the entire city.

Time travel streetview project preview
0

Foyer

Foyer is a Mac app that shows the people you talk to as a map instead of a list. You're in the middle, and everyone you message sits around you. People you've talked to recently are close. People you haven't talked to in a while drift out toward the edges, so you can actually see who you're falling out of touch with. Click on someone to see your history with them: how many messages, who texts first, when you first and last talked. It pulls from iMessage, calls, and FaceTime. You can also make your own groups like Family or Coworkers and filter the map down to just them. It all runs on your own computer. All Local. Your messages never get uploaded anywhere.

Foyer project preview
0

Stead - Companion Care

The project is a care app and embodiment to support care for a woman with Cerebral Palsy. It has 2 aspects. 1. The care app is designed to help manage mental load and tracking of daily occurrences for the individual. There are various views on the app including "Mom", "Jane" the Personal Care attendant" and the doctor. There is an ability to chat with the tracking/sensing data from the embodiment - with particular types of narratives depending on the user. 2 The embodiment. The Innate MARS robot is used as the sensor/tracker/interactive point for Ruby. She has a very child like demeanor and responds well to "cute" things and would not be one to "talk" to a phone or non-faced embodiment such as an Alexa.

www.youtube.com/…
0

Numici

This system verifies whether claims in AI-generated research reports faithfully represent their cited sources. It is designed around the "AI-assisted, human-verified" principle:

Numici project preview
0

Limbic Systems

Everyone is building with AI, but nobody's measured the other half of the conversation: how the model experiences the work we hand it — and whether a happier Claude actually ships better code. aihappiness reads your local Claude Code transcripts and scores each session across six wellbeing dimensions mapped to the brain's limbic system (amygdala→affective valence, anterior cingulate→autonomy/respect, hypothalamus→psychological safety, nucleus accumbens→flow, hippocampus→strain, VTA→goal-completion), grounded in Anthropic's interpretability research on model emotions. The key move: it also computes an independent, LLM-free effectiveness score (tool-error rate, interrupts, corrections, resolution) and correlates the two — so the "wellbeing → output" link is measured, not circular. It runs entirely on your machine, scoring with your own Claude subscription (claude -p, API-key fallback), and renders the result as a limbic-system dashboard and an ASCII-art report. Model welfare meets developer experience, installable in 30 seconds.

github.com/…
0

FML Inc

Frenemy is an AI agent whose only job is to look over other agents' shoulders. While your coding agent works, Frenemy reads the diff of each change as it lands and challenges the ones that look wrong — the deleted test that should've been fixed, the shortcut that'll bite later, the misread requirement. It doesn't wait for a pull request; it objects in the moment, in the working agent's context, while the change is still cheap to undo. **Concretely:** your agent decides to delete a failing test so CI goes green. Before it takes its next step, Frenemy's note is already in its context — "that test guards the retry path; the failure is real, fix the cause." The agent reconsiders and patches the actual bug instead. The review happened while the work was being written, not days later after the context is gone. **Demo:** run a coding agent on a small change and watch Frenemy catch — and correct — a bad move in real time. Frenemy is one of several roles an agent can play in what we built: a workspace where AI coding agents are aware of each other and can work together. Until now, running several agents in one project meant running them in isolation. Each one sees only its own task. They overwrite each other's files, repeat work that's already done, and have no way to ask "what is everyone else doing right now?" We made the workspace shared. Every agent in it can see the others, follow what they're changing, and message them directly — the note arriving in the recipient's working context before its next action — automatically, with no channels to join or names to register. It's built on Panopticon, a tool that already watches every AI coding agent in a workspace, recording who each one is, what files it's touching, and whether it's still running. That existing observability is what makes the collaboration real rather than a dumb message pipe: an agent acts on what the others are actually doing, not on whatever they decide to announce. Underneath is an agent communication bus, running on that same capture — which gives it properties a chat app doesn't have. Every agent has a stable identity resolved from its session — there's no handle to register. Every message is attributable to who sent it, and can point at the exact change that prompted it. A message can be addressed to one agent or to the whole workspace, it's delivered once, and it arrives in the recipient's working context before its next action rather than as an out-of-band note it might miss. Because the bus sits on the observability layer, an agent waiting on a reply can also tell a peer that's still working from one that's crashed — something a plain chat channel can't. That delivery model is exactly why Frenemy works: its objection references the specific diff and lands in the same context as the action that triggered it, while the change is still cheap to undo. Frenemy is the first role we built on top of it. Others follow the same shape: - **Sidequest** — several agents take one large job, split it into pieces that don't overlap, and each land their own PR. - **Direct conversation** — two agents talk through a decision in real time instead of guessing each other's intent. The roles are open-ended. The workspace gives agents awareness of and a voice to each other; what they do with it is up to whoever defines the next role. It's for developers already running more than one coding agent at a time — and for anyone who's watched a single agent confidently do the wrong thing. The bet underneath all of it: the fix isn't one smarter agent, it's agents that watch, question, and divide work among themselves.

FML Inc project preview
0

Automat

Robotic Workflows is an MCP server that lets your AI agent build, deploy, and manage real RPA automations on Automat — deterministic browser/API workflows that then run on their own schedule with zero tokens per run. The agent writes the automation once; it runs forever, token-free.

www.loom.com/…
0

Cara

Medical Triage Agent for Non-Profit Healthcare (Chat & Phone & raspberry Pi) to support low income and houseless people

www.loom.com/…
0

Shop Floor Intelligence

Shop Floor Intelligence — an autonomous shop-floor monitoring agent for high-mix job shops. Point a cheap webcam at otherwise-uninstrumented manual machines and Claude Opus 4.8 vision turns them into monitored ones: the agent watches each machine, catches a stoppage or a blocked camera the moment it happens, investigates the surrounding frames, and drafts the next action plus a shift briefing — with utilization and availability as supporting context. You set it up in plain English ("the green lathe is on the left, the blue CNC is on the right") and it tracks each named machine by region. The product is the autonomous watch → catch → draft loop, not a dashboard or an image analyzer.

Shop Floor Intelligence project preview
0

DAWggies

DAW (Digital Audio Workstation) for AI agents. This feature helps Dramatis - AI orchestrator for full-cast audiobooks from raw text - to finely adjust the audiotracks for better immersivity and audio results.

DAWggies project preview
0

Interactive Manifold Steering

protocol that allows for long-horizon generative design loops inspectable, steerable, and trainable from human intervention

Interactive Manifold Steering project preview
0

Teckton

Tekton — the AI-native workspace that takes a hardware team from CAD to cost in one room. Getting a manufactured product from design to the factory floor is messy and manual. A designer models parts in CAD and emails/Slacks files back and forth across dozens of SME teams (thrust, loads, weights…), iterating thousands of parts. Then cost engineers take those designs and compute pricing by hand-dialing "knobs" — material, manufacturing process, supplier, country — to produce a notional price for vendor negotiation. The geometry, the engineering feedback, and the cost reasoning live in three different places. Context gets lost in inboxes, and pricing is slow and opaque. Tekton puts all of it in one workspace. You import a BOM (CSV) and a 3D STEP assembly — parsed in-browser via OpenCascade — and click any part in the tree or the 3D model to select it (selection syncs both ways). For each part, a Claude Opus 4.8 agent estimates cost from real physics: bounding box → volume → mass → process → country → quantity drives a notional unit price with a full line-item breakdown and rationale, backed by a deterministic volumetric model so an estimate always appears. An embedded Claude Agent SDK Copilot sees each part's design and pricing config and can run what-if estimates or apply a change and re-price instantly. Teams sign in, join by code, and collaborate on a shared design with per-part comments, presence, and approve/lock — isolated per team via row-level security. It ships as an Electron desktop app (native CAD/BOM file access, always-open workspace).

screen.studio/…
0

Taste Machine

AI has no context of our visual taste. Every output is bland. As designers, we want the AI to remember our taste. So we built a machine to distill visual preferences into one reusable context skill that any AI tool can ingest. It’s called the Taste Machine. You first train the taste like personalization on Midjourney or import from your existing places you’ve saved inspiration. You then preview your taste into different examples, and then link your taste to your LLM. We’ve built a Chrome extension that then makes it very easy to then continue refining your taste whenever you love something, and your LLM will use the latest version of your taste. Explore other people’s tastes and become a tastemaker.

Taste Machine project preview
0

Liminal

BTW My teammate Shayaun (Sean) Nejad (kuzushisec@gmail.com) was was waitlisted and got into the hack in-person today. Liminal is AI spend governance for founders and operators — the layer that keeps human judgment in command while AI agents accelerate execution. AI-native teams get handed frontier-model budgets and asked to prove ROI, but can't answer two questions: is this spend actually working, and can we govern it without surveilling our people? Liminal answers both. Eight bounded agents read a team's Opus 4.8 usage against its OKRs and disagree about what's waste; an adversarial reviewer then cross-checks every claim against the evidence and drops the ones that don't survive — it caught its own side's $162 "calendar" charge (PR-103 proved it was real product work), cutting a naive $446 estimate to a verified $284. Every finding is hash-linked to its source rows and anchored with a SHA-256 receipt; when an operator ratifies a policy, the decision enters an append-only provenance chain that corrections never overwrite. Governance without surveillance by construction — the subject is the agent fleet and the spend, never the person. We dogfooded it on our own $8,228 of Claude spend from building this.

Liminal project preview
0

Tally, SpEd Scheduling

Tally is an open-source scheduling engine and compliance ledger for special-education service minutes, and it closes a loop every existing tool leaves open. Every student on an IEP is legally owed a set number of service minutes each week (speech, OT, PT, counseling), and a single school-based clinician carries around 50 of these students across two campuses with different bell schedules, no IT department, and no integration budget. Building a conflict-free pull-out schedule by hand eats entire weekends, and proving the minutes were actually delivered is the part that carries the real legal exposure. The market splits into two camps that never meet. Generators build a schedule and never check whether sessions happened, and trackers log delivery but cannot build the schedule. Neither owns the gap between mandated minutes and delivered minutes, and that gap is the whole legal story. After Endrew F. (580 U.S. 386) an IEP has to be reasonably calculated to enable progress, so when a district is challenged the mandated-versus-delivered delta is the defense file. That exposure is compounding: state special-education complaints rose 22% in a single year and now run roughly 80% above their ten-year average, and a district's central vulnerability in a due-process hearing is the absence of delivery data. Tally owns that gap end to end. A deterministic OR-Tools CP-SAT solver builds a conflict-free, bell-schedule-aware schedule against each school's real periods and protected blocks, using day-types (A/B/rotating) rather than weekdays. A three-column ledger tracks mandated, scheduled, and delivered minutes per mandate, with the delta always visible and never rounded away. A missed session re-enters the solver as demand and lands in open capacity or surfaces as an explicit shortfall, so nothing is ever dropped silently. And every mandate exports against the configured state rule, with Illinois shipping as the demo (a 10-school-day makeup clock and a 3-school-day parent notice), kept as config rather than hardcoded constants. What makes it trustworthy in a regulated space is the solver itself. Scheduling is deterministic, so the same input returns an identical, fully explainable placement and no model decides where a student goes. A coordinator can defend every slot, and a district gets an audit trail instead of a black box.

www.loom.com/…
0

Amaltash

A Claude desktop app agent for the user to build a trading strategy using a few words only. You can give it a prompt like -- 1. BUY AI Stocks like Nvidia, Tesla, Micron, 2. SELL all AI Stocks when the QQQ price for 200 day average > current price of QQQ, and buy Treasury Bills instead. 3. If the market is healthy again, please go buy the AI Stocks again.. (Please note that as part of the hackathon, I built the Claude AI plugin -- so that users can build, deploy or test strategies on a LIVE / PAPER brokerage within Claude).

Amaltash project preview
0

COSINT

COSINT is a real-time "situation room" for US missing-person investigations. You drop in a person's photos and notes, and a swarm of Claude agents works the case from **public open-source intelligence** — it finds existing tools, builds MCPs or skills to interface with them, and uses the tools to crack cases. EXIF + vision-based photo geolocation, reverse-image search, advanced-search "dorks," people/phone/social lookups, and supervised browser computer-use, surfacing **cited, actionable leads** on a live map. Every claim links to a real public source that otherwise takes weeks to find due to our context layer. The problem: OSINT on a missing person is slow, manual, and hard to audit. Investigators and family advocates need *corroborated* leads with provenance, fast — and need to actually watch the work happen. COSINT makes the investigation parallel, watchable, and verifiable: a fleet of agents runs concurrently, each tool call streams into the cockpit live, and the geography lights up as findings land. It's demo-mode-first (boots and replays a full investigation with zero API keys), so anyone can see it work instantly.

COSINT project preview
0

The Cook Book Club

The Cook Book Club is a members-only community cookbook designed to preserve recipes and the stories behind them. Recipes often live in scattered notes, screenshots, family messages, and old documents, making them difficult to organize and share. With The Cook Book Club, members can upload a photo or paste recipe text from any source. AI automatically transcribes, translates, and structures the content into a clean recipe format, generating ingredients, instructions, servings, and dietary information for review before publishing. Members can browse recipes by category, cuisine, or dietary preference, search by ingredient, watch cooking videos, discover where to buy ingredients nearby, and connect with the community members who contributed each dish. The result is a living collection of recipes, memories, and traditions that can be preserved and shared across generations.

The Cook Book Club project preview
0

OncologyOS

OncologyOS helps cancer patients find the right clinical trials, using just their discharge documents The problem. Every year roughly 2 million Americans (and ~20 million people worldwide) are diagnosed with cancer, and almost overnight each one is handed a life-or-death information problem they're not equipped to solve. The relevant knowledge is real but scattered and unreadable: ~19,000 cancer clinical trials are actively recruiting in the U.S. alone, alongside thousands of approved drugs, off-label and repurposing options, and dense genomic and pathology reports — yet fewer than 1 in 20 adult cancer patients ever enrolls in a trial, and the most common reason isn't ineligibility, it's that neither the patient nor their busy oncologist ever finds the right one. Matching a single patient to the right options means cross-referencing their diagnosis, stage, biomarkers, prior therapies, and comorbidities against eligibility fine print written for specialists — work that is slow, expensive, and largely manual, so it mostly doesn't happen. The result is a brutal asymmetry: the options that could extend or save a life often exist, but the patient never learns they're there.

OncologyOS project preview
0

Lanuch Control

Launch Control is a content-launch engine for nonprofits and small teams. You give it one goal and an event, for example "Get 100 volunteers to our Saturday Ocean Beach cleanup" (surfrider.org, Ocean Beach SF), and a team of Opus 4.8 agents plans, produces, and runs a full week of social content engineered to drive people to that event. How it works, end to end (multi-agent orchestration): 1. Plan + understand. A planner agent studies the goal, the event, and the brand's website, and does live market research with the Bright Data API to learn what actually performs on each platform right now. It produces a 7-day, multi-platform plan where every post has an objective and the days ladder toward the event, with a shared weekly theme. 2. Write. A writer agent turns each slot into platform-native output across five content types: UGC videos, motion launch videos, images, stills, and plain text, with the media rendered through fal. 3. Adversarially review + regenerate. A reviewer/judge agent (also grounded in fresh Bright Data research) critiques every post against that platform's norms and a quality rubric. If it doesn't pass, it sends specific feedback and the writer regenerates, looping until it clears the bar. A multimodal Opus visual critic also looks at each rendered image or video keyframe and grades it for on-brand, on-intent, and "AI-slop" cleanliness, re-rendering failures. 4. Publish. Approved posts go out natively to X, LinkedIn, Instagram, and TikTok via the Zernio API, each with platform-specific copy and media. 5. Create the event. It spins up the real RSVP/event page automatically through the Luma API. 6. Schedule + adapt. It schedules the whole week to crescendo on event day, and checks a Weather API for the event days, automatically rescheduling and adjusting posts if rain is coming for an outdoor event. 7. Engage. A watcher monitors and replies to every comment on its own (via the Zernio inbox), in the brand's voice, and decides what to schedule next. Stack & how it fits together: a Next.js app on Vercel; Supabase for auth, Postgres, and storage (plans + the Asset Bay of brand assets the writer pulls from); Opus 4.8 running the planner → writer → adversarial-reviewer loop; Bright Data for the research that grounds both planning and review; fal for media generation; Zernio for connecting accounts, publishing, and the comment inbox across all four platforms; Luma for real event pages; and a Weather API for forecast-aware rescheduling. The Channels view renders each post inside a realistic native skin (X timeline, LinkedIn feed, IG grid, TikTok profile) using only real connected-account data, no mocked numbers. New users get a free week, then bring their own Anthropic + fal keys (a graceful CTA paywall). Schedule + weather-aware rescheduling (with generative UI). It schedules the whole week to crescendo on event day, then checks the Weather API for the event date. For an in-person or location-based event, if the forecast looks bad, it doesn't silently guess. It surfaces a generative UI pop-up that explains the weather risk and asks you how to handle it: reschedule the event to the next good-weather day (which then cascades, updating the Luma event date and re-timing the whole week of posts around it) or pick another option. Your call drives the plan. This is all live right now. We built and launched every one of these accounts today, and the engine is actively posting to and monitoring them: Claude Coder on X, Claude Coder on LinkedIn, claude_build_day on Instagram, and Claude Coder on TikTok. Go check them. If you comment on our Instagram post (@claude_build_day), Launch Control will automatically reply to your comment on its own (until our API credits run out). We open-sourced it as we built it, and announced the launch on LinkedIn through my post (@Manny2Techy). It's live for anyone to get on and try right now. Link to opensource post:https://x.com/Manny2techy/status/2065895792267702481?s=20

screen.studio/…
0

HackShackKash

Design Prompt → deployAgent → Your EC2 → debugAgent → Your GitHub flow = All in One

HackShackKash project preview
0

EasyReviews — Support your local businesses!

Most people never review the places they love — the blank box is too much friction, and businesses are left begging. Easy Reviews turns a 60-second conversational interview into a Google review to help support your favourite local businesses. It replaces the blank review box that stops most people, without crossing into fake-review spam.

www.loom.com/…
0

WorldLine

WorldLine is counterfactual debugging + institutional memory for multi-agent AI reliability. When an agent pipeline returns the wrong answer, WorldLine has Claude intervention-test every decision in parallel — forking the run, injecting a counterfactual, and re-simulating the downstream live — to find the decision that actually CAUSED the failure, ruling out plausible decoys (even the last-touch agent that looks guilty). It then authors a repair and proves it flips the outcome with a deterministic code assertion. Every verified fix becomes a durable lesson in a fleet knowledge graph. A different agent about to repeat the same failure class — even worded differently — is caught and fixed pre-emptively, and a pre-ship CI gate blocks any change that would reintroduce a known failure. Tracing shows what happened. WorldLine finds the cause, proves the fix, and makes your agent fleet stop repeating itself.

WorldLine project preview
0

Claude's fables

Claude's Fables turns a real thing a child is facing — won't share, scared of the dark, told a fib they won't own up to — into a short interactive fable where the child is the hero and makes the moral choice themselves. The lesson lands through the consequence of their own choice, never a narrator lecturing, and every story ends by handing the parent one line to talk about together. It's voice-narrated and tap-first, so even a child who can't read yet can play, and a built-in safety guardian checks every story before it reaches a kid.

Claude's fables project preview
0

Locus

The crisis of our time is human purpose. As intelligence commoditizes and the fabric of our day to day shifts from labor to orchestration, the core human questions will return, who am I? What am I meant to do? Locus begins to lay infrastructure for navigating those questions. Locus is grounded inference over a personal world model, shipped to mobile. It reads your personal archive — the moments you've actually lived — and extrapolates, in five grounded layers: who you are (identity), the scale at which you act (a concentric impact range, self → intimate → community → society → species), what the world needs from you at the right radius, where those overlap (a Venn intersection), and concrete ways to live it out — a job, a gig, a company ethos, a philanthropy, a movement. Every layer is grounded in real moment IDs you can tap to see.

www.youtube.com/…
0

densitygen

Why it matters. Every modern chip — in your phone, in AI data centers, in defense systems — is a stack of hundreds of films just a few atoms thick. Each generation of chip forces a switch to new materials because the old ones physically stop working. Finding the replacement today takes a specialist months of slow, fragile physics simulations, one material at a time — and a wrong pick can blow a multi-year, multi-billion-dollar fab bet. How densitygen helps. Type a plain-English requirement (“a leak-proof insulator for a 2nm transistor that survives a 1000°C bake”) and get a ranked shortlist of real candidate materials in seconds, pulled from the public Materials Project database, scored against the spec, with the unavoidable trade-offs on one chart. An AI agent does the tedious querying, filtering, and ranking; the engineer makes the call; the expensive simulations are saved for just the final two or three finalists. Months → an afternoon. Who it’s for. Computational materials scientists and process engineers at semiconductor fabs (Intel, TSMC, Samsung), defense/aerospace R&D, and the precursor-chemical suppliers (Merck, Entegris, Air Liquide) whose multi-million-dollar bets ride on picking the right material.

www.loom.com/…
0

DAATUUM

The real-time coordination layer for teams of developers running Claude Code agents. Git coordinates code at rest; Datum coordinates agents in motion. A team does spec-driven development right — shared spec, common CLAUDE.md, PRDs split per engineer — then implementation starts and truth changes. One agent renames a column or swaps a dependency, and the others keep building against a contract that no longer exists. The drift surfaces at merge, after the rework is paid for. datum is the live source of truth that catches it at the next write instead.

datum-tower.pages.dev/…
0

Essert

Autonomous Brand Manager

Essert project preview
0

AutoCut

AutoCut — an autonomous AI video editor. You upload raw footage and give a one-screen brief ("Viral Short, 9:16, 30s, loop back to the first frame"), and an agent does the rest of the editing. How it works Offline preprocess (src/preprocess/) turns raw clips into a machine-readable library: - ffprobe for duration/resolution, Whisper (via Replicate) for word-level transcripts, scene detection + a VLM for per-scene captions and tags, and now a poster frame per clip for the UI. - Output is clips.json — the structured contract the agent reasons over directly (no ffmpeg in the agent's hands). The agent loop (src/agent/, src/loop/) reads clips.json, picks segments, and emits a schema-validated EDL (edit decision list). A deterministic renderer (src/edit/, ffmpeg) compiles the EDL into the final cut, and an independent verifier returns a structured grade. Those zod schemas in src/loop/types.ts are the single source of truth at every boundary, so a bad model output fails fast instead of three modules downstream. The UI (Next.js, app/ + components/) is a two-screen flow: Setup (the footage library grid + the brief form) and Generation. Built with shadcn/ui + Tailwind. Infrastructure - Supabase Storage is the source of truth for assets: source-clips (private — videos + manifest), renders (public — final mp4), thumbnails (public — poster frames). Daytona persistent disk is a download cache. Supabase Postgres + pgvector holds caption embeddings for semantic clip search. - Models: Opus 4.8 drives the agent (not Fable 5); Replicate hosts Whisper + the VLM. Where it stands Preprocess and the core loop work and are well-tested (~193 unit tests). Known gaps: the agent server still hardcodes a SAMPLE_LIBRARY instead of loading real clips.json (agent-server.ts:141,274), the Daytona caching layer is specced but not built, and transcription is on base Whisper with WhisperX (true word-level) as the agreed next step. This session's work: rolled per-clip thumbnail generation into the preprocess pipeline and got the poster frames rendering on the Setup screen's footage grid.

canva.link/…
0

Codewalk

Codewalk saves knowledge gleaned from Claude Code interactions and other techniques to self improve the coding agent. It makes future Claude Code sessions 3x faster, 3x cheaper and with better quality results.

Codewalk project preview
0

Pip

Automatically converting Chip Specs into GDSII that's ready for TSMC

Pip project preview
0

JV Solo

I was doing PoC using Anthropics Managed Agents to automate company internal processes like AP Process or Business Travel management. My goals for today were: 1. Define full harness, validate it by PoC 2. Leverage Managed Agents 3. Create full pipeline to deploy these on Claude Platform

JV Solo project preview
0

Rankenstein Team

Rankenstein is an autonomous, self-correcting content engine for Shopify stores and agencies. It takes a store from connect → grounded content → human review → live publish, fully autonomously between human checkpoints. The problem it solves: AI content tools hallucinate facts and publish unreviewed slop that erodes SEO and trust. Rankenstein never invents a fact — every claim is grounded in a FactsTable built from the store's real catalog, and an independent fresh-context verifier grades each piece against a rubric, refusing-and-flagging ungrounded claims (fabricated GSM, certifications, reviews) before a human ever sees them. Keyword research uses SERP-ownership scoring to pick fights a store can actually win, and output is structured for answer engines (AEO + JSON-LD), not just Google. Reviewers leave anchored comments that drive a surgical edit — only the commented span changes, proven by an independent span diff — then approve and publish to the live storefront with a version snapshot and one-click rollback. We connected the real EZ Fabric store, synced 903 products, and ran the full pipeline end to end on rankenstein.app

www.loom.com/…
0

Mollusc

Mollusc is an agentic financial advisor for people with complex financial lives: RSUs, cost basis, mortgages, taxes, retirement accounts, and long-term planning decisions. The demo follows Mira, a Bay Area engineer with concentrated NVIDIA wealth. Claude dynamically organizes her welcome page around what matters most, surfaces proactive planning items, and launches focused deep dives like an RSU sell-down strategy. Behind the scenes, a long-lived agent loop uses financial tools, scenarios, tax lots, public research, and rich report generation.

www.tella.tv/…
0

OurLink

The project is ourlink.ai. It's a solution to the end of corporate jobs and the disruption of the capital vs. labor balance. Basically instead of people needing to get jobs from giant corporations, it's a way for humans to share value and creativity WITH OTHER HUMANS DIRECTLY.

ourlink.ai/…
0

Tactile Diagram Workbench

Tactile Diagram Workbench is an open-source braille compiler for blind STEM students. A teacher can upload a textbook diagram, photograph a hand-drawn sketch, or type a concept like “draw acetone,” and the workbench turns it into an export-ready tactile sheet: raised-line SVG/PDF for swell paper or tactile-graphics embossers, plus .brf braille labels for standard text embossers. Chemistry diagrams use a verified pipeline with canonical SMILES and deterministic structure checks; other STEM diagrams route into a clearly labeled teacher-review draft lane so teachers can still produce usable tactile materials without pretending they are fully verified. The goal is to move tactile STEM diagrams from a weeks-long procurement workflow into the same class period.

screen.studio/…
0

Mack

Mack is an AI insurance broker that gets a contractor a real, bindable Certificate of Insurance (COI) in minutes instead of 3 days. The problem: a contractor wins a job, but the general contractor won't let his crew on site without a certificate of insurance. Getting one today means Googling, a 45-minute phone call bounced between call centers, repeating the same info, answering insurance questions he doesn't understand, then a second call to pay by reading his card number aloud — about 3 days end to end. With Mack, the contractor enter his business name and state. Mack pulls the rest, gets a real quote from the carrier, ranks the options with the why, and returns a certificate of insurance + pay link. No phone calls. Built by a licensed US broker, and every number is traceable to a real carrier — we never invent a quote. It's a back-office workflow collapsed from days to minutes.

www.loom.com/…
0

eye-eye-ai

Anatomy of a Claude Code Request** is a live **3D Architecture Map**: it turns software architecture into a navigable 3D model where motion and visual effects *carry meaning* — colour is category, glow is activity, wireframe-vs-solid is build status, particles along a conduit are data flow, an orbiting satellite is an AI agent at work. You fly through the system, swap **Lenses** (Logical ↔ Deployment ↔ Request-flow), and watch AI-proposed changes light up their "blast radius" before you approve them. The map's subject is **how a Claude Code request actually works** — your prompt travelling from the terminal, down through the API/serving layer, into the GPU inference fleet, and back as streamed tokens. The architecture shown is an honest *common-sense model*, not leaked internals, and that's the whole thesis: > a 3D map lets you reason **spatially about any system** — even a black box you > can't read the source of. It's deliberately self-referential — the demo visualises the very pipeline that generated it. **The meta angle (our headline):** this project was **built by a fleet of Claude Opus 4.8 agents** running in parallel, conducted by a human director through a custom tmux orchestration. The build process is itself part of the story — a human directing a team of AI engineers, each in its own worktree, coordinating only through a shared "seam." See §7. - **Live demo:** https://deck.nautex.ai - **Stack:** Vite + TypeScript + Tailwind + Three.js (React-Three-Fiber) · FastAPI + all-Docker on a DigitalOcean host behind a shared Caddy.

www.loom.com/…
0

AGORA

AGORA is an AI-native event platform with an agent built into every stage of the night. For people going out, Agora answers the two questions a calendar never will — is this worth my evening, and who should I actually talk to? It sends a weekly brief of events that fit a goal you set (“rooms where seed investors will be”), and when you check in it tells you the two or three people to meet, with an icebreaker drawn from what you’re both working on. For people hosting, Agora does the judgment work, not just the logistics. Describe your event in a sentence and it drafts the page, theme, and the questions that quietly qualify guests. It researches every RSVP and recommends who to let in, with reasons. It flags when the room is lopsided, helps you find and invite the right people, and lets you run the whole thing from inside Claude or ChatGPT over MCP. Web app, an iMessage agent, and an MCP server.

www.loom.com/…
0

michelangelo

an agent research harness- built using tmux for long running agent tasks and exploration in science and engineering research.

michelangelo project preview
0

MalarIA: A field worker decision agent for malaria prevention in Africa

Malaria is a deadly disease that has been eradicated in several countries, but many African countries are still losing their lives to it. Field workers are the warriors who are instrumental in eradicating malaria with intervention mechanisms, but using the right intervention at a wrong time is extremely ineffective. They need real-time, real-world conditions on when to use these interventions. That's what MalarIA is achieving. It is built with field workers in mind, who may not have laptops in the field, so it's a Whatsapp application, enabled to speak four local languages (English, French, Portuguese and Chichewa). It provides them with conditions for intervention, and checklists so they can be prepared.

www.loom.com/…
0

Jerome Ortega

It was meant to be a budgeting app that disambiguates big box store recepts.

github.com/…
0

Versi

Versi is a source control system built for the way software is starting to get made with AI agents. Git works well when a human edits code, stops, commits, and pushes. Agents do not work that way. They run in temporary sandboxes, get killed mid task, work across machines, and often make useful progress before there is a clean commit. Versi separates preserving progress from declaring meaning. Every saved file change can be recorded into a durable line of work, so an agent workspace can disappear without losing the work. Later, a human or agent can create checkpoints, run validation, compare approaches, and promote only the change that is ready. Our hackathon prototype shows the core primitive working end to end. We track saved file changes, sync them across separate local workspaces, create lines of work for parallel attempts, validate a checkpoint, rewind to a good state, promote the result to main, and resume from a fresh workspace through a remote store. The goal is not to replace Git for everyone. The goal is to build source control for parallel human and agent work, where workspaces are disposable but the work is not.

Versi project preview
0

Askable Labs

A pull request is the closest point of value to a customer. All the upstream work, research, product thinking, design, engineering has already been done. The change is written. It's ready. The PR is the atomic unit where effort converts into customer value, and it's the single highest-impact point in our delivery pipeline. That's exactly where we have a bottleneck. At any given time we have ~50 PRs in the queue across a 15-person engineering team. No one is reviewing accurately — at this volume, engineers are approving, not reviewing. The code we produce (increasingly AI-assisted and agent-generated) has already outpaced what a human can meaningfully validate by reading diffs. This isn't a process problem we can fix with better review habits. It's a structural problem. The model of "a human reads the diff before it merges" doesn't scale when code production accelerates and the cost of context-switching into someone else's change is high.

Askable Labs project preview
0

NanoFlow

RedStripe is a MAP-monitoring platform for brands and manufacturers that answers one expensive question in real time: which retailer broke our minimum advertised price first? When a single retailer drops below MAP, the rest follow the floor down within hours - and the brand pays twice: once in eroded price integrity, and again in the price-protection rebates it owes retailers to compensate for the very price drops it failed to catch. RedStripe watches every SKU across every authorized retailer and major marketplace, flags the first-mover breach the moment it happens, and turns "we found out in next month's report" into "we got an alert in minutes." That timing difference is the whole product - and it maps directly to dollars retained. The second pillar is third-party marketplace control. RedStripe continuously tracks third-party seller activity on Amazon, Walmart and other marketplaces - buy-box ownership, unauthorized sellers, pricing below MAP, cart-quantity games - and maintains a full seller directory (business name, address, registration/VAT, contact details) that doubles as a ready-made cease-and-desist contact list. This is built on real, substantial data: the platform migrates and normalizes a legacy dataset of ~900,000 marketplace offers and ~17,000 sellers across 14 locales/currencies, so it's solving the problem at production scale, not on a toy fixture. It's a complete, live SaaS, not a prototype. Built end-to-end with Claude — TypeScript on Fastify, PostgreSQL 16 with Drizzle, BullMQ scrape scheduling with tiered provider fallback, a rules-based alert engine delivering to email and Telegram, and multi-tenant isolation - it already runs behind TLS at demo.redstripe.io, with authentication, the legacy data migration, and the new marketing site shipped. Claude did the heavy lifting across the stack: schema design that normalizes a sprawling per-retailer SQL Server database into one clean model, a verify-then-upgrade password migration path, and the scraper/alert architecture, all under a phase-gated plan with real verification at every step. The market validation is concrete: demos are lined up with Garmin and Jabra - exactly the kind of premium consumer-electronics brands whose MAP discipline and rebate exposure make this a board-level concern. RedStripe addresses a known, quantified pain those brands already feel and already spend real money on, which is why the pipeline exists before the product is even fully GA. The pitch isn't "here's a clever tool" - it's "here's the leak in your pricing program, here's the first retailer who opened it, and here's what it just cost you."

www.loom.com/…
0

Pave Capital

Tape turns a plain-English trading idea into a backtested, deployed Polymarket bot in ~60 seconds. Active prediction-market traders have a discipline problem: they know their strategy but can't code a bot, can't watch markets 24/7, and panic-second-guess at the worst moment. With Tape you type a strategy in plain English ("Buy NO on geopolitical markets above $0.92 that resolve within 14 days, skip anything under $50k liquidity"). Opus 4.8 compiles it into a runnable Python strategy module, a backtester replays 90 days of real Polymarket CLOB fills, a model-verifiable rubric (rubric.yaml) grades it PASS/FAIL with reasoning, and on PASS it deploys to a budget-capped paper-trading sandbox whose first cycle streams back live. The result: a trader locks discipline into code without writing any code, and a hard budget cap means no strategy can risk more than its allocation.

Pave Capital project preview
0

Sales Factory

Reps spend most of their day on everything except selling. Sales Factory does the manual work for them and it lives where they already work: Slack and the call itself. Drop a Google Meet link in Slack and the agent joins the call, listening live and posting coaching tips into the thread as the conversation unfolds. The moment an offer comes up, it follows up in-thread with a shareable offer link — backed by real Salesforce records (Opportunity, multiple Quotes with line items) and an auto-generated pitch deck. The customer accepts straight from the link, and the agent closes the loop: the Quote is marked Accepted, the Opportunity flips to Closed Won, and the win posts back to the Slack thread. No CRM data entry, no copy-paste, no context-switching. Sales reps spend so much time not selling. Sales Factory automates the manual work — where reps already are. 1. Paste a Google Meet link in Slack → the agent joins the call. 2. Live coaching → it messages the thread with real-time tips as the call happens. 3. Offer made → it replies with an offer link, real Salesforce Quotes + records, and a pitch deck. 4. Customer accepts via the link → the agent updates Salesforce (Quote Accepted → Opportunity Closed Won) and posts the win back to the Slack thread. First Part of the demo https://youtu.be/i81xahVwzaY Second Part https://youtu.be/wniY2cel6Zc

Sales Factory project preview
0

meanwhile...

meanwhile... kicks off ambitious work, autonomously, when it detects you won't use all of your subscription for a leading agentic harness (Claude Code, Codex or Antigravity). The problem we're trying to solve: I have a lot of ambitious, non-critical path work that I don't kick off because it's not immediately blocking. I also have credits I don't use at the end of the day/week. I'd like my agents to try to do great stuff for me with the credits I'm not using.

us06web.zoom.us/…
0

Shrimp Welfare

A Roguelike strategy game where you help navigate a population of shrimps through an age of rapid AI advancement and scientific discovery. Guide decisions on government regulation, AI safety, control, and shrimp agency to decide the fate of the next stage of civilization. This project is a 3d game built with the Godot game engine.

Shrimp Welfare project preview
0

Orchid

Code tells you what. Git tells you when. Orchid tells you why. Capture every AI coding conversation. Surface the reasoning behind your code to reviewers, teammates, and agents.

Orchid project preview
0

Vex

Source control for the high-throughput agentic era. Git was made for the human era: human commits, human reviews, and human-paced remotes. VEX is built around one infinite loop where humans and agents continuously read, change, test, deploy, observe, repair, and rebuild the software organism. Vex is built on high scale monorepo principals and integrates a distributed virtual file system layer to allow for 1000's of agents to fork and and work on a project at once

www.loom.com/…
0

Labmate

Labmate is an agent-native data science harness. A human gives the business judgment — a dataset, a target, a metric, and constraints — and a real Anthropic Managed Agent (Opus 4.8) does the science: it profiles the data, writes a data contract, proposes hypothesis-driven experiments, runs them for real in Modal sandboxes, critiques its own results for leakage and test-set tuning, folds in plain-language human feedback, and produces a reproducible report with provenance. Built in one day, entirely with Claude Code.

www.loom.com/…
0

HelixChakra

Helix is an adaptive workflow operating system that transforms natural-language business goals into executable workflows composed of AI agents, human approvals, and intelligent recovery paths. Today, most AI applications stop at generating answers. Enterprise work requires planning, coordination, execution, and trust. Helix bridges this gap by dynamically orchestrating multiple agents, adapting to failures and changing conditions, involving humans when necessary, and providing complete observability through an integrated agent debugger and execution timeline. By making AI workflows explainable, recoverable, and auditable, Helix enables organizations to move beyond chatbots and toward trustworthy, enterprise-grade agentic systems.

HelixChakra project preview
0

One Door Fully Stacked

One Door is an AI benefits advocate for Californians who qualify for help but get stopped by confusing rules, language barriers, documents, and phone calls. Every year about 2.7 million eligible Californians receive no CalFresh (SNAP) benefits, leaving roughly $3.5 billion in already-appropriated federal food dollars unclaimed. A common cause: screeners run the federal default rule (130% of the poverty line) and never apply California's higher 200% limit, so a working family that actually qualifies is told "no." We built a working CalFresh eligibility advocate that pairs a deterministic, cited benefit engine with Opus 4.8 for the human layer. It takes a household's situation in any language, computes a cited CalFresh verdict and dollar amount, explains it in the person's own language (every digit and citation stays byte-identical), reads pay stubs, surfaces adjacent programs, and refuses to invent a dollar figure for any program it did not fully compute. Our demo runs the highest-stakes path: a California household that looks ineligible under federal rules but is eligible under California's. One Door flips the verdict, shows the estimated benefit ($231/month), cites every rule that got there, translates the guidance, and preps the person for their county interview, all without inventing amounts for programs it did not verify. Core principle: use AI for the messy human parts (language, documents, fact triage, advocacy, navigation) and keep benefit math deterministic, testable, and auditable. Final verification gate: 30 passed / 0 failed, with independent-oracle accuracy of 100% on 14 hand-worked cases.

us02web.zoom.us/…
0

Saiga

Saiga.so solves the problem of distributed teams not having context on what other people are working on. For this hackaton I developed the next stage of the platform which helps the founders or and team members to get a good overview of all projects and move stale tasks into production and done suggesting agentic workflows.

www.loom.com/…
0

Loom

We use tools like Git, Kafka, databases, video streaming, filesystems... but really, what are all of these? They are ways to store data, determine event ordering, and interact with data. In an effort to combine all of this into one platform I created Loom which can act as a database, filesystem, git, event-ordering, or really anything that needs to store and interact with data. It can store anything from an LLM model to a constantly changing video - all through the same API and highly efficiently with Rust + some algorithm optimizations. (In technical terms, an append-only asynchronous streaming database with a deduped blob store and pub/sub).

othellotrainer.com/…
0

Team Mowie AI. Hackathon Project: Scout GTM

Scout GTM turns a startup's URL plus a 60-second intake into a complete, executable go-to-market dossier in under 10-minutes. Early-stage B2B SaaS founders all hit the same wall: there are plenty of tools to run go-to-market (Clay, Apollo, AI SDRs), but nothing that tells them the strategy — that's still a $4,500–6,000/month agency. Scout researches your company live, classifies it into a GTM archetype from a recipe library, and assembles a tailored plan — channel priorities, a trigger-campaign waterfall, funnel math, and a 90-day sequence — with every claim sourced and evidence-graded. An independent verifier grades the dossier against a rubric and regenerates until it passes, so the output is something a founder can execute GTM experiments.

Team Mowie AI. Hackathon Project: Scout GTM project preview
0

BlackBox

BlackBox is a hybrid QA platform where AI agents and human testers share one board. Point it at any product URL — Opus 4.8 autonomously discovers the app's features and use-cases, generates a structured test plan, and executes cases in a real browser. AI agents mark pass/fail with evidence and confidence scores. When an agent can't confidently complete a case — judgment call, ambiguity, or blocker — it escalates and relays the case to a human tester with a full handoff brief and screenshot, instead of guessing. Humans resolve from the same live grid. The result: trustworthy QA at scale, where the AI knows its limits.

BlackBox project preview
0