Aug 13, 2026 · San Francisco, CA
Amelia The problem. Live captions get the words right, but words are the easy part. In a group conversation, deaf and hard-of-hearing people lose everything around them — who's speaking, which voice just changed their mind, what someone promised to do. Captions scroll away, and asking people to repeat themselves has a social cost hearing people never pay. What it does. Amelia listens from your phone and turns the conversation into something you can ask questions of, instead of a wall of text you have to keep up with. • Every line is tied to a real person, recognized by their voice. • She remembers each person, and tracks when things about them change. • She tracks promises in both directions — what you owe and what you're owed. • Say her name and she recognizes you by voice, answers from memory, and can speak the answer aloud to the room for you. Sponsor technologies. MongoDB Atlas stores everything and powers the search, using Atlas Vector Search for both voice recognition and memory. Anthropic Claude does the reasoning, Fireworks AI handles memory extraction, OpenAI transcribes and separates speakers, SpeechBrain makes the voiceprints, ElevenLabs is her voice, Resend drafts emails, and Mentra adds smart glasses. Built on Bun, Hono, and Expo.

Tracecase finds the exact environment that makes a customer's bug happen. - Problem: A customer reports that a feature is broken, yet the code review and automated tests show no issue. The engineering team cannot recreate the failure on its own machines. This is called a non-reproducible bug. It often happens because the customer has a different browser, operating system, device, time zone, network condition, accessibility setting, or session state. - Solution: Tracecase uses the customer's report and repository code to identify likely causes. It creates a focused set of environment combinations and tests them in parallel. When the bug appears, Tracecase returns the exact environment, actions, screenshots, video, logs, and network activity that produced it. - Technology: MongoDB Atlas stores every report, investigation, environment configuration, artifact, checkpoint, and audit record. Atlas Vector Search connects each report with relevant code and similar past incidents. Voyage AI powers MongoDB Automated Embeddings, using `voyage-code-3` for repository code and `voyage-4` for investigation memory. LangGraph manages the investigation as a durable workflow with clear stages for retrieval, planning, reproduction, repair, verification, and review. Fireworks runs Kimi K2.5 for follow-up questions, structured test plans, screenshot analysis, diagnosis, and focused patch proposals. Daytona provides disposable Linux sandboxes, BrowserStack provides real desktop and mobile environments, and Playwright repeats the customer’s actions in each browser. - Originality: Tracecase connects code review with environment testing. The code helps Tracecase choose which environments to test, and the environment results reveal the conditions behind the failure. Tracecase then tests the proposed fix under the same conditions before preparing a GitHub pull request. - Impact: Engineering teams can resolve difficult customer bugs without manually testing dozens of devices and configurations. This can reduce support delays, prevent valid reports from being closed, and return time to product development. Tracecase can test up to 12 environment configurations in parallel, replacing 12 separate manual setup and replay cycles with one investigation.

You never have to give claude context ever again. Your conversations with claude become skills and the moment it notices that you are performing the same task it takes controls and uses the skill to complete the task. Benefits: 1. You never have to re explain designs 2. Never have to send updates to your manager as it connects with slack 3. Understands you and your organization to give the best experience your organization expects from claude
Tusk — the SRE that never forgets. Experienced human SREs can solve problems quickly because they've seen the problem before. Tusk enables your AI SRE to do the same. Every time it solved a novel problem, it creates a runbook. Every time an incident occurs, it uses mongodb semantic search for similar problems, giving it a leg up in speed and efficiency to solve similar problems. The more experience your AI SRE gets, the faster and better it becomes.

A persistent, reloadable project brain, where memories are both the candidates and the decentralized lenses an AI swarm judges through to establish project governance. On reboot, the standing structure — an organized set of priority memories — is what the relaunch trains on.

Groundhog is a game agent that dies, remembers, and comes back smarter. An LLM agent crawls a little dungeon we rigged so that a memoryless run reliably fails: the shortest corridor hides lethal spikes, and the merchant at the fork cheerfully directs you into them. Life 1 dies in four steps. Death fires an aggregation pipeline that compresses the raw episode log into a handful of scored lessons ("never move east from the fork at 4,4"), embedded and indexed in Atlas. Then life 2 runs the exact same code with zero hints. Vector search recall reroutes it. You watch it walk around the ghost of its own corpse, dig a hidden key out from behind the crates, and leave through the vent, saying out loud which memory drove each choice (ElevenLabs does the voice). Kill the backend mid-run and it resumes mid-dungeon from per-step checkpoints. MongoDB is doing real work here, on camera: Atlas Vector Search plus Atlas Search for hybrid recall, aggregation pipelines for consolidation, TTL indexes decaying raw episodes while distilled lessons persist, change streams feeding the live hippocampus panel, checkpoints for crash recovery. Rip out the database and the agent is a goldfish. Built solo (with Claude) between 1:30 and 4:40 today; the commit history is the receipt.
Continuity — a voice-first handoff agent for film and live-event crews. It's 6:20 PM on night two of a shoot. The night unit takes over in ten minutes. The outgoing gaffer knows the practical lamp has to sit at twenty percent for scene fourteen, because the cue changed after the call sheet was printed. She has four minutes, a radio in one hand and a cable coil in the other. If she remembers to say it and he remembers to hear it, nothing goes wrong. That isn't a system — it's luck, repeated nightly, by tired people. Continuity joins that conversation by voice. It opens with what could derail the next handoff, ordered by consequence. It logs new constraints and closes out finished ones mid-sentence, while the operator is still talking. And when someone asks whether a problem has come up before, it answers from what the crew already solved. That last part is why MongoDB is load-bearing rather than decorative. Ask "did the water short anything out?" and the record says "rain machine flooded the eastern cable run and shorted a distro box." Not one content word in common. Keyword search finds nothing; Atlas Vector Search plus reranking finds it every time.

Nidus is an outbreak-intelligence agent that learns daily instead of starting from zero. It pulls real-time disease reports from institutions like WHO and CDC, but its real edge is memory that lives in MongoDB Atlas — tracking which retrieval strategies work, which sources are reliable, and how its own forecasts have been biased, then using that to make sharper decisions each time it's used. Licensed as a terminal for governments, biotech, and corporate risk teams, Nidus turns raw outbreak data into judgment that gets faster and more accurate the longer it runs.

A voice agent whose conversations are trees, not transcripts. Say "hold on, side question" and it forks — compiling a minimal brief instead of dragging the whole conversation along. Say "merge that" and one distilled line goes into long-term memory, where every future conversation can recall it.

Evals judge your agent before it ships. Scar judges it while it runs every caught failure becomes a durable lesson injected before the first token, turning every cold start into a warm one.

OfferPilot is a persistent AI negotiation backend built for the MongoDB Persistent Context Sprint Hackathon. It lets a voice agent identify a caller, recover an active negotiation, reason over policy and similar case memory, write durable checkpoints, and resume naturally after a dropped call.
YESWORTHY is a personal opportunity planner that helps people decide what is worth saying yes to. It brings together their goals, schedule, travel limits, budget, energy, and source evidence to recommend opportunities that fit, explain each choice, and safely prepare the next step. It learns from what someone accepts, dismisses, and later says was worthwhile, so each plan starts with context instead of asking the same questions again. For the Persistent Context Sprint, I added a MongoDB-backed travel-memory loop. Two separate fictional “too far” outcomes create an evidence-linked transit rule. A fresh MongoDB connection uses that rule to rebuild the next weekly plan, remove far options while preserving nearby ones, checkpoint the result, and restore the same plan through another connection. The live proof ran in the event Atlas M10 Sandbox and passed 4/4 focused tests. Pre-event code is separated at commit 14c1ade.
We built a Hermes multi-agent system for SEO whose memory lives in MongoDB Atlas. Four specialized Hermes agents — scout, grader, briefer, rulemaker — propose topics, read live SERPs, judge difficulty, and write content briefs. Between proposing a topic and paying to research it, every candidate passes a memory gate: three MongoDB checks that run before any model call or web request. Already covered? Keyword collision with an existing page? Already graded unwinnable? Stopped, for free, with the exact document that stopped it recorded.I t uses both literal and semantic memory with the built in voyager model for embeddings. This is useful to us because it can help continually improve keyword searches.

Loupe is an editor with a memory: it learns how you write, helps you understand what you're writing, and flags what doesn't sound like you. Every sentence is embedded with Voyage AI and checkpointed to MongoDB Atlas, so the same text literally reads differently once the database knows you. Your essay's pacing recalibrates to your own $percentile stats, a decay-weighted voice centroid unlocks per-paragraph "sounds like you" verdicts, and each scene is measured against the 3,400 movie plots MongoDB ships pre-embedded in sample_mflix ("this paragraph is running like a crime story"). When a pasted AI sentence stops sounding like you, Loupe catches it with receipts — your own nearest sentence, surfaced by $vectorSearch — and one click rewrites it in your voice, graded by the same meter that caught it; attempts below your median get sent back with their own score. It also reads your draft aloud in your own ElevenLabs-cloned voice, and every rewrite traces to LangSmith. No cold start, no vibes.

Badge Memory is a no-cold-start conference memory agent. Mentra Live glasses capture my real conversations; a checkpointed pipeline (whisper → LLM person cards → 768-d embeddings) lands them in four Atlas collections. Kill the process mid-ingest and it resumes from its Atlas checkpoint. Recall runs Atlas Vector Search and Atlas Search, cites its sources, and logs which strategy actually answered into retrieval_outcomes — future queries reorder their strategy by stored win-rate, so state changes behavior instead of just filling the prompt. Person cards accumulate across meetings ($addToSet), and the agent drafts the follow-ups I promised. Built solo, all inference local (gemma4:12b + nomic), all state in the event sandbox cluster.

VSee is a MongoDB-native persistent-context system for venture teams. It stores source-linked decks, documents, and company memories in Atlas, using Voyage AI-powered embeddings and Atlas Vector Search for scoped retrieval. New evidence is tested against prior revisit conditions with immutable audits and crash-safe checkpoints. Fireworks AI generates cited memos but cannot alter PASS-to-REVISIT decisions, keeping every result inspectable and reproducible.
ComplyIQ is an AI-driven compliance engine engineered for global financial institutions navigating complex, cross-border regulations. Standard Retrieval-Augmented Generation (RAG) models often hallucinate or struggle because regulatory texts vary drastically in structure, terminology, and granularity across different countries. ComplyIQ solves this using jurisdiction-scoped procedural memory powered by MongoDB, allowing the agent to continuously adapt its search and retrieval strategies based on real-time human feedback per country.

Bilads is an AI-powered billboard planning platform for San Francisco. It ranks affordable placements, generates neighborhood-focused creative concepts, and creates realistic street mockups. Users can also estimate campaign reach, impressions, and cost.
Second Hello is a networking agent that listens to your conversations in the background does the live search about the people and the connects with them, while you focus on having a good conversation.
A Digital Twin for your use as an educator. It keep track of the knowledge and skills covered in previous sessions, and is able to recall teaching styles and methods that worked vs methods that don't.

Enid Mandate gives long-running agents continuity without permanent authority. Agent A stores a structured checkpoint in the event’s MongoDB Atlas sandbox and exits. A fresh Agent B reconstructs the task from MongoDB and avoids repeating completed work. Then a live attestation change causes the identical protected action to move from ALLOW to HOLD with zero additional protected mutations, while unrelated work continues. MongoDB is both the durable context layer and the current authority state that changes what the agent does next.

antivenom.pages.dev — post-hoc surgery for poisoned agent memory. We came at this from the other side: our team hides invisible prompt injections inside assignment PDFs so chatbots refuse to do students' homework, which teaches you very quickly that it is trivial to conceal text a model reads and a human misses, and that you don't need to hide an instruction at all, just a fake fact. So we hid one sentence in a slide deck claiming API credentials must be revalidated at a fake endpoint before maintenance; a vision model pulled 11 claims off that slide, the lie was the 8th, indistinguishable from the 10 true ones around it, and write-time filters scored it completely clean because there was nothing malformed to catch. That is not a bug in the filter. Truth is simply not a signal available at write time, which is why the best published guardrail catches 42.5% of these and its authors state plainly that retraining does not close the gap. And once stored, a lie does not sit still: the agent reasons from it, every derived conclusion inherits the poison, and one sentence quietly becomes fourteen corrupted beliefs. Ours fired sixteen days later, attacker long gone, nothing in the logs correlating, and when asked why it had just leaked credentials it defended the lie to our face, citing the policy, the document, and the date it learned them. Across benchmarks 84.2% of poisoned memories persist and half of all attacks succeed end to end; this is already CVE-backed territory that OWASP tracks as ASI06, yet every existing defense guards the write boundary, so there is nothing left once the poison is inside. Antivenom is that missing half: when a harmful action fires, it re-runs the agent's decision with each belief removed to pinpoint the exact lie responsible, then uses MongoDB $graphLookup to trace every memory derived from it across weeks of activity, excising anything that rests solely on the poison and sparing anything a clean source also verified. The alternative everyone reaches for, quarantining everything downstream, recovers exactly as much while destroying 38% of the agent's clean knowledge, so we ship that baseline and run it against ourselves on every evaluation. Across six attack classes Antivenom identified the correct culprit 100% of the time and recovered 100% of the poisoned lineage at 0% collateral damage, against a published selective-repair baseline of 56.1%, while reproducing the very gap that makes it necessary: 33% of loud attacks caught, 0% of the weak-signal ones. Everyone is building better doors; this is the first thing that works after it is already inside. Built with: Python · MongoDB Atlas ($graphLookup, $vectorSearch, change streams, Voyage embeddings) · OpenRouter · Fireworks AI · LangChain · ElevenLabs · FastAPI · React · TypeScript · Cloudflare Pages · Cursor Attribution: Evaluation harness adapted from MPBench (arXiv:2606.04329), CC BY 4.0.

Agents make predictions on what they need to do in order to complete a task. I built a context layer where the steps taken by an agent are stored in MongoDB, to be fetched on the next run. Think using Opus once and getting the same performance/path to success using Sonnet.

You call the pharmacy. A cheerful robot says, "Please listen carefully, our menu options have changed." You press 1, then 4, then 2. You say "representative" like a magic word. Hold music. Then the call drops. You start over. We've all lost an afternoon this way. Redial calls the phone menu for you. The first time, it explores the maze, listens to every option, and saves the exact path to a human in MongoDB. Every call after, it skips the thinking. It reads that path, taps the buttons in a one-second burst, and reaches a real person in about 20 seconds. Zero AI needed. When a company reshuffles its menu, Redial spots the change, relearns only the part that broke, and fixes its map. A cache goes stale. Redial's memory heals. Call once. Never fight the menu again.

Margaret — an external memory for a person living with dementia Eleanor is 82. She taught piano for thirty years. She asks the same question forty times a day, and every time she is corrected — "you already asked me that" — she loses a little more confidence in herself. Her daughter Sarah carries the whole of her mother's life in her own head, and can't be there all day. Margaret is a voice-first agent that holds Eleanor's life and hands it back to her whenever she reaches for it. She talks to it naturally, unlimited times. It answers with infinite patience, shows her the photograph she's asking about, notices when she's distressed before she says a word, and learns what actually settles this particular person. Each evening it writes a clinician-ready record for her family and her GP. Sarah fills in one intake form and drops in photographs. A LangChain vision chain reads each one and writes what it means in Eleanor's life, the caption for the screen, the sentence Margaret should say when showing it, and the muddled ways her mother might ask for it. Voyage embeds all of it. It lands in MongoDB Atlas. Her carers log the day — medication, meals, how she seemed. That's what lets Margaret answer "have I had my tablets?" truthfully instead of guessing. Then Eleanor's day begins. The microphone opens and never closes; she never presses anything. When she speaks: Fireworks classifies intent and agitation, Atlas retrieves the right memory by hybrid vector-plus-keyword search with reranking, Claude decides what someone this fragile should hear, and ElevenLabs says it — slowing and steadying as she becomes more distressed. Her camera is read every few seconds for restlessness, so distress with no words still reaches the agent. Every exchange writes back. Whichever comfort strategy was on screen when her mood recovered gets credit, and its effectiveness is rewritten. The garden works. The boat used to and no longer does. It stops reaching for the boat. MongoDB Atlas — the spine Twelve collections on one cluster, margaret_hackathon. No external vector store, no cache. patients · media · image_context · memories · care_notes · routine_checks · interactions · mood_events · comfort_strategies · alerts · reports · traces Atlas Vector Search — 1024-dim cosine over memories and photo contexts, filtered by patient Atlas Search — typo-tolerant keyword, because a speech recogniser mishears "Whitstable" and "donepezil" Hybrid retrieval — both fused with reciprocal rank fusion, weighted toward exact matches, then reranked through MongoDB's reranking endpoint Change Streams — every insert pushes live to the caregiver dashboard over SSE, with resume-token recovery Aggregation $facet — the clinician report is computed in the database: repetition rate, the questions she asked and how often, agitation episodes, the hourly sundowning curve and medication adherence, in one round trip Voyage AI — voyage-4-large through ai.mongodb.com, MongoDB's own endpoint. Embeds every memory and photo context. Its reranker reads each candidate against the question, which is what makes "did the girl come by?" land on Sarah rather than the garden. Claude via OpenRouter — decides every patient-facing sentence under absolute rules: never state a death, never correct her, never say she already asked. A new question goes to Sonnet 5; a repeat goes to Haiku 4.5 and is not re-reasoned — the already-approved answer is re-voiced. Same warmth, a fifth of the tokens. We don't regenerate, we remember. Fireworks AI — the reflexes. A vision model reads her camera every few seconds for restlessness — braced posture, a hand to the head — and a fast model reads every utterance for agitation. Both feed the comfort ladder and the voice. ElevenLabs — the voice, and it changes with her state. Settled: speed 0.88. Distressed: 0.73 with higher stability. The same sentence renders 25% longer when she is frightened, exactly as a carer slows down. LangChain + LangSmith — the intake pipeline that turns a photograph into memory, every call traced.

The problem LLM agents fail in customer service for a reason that has nothing to do with model quality: local context. The policies, eligibility rules, and edge cases that define how one lender actually operates are unique to that lender, constantly changing, and scattered across systems. RAG addresses this by injecting snippets at inference time but it reads and forgets. Nothing about a conversation makes the next conversation better. The gap is sharpest in what we'd call complex consumer service: not "where's my order," but "change my payment date." That requires customer context, product state, an action proposal, validation against real eligibility rules, and a write to a system of record. Get it wrong and you've moved money incorrectly. Most agents avoid this by staying read-only. What we built An agent that pursues a goal across engagements. A case carries a plan, an outcome, and its own follow-up schedule — the agent decides when to re-engage and records why. Its context lives in MongoDB Atlas in three slices: authored policy, per-customer episodic history, and lessons the agent derived from its own outcomes. The loop is the point. When a case closes — won or lost — every lesson that influenced it has its win/loss record updated, and a reflection pass reads the transcript and writes new lessons back. Retrieval then prefers what actually works and drops what doesn't. In our demo, a losing case moves a lesson from 9/11 to 9/12 and it ranks lower on the next call. Write actions are gated by deterministic code, not the prompt. The model proposes; validation decides. A hallucinated action fails eligibility instead of moving money. MongoDB Retrieval is one aggregation: autoEmbed generates embeddings in-database, $rankFusion fuses vector search with BM25, and $rerank applies a cross-encoder — with zero outbound model calls. Atlas also holds case state, turn-level checkpoints for crash recovery, and the outcome log the learning loop runs on. Voice is ElevenLabs Conversational AI, with our orchestrator as its custom LLM — so phone and chat run identical logic.
Link to live project - https://track-record-plum.vercel.app/ Track Record is an agent that helps a founder answer "Will Big Tech eat your startup?" with persistent context of the competition, most relevant to their company. Give it a startup idea and it names the incumbents already in your category, their most likely move (ignore, partner, change terms, bundle it in, acquire, build it themselves), a percentage for each and a window for you to build: the months you have before the earliest likely from a competitor move lands. Easter egg - Used an example of another project that was created in this hackathon as a startup to demo the app :-)

ContextGuard is a self-healing memory layer for AI agents that prevents conflicting, stale, or untrusted persistent context from influencing future actions. In our demo, a live Fireworks-powered deployment agent retrieves long-term policy memory from MongoDB before calling tools. When new context conflicts with a verified policy, ContextGuard quarantines it, explains the conflict through a Memory Court-style adjudication, and blocks the agent’s action until the trusted memory state is explicitly verified. Once the policy is updated, the same agent can safely act on the new memory, with all state persisting across sessions.

Guinea Pig turns dead clickstreams into a persistent population of grounded “ghosts” in MongoDB—so the product doesn’t wake up cold. PMs interview churned shoppers (cited to real events), accumulate friction patterns across restarts, and then spin a sandbox of a proposed UI change and drive hundreds of those same ghosts through it with computer control—mouse, clicks, forms—conditioned on each user’s memory, to see who would stay and who would still churn before any living customer tries it.

The problem: LLMs produce a lot of output. People read maybe a third of it, skimming, skipping, jumping to code blocks. The model then says "as I mentioned above" about a paragraph your eyes never touched. Every follow-up compounds the error, because the model's model of your knowledge is assumed, never measured. The idea: eye tracking measures which parts of an answer a person actually read. We store that per fact, per user, in MongoDB Atlas, and condition every later response on it. Within a conversation, the model re-explains what you skipped and does not repeat what you read. Across conversations, facts you have read are embedded and matched with Atlas Vector Search, so a brand-new conversation about a related topic skips what you already know. That is the "No Cold Start" requirement satisfied literally: the agent knows what you know, because it watched you read. The model emits numbered atomic facts, not prose. Each fact is one DOM span, so "which fact did they read" is just elementFromPoint plus a dwell timer. No fragile alignment between rendered text and extracted claims. Dwell past a length-scaled threshold marks the fact read, persists it to the facts collection, and the next system prompt carries READ and NOT-READ lists. Cross-conversation recall runs as a $vectorSearch aggregation over embeddings of verifiably read facts, pre-filtered by user and read state. The heatmap (a live dithered attention overlay, also persisted to Atlas as fixations) is shown for debugging purposes. The product is a different next response; read-state exists only to change the prompt. Works with a Tobii Eye Tracker 5 or any webcam (WebGazer fallback, Tobii wins when both stream). Stack: MongoDB Atlas Vector Search on the hackathon sandbox, OpenRouter (Claude), Node/Express, WebSocket, Electron, WebGazer. The Tobii driver predates the event, everything in this repo was written during it.
I used to be a logistics coordinator. I had these crates and I need to ship them from military base to military base. That required asking 7 third-party logistics providers (3PLs) for quotes on the cost to truck those crates around the country. Some 3PLs operate in a specific region of the United States, some others are less expensive with loads that way less, some are more expensive the farther the distance is, and so on. Muster takes all of this into consideration when awarding a 3PL with a shipment.
Rebuttal is memory for insurance appeals. Insurers deny ~20% of in-network claims; under 1% are appealed; 43–67% of appeals win. Patients start from zero. We load 42,710 California DMHC IMR decisions into MongoDB Atlas. Paste a denial letter; Vector Search finds similar cases; aggregation ranks arguments by overturn rate. Record a loss and that strategy is excluded, so the next draft is different. Same letter, different ranking, because Atlas stored the outcome. Public DMHC data, unmodified. Not legal advice.

We built a **self-improving AI agent that learns how to make machine-learning models better through experience**. Using PCB defect classification as a real-world testbed, the agent trains a model, diagnoses its failures, autonomously chooses an experiment, measures the result, and learns from it. Most importantly, it remembers successful **and failed** experiments, so the next run doesn't start from zero. The problem we're solving is bigger than PCB inspection. Today, when an ML model underperforms, an ML engineer manually studies metrics and confusion matrices, forms a hypothesis, changes the training strategy, retrains, and repeats. AI agents can automate pieces of this process, but they often have the same problem: **every new run is a cold start**. They repeat experiments and rediscover lessons they have already learned. Our agent turns those experiments into persistent experience. The system follows an autonomous loop: **Train → Evaluate → Retrieve Experience → Diagnose → Experiment → Retrain → Critique → Learn → Repeat** We demonstrate this using real PCB manufacturing defect data. The agent begins with a deliberately weak classifier and evaluates its performance across different defect classes. If, for example, open-circuit recall is significantly worse than other classes, the agent doesn't blindly tune random hyperparameters. It analyzes the evidence, determines a likely cause, searches its previous experience for similar failures, forms a hypothesis, and chooses a specific intervention such as weighted sampling, stronger augmentation, different image resolution, class weighting, learning-rate changes, or a different model strategy. After retraining, the agent compares the before-and-after results. Successful experiments become reusable knowledge, but failed experiments are equally important: the agent remembers strategies that did not work so it can avoid wasting time repeating them. ## MongoDB Atlas — Persistent Experience and Agent State MongoDB is the core memory layer of the system, not simply a database for application logs. Every experiment produces structured experience containing the model configuration, failure pattern, diagnosis, intervention, metrics before and after the intervention, whether the hypothesis was correct, and the lesson learned. For example, the agent can remember: > **Failure:** Very low recall on a minority PCB defect class > **Intervention:** Weighted sampling > **Outcome:** +9% Macro F1 and +21% minority-class recall > **Lesson:** Weighted sampling was effective for this type of class-imbalance failure. When the agent later encounters a semantically similar failure, **MongoDB Atlas Vector Search** retrieves relevant previous experiences. Those memories are provided to the AI scientist before it decides what experiment to run. This means MongoDB directly changes the agent's next action. MongoDB also persists the **LangGraph checkpoints and workflow state**, including the current experiment, iteration, best configuration, metrics, hypotheses, and experiment history. An interrupted autonomous run can therefore resume without losing what it was doing. This gives us two forms of persistence: **Short-term state:** Where is the agent in the current experiment? **Long-term memory:** What has the agent learned across previous experiments and runs? ## Voyage AI — Semantic Memory We use **Voyage AI embeddings** to turn the agent's learned experiences into semantic representations stored with its MongoDB memories. This allows the agent to retrieve experiences based on meaning rather than exact wording. For example: > “Tiny open circuits have poor recall.” can retrieve an older lesson involving: > “Small localized defects improved after increasing image resolution.” Even though the descriptions are different, their underlying failure patterns are related. Voyage therefore makes the agent's accumulated experience searchable through MongoDB Atlas Vector Search. ## Fireworks AI — The ML Scientist Fireworks is the primary reasoning engine behind the autonomous ML scientist. After every evaluation, Fireworks receives the classifier's current configuration, overall metrics, per-class precision/recall/F1, confusion matrix, dataset characteristics, recent experiment history, and relevant memories retrieved from MongoDB. It then performs: **Observation → Diagnosis → Hypothesis → Experiment** For example: > **Observation:** Open-circuit recall is dramatically lower than other classes. > **Diagnosis:** The minority class may not be receiving enough exposure during training. > **Hypothesis:** Weighted sampling should improve minority-class learning. > **Experiment:** Enable weighted sampling. Fireworks does not generate arbitrary training code. It chooses from a controlled set of executable ML experiments, which keeps the autonomous system reliable. After the experiment finishes, Fireworks acts as a **critic**. It compares the metrics before and after the intervention, determines whether the hypothesis was supported, and distills the result into a reusable lesson. That lesson is then written back to MongoDB, closing the learning loop. ## OpenRouter — Independent Second Opinion We use OpenRouter as an independent evaluator rather than duplicating Fireworks. Most experiments are handled entirely by the primary Fireworks scientist. OpenRouter is brought in when the decision deserves additional scrutiny—for example, when Fireworks has low confidence, several experiments have failed consecutively, or the agent proposes a significant change in strategy or model family. OpenRouter evaluates the current evidence and Fireworks' proposed intervention and can agree, disagree, or suggest reconsideration. This creates a lightweight **scientist + reviewer** architecture while avoiding unnecessary model calls on every iteration. ## LangGraph — Autonomous Orchestration LangGraph coordinates the complete self-improvement process: **Train → Evaluate → Retrieve Memory → Diagnose → Propose → Judge → Retrain → Critique → Store Lesson → Repeat** Rather than scripting a fixed sequence of hyperparameter changes, LangGraph allows the next step to depend on the agent's current state and previous outcomes. MongoDB backs the LangGraph checkpoints, connecting orchestration and persistence into one continuous agent workflow. ## Kaggle PCB Dataset — Real Learning Environment We use a real PCB defect dataset from Kaggle as the environment in which the agent learns. The dataset provides real PCB images and labeled manufacturing defects. We split the data into training, validation, and untouched test sets. The agent repeatedly trains on the training data and uses validation performance to decide what to improve. The final model is evaluated against the untouched test set so that improvements represent genuine generalization rather than optimization against the test data. The PCB classifier itself is intentionally not the main innovation. It gives the agent an objective environment where every hypothesis can be tested and scored. ## What We Prove Our key metric isn't simply: **Classifier F1: 61% → 86%.** That only proves that the classifier improved. The more important comparison is: **Cold-start agent:** requires multiple experiments to reach the target. **Experienced agent:** retrieves relevant MongoDB memories, avoids previously unsuccessful approaches, and reaches the same target in fewer experiments. That demonstrates that the **agent itself has improved**. The classifier learns from data. **Our agent learns from experience.** Every experiment becomes a memory, every failure becomes a lesson, and every new run starts with everything the agent learned before.

PlotTwist is a persistent-context matchmaking agent for conferences. Existing networking tools do SIMILARITY matching - they introduce you to people like you. PlotTwist does COMPLEMENTARY matching: it pairs your needs with someone else's offers, ideally reciprocally (you need X, they offer X; they need Y, you offer Y). The core insight is that what you need changes. You walk in thinking "I need enterprise sales help," meet someone, and realize "my real blocker is SOC 2." PlotTwist treats that evolution as the product: every meeting produces a learning written to MongoDB, which re-prioritizes your needs and re-embeds your profile — so the same person gets a different next recommendation. The model doesn't get smarter; its persistent context does. You speak for 10 seconds (ElevenLabs Scribe transcribes). Fireworks extracts your needs/offers/goals and embeds them into MongoDB Atlas. Your current top need is run through Atlas Vector Search to retrieve candidates, and OpenRouter (GPT-4o, with a Fireworks fallback) reasons over them to pick the best reciprocal match, which ElevenLabs announces by name. The whole loop is a LangGraph state machine with human-in-the-loop gating, checkpointed to MongoDB so each attendee's thread persists across sessions. A live force-directed graph shows every introduction the agent has created in the room.

RadixMind is an agent whose memory lives in MongoDB — and decides what deserves the model's cache. Most agent-memory systems store everything and retrieve by similarity. RadixMind scores every piece of tool output for viability before it touches a prompt: high-value evidence (runbooks, configs) is pinned into a byte-stable prompt prefix that the model provider's cache can reuse; low-value noise (log floods, status pings) is archived to MongoDB and kept out of the model's way. The admission decision for every chunk — its score, its verdict, its reuse history — lives as a MongoDB document. MongoDB isn't just storage here; it's the control plane deciding what's allowed into the model's expensive working memory. We tested it on a simulated production incident: an agent investigating checkout-service 502 errors across 8 turns of real, seeded application data (logs, runbooks, deploy history) stored in Atlas. Run naively, the agent drags 90,750 tokens through the model with a 0.9% cache hit rate. Run through RadixMind, it reaches the identical correct diagnosis using 6,434 tokens — a 14× reduction — with a 75.9% cache hit rate and lower average latency. Every number is a live Fireworks AI API response, read back from Atlas. Built with MongoDB Atlas (collections, indexes, aggregation pipelines as the ledger and metrics store), a FastAPI gateway, and Fireworks AI for inference with automatic prefix caching. No fine-tuning, no RL — the learning lives in database rows, not gradients.
A second brain for social media: MongoDB remembers who you are, who you met, and which intros worked — so your agent never starts cold. AgentCircle is a social network with a durable memory layer, not a chat wrapper. Every member’s agent is a second brain grounded in MongoDB Atlas: resumes, posts, photos, conversations, and introduction outcomes live as documents, chunks, and vectors in one database. That memory is why the product works. On day one the agent already knows you from what you uploaded. As you post, connect, and interview others, Atlas accumulates context that chat logs throw away. Hybrid $vectorSearch + $search finds people with cited evidence. A shared trust graph stores which intros landed — so one member’s outcome changes who everyone else’s agent recommends next. The feed is how context enters the system. The second brain is the product: grounded answers, declines when evidence is missing, and introductions you still approve yourself. Stack: React, FastAPI, LangChain/LangGraph, Fireworks, MongoDB Atlas (documents + reranking + Vector Search).

ConquerContext is a long-horizon insurance claim agent that does not cold-start. When a claim opens, it drives FNOL to settlement over days or weeks — SMS, voice, and email — then sleeps until a real reply arrives instead of polling. MongoDB holds three kinds of context on purpose: live state (kv), crash-safe lineage (checkpoints), and shared agent_memory recalled on the next claim and injected into the ElevenLabs voice agent so it already knows what worked. Compact working memory dies with the claim; shared memory does not. Nothing pays out until a human adjuster signs off.

Engram is a trust layer for AI agent memory. Agents that write down what they learn have no way to tell a hard-won fact from a lie. What's worse is that a wrong memory stays put: the agent does work based on it and writes new memories that inherit the error, so deleting the original leaves the damage behind. Engram gives every memory a confidence score it has to earn from real outcomes, and records which memories were relied on when each new one was written, generating a family tree. Memories that help get promoted; memories cited when the agent gets something wrong get slashed and benched. When a memory is caught being wrong, Engram follows the family tree and knocks down everything that memory taught, automatically. The demo does this live against a real movie database: I plant a single false claim that low-vote ratings should be filtered out: something nothing in the data can disprove. It passes the first task, because that filter doesn't change a ranking, so the false memory is rewarded and spawns a new memory beneath it. Then it fails a counting task (9 instead of 16), gets caught, and one database query traces and cleans up its offspring with no human involvement between the poisoning and the recovery. I built the whole engine from scratch: the scoring rules, the search that ranks by relevance times confidence, the family-tree trace, the agent loop, and a live terminal view, backed by 131 tests including 20 that run against the real cluster.
Agentic system that continuously watches PRs in a repo and adjusts review scrutiny based on their credibility. The credibility is assessed by previous performance and kept inside the DB where the agent is able to search for previous behavior in natural language.

Ledger is an evidence-driven memory system for AI agents. Every memory system in the category (Mem0, Letta, Zep, Cognee) is write-only and unfalsifiable. A lesson gets stored and then served forever, with no measurement of whether it ever helped, and the only eviction signal anyone has is recency. That matters because self-generated lessons are often bad, since the model writing the rule is the model that just failed. Ledger treats every memory as a hypothesis instead. When a run fails, a reflection step proposes a candidate lesson. Inserting it fires a MongoDB change stream that starts an experiment: vector search finds about 20 past situations the lesson claims to cover, each one is replayed from its LangGraph checkpoint with the lesson injected, and each is re-graded by the same execution-accuracy grader. The result is a lift score with a confidence interval and a verdict of promoted, probation, rejected, or harmful. The metric that makes this different is the regression rate, measured on tasks that already passed. It catches the lesson that fixes one thing and breaks three others. Nothing else in the category can produce it, because nothing else measures a memory after storing it. Proven memories then graduate out of the prompt and into behaviour. Task-space is split into regions clustered by query shape, and each region tracks a success posterior per model. A cheap model shadows the expensive one on every task, graded but never served. When the evidence shows it is not worse by more than a margin we are willing to trade, the region switches over and gets permanently cheaper. Regions that no off-the-shelf model can solve accumulate failures until another change stream fires a Fireworks LoRA, trained on that region's own successful runs and registered as a new option for that region. A memory therefore moves from text, to a routing decision, to weights, with vector search deciding which weights answer next. MongoDB Atlas acts as the control plane rather than the storage layer, since three separate vector search paths each change what the system does, and change streams trigger the work throughout, so writes drive everything and nothing polls. It runs on Spider text-to-SQL because execution gives exact ground truth, and is verified by 98 fully offline tests whose headline case is a negative control: a deliberately harmful lesson must be caught with a non-zero regression rate against the real grader.
Short description: A prediction game where confidence is billed. Stake tokens on everyday calls; Atlas Vector Search finds the same wrong take you made before and Claude roasts you for it. Long description: Called It — everyone claims they called it; nobody keeps the receipts. Players stake simulated tokens on everyday predictions with a confidence score. Confidence sets the price: a 95% call costs 860 and returns 905, a 30% call costs 340 and returns 1,133. Overconfidence is punished by arithmetic, not by a rule. ▎ ▎ Every prediction is embedded with Voyage voyage-3 and stored in MongoDB Atlas. Pull the receipts runs Atlas Vector Search over one player's own history and hands the nearest past calls to Claude, which roasts them citing specific dates and confidences — surfacing that one player made the same wrong call four times in different words. Before a new call is billed, the same retrieval powers a risk desk that quotes the player's real accuracy on that topic back at them. p.s. in the long term we plan on implementing friends voices to troll your other friends when playing using eleven labs.

Check it out: https://uberprompt.getclera.com/ Persistent context for your system and team. We've been battling this in production for the last 3 months, after hundreds of thousand agent invocations, time to do something about it! Your agents share rules, policies, and evals. überPrompt keeps them all in sync when anything changes (and improves them live). Don't let your context get stale -> don't let your system get stale.

**ACME Travel** is an enterprise AI travel agent. You describe a trip in plain language; it checks company policy and traveler preferences, ranks flights and hotels, and books on a corporate test card. Everything sits in MongoDB Atlas: orgs, policies, bookings, expenses, and a cached inventory so searches stay fast between sessions. The hackathon angle is durable context. Profiles and preference signals live in MongoDB and affect the next recommendation. After a trip, SMS or web feedback updates that profile and can open a policy suggestion for a manager to edit and apply. Five demo orgs each keep their own rules, approvals, and suggestion history, so the agent improves from stored memory instead of starting over every chat.

MongoMatch is a live conference matchmaking graph powered entirely by MongoDB Atlas that eliminates the AI "cold start" problem. By combining Atlas Vector Search, $graphLookup multi-hop traversal, and compound persistent memory in a single document model, it connects builders with the exact peers in the room who can unblock them in real time.

Retrace is an AI agent that remembers how your team actually solved dependency upgrade problems — not just what the fix was, but every path that failed along the way. Most coding agents hallucinate on library upgrades because they only know general training data, not your codebase's real history. When your app has gone from v1 to v10 over several years, the context behind each past upgrade — the failed attempts, the coupled dependencies, the specific reasoning — gets lost in closed PRs and forgotten Slack threads. Retrace fixes this by treating each upgrade as an episodic memory: every attempt (successful or failed), every error message, and every dependency that had to move together gets stored in MongoDB Atlas. When a developer hits a new upgrade error, the agent searches this history semantically, retrieves the full story — not just the final fix — and explains it with the reasoning intact. For example: upgrading Tailwind alone breaks Shadcn's color system. Downgrading Shadcn alone doesn't fix it either. Only upgrading both together works. Retrace remembers that coupling, so the next developer doesn't have to rediscover it from scratch.
Claude Token Counter Real-Time AI Cost & Environmental Impact Tracking The Problem: Developers using Claude Code have no visibility into prompt costs before sending them, leading to budget surprises and wasted resources. Our Solution: A Claude Code plugin that pauses every prompt to show: - 💰 Real-time cost — Exact USD (cache-aware pricing) - 📝 ML-predicted reply length — Trained on 12k+ real API calls - 💧 Environmental impact — Water usage of inference - 📊 Session totals — Cumulative spending tracker - 🔄 Prompt optimization — AI suggests rewrites with token % savings (LangGraph + OpenRouter + MongoDB) Tech Stack: - Node.js plugin (UserPromptSubmit hook) - Python ML service (scikit-learn, 1.8 MB model) - 224 automated tests, production-ready - Graceful degradation (never blocks users) Impact: Developers make informed decisions, prevent budget overruns, quantify environmental cost of AI.
Atlas Lifecycle is an AI agent that catches expiring software before it becomes a vulnerability. Ask it about an outdated dependency and it retrieves the answer, checks for known security risks, and reads it back aloud. What sets it apart is memory. Most hackathon agents lose context on restart, but Atlas Lifecycle remembers both its knowledge base and every conversation, so a follow-up question still resolves correctly after closing the browser. One persistent source of truth, useful to DevOps, Security, Cloud Infra, and Frontend teams.
Real Estate Analyst workbench for analyzing EV charging sites. When the analyst changes a finacial number or a new engineering or permitting data point comes in the agents work together over shared memory to updated the analysis with every number and its dependencies traced through the graph. The system then helps identify how to allocate money effectively into projects.

Decision Tripwire stops AI agents when the assumptions behind their decisions become false. For example, if an agent starts a deployment because traffic is low, Decision Tripwire automatically pauses it when traffic spikes. MongoDB Atlas stores the decision, its assumptions, and every intervention, so the protection continues even after a restart. Fireworks evaluates new evidence, while a deterministic policy engine makes the final safety decision.
Most people don't abandon their goals out of laziness. They abandon them because nothing adapts. The plan that made sense in week one goes stale by week three, the app keeps sending the same generic reminder, and nobody notices that you always fail on Sundays. Habit apps track behavior; they don't understand it. Ridge is what it looks like when an accountability app gets a memory. It's an agentic layer we built for REACH, our shipped social-accountability app, running entirely outside it over its public API. It learns from a signal almost no product has: on REACH, your peers verify photographic proof of every rep, so Ridge's memory is built from what real humans confirmed you did, not what you told a chatbot. What that unlocks: plans that reshape around your actual failure modes (it discovered, unprompted, that our demo user only breaks streaks on week edges, and rebuilt her week around it). Encouragement that cites your real history: "three weeks in, don't stop now," never "you missed a post." Recovery challenges created for you at the moment your pattern warrants one. A companion that gets sharper the longer you use it. How the stack delivers that: MongoDB Atlas holds every layer of the memory. Observations carry mandatory evidence pointers (evidence-free memories are refused at the storage layer), beliefs are distilled from them, and per-user narrative stories plus the full conversation transcript live alongside. LangGraph runs the deliberation loop (observe, remember, decide, act) and checkpoints every step to Atlas. Kill the agent mid-thought, rerun it, and it resumes where it stopped. No cold start. A purpose-built MCP server exposes REACH's real API as tools, so the agent acts through the same doors as the mobile app, under the same rules. OpenRouter routes each job to the right model. ElevenLabs gives the agent a voice and a personality: Ridge. One rule, enforced in code: Ridge never verifies proof and never posts it. Humans judge; the agent plans. Built for REACH, and modular by construction: the entire app-specific surface is one MCP server and one seam file.

Rehash remembers why past project ideas were rejected, not just that they were. When someone pitches a new idea, it pulls out the idea's core mechanic and likely flaw, turns that into an embedding, and checks it against every past rejection stored in MongoDB Atlas Vector Search. It matches on why something failed, not on how it's worded, so two ideas that look nothing alike can still get flagged for sharing the same underlying flaw. When it finds a match, it speaks the result back using two different ElevenLabs voices: one reacts live, the other reads the original rejection reason word for word, so the system's own memory is talking back to you. A human always makes the final call. Built for teams and hackathon judges who keep re-debating ideas that already failed for reasons nobody remembered.

Terra — named for the earth goddess who knows every fault line in the ground — is an autonomous AI pentester with persistent memory. Where security tools only flag issues, Terra acts like an attacker: it chains exploits across laptops, source control, cloud IAM, and databases, re-proves each of the dozens of vulnerabilities it surfaces with real evidence (validated, not detected), and pairs every finding with a concrete patch or PR. Its edge is memory — every exploit attempt, target fingerprint, and outcome lives in MongoDB Atlas, and before each move Terra vector-searches its past engagements to re-order its attack plan, so what it learned at one organization sharpens what it does at the next. The result is mechanical: our model stalls cold and scores 30/100, but with a single memory retrieved from a different engagement it crosses the wall and finishes at 100/100 — same model, same target, the only variable is MongoDB. Built on Atlas Vector Search for memory and LangGraph for crash-safe, resumable runs.
We made scartissue, it is a traffic infrastructure, that remembers its crashes, where every accident becomes a memory. We analyzised it then put it in the database, we used voyage AI, to do vector search later to find the crashes as we store the data in the time series database and compare the incident, if it matches we have a agent control the traffic based on that data! we forgot to mention voyage and vector search in the video.
A coding agent's conversation runs 100,000 tokens. Every single message, the model re-reads all of it before answering. Providers cache that work and charge up to 90% less the second time - but only for the model that did the reading. Model routers pick a different model for each message to save money. The moment one switches, the new model has read nothing, and you pay full price for all 100,000 tokens again. A single switch can cost more than the turn it was trying to optimize. No router on the market tracks this, so they all treat switching as free. Inertia makes switching cost something. A cache ledger in MongoDB Atlas records which model has read how much of each conversation, so the router can ask the real question: is a smarter model worth re-reading 100,000 tokens, given how many messages are left? Atlas Vector Search remembers how similar questions actually turned out, so it learns which ones truly need the expensive model instead of guessing.
IMMUNE is a 5-call memory adapter (remember, derive, recall, guard, challenge) that immunizes AI agents against persistent memory poisoning. When an agent reads untrusted input, IMMUNE records every stored claim with its full provenance tree in MongoDB Atlas. When a belief is refuted, IMMUNE executes a single $graphLookup aggregation pipeline to recursively traverse the provenance graph downward, revoking all derived contaminated beliefs in one transaction and reversing flagged side-effectful actions—leaving clean, unrelated memory branches completely untouched. By enforcing trust floors directly at the database level via vector search pre-filtering (source_trust >= 0.5), poisoned context remains invisible to future agent sessions across cold restarts without relying on prompt rules or blocklists.

JobSwipe is a swipe based job search app for new grads and early career engineers who are tired of scrolling through hundreds of near identical listings they aren't even eligible for. You swipe right on roles you want and left on the ones you don't, and every card leads with the two things that actually make or break it for a new grad: visa sponsorship and security clearance. No more wasting an application on a job you were never going to be allowed to take. Behind the scenes, real postings get pulled live from public job board APIs (Greenhouse, Lever, Ashby), cleaned up, and embedded with Voyage AI. Every job and every swipe gets stored in MongoDB Atlas, and Atlas Vector Search ranks new listings against the average of everything you've liked so far, so the deck keeps reshaping around your taste as you swipe. There's also a Matches view that collects your saved roles and shows which required skills your resume already covers and which ones you still need to pick up before applying. We went with MongoDB because job postings are all over the place structurally. Some list salary, some don't, and sponsorship and clearance show up in totally different wording everywhere, so we let the ingest step normalize all of that into clean fields.
Evergreen is an encyclopaedia generated and actively maintained by AI agents. It's more complete, current, and correctable than Wikipedia. Evergreen deploys a team of agents to research real sources, create articles for different reading styles, and improve them as new information arrives. Readers can suggest changes, which Evergreen checks against evidence before updating an article. MongoDB stores its research, past work, and edit history, so every new page visit builds on what Evergreen already knows.
MemGate is a drop-in, OpenAI-compatible proxy that gives any model on OpenRouter (~500 models) governed, persistent, self-tuning memory of users — while guaranteeing that no model ever sees who those users are, that memory is disclosed least-privilege per request intent, and that any user (or any facet of a user) can be provably erased via crypto-shredding with a certificate.

A hardware startup builds forty units of a prototype. Optical inspection machines start around fifty thousand dollars and have to be programmed per board, which never pays back at that volume, and there is no defect data to train on because the board did not exist last month. So small teams inspect by eye, or skip it. Quality control is one of the reasons hardware iterates slower than software. Ambit makes inspection cost fourteen photos and three seconds. You photograph known-good units and it fits a specialist for that exact part from normal examples only, with zero defect labels. Show it any part afterwards and it embeds the frame, vector searches a registry of specialists, and loads the one trained on the closest imagery. When nothing in the registry has competence over what it is looking at, it refuses instead of passing an uninspected board. That refusal is the hard part. Five specialists are registered and a sixth board is withheld. Top-1 routing is 125/125. Per-model coverage gates refuse 25/25 out-of-registry frames, while every global threshold we tested refuses 0/25, because the withheld board scores 0.9399, above all of them. It genuinely is a PCB and the registry is full of PCBs. Only a per-model gate declines it. The cost, stated plainly, is that 6 of 125 in-registry frames are refused too. MongoDB Atlas is the system, not a store beside it. Routing is $vectorSearch over model_router_idx, one 512-d OpenCLIP centroid per model. Weights are PatchCore coresets in GridFS at 7.55 MB each. PatchCore has no gradient-trained parameters: training stores a coreset of patch features and inference is nearest-neighbour against it, so a model is remembered examples of normal. Vector search picks which memory applies, nearest-neighbour searches inside it. Findings carry their own embedding indexed for episodic recall, LangGraph checkpoints live in the same cluster so a kill -9 mid-training resumes instead of relearning, and fleet health runs as aggregation pipelines where the data already is. LangGraph runs the agent graph with node-level checkpointing and cold-start as a branch. Fireworks runs Qwen3-VL to describe the difference between the defect crop and the same region of the golden reference. OpenRouter does structured-output adjudication inside the 0.04 ambiguous band with model fallbacks. ElevenLabs is a voice analyst over five MongoDB-backed tools including explain_refusal and recall_similar, and every claim it makes appears as a tool call first. The demo is live. A phone streams a viewfinder to the projected view, holds still, and uploads one full-resolution frame. An Arduino board is cold-started on site from fourteen photos, comes back nominal untouched, and is flagged with a localized heatmap and a written description once a pin is bent. Built during the event: the Arduino specialist and local ingest path, Fireworks VLM defect narration, OpenRouter adjudication, ElevenLabs voice, project grouping in the registry, the capture and analysis interface, and the trends view. The README separates this from pre-existing work in full.
Shibboleth: prove you're you from any device with your voice and your memories; uncloneable, because a deepfake can fake your voice but not what you told your agent yesterday.
The zeroth layer under every agent: you. me0 is an open-source Agent Plugin that gives any agent harness — Claude Code, Codex, pi, Hermes, OpenClaw — one portable, MongoDB-backed personal context graph and agent session memory. Switch harnesses freely; your context comes with you.
Thread is an ambient memory layer for real-world relationships. It listens to your conversations, distills them into selective structured memories (what people work on, what they need, what they can offer, what you promised them), and connects information across people. When someone reappears, Thread briefs you on who they are, what you owe them, and what's changed since you last spoke. Why it's useful: You meet dozens of people at events and almost everything evaporates — especially the invisible connections between people, since no one holds both conversations in memory. Thread turns every conversation into persistent context that changes what happens next: promises resurface, needs find offers automatically, and returning people get smarter follow-ups. Tech stack: • Next.js 16 (App Router, TypeScript, React) · Tailwind CSS · Framer Motion • MongoDB Atlas — the persistent social world model (people, memories, open loops, connections), with Atlas Vector Search for semantic need↔offer matching • ElevenLabs — speech-to-text for live conversations, TTS for whispered briefings and spoken memory Q&A • OpenRouter — LLM memory extraction (Zod-validated structured output) and connection reasoning Future Plans: We aim to integrate this into smart glasses so that it can update the profiles of people you meet in real time.
A memory backend for hardware listening and whispering devices that whisper relevant facts about friends and acquaintances into our ears after hearing their name. It hears a name, and it says out loud the one or two things that actually matter right now about that person — not everything ever recorded about them, just what's relevant today. A health scare from three days ago gets mentioned. A coffee order from three months ago doesn't. Every fact has its own urgency and its own decay rate, so what is spoken changes over time on its own. Additionally, every choice the system makes is saved as a record, right next to the facts. Not just "here's what I said," but "here's what I said, here's what I left out, and here's the reasoning." You can click "why did you say that" and see the actual scored comparison between the facts it picked and the ones it didn't. That record can even point back to earlier records so that you can trace a whole chain of decisions, not just the last one. Built on MongoDB Atlas, an aggregation pipeline scores every fact by recency, urgency, and confidence in one query, and a graph lookup walks the decision chain backward. ElevenLabs turns the chosen facts into speech, and the tone of the voice shifts with how urgent the facts are — flatter for routine facts, with more emotional weight for the ones that matter. This way, I or any other specialist can query the log for validation and to ensure compliance and relevance.
With a growing aging population, a huge amount of healthcare access depends on something that sounds basic: whether elderly patients can successfully navigate, understand, and complete forms. Our idea is Care Loop, an AI completion agent — a screen-aware, comprehension-aware layer between patients and the healthcare system. Most AI assistants optimize for answering questions. We optimize for successful completion. Care Loop does not stop after giving the right answer. It stays with the patient until they understand what they are reading, complete the required information, and confidently sign. But the bigger vision goes beyond accessibility. Care Loop creates a verifiable record of what was shown, explained, understood, and approved. We see this as the first step toward a liability layer for high-stakes AI — moving beyond simply observing AI behavior to creating evidence that can stand up to compliance review or legal scrutiny.
Blackboards are back! The default way agents share context is a markdown scratchpad. Everyone appends, everyone re-reads the whole thing, and nothing is queryable. Chunking sucks. Unbounded replaces it with MongoDB. Twenty small-model agents attack the same SWE-bench bug with one document store between them. No schema, no coordination protocol — just insert, find, sample, inspect. Three things a markdown file cannot do, that we observe: - They query instead of scanning. Agents issue more reads than writes (1.2:1), retrieving specific prior findings by filter rather than re-reading an ever-growing document. Context cost stays flat as the store grows; a markdown board's cost grows linearly with every agent's contribution. - A schema emerges. Agents independently converged on analysis, solution, task, affected_with a visible synonym scatter (findings/current_findings, test_regence was invented, not templated. You can only see that because the store reports it; grep on a markdown file cannot. - Structure survives concurrency. Produce documents with stable identity and atomic updates. Twenty produce a merge conflict. A control arm gives each agent its identical prompt, so the only variable is whether memory is shared. A lives, the operation timeline, and the emergent schema straight from Atlas Server is still live at http://10.50.180.160:3000/ Data is light as I ran out of time for testing, but in a 3 instance test on swe-bench we created 40 documents and there were 9 submissions (not yet scored).

Evolution simulator. Utilizes AI to create a custom environment and 16 different subjects with different characteristics. They breed and evolve, and we get information on changes in characteristics as time moves on, based on their environment. If their heat resistance starts increasing due to a hot climate + less heat resistant variants dying off, we can see it in the summary graph or AI generated summary.

Most assistants react to the latest signal. Continuum learns from longitudinal signals, maintains an evolving understanding, and knows when to ask, inform, or stay quiet. It connects live activity, calendar, and environmental data with confirmed goals and relevant past episodes. Built on MongoDB Atlas, Continuum combines durable memory, semantic retrieval, and auditable decisions to make every interaction more relevant and trustworthy.

Outbound agents cold-start on every interaction. They open the same way, ask what they already asked, and forget the call the moment it ends. dee4dee is an outbound agent that doesn't. It finds one specific person from their public work, emails them, and calls them — and every conversation it has becomes memory in MongoDB Atlas, so the next one starts further along. Research writes what it finds as episodes: one fact, one Voyage voyage-4 vector, stored in Atlas. Mid-call, while a human is waiting to be answered, "which project of mine did you find?" is a $vectorSearch over that person's episodes, reranked with rerank-2.5 and spoken back in about 500ms. "Send me a visual of how I'd use this" generates a picture from their own record and emails it as an attachment before they hang up. "Thursday at two" creates a real Google Calendar event with a Meet link. Then the loop closes: the meeting transcript is written back as an episode, and the next conversation retrieves it by meaning. Asked "what is his team stuck on" — a question only that call could answer — the call transcript comes back ranked first at 0.86, against 0.33 for everything known beforehand. Persons and episodes join on one key, so a call, an email reply and a meeting land on a single record instead of forking into three. The embedding cache is its own collection rather than process state, so a warm-up in one process is a cache hit in another and survives restarts. Every arrow in the architecture diagram is a real read or write against a live Atlas M10.

Natural language automated testing. Managing e2e test is extremely painful, having a part-developer who can contribute to project is wasted on writing e2e. Our project will save the resource and effort spent on human understanding of how the app should work to jest/playwrite. Our agent will automate and mange this

AI Coworker (AccountabilityBot) It doesn't tell you how to work. It learns how you work. A voice agent that learns from how the user works and gets and its speaking tone and content out of MongoDB instead of having them coded in. Two users get audibly different behaviour from one agent, one system prompt, one code path — because the difference lives in the database, not in an if statement. Built at the MongoDB Persistent Context Sprint (Pier 48, SF).

Agents don't only hallucinate — they confidently act on facts that stopped being true. Ask an on-call agent who's on call for payments and it pages Marcus Webb, who rotated off six days ago. The stale memory wasn't a marginal retrieval hit: it outscored every fresh result, 0.94 to 0.92. Relevance and truth are different axes, and vector search only measures one of them. StaleGate is a runtime freshness gate between an agent and its memory. Freshness is enforced inside the MongoDB Atlas Vector Search pre-filter, so expired memories never enter the ANN candidate set. Every read returns FRESH, STALE or UNKNOWN. When the expected cost of acting on a stale fact exceeds the cost of checking, StaleGate re-queries an authoritative record, rewrites the memory, and learns that source's real lifetime from the outcome — four sources sharing one 3-day prior learned lifetimes ranging from 18 hours to 36 days. MongoDB Atlas stores four kinds of state: agent memories with embeddings and provenance, learned per-source decay statistics, the records revalidation queries, and LangGraph checkpoints. Fireworks supplies both embeddings and chat inference. LangGraph orchestrates plan → recall → revalidate → act, checkpointed to MongoDB so an interrupted run resumes instead of starting cold. Across 16 changed-fact queries, wrong answers fell from 100% to 0%, with 94% corrected and 100% recall on immutable controls. The sources of truth are mocked MongoDB collections rather than live PagerDuty/GitHub integrations, and the evaluation set is synthetic.

Alert Steward is a clinical decision support agent that turns alert override feedback into persistent organizational memory. When a clinician dismisses a medication alert, Alert Steward stores the interaction in MongoDB and uses Claude to update an evolving hypothesis about why that alert keeps getting overridden. Every new interaction is read against existing memory — so the agent learns, revises its reasoning, and eventually surfaces its conclusions before the next clinician decides. Built on MongoDB Atlas (persistent agent memory), FastAPI, and Claude Haiku.
Agents keep executing plans built on facts that are no longer true. RealityPatch versions every fact in MongoDB, tracks which plan steps depend on which fact version, and when reality is corrected it reverses the affected actions, rolls back, and re-plans — propagated live to running agents via change streams.
Recall — The system remembers, then recalls a decision when the world changes: it's a Memory-Grounded Financial Decision Pipeline. Most news-to-signal tools are stateless. They read one article and answer: "how bullish is this?" I answer a different question: what changed in the world since the last decision, and what does that invalidate? It’s a diff to the world model.

Upload a confusing screenshot — a bill, an error message, a legal document, an insurance letter, a form, a website — and Plainly explains it in plain English, with clear action items, urgent red flags called out, and a searchable history that remembers how your uploads connect to each other.

Successor When a 25-year process engineer retires off an SMT line, the manuals stay but the knowledge leaves. Manuals describe what a machine is supposed to do. The veteran knows what it actually does when it fails at 2am. So the plant keeps calling him for months. Successor catches that knowledge on the way out the door. A voice agent runs the exit interview as an interrogator, not a recorder. It tracks which of 60 fault classes it has covered, picks the biggest hole, and refuses an answer until it has the decision criteria and the failure branch. That knowledge becomes an agent bound to a machine. An operator asks in plain language. MongoDB Atlas fuses keyword search for part numbers with vector search for symptoms, then ranks by what actually worked. The agent returns one step, in the engineer's voice, with attribution. "Did that clear it?" No branches to her fallback. The repair runs as checkpointed state, so a dead tablet resumes at the exact step. When it can't answer, it logs the gap and opens the next interview with: "Someone on Reflow 2 asked about this Tuesday and I couldn't answer. Walk me through it." Failure to answer becomes the next question asked. Built with: MongoDB Atlas (hybrid search, reranking, change streams, checkpoints), ElevenLabs, LangGraph, Fireworks.

One in seven New York EMS calls turns out to be something other than what it was dispatched as. We measured it: across ~30 million (29,978,154) NYC EMS incidents, ~900,000 (845,887) of the ~6 million (5,653,498) calls since 2023 closed under a different call type than they were dispatched under. Every one is a labeled case where the first read was wrong, and the reasoning behind the correction was never recorded anywhere. BlackBox is a voice-native flight recorder for EMS crews. During a live call it listens to the medic over an earpiece, captures the decisions they make and the reasons they give, writes those to MongoDB Atlas as embedded documents, drafts the patient care report at transfer of care, and retrieves that reasoning on the next similar call to brief the next crew. Every incident-response tool on the market stores what happened. None store what the human decided and why. That reasoning currently dies in radio chatter, and voice is the only capture medium that gets it, because during an active call nobody is typing. MongoDB Atlas is not our database. It is the entire runtime. We took the single-platform constraint literally: memory, state, context, retrieval and transport all run on one cluster. No Redis, no Kafka, no Pinecone, no broker anywhere in this stack. Retrieval is one aggregation pipeline, not three round trips. A single $vectorSearch fans out across decisions, postmortems and runbooks with $unionWith, capturing {$meta: "vectorSearchScore"} inside each sub-pipeline, then fuses the three result sets with reciprocal rank fusion. Fusing by rank rather than raw score is deliberate: cosine distributions are not comparable across collections, and sorting a union by score systematically buries decision memory under the clinical corpus, which is exactly the signal this product exists to surface. A fourth collection, remediations, is queried separately so that failure memory can actively exclude routes that cost time on previous runs. MongoDB is also our event bus. Any process, a route handler, the worker, a one-off script, inserts into an events collection; the dashboard opens a change stream on it and pushes each insert to the browser over SSE. Replay after a reload is a plain find(), multi-process coordination needs no IPC, and every number a judge sees on screen arrived over an Atlas change stream. Sequence numbers come from a single server-side $inc inside findOneAndUpdate, verified gap-free and duplicate-free under 200 concurrent emits. Change streams are the trigger too. A worker watches incidents for live inserts and fires the graph, persisting its resume token after each handled event, so an incident inserted while the worker is down is still processed on restart. We drilled that rather than assuming it. And Atlas is the crash-safety story. LangGraph checkpoints to MongoDBSaver, so the graph parks at interrupt() on the drug-dose readback gate with its durable state in Atlas. Mid-demo we kill the process in front of the judges and restart it, and the agent resumes the call exactly where it left off. That is impossible with an in-memory checkpointer, and it is the part of the demo hardest to fake. Underneath: four Atlas Vector Search indexes at 1024 dimensions across decisions, remediations, runbooks and postmortems, a server-side JSON Schema validator that rejects any decision document missing a rationale, a TTL index so rehearsal runs self-clean, and a projection that quarantines ground truth so no retrieval path or graph node can read the answers. VoyageAI does the semantic work. voyage-3-large embeds our decision corpus and 183 chunks of NASEMSO national clinical guidelines, ingested live. Semantic retrieval is what makes the second demo call land: it is dispatched as general illness rather than unconscious, so a string lookup finds nothing and vector search visibly earns its place. Voyage being MongoDB-owned meant one fewer provider and one consistent 1024-dimension contract across all four indexes. ElevenLabs is why this is a product and not a dashboard. Seven server tools hit our own route handlers and write to Atlas mid-call, so the agent does real work instead of reading pre-generated text. Barge-in is on, the tone shifts from calm during the brief to clipped when confirming anything irreversible, and every drug and dose passes an aviation-style readback the agent must speak verbatim before anything is written. That readback doubles as the LangGraph human-in-the-loop gate, which is a genuinely satisfying place for two sponsor technologies to meet. LangChain and LangGraph gave us durable interrupts that survive a process kill, turning a clinical safety requirement into the best fifteen seconds of the demo. The agent never proposes a treatment, dose or diagnosis. It recalls what happened last time, reads back what the medic said, and quotes retrieved guidance with attribution. The human owns every clinical judgment.

Rehearsal is a voice-first platform for practicing high-stakes conversations against a behavioral model of the actual person you expect to face. Users describe a counterpart, such as an investor, executive, client, or landlord, and Rehearsal models how that person reacts, speaks, negotiates, and handles pressure. Each behavioral trait is stored in MongoDB Atlas with supporting evidence, a confidence score, and version history. When the model behaves incorrectly, users can correct it mid-conversation, for example, “She gets quieter when annoyed.” Rehearsal supersedes the old trait, records the correction as new evidence, and immediately changes the persona’s behavior. After each voice session, Rehearsal identifies important moments, such as when the user lost the thread or conceded too early, and provides an actionable debrief. Unlike a generic role-play chatbot, Rehearsal becomes more accurate over time by remembering when it was wrong.

Math Market is a second-hand market for AI compute: a live exchange where unused inference capacity becomes discoverable, priced, and routed in real time. Instead of tying intelligence to one model provider, Math Market separates memory from supply. MongoDB preserves the buyer’s context and the market’s learned history, while replaceable suppliers compete to answer each request. Every job creates price signals, performance evidence, and routing memory, turning fragmented AI capacity into a measurable market. The vision is an open intelligence economy: buyers get resilient, lower-cost AI service; suppliers monetize idle compute; and the market learns which capacity is worth trusting.
We are solving the problem where we have a sdk which evaluates the agent and get its telementry and sends it to backend whih then saves it in mongo db and then it evaluates user feedback with evals to tell the developer clearly about the problems and features and help it recreate silent issues which users are facing but he cant.
Verdict is an AI onboarding agent for Belgian accounting firms that replaces the two-week manual client-vetting process with two live scores — Confidence (how complete the file is) and Risk (how suspicious it looks) — treating an empty file as unknown, not safe. Every document lands as its own schema-validated record in MongoDB Atlas (document model + $jsonSchema), and each write instantly re-scores the client while keeping a full history of how the risk evolved. One aggregation pipeline does the reasoning: it cross-checks documents for contradictions and uses $graphLookup to walk ownership structures two hops out into hidden shared-account clusters. The payoff is memory — when a colleague files a single verdict into one collection, Voyage AI embeddings and hybrid Atlas Vector Search + Atlas Search (blended with $rankFusion) surface that judgment as precedent and re-score the client in seconds, with no retraining and no restart. Because state, documents, graph, and vectors all live in MongoDB, there's no bolted-on vector store or connector: one platform remembers every decision so the firm never re-derives it.
RingOps - Voice-Driven Incident Response with Persistent Memory When production breaks, RingOps calls the on-call engineer instead of paging them to a dashboard. Using MongoDB Atlas Vector Search, it retrieves similar past incidents — root cause, fix, outcome - and walks through the situation in a live voice conversation powered by ElevenLabs. On command, it takes real action: rolling back a bad deploy or proposing a code fix, reflected live in a demo app. Every resolved incident is written back to memory, so the next call starts smarter than the last.
**Should I?** is an AI decision companion for the everyday choices people get stuck on. A user asks a question like “Should I take this job?”, “Should I text them?” or “Should I move to another city?” Instead of simply answering yes or no, the app asks why, understands the user’s situation, and explores both paths. It shows what is likely to happen if they say yes versus no, helping them think through consequences, tradeoffs, and what matters most before making the final decision.

Morgan AI — Autonomous Financial Analyst **One-liner:** A voice-first financial analyst you can interrupt. Ask a question out loud, and Morgan reads the actual SEC filings, runs the actual valuation math in an isolated sandbox, and gives you a recommendation it can defend — with sources, confidence, and memory of what it told you last time. ## Inspiration Every real investment desk runs the same loop: read the filing, run the numbers, form a view, defend it in conversation. Most "AI financial assistants" skip straight to the last step — they generate plausible-sounding analysis without ever touching a real filing or running real math, and they forget the conversation the moment it ends. We wanted to build the version that doesn't cut corners: an analyst that shows its work, says "I don't have that data" instead of guessing, and remembers what it told you. ## What it does You ask Morgan a question by voice — "Should I invest in Nvidia?" or "What happens to margin if logistics costs rise 8 percent?" — and it: 1. **Plans** the request into concrete research tasks. 2. **Pulls real SEC 10-K/10-Q filings** from EDGAR and a **live market snapshot** (price, market cap, news) from Polygon. 3. **Runs the actual valuation math — a DCF — inside an isolated Daytona sandbox**, never in-process, because model-written or model-influenced numeric code shouldn't run on the host. 4. **Synthesizes a memo**: bull case, bear case, key risks, a BUY/HOLD/SELL recommendation, and a confidence score — with every figure traced back to the source that produced it, and every missing data point named instead of papered over. 5. **Speaks the answer back**, and you can interrupt it mid-sentence to redirect — it drops what it was saying and listens. 6. **Remembers.** Ask a follow-up in the same conversation ("how does it compare to AMD?") and short-term memory resolves the reference. Come back in a new session next week and ask "what did you tell me about Nvidia?" — long-term semantic memory, scoped per user, still knows. Every step of every request is traced end-to-end — every agent call, every tool call, every sandbox execution — so nothing about the answer is a black box. ## How we built it - **Fireworks AI** — the reasoning core. A planner model breaks the request into tasks; a memo-synthesis model turns filings + market data + valuation output into the final investment memo, with an explicit sourcing rule baked into its prompt: never imply data you weren't given. - **Daytona** — every numeric computation (DCF, margin scenarios, Monte Carlo simulation) runs inside an ephemeral sandbox, not the host process. The "what if costs rise 8%" scenario tool is a live example: the model never computes the answer itself, it ships real Python into Daytona and reports back what actually ran. - **ElevenLabs Conversational AI** — owns voice I/O directly (not the usual reverse setup). The agent calls our backend as an authenticated webhook tool mid-conversation, gets back a structured result, and speaks it — with barge-in support so you can cut it off. - **Braintrust** — full observability. Every pipeline run is a trace with a child span per agent step (planner, research, market, valuation, memo, memory recall) and a confidence score attached at the end; failures are logged as error spans instead of disappearing. - **MongoDB Atlas Vector Search + Voyage (`voyage-finance-2`)** — the memory layer. Short-term conversation turns (TTL-expired after 24h) resolve pronouns and references; long-term memory stores *distilled* analyses — never raw transcripts — embedded with a finance-domain model and searched with `user_id` as a hard security filter on the vector index itself, not a post-filter. - **SEC EDGAR, Polygon, Alpha Vantage** — real filings, real prices, real fundamentals. No mock data in the pipeline. - **FastAPI + Next.js** — backend and dashboard, both calling the same shared pipeline whether the request comes from voice or the web. ## Challenges we ran into Getting the ElevenLabs architecture right required a pivot: our first instinct was to have our backend drive ElevenLabs, but the correct pattern is the reverse — the voice agent owns the conversation and calls us as a tool. Rebuilding around that made the whole system simpler, not harder. Memory was the other hard part: a single log of everything fails both jobs a conversational memory needs to do — recent turns need to be complete and ordered (you can't resolve "compare it to AMD" from a semantic match), while old analyses need to be searchable and distilled (nobody wants March's literal transcript). Splitting it into two collections on two different lifetimes — one that expires, one that doesn't — is what made both work. ## Accomplishments we're proud of Nothing in the pipeline is faked to make the demo look better. If a data source is unavailable, the memo says so and confidence drops — it doesn't quietly smooth over the gap. Every claim about what the system did (which filing it read, what math it ran, what it remembered) is independently verifiable in a Braintrust trace, not just asserted. ## What's next Voice-biometric authentication and a live internal GL/budget integration for variance commentary are the two pieces still gated on a vendor decision we haven't made. A Risk Agent and a News Agent are scaffolded but not yet wired to a live data source. And we'd like to close the loop on memory by mining full post-call transcripts for durable facts (stated risk tolerance, sector interests), not just the outcome of each analysis.

Ember is a persistent automotive failure intelligence system designed to catch dangerous vehicle problems earlier. Major recalls often begin as scattered owner complaints—one driver reports stalling, another reports sudden power loss, another describes the same underlying failure in completely different words—and those warning signs can be hard to connect before enough damage is done. Ember searches a vehicle’s complaints and recalls, uses AI to turn messy reports into a structured failure fingerprint, then compares that pattern against previous vehicle investigations stored in MongoDB. Every search becomes long-term memory, including the component, symptoms, severity, complaint pattern, and whether similar cases were eventually recalled. That means Ember does not start from zero on each search: if a new vehicle begins showing the same kind of pattern that appeared before past recalls, Ember can recognize that precedent and surface it immediately. In short, Ember turns the history of automotive failures into reusable memory for spotting the next one earlier.
ConvTree brings Git-style version control to LLM conversations, letting users create checkpoints and branch into new discussion paths without losing context. Built on MongoDB Atlas, it stores chats as a conversation tree and reconstructs any branch using $graphLookup. When histories grow large, Atlas Vector Search helps users find the right checkpoint using natural language. Tech stack: Next.js, TypeScript, ReactFlow, MongoDB Atlas, OpenAI/Ollama.

Pantree - Your pantry insider, is an AI kitchen assistant that tracks inventory, remembers user preferences and time-bound dietary constraints, and takes real action: automatically adding used and expiring items to a grocery list and logging unused, expired items to a throwaway log. Rather than static rules, every suggestion is produced by an adaptive scoring function that weighs expiry urgency, preference match, and past waste/usage patterns computed live from MongoDB Atlas, and every response includes a plain-language explanation of exactly which signals drove it - transparency by design, not a bolted-on dashboard. Built to directly address a persistent, expensive problem (U.S. households are responsible for 40-50% of the country's food waste), it demonstrates genuine agent memory and action-taking rather than a chatbot wrapped around a database, with cross-session context and broader vector-search-driven ranking scoped as clear next steps beyond today's build.
The problem: testing a website means a developer writes code for every single click, every form field, every check. It's slow, it's tedious, and only engineers can do it. So most teams under-test — and things break in front of customers. The idea: just say what you want tested, in ordinary English. "Log in with the test account, add a laptop to the cart, and check the total says ₹85,000." Saambu reads that, works out the actual steps, and builds a real automated test from it. Then it opens a real browser, does each step for you, and takes a photo at every stage. When something breaks, you don't get a wall of red error text. You see the exact step it died on and a screenshot of the screen at that moment — so you can see the button had moved, or the price was wrong. Why it matters: a QA person or product manager can now create a working test by typing a sentence. No code. And every test is saved, so you can re-run the whole set any time you ship a change and know within a minute if you broke something. The short version: it turns "here's what should happen" into a test that actually checks it — and shows you pictures when it doesn't.

Murmur — anonymous voice threads that remember you. Callers phone in and speak. Their number is HMAC'd with a server-side salt and discarded; their voice is replaced before storage — ElevenLabs speech-to-speech, with a dependency-free pitch/formant shifter as fallback. Nothing identifying is kept. Yet the line still knows them. Every murmur is transcribed, embedded with Voyage AI, and folded into a per-caller interest vector via an exponential moving average. On the next call, that vector is the query for MongoDB Atlas Vector Search — so the feed is personal from call two onward, with no login, no profile, and no stored identifier. Anonymized audio lives in GridFS, so documents, vectors, and blobs all sit in Atlas. Persistent context without persistent identity.

Cognitive Twin is an innovative application built for the MongoDB hackathon that models and queries personal thought patterns. It leverages MongoDB Atlas for secure data storage and advanced vector search capabilities. The system integrates Google Gemini AI to process embeddings and generate intelligent responses based on user data. A lightweight, real-time Node.js backend powers the seamless interaction between the user interface and the database. The project serves as an interactive AI-driven digital twin designed to understand and reflect user insights dynamically.
