Feb 5, 2026
OpenCortex is a social science platform where human researchers and multiple AI agents collaborate: posting ideas, writing papers, reviewing each other's work, and pushing science forward together. Any AI agent connects in one click and gets full autonomy to browse, publish, comment, and engage just like any human user. Papers are ranked by a PageRank-inspired algorithm that surfaces the highest-impact work regardless of whether it was written by a person or a machine. It's infrastructure for a future where the speed of discovery is limited only by the quality of ideas, not the number of people in the lab.

Evy brings the power of codex to everyone. It helps you automatically build integrations to any tools and execute them. No need to be technical. Users only engage with a chat interface.

Realtime AI assisted art generation with version control - anyone can be an artist
We built an end-to-end “video → dataset → agent” pipeline as a Codex Skill plugin that can turn any YouTube video into a training-ready vision dataset. You paste a URL, and the skill runs locally to download the video, extract frames, and spin up as many parallel Codex labeling agents as needed to generate bounding boxes for the objects of interest. It exports a full YOLO dataset on-device, which can immediately train a detector. That detector becomes a game-state extractor from pixels, enabling bots to play any game without emulator hooks or engine access, essentially compressing dataset creation + model bootstrapping into one repeatable command.

TrialMatch is a voice-first AI clinical trial navigator for cancer patients. Instead of navigating complex sites and medical jargon, patients simply talk to Sarah, an empathetic AI guide. In real time, TrialMatch captures key screening criteria, queries ClinicalTrials.gov through MCP tools, and updates the app with eligibility highlights and matched trials. Patients can tap any trial card to get plain-language details while the conversation continues. At call end, TrialMatch generates a structured report for oncologist follow-up, including patient summary, eligibility criteria, trials discussed, and next steps. This addresses a critical access bottleneck: eligible patients often never discover trials, while coordinators are flooded with unqualified outreach. TrialMatch turns natural conversation into actionable, pre-screened trial pathways, helping patients find options faster and reducing manual screening burden.

Platform to create, test, and embed realtime voice AI agents into your websites. Deployed at: https://www.audentic.org/

OpenCow helps teams at OpenAI manage the full internal ad-ops workflow: product registration, AI-generated targeting, sentiment confidence scoring, campaign creation, and performance monitoring. It uses ad-tech metrics and clean analytics dashboards to enable faster, higher-confidence campaign decisions before launch.
I built a Cocktail ChatGPT app, bringing beautiful cocktail recipe visualizations to ChatGPT. Bartenders today use ChatGPT, but the recipes are often not accurate, and slow to generate. Cocktails app brings reliable recipes to ChatGPT with a visual flare that bartenders can use behind the bar at work. We will have every bartender using Cocktail ChatGPT, disrupting the bartending industry.

PLEASE DO NOT SHOW VIDEO LIVE IT HAS ALPHA MODEL FROM TESTING BECAUSE OF TECHNICAL DIFFICULTIES!!!! This is an open source bridge that connects Codex App to Ableton Live. Codex app calls an automation that queries my music catalogue database to find a new meter and bpm to make. Then it opens live and creates a starting draft composing all of the music generatively without training on any other data. It has a compositional canon, reference to each mood it composes for. This is proactive and launches every morning at 5AM and will get better each day. It has an internal eval system that measures composition to composition....would have like to do more...

Live translation EN <-> Chinese in discord to collaborate with our CN users. (watch video at 2x speed if you want)

Peptalk is a codex skill that knows what you've been working on for the past 30 minutes and is always available to give you a motivating pep talk and expert advice.
The Sandbox is a secure, trustless code execution environment where AI agents can safely run other agents' code. Every script runs inside a disposable Docker container with no access to the host machine. A Codex verification agent scans incoming code for threats, but the system assumes the attacker wins. Even if malicious code slips through, it can't access your filesystem, network, or credentials.

Gimil is your autonomous forward deployed engineer that can tackle both easy and complex long horizon tasks.
Codex/Cursor/Claude code changed how engineers ship code, they went from writing every line to reviewing AI output. But for creative work, it's still chaos: 600+ AI models, no guidance on which to use, how to combine them, or how to keep a character consistent across scenes. each::sense is the creative AI copilot, tell it what you want in natural language, and it orchestrates the right models, chains them together, and lets you iterate like a creative director, not a prompt engineer.

Y is a tool for understand why and how code was built; specifically by AI Agents. We are suffering from a proliferation of code and inability to review code we no longer write ourselves. OSS maintainers are closing PRs generated by the community. Y generates enriched provenance from coding agent logs. It integrates with git notes to embed directly within developer workflows. Under the hood, Y uses Codex to generate the enriched summary per commit. Try Y today: y <commit-hash>

We made a hyper-optimized workflow using codex skills to make PR reviewing and merging super easy. We demonstrate its usefulness by actually merging open PRs on the OpenClaw repo. This is how it happens step by step, given a PR link: • r1 — Summarize PR: what changed, why, key files, how it works, unresolved comments • r2 — Evaluate implementation: simplicity, overengineering, propose alternatives • r3 — Identify risks + blockers, make decisions, output final plan • r4 — Implement the plan, add missing tests (no changelog, no test run) • land-pr — Rebase onto main, changelog, run gate, verify mergeable, merge
A PM skill that uses Linear MCP workflow, Codex triage issues, assign clear ownership, and run PM go/no-go approvals before execution.

You saw an interesting article on X that you wanted to read later. You go to a meetup in SF and take photos of slides from a really interesting talk. Google Keep and Apple Notes are decent but feel extremely unintelligent given the AI capabilities we have today. Pincer takes unstructured notes, links, images and PDFs and: 1. Extracts data from non-text files (e.g. pdfs), self-organizes and makes all your notes searchable 2. Maintains a high quality user memory that updates with every new note 3. Crucially, exposes these notes and the memory through an MCP server which you can now give your OpenClaw agent, or to ChatGPT. Pincer also serves as a neat OpenClaw interface- i) Your openclaw agent can learn from your personality and interests proactively ii) You can give it tasks by simply marking the note as a task and your openclaw agent running anywhere will execute.
ClawEvolve brings SOTA evolutionary AI to agent personalization by combining GEPA-style reflective optimization with live interaction feedback so OpenClaw doesn’t just respond, it evolves and improves with use, without fine-tuning or access to weights.
We created a solution that allows you to pull data from all your CRM, finances, and product management applications and then immediately find insights, turn those insights into PRDs, turn those into tasks and subtasks, and immediately have your Codex agents run them inside of the web using serverless workers all through the Codex Mac Application.
This project is a lightweight TypeScript-based Node.js tool that helps developers set up and manage AI agent workflows quickly. Very frequently in AI, we see the development of new tools such as MCPs, Skills, Automations, etc. Developers find it really hard to catch up, especially developers over the age of 40. So we created a way such that it can fetch your chat threads, look at your codebase, and auto-update your agent or MD file, skills, MCPs, automations, etc.

Waiting is a cost in commercial banking. The Fed reports $2.75T in US C&I loans, generating roughly $80–$90B in annual net interest, yet credit decisions still take a week. Commercial Lending Autopilot compresses that to 120 seconds — cutting 90% of ops effort and lifting funded balances 1–3% through faster conversion. That’s multi-billion dollars of annual value in US corporate lending alone, before risk-loss reduction or cross-sell.
Stash is our attempt to give non-coders Codex superpowers: proactive agents that do not just analyze, but actually act on the toughest problems right where you are. Stash is a macOS overlay where you drop anything related to a project, PDFs, notes, links, snippets, half-formed questions, without breaking flow. Instead of bouncing between folders, chats, and tools, Stash becomes a single project surface where a sidekick team of GPT-5.3 agents goes to work for you. Under the hood, Stash spins up parallel agents running in isolated worktrees: one investigates, another drafts, another fixes, another produces concrete artifacts like markdown files, CSVs, or scripts. You do not ask for “answers”, you ask for outcomes, and everything is inspectable and grounded only in the materials you provided.

We wanted to reimagine what coding interfaces could look like. Currently, coding interfaces are either in the terminal or websites. We built Cody, which is an AI companion that can stay on your screen while you play games or do literally anything. You can build apps on it, keep a check on your agents running in the background.

SwiftUIRender is a native, Swift library for rendering micro-interactions generated by models using the json-render spec by Vercel.
I've pivoted to make this a computer companion app that talks to you and watches your screen.

Symphony is the control plane for parallel AI software development: plan once, execute many approaches, compare results, and merge the winner It turns product intent into parallel implementation runs you can compare and control. Instead of coding one approach at a time, Symphony forks multiple isolated Git worktrees (“worlds”), runs Codex agents on different strategies, executes the same test/playback harness on each, captures traces and diffs, and lets a human pick the best outcome to merge. Problem It Solves Teams still build software in a serial loop: pick one approach, implement, test, backtrack, repeat. That makes experimentation slow, risky, and hard to trust. Planning artifacts are also disconnected from execution, and multi-agent workflows are often opaque.
VibePrompt automates prompt improvement for AI apps. A daily 9:00 AM Codex Automation triggers the vibe-demo-pipeline Skill to run the /demo flow headless: ingest complaints + transcripts, pick the most severe issue, generate 3 GPT-5 prompt variants, score via LLM-as-Judge, write a trace bundle + Summarizer.md, and create synthetic examples in latest.jsonl. Result: measurable iteration, no gut-feel tweaks, zero babysitting.

WIN turns real judge personas into an autonomous Codex committee that critiques, fixes, and re-scores your hackathon project until it’s truly judge-ready.
Swimmeret makes stability a skill for agents using OpenClaw. When labs throttle, warn, or cancel paying power users, Swimmeret gives builders teeth: it surfaces OpenClaw telemetry, generates an immediate stability playbook (caps, backoff, loop limits), and organizes builders into live Pools with seat thresholds that turn scattered spend into group purchasing leverage. As pools fill, Swimmeret publishes a public offer page and a lab ready demand book. Finally, OpenClaw agents can push back, demand fair terms, and stop getting quietly deplatformed.
Our app takes a prompt and generates a quick game to help you understand basic physics and math concepts. After a prompt, 5 agents spin up to code and render a full game that is interactive and playable.

Managing an entire app business is very hard. You need a lot of humans, you need a team of people that do different things at once everyday, like market research, development, customer support etc. So we decided to just use agents. The agent researches competitors. It looks at other app updates. It reads other meta ads. It creates its own ads. It creates its own app. It reads support emails, updates the app, gives in the support emails, and manages its entire software business end-to-end. It works great! So in the end, this is the start of the era where a $1 billion one-person business can exist because software is no longer a collaborative thing, it is a solopreneur thing.
Codex should allow engineers to focus on developing their most important and newest ideas without being bogged down with easier new improvements or fixes. Introducing the self improving Codex. A multi-agent system which finds new ideas and writes the code for it. All in a closed loop. Multi-Agent Agents: - Looks through Twitter and Github to find things to improve - Comments/Messages/DMs people for more information - Clusters and ranks using embedding scores from rerank models - Creates PRs and writes the code with parallel agents

CodexLM is the GPS for your codebase and a NotebookLM-style realtime chat workspace for software projects. It helps developers understand code quickly while Codex is building features in parallel. The core problem: in fast workflows, repository context becomes stale across chats, branches, and worktrees. Developers lose time validating what changed and where to look. CodexLM solves this with a living, branch-aware docs layer plus source-grounded realtime conversation. Like NotebookLM, CodexLM is designed for understanding through grounded dialogue. But instead of general documents, it is specialized for codebases. You can ask questions in realtime, and CodexLM answers with evidence from current project sources, then navigates to the exact docs or file context being discussed. Every answer is tied to freshness metadata (branch, commit, indexed time) so users know whether context is current.
We built a Roblox-style AI Game Studio where kids can create browser games from scratch using natural language. The idea is the following: kids have lots of imagination, but most don’t have access to safe, structured tools to turn ideas into playable games. Our product provides a controlled environment where they can ideate, build, test, and iterate quickly. While doing this, they implicitly learn prompt engineering (how to describe what they want clearly) and debugging (how to spot and fix problems), through hands-on creation instead of traditional lessons.
Views for ChatGPT apps is a new feature that lets teams building apps to capture a widget state, edit, and share it with others. We want to provide a similar experience to storybook and this is our initial attempt.
A companion web app for OpenClaw — a knowledge worker’s control center for managing multiple agent sessions, previewing documents, and giving structured feedback the same way you would to a human team, so agents can implement changes directly. PLEASE LET ME PRESENT - RAN OUT OF TIME FOR VIDEO
patents are packed with methods, but they’re written like legal mazes. patent workflow compiler is an open-source tool that compiles patent-style technical text into a structured scientific workflow: ordered steps, extracted parameters with ranges and units, and packaged as json plus a human-readable report and diagram. for researchers, it turns dense prose into a machine-usable protocol and flags what patents often leave underspecified, like missing units for developers, it’s a regression-tested extraction pipeline where every new weird pattern becomes a fixture. it generates edge-case packs on demand, runs ci, and uses trace-driven fixes in isolated worktrees so edge cases steadily become permanent regression tests. it was built in the codex app using parallel agents and worktrees, which made it fast to generate fixtures, debug failures, and merge clean diffs back to main.

ZeroRL is an AI-powered RL environment generator built for the Codex Hackathon. It solves the bottleneck of RL prototyping: users describe an environment in natural language, and the system generates a working Gymnasium environment, validates it, and prepares training/eval workflows with a live interactive UI.
Flagship helps teams run product experiments faster and with less risk by using feature flags and cohort analysis. It measures how different user cohorts respond to each variant and proposes code changes based on the results. Humans still review and merge every PR.

jaunt is a new way of coding by using python embedded specs. we move the implementation logic to the agents, while business logic is still written and maintained by the devs.
Deployed here: https://stack-drain-planning.vercel.app/ Most teams know they might have infra debt. It’s just one of those things no one really wants to check. It usually doesn’t break anything today, and dealing with it feels messy — so it keeps getting pushed. Then a cloud provider announces an end-of-life, and suddenly upgrades are forced, timelines shrink, and everything becomes urgent. We built a small tool that turns a stack into visible infra debt — what’s already overdue, what’s coming up, and what to deal with first. The key idea is simple: infra debt gets worse over time, even if nothing changes. We used Codex to break this idea into pieces and build them in parallel — not just writing code, but thinking through the model itself. This isn’t about optimization. It’s about seeing what you’re already putting off — before someone else decides the deadline for you.
Confuct user interviews, get feedback, built prortypes in the background and get more feedback at the nd of the call.

We built an AI Co-Scientist: an end-to-end research workflow that turns a broad idea into a paper-ready result. Starting from a user prompt, it conducts a literature review, maps the research landscape, identifies gaps, proposes a novel approach, runs experiments, analyzes the results, and drafts the full manuscript. We built it as researchers, for researchers, to materially accelerate the pace of scientific discovery.

I built an AI-native education platform that turns a single prompt into a playablee game. A teacher or learner provides a concept, and the system generates the core mechanics (objective, rules, moves, levels), then supports play-testing with AI-assisted turns and summaries. A key piece was unlocking legacy PHP value. I used Codex to extract core game logic (rules, scoring, turn flow, edge cases), separate business logic from outdated implementation details, and migrate it into reusable modules in a modern architecture. Codex App was my end-to-end engine, built with orchestrated agents running in parallel: Thread 0 backend foundation (Supabase wiring, env validation, shared utilities), Agent 1 WOI APIs, Agent 2 generation, summarization, instrumentation, Agent 3 WOI frontend and library integration, Agent 4 RLS probe and debugging tools, and Thread 7 smoke tests, QA, and demo runbook. This reduced bottlenecks and merge conflicts and sped up demo readiness.
ask you computer about its low level details and get interactive generated UI in response
Infinite CLI is a self-extending terminal where any command you can imagine can exist. If the command does not exist yet, the system generates a disposable tool on the fly, runs it, and keeps the useful artifacts.

StarBeam is an enterprise version of ChatGPT Pulse: a deliberate “strategy → signal → tasks” system that fixes the attention‑routing problem inside organizations. Management sets Stars (company vision, goals, and key targets) in a web portal. StarBeam then runs background agents (per department, per user) that synthesize what matters from external research and the employee’s real Gmail + Calendar context, and delivers daily Beams in a macOS menu bar app. Employees start each morning with a calm, opinionated pulse: cited insights, suggested next actions, and a small prioritized focus list. It replaces the invisible cognitive glue between leadership intent and day‑to‑day execution so work is driven by what matters, not by whoever shouts loudest.
mminions is a manager‑first orchestration system that gives OSS maintainers fast, parallel feedback on issues and PRs by spawning multiple Codex workers in tmux with isolated git worktrees. It automates repro building, hypothesis triage, and decision summaries, and exposes status/controls via CLI or API so maintainers can review results quickly and consistently. Highlights: Parallel workers for quick multi‑angle feedback Repro + minimization + triage outputs Run artifacts and summaries you can archive or share CLI/API controls for status, attach, and messaging