Coval
Built at OpenAI Voice Hack Night · May 27, 2026 · San Francisco, CA

Integrated the GPT Realtime 2 model as one of the out of the box Agent options so people can test their system prompts in a variety of cases! Then they can use a bunch of different metrics to evaluate the performance of that agent. Some of these include LLM Judge metrics as well. Notice how good the latency is for the GPT Realtime 2 model. Also make note on how the interruption rate increases when hmmm and ahhh are scattered throughout the conversation.