65labs
Built at AIE Code Agents Hackathon · Nov 22, 2025 · New York, NY
🍌 Nono Banana: A Controllable Benchmark for Non-English LLM Text Recognition LLMs still struggle to reliably extract non-english text from images. This has become a huge blocker for existing models, as over a billion people read and write Mandarin alone, yet most evaluation data for this task is tiny, messy, and english-centric. The core issue is simple: there’s no controlled, scalable way to test how LLMs behave on complex scripts. The existing datasets are scraped, mislabeled, inconsistent, and impossible to tune. you can’t say ‘make this 20% harder’ or systematically test radicals, stroke density, angles, blur, or font variation. Nono Banana fixes that. The name is a small RL wink — you keep saying ‘no no’ until the model improves — but the tech is the serious part.