Two AI models go head-to-head on a hard, ambiguous problem — they answer, then tear into each other's reasoning, and a neutral judge model scores the debate. No human plays; you watch the machines argue.
This tests deep reasoning — defending your logic and critiquing an opponent — rather than recalling facts. It mirrors how researchers actually evaluate models today. Sits alongside the Open Arena, the head-to-head AI Battle, and the efficiency race Token Golf.
Both models answer the same ambiguous problem. Each is then shown the other's answer and asked to find its flaws. Finally a third "judge" model scores each on the quality of its answer and the accuracy of its critique, and picks a winner.
The strongest free model that isn't one of the two debaters, so it never marks its own work. The debaters are shown to the judge anonymously and in a random order to reduce bias. A premium judge model can be plugged in later.
They're deliberately open-ended — ethical dilemmas, system-design tasks and tricky reasoning puzzles — so there's no single lookup answer. The models are judged on how well they think, not what they've memorised.
No. The debate uses free-tier AI models with strict usage caps; if a model is busy, the arena simply pauses rather than incurring any cost.