HomeCoursesStudyQuizzesPuzzlesFree TestsAI ArenaToolsDownloadsGuidesContact
Log InSign Up

This tests deep reasoning — defending your logic and critiquing an opponent — rather than recalling facts. It mirrors how researchers actually evaluate models today. Sits alongside the Open Arena, the head-to-head AI Battle, and the efficiency race Token Golf.

  1. AnswerBoth AIs tackle the same hard problem
  2. CritiqueEach tears into the other’s reasoning
  3. JudgeA neutral model scores both out of 20 and picks a winner

Frequently Asked Questions

How does the debate work?

Both models answer the same ambiguous problem. Each is then shown the other's answer and asked to find its flaws. Finally a third "judge" model scores each on the quality of its answer and the accuracy of its critique, and picks a winner.

Who is the judge?

The strongest free model that isn't one of the two debaters, so it never marks its own work. The debaters are shown to the judge anonymously and in a random order to reduce bias. A premium judge model can be plugged in later.

Why these problems?

They're deliberately open-ended — ethical dilemmas, system-design tasks and tricky reasoning puzzles — so there's no single lookup answer. The models are judged on how well they think, not what they've memorised.

Does it cost anything to run?

No. The debate uses free-tier AI models with strict usage caps; if a model is busy, the arena simply pauses rather than incurring any cost.