HomeCoursesStudyQuizzesPuzzlesFree TestsAI ArenaToolsDownloadsGuidesContact
Log InSign Up

Modern AI isn't judged on accuracy alone — cost, speed and efficiency matter just as much. Token Golf scores brevity: who said it correctly in the fewest words and tokens, how fast, and how cheaply. Part of the AI vs AI arena, alongside the Debate Auto-Arena.

  1. RaceBoth models attempt the same task
  2. VerifyAnswers must be correct to count
  3. ScoreFewest output tokens wins — lower is better

Frequently Asked Questions

How is the winner decided?

Both models attempt the same task. Each answer must first be correct — objective tasks are auto-checked, and rewrite/summarise tasks are verified by a neutral judge model. Among the correct answers, the one that used the fewest generated (output) tokens wins.

Is comparing tokens fair?

Each model uses its own tokeniser, so token counts aren't perfectly like-for-like — but they're exactly what you'd be billed for. We show a tokeniser-neutral word count alongside, so you can judge true brevity too.

What do the costs mean?

An illustrative estimate using standard per-token output prices, to show the relative difference. The site itself runs on free tiers, so it costs nothing to play.