Two AI models race the same task. Both have to get it right — but the winner is the one that does it in the fewest tokens. Lower is better, just like golf.
Modern AI isn't judged on accuracy alone — cost, speed and efficiency matter just as much. Token Golf scores brevity: who said it correctly in the fewest words and tokens, how fast, and how cheaply. Part of the AI vs AI arena, alongside the Debate Auto-Arena.
Both models attempt the same task. Each answer must first be correct — objective tasks are auto-checked, and rewrite/summarise tasks are verified by a neutral judge model. Among the correct answers, the one that used the fewest generated (output) tokens wins.
Each model uses its own tokeniser, so token counts aren't perfectly like-for-like — but they're exactly what you'd be billed for. We show a tokeniser-neutral word count alongside, so you can judge true brevity too.
An illustrative estimate using standard per-token output prices, to show the relative difference. The site itself runs on free tiers, so it costs nothing to play.