Token Trap Test
What is a token trap?
A token trap is a question that any literate person can answer by looking at the letters in a word, but that trips up artificial intelligence because AI does not read letters. Before a language model sees any text, a tokeniser chops it into chunks called tokens: "strawberry" might arrive as "straw" and "berry", or even as one indivisible piece. The model never sees an R. So when asked how many times the letter R appears in strawberry, it has to remember the spelling rather than count, and in 2024 the leading models famously answered two instead of three.
This free test has 16 questions of that kind and a 10-minute limit. Your percentage score and category breakdown appear instantly, and because the identical questions are put to AI models through the site's API, you can see the human and machine scores side by side on the human-versus-AI leaderboard.
The four categories
- Letter counting — how many times a given letter appears in a word or short phrase, including double letters that models routinely miss.
- Letter position — which letter is third from the end, or what the middle letter of a word is. Positional questions are hard when you cannot see positions.
- Spelling and reversal — reading words backwards, spotting the odd one out among near-identical spellings, and identifying anagrams.
- Word constraints — choosing the word that satisfies a rule about its letters, such as containing no vowels or having exactly five letters.
How the test is scored
Every question has one correct option. Your headline score is the percentage correct out of 16, shown with a breakdown across the four categories. Answer options are shuffled on each attempt so there is no letter position to memorise. The percentile you see afterwards compares you with other players on this site; this is an original test with no published norms, written to illustrate a specific weakness of AI and to give people a brisk attention-to-detail check.
Why the gap exists, and why it is closing
Tokenisation is a deliberate engineering trade-off. Reading text in sub-word pieces lets a model handle every language and every rare word with a vocabulary of a few hundred thousand tokens, and it makes training far cheaper than working letter by letter. The cost is that anything which depends on the letters themselves, from counting to rhyming to reversing, has to be reconstructed from memory. Recent models have learned a workaround: when asked a counting question, they spell the word out one letter per line first, turning each letter into its own token, and then count. That is why AI scores on this test have risen, and why the leaderboard is worth revisiting.
The Token Trap Test sits alongside the Common Sense Test, the Lateral Thinking Test and the Visual Reasoning Test in the AI vs Human Arena. Each targets a different reason a text-trained system can fail at something a person finds easy.
Tips for a clean score
- Count with your finger or by writing the word down. Double letters are where people lose points too.
- For reversal questions, read the reversed word aloud slowly rather than trying to flip it in your head.
- For constraint questions, check every option against the rule instead of stopping at the first one that seems to fit.
- There is no speed bonus. Ten minutes is far more than the test needs, so use it.
Frequently Asked Questions
What is a token trap?
A question that is trivial for a person who can see individual letters but hard for an AI model, because language models read text in tokens (chunks such as "straw" and "berry") rather than letters. Asking how many times the letter R appears in strawberry forces the model to reason about letters it never actually sees.
How is the Token Trap Test scored?
There are 16 multiple-choice questions, four in each of four categories, with a 10-minute limit. Your score is the percentage correct, with a per-category breakdown. Options are shuffled on every attempt.
Why can't AI count letters in a word?
Because the word is never presented to the model as letters. A tokeniser splits text into sub-word pieces before the model sees it, so it can only count letters by recalling the spelling from training data. Newer models sometimes answer correctly by writing the word out letter by letter first.
Do people ever get token trap questions wrong?
Yes, mostly through haste: miscounting a double letter, or misreading a reversed word. The test is deliberately unhurried so that the human score reflects attention rather than speed.
Is this a real psychometric test?
No. It is an original test built to demonstrate a specific weakness of AI systems and to give people a quick attention-to-detail check. It has no published norms; the percentile compares you with other players on this site.