Every test on this site is a browser adaptation of an idea from research β not the instrument itself. This page says exactly where each task comes from, what it cannot tell you, and every correction we have made since launch.
And the distinction that most sites blur: the research belongs to the idea of the task, not to our version of it. Citing a 1935 experiment does not make a browser page a validated instrument. It only explains where the idea came from.
| Test | Origin of the idea | What it does not measure |
|---|---|---|
| Reaction time | Choice reaction time and the HickβHyman law (Hick 1952; Hyman 1953): choice time grows with the logarithm of the number of options. | Not simple reaction time, and not reflexes outside a screen. Display and input lag are inside the number. |
| Visual memory | Block-tapping span tasks in the tradition of Corsi (1972): positions shown in sequence, then reproduced in order. | Not memory for faces, words, or anything you need for more than a few seconds. |
| Matrix reasoning | Matrix completion, the format used by many reasoning tests. Our grids are generated from Latin squares β they are not Raven's matrices. | Not an IQ score, and not knowledge-based reasoning. |
| Stroop | Stroop (1935): naming the ink colour of a conflicting colour word. | Not attention in daily life, and the colour requirement limits who it can describe fairly. |
| Task switching | Task-switching paradigms in the tradition of Rogers & Monsell (1995): switch cost as the difference between switch and stay trials. | Not multitasking, and not how you handle real competing demands. |
| Mental rotation | Shepard & Metzler (1971): response time rises linearly with angle of rotation. Our shapes are flat, not the original 3D figures. | Not navigation, and not comparable with scores from the published task. |
| Delayed recall | Recognition memory after a filled delay: shapes seen, interference, then old-versus-new judgement scored as hits minus false alarms. | Not clinical memory testing, and not a screen for memory disorders. |
Saying you are faster than 87% of people is a claim about other people, and it needs a published norm collected on the same task under the same conditions. We looked for one. The closest usable data covered simple reaction time in a laboratory, measured with standardised equipment, split by age β and it does not fit this site for three reasons together:
So instead of inventing precision, we compare you with the only fair reference available: your own earlier attempts, on your own device.
The door is left open in code rather than in intention: a sourcing gate marks any scale that gains a real published reference, and percentiles would appear only for a test that passes it. None does today.
The six-sided chart needs one shared scale, because its axes are measured in different units: milliseconds, positions, correct answers. These are the bounds we convert with. They are our choice, not a standard β a generous range on one axis and a tight range on another would make the same person look strong on one and weak on the other.
| Axis | Measured as | 0% end | 100% end |
|---|---|---|---|
| Speed | median search time | 1600 ms | 500 ms |
| Memory | longest sequence | 2 positions | 9 positions |
| Logic | share correct | 0% | 100% |
| Focus | Stroop interference | 320 ms | 40 ms |
| Flexibility | switch cost | 800 ms | 90 ms |
| Spatial | share correct | 50% (chance) | 100% |
This is why your raw number is always printed next to the chart: the chart is a shape, the raw number is the fact.
No puzzle on this site was drawn by hand. Each one is generated when you ask for it, and checked in code before you see it: a matrix is built from a rule so its answer is derived rather than typed; a rotation shape is rejected unless its mirror differs from all of its own rotations; recall shapes are regenerated until every pair differs by a wide enough margin.
That is why the tests never repeat and never need a content update β and why a bug in a generator is a bug in every puzzle, which is the trade we accept and test for.
Corrections are published with their date and reason. This is the part of the page we are least comfortable writing and most convinced is necessary.
The flexibility range was 420β60 ms. The first real session produced a switch cost of 567 ms β outside the range β which the chart clamped to 0%. A zero on a scale whose bounds are our own choice is a verdict on the player caused by our error. The range became 800β90 ms, and that result now reads 33%.
The closest pair of shapes in a set differed by about 7.6 pixels on screen, which made the task closer to an eye test than a memory test. Radius levels were widened, the stage was enlarged, and display time was raised to 2.2 seconds. The closest pair now differs by about 11.6 pixels on a phone and 15.2 on a desktop.
The tests existed only in Arabic. The engine was language-neutral from the start, so this release adds English pages written natively rather than translated, with the same limits stated in both languages.
If you find an error in the reasoning, the wording or the code, tell us. A correction that holds up will be applied and recorded in the changelog above, with credit to whoever reported it.