Methodology & Sources

What we measure, and what we refuse to claim

Every test on this site is a browser adaptation of an idea from research β€” not the instrument itself. This page says exactly where each task comes from, what it cannot tell you, and every correction we have made since launch.

1. What these tests are not β€” stated first, not last

And the distinction that most sites blur: the research belongs to the idea of the task, not to our version of it. Citing a 1935 experiment does not make a browser page a validated instrument. It only explains where the idea came from.

2. Where each task comes from

TestOrigin of the ideaWhat it does not measure
Reaction timeChoice reaction time and the Hick–Hyman law (Hick 1952; Hyman 1953): choice time grows with the logarithm of the number of options.Not simple reaction time, and not reflexes outside a screen. Display and input lag are inside the number.
Visual memoryBlock-tapping span tasks in the tradition of Corsi (1972): positions shown in sequence, then reproduced in order.Not memory for faces, words, or anything you need for more than a few seconds.
Matrix reasoningMatrix completion, the format used by many reasoning tests. Our grids are generated from Latin squares β€” they are not Raven's matrices.Not an IQ score, and not knowledge-based reasoning.
StroopStroop (1935): naming the ink colour of a conflicting colour word.Not attention in daily life, and the colour requirement limits who it can describe fairly.
Task switchingTask-switching paradigms in the tradition of Rogers & Monsell (1995): switch cost as the difference between switch and stay trials.Not multitasking, and not how you handle real competing demands.
Mental rotationShepard & Metzler (1971): response time rises linearly with angle of rotation. Our shapes are flat, not the original 3D figures.Not navigation, and not comparable with scores from the published task.
Delayed recallRecognition memory after a filled delay: shapes seen, interference, then old-versus-new judgement scored as hits minus false alarms.Not clinical memory testing, and not a screen for memory disorders.

3. Why you will never see a percentile here

Saying you are faster than 87% of people is a claim about other people, and it needs a published norm collected on the same task under the same conditions. We looked for one. The closest usable data covered simple reaction time in a laboratory, measured with standardised equipment, split by age β€” and it does not fit this site for three reasons together:

So instead of inventing precision, we compare you with the only fair reference available: your own earlier attempts, on your own device.

The door is left open in code rather than in intention: a sourcing gate marks any scale that gains a real published reference, and percentiles would appear only for a test that passes it. None does today.

4. The chart scale is a choice, and here it is

The six-sided chart needs one shared scale, because its axes are measured in different units: milliseconds, positions, correct answers. These are the bounds we convert with. They are our choice, not a standard β€” a generous range on one axis and a tight range on another would make the same person look strong on one and weak on the other.

AxisMeasured as0% end100% end
Speedmedian search time1600 ms500 ms
Memorylongest sequence2 positions9 positions
Logicshare correct0%100%
FocusStroop interference320 ms40 ms
Flexibilityswitch cost800 ms90 ms
Spatialshare correct50% (chance)100%

This is why your raw number is always printed next to the chart: the chart is a shape, the raw number is the fact.

5. Generators, not libraries

No puzzle on this site was drawn by hand. Each one is generated when you ask for it, and checked in code before you see it: a matrix is built from a rule so its answer is derived rather than typed; a rotation shape is rejected unless its mirror differs from all of its own rotations; recall shapes are regenerated until every pair differs by a wide enough margin.

That is why the tests never repeat and never need a content update β€” and why a bug in a generator is a bug in every puzzle, which is the trade we accept and test for.

6. Changelog

Corrections are published with their date and reason. This is the part of the page we are least comfortable writing and most convinced is necessary.

27 August 2026 β€” calibration error, flexibility axis

The flexibility range was 420–60 ms. The first real session produced a switch cost of 567 ms β€” outside the range β€” which the chart clamped to 0%. A zero on a scale whose bounds are our own choice is a verdict on the player caused by our error. The range became 800–90 ms, and that result now reads 33%.

27 August 2026 β€” delayed recall shapes were too similar

The closest pair of shapes in a set differed by about 7.6 pixels on screen, which made the task closer to an eye test than a memory test. Radius levels were widened, the stage was enlarged, and display time was raised to 2.2 seconds. The closest pair now differs by about 11.6 pixels on a phone and 15.2 on a desktop.

2 October 2026 β€” English pages published

The tests existed only in Arabic. The engine was language-neutral from the start, so this release adds English pages written natively rather than translated, with the same limits stated in both languages.

7. Found a mistake?

If you find an error in the reasoning, the wording or the code, tell us. A correction that holds up will be applied and recorded in the changelog above, with credit to whoever reported it.