How the score is computed

Every step from your answers to the number on the results page, including where the method is weakest.

Framework

The Cattell-Horn-Carroll model is the taxonomy of cognitive abilities that contemporary clinical batteries are organised around. It separates general ability into broad abilities that are related but measurably distinct.

This battery samples six of them: fluid reasoning (matrix problems), quantitative reasoning (number series and rule grids), verbal reasoning (analogies, vocabulary, deductive logic), visual-spatial processing (mental rotation, paper folding, mirror discrimination), working memory (forward, backward and re-ordered span) and processing speed (timed symbol search).

Sampling several broad abilities rather than one is what makes a profile possible. A single-domain instrument can give you a number; it cannot tell you that your verbal reasoning outruns your processing speed.

Items

There are 70 scored items, and difficulty rises within every section rather than across the test as a whole.

Matrix and spatial items are generated from explicit transformation rules, which means every keyed answer is correct by construction and no distractor duplicates a visible cell. Series items are defined by generator functions. Verbal items are hand-written, each with a single defensible key.

The scored items are not published, because an item only measures anything the first time you meet it. The sample questions are written separately for illustration.

Scoring

Item response theory scores a test by modelling each item separately rather than counting correct answers. This battery uses a three-parameter logistic model, where each item carries a difficulty, a discrimination (how sharply it separates people near that difficulty) and a guessing parameter.

The consequence is that two people with the same number of correct answers can receive different scores. Solving a hard, highly discriminating item moves the estimate a long way; getting an easy one right moves it very little; and a correct answer on an item with a high guessing parameter is discounted rather than credited in full.

Ability is estimated by expected a posteriori (EAP) estimation over a standard-normal prior — the mean of the posterior distribution of ability given your response pattern, rather than the single most likely value.

The scale

A deviation IQ places an ability estimate on a scale with a mean of 100 and a standard deviation of 15. The constants used for that conversion here are derived from a 20,000-person simulated normative population.

The calibration is re-run whenever an item changes, and an automated test suite checks four things: that the population mean and spread stay at 100 ± 1.5 and 15 ± 1.5, that the scale is linear, that random guessing scores below average, and that answering more items correctly never lowers a score.

Confidence interval

The reported 95% interval comes from the posterior standard error at your ability level. It is the honest form of the result: a best estimate plus the range a repeat sitting would plausibly fall in.

Precision is not constant across the scale. Mid-range scores are estimated from many items of roughly the right difficulty and are therefore tight. Extreme scores rest on the few items extreme enough to be informative, so their intervals are wide. Any fixed-length test behaves this way. See IQ to percentile for what a given band is worth in rank terms.

Profile measures

The profile panel is self-report, not ability testing, and is scored separately from the IQ estimate. Personality uses a ten-item Big Five inventory with two items per trait, one of them reverse-keyed. Cognitive style and perception items are preference measures. They describe tendencies; they are not scored against a norm and they do not feed into the IQ number.

Limits

This is a self-administered, unsupervised online assessment. Nothing verifies who sat it, whether it was taken in one uninterrupted session, or whether outside help was used, and a ten-item Big Five inventory is a coarse instrument by design.

It is not a substitute for a supervised clinical evaluation and should not be used for diagnosis, placement or employment decisions. For how it compares to instruments that can be, see types of IQ test.

Keep reading

Find out where you actually land.

70 calibrated items, about thirty minutes, no account.

Start the test