← Results

Reliability

How consistently each test measures: whether its items agree with each other, whether people who take it twice get the same score, and whether short or adaptive forms agree with the long one.

Method

Internal consistency: Cronbach's α of each scale among the people who answered all its items, first attempts (for right-or-wrong items this is KR-20). Where each person gets a random draw of the questions (the citizenship tests), the average correlation between questions answered together, stepped up to the number a person answers (standardised α). A test that gives people different items (adaptive, or scored by IRT) has none; it gets the empirical reliability of its ability estimates instead: their variance over that plus their mean squared standard error.

Retest: each person's first and second completed sitting. Retakes on this site are the person's own choice and mostly the same day, often to see how a different answer changes the result, so this is a lower bound on the test's retest reliability, not a planned retest. Changes on the main score more than 3.5 robust standard deviations from the typical change are marked; large gains are the pattern of answers looked up between sittings. No one is identified.

Part and whole: the correlation between an adaptive or short form and the full test among people who took both, first attempts.