Item statistics
How each item of a cognitive test behaves, from everyone's first attempt: how many answer it correctly, how well it separates people who do well on the rest of the test from those who do not, and its loading on the ability the test measures. Items that measure little are flagged.
What the columns mean
Correct: the share who answered the item correctly, of those given it. Biserial: the correlation of the item with the number correct on the rest of the test (on an adaptive test, with the ability estimate), corrected for the item being correct or not; below about .2 the item tells little about the rest. Loading and difficulty: from a one-factor two-parameter logistic IRT model fitted by marginal maximum likelihood, with items a person was not given (adaptive tests, time limits) left out rather than counted wrong; the loading is on the correlation scale, the difficulty is the ability (in standard deviations) at which half answer correctly.
Items given to fewer than 30 people get no statistics beyond the share correct. Flags: broken? loading or biserial below .1; weak loading below .3; too easy above 97% correct; below chance fewer correct than guessing would get. Answer keys are not shown. The takers can be limited to English-speaking countries (US, UK, Canada, Australia, New Zealand, Ireland) or the rest, by the country of the internet address: the tests are in English, and a verbal item can look broken only among people who are not native speakers. Time: the median time people spent on the item, first attempts. A difficulty whose standard error is above 1 (hover for it) is marked unstable: it happens when an item barely loads, since the difficulty divides by the slope.