More…
8. Dick: The percentages don’t form a linear scale because to get a hard item correct depends both on an increase in item difficulty and an increase in student ability.
8. John: I think your point here is that where the test is designed correctly, you won’t have a situation like I suggested in paragraph 6, where the difference in difficulty is so small between some questions that it doesn’t really take greater ability to get the slightly harder one correct. In other words, if your norming process showed that 64% got the hardest problem correct, and 63% got #2 hardest correct, you have not selected the two hardest problems correctly—you need a more difficult "hardest" problem that substantially fewer than 63% get right, so that there is a significant ability difference or gap between the smartest kids who get the hardest one right vs the next-smartest who get nine correct but get the hardest one wrong.
It doesn’t matter whether two items are close together in difficulty. The analysis will show their difficulty level as close together. For students, since the items are about the same difficulty, the chances of answering either item correctly won’t be different. Getting the slightly easier one wrong and the harder one right, might occur often with students if the probabilities of answering each one correctly overlap so that there isn’t a significant difference between them. In other words, since they are close to each other on the scale, we expect some minor inconsistency.
But we can have many items very close together at one point on the scale if we want increased accuracy. If we have big gaps between items, then we can measure students accurately. So the state actually does pack in many more items at each of its cut scores so as to reduce the error of measurement at that point.
Remember, we aren’t using a norming process. We are using a different model, one based on determining the probabilities of correct responses for each item and converting those to a scale. There are no norms unless we then use our test and scale to see how many students score at various points, or what the average might be for a group of students, or whatever. But these data have nothing to do with the test calibration and design. By freeing the test from dependence on norms, we can create a measurement that is independent of the students or the items we use to develop the test. It simply doesn’t matter how difficult or how easy or how many of each kind of items we have. They simply show up as points on the scale. Of course, in general, we probably want a nice even distribution of the items along the scale so that the scale is equally precise all along it. But we can add or remove items as we like to get the kind of precision we want.