Sunday, June 12, 2005

Rasch Response 6



Here's more in the Rasch discussion…

6. Dick: Since the items vary in difficulty and students vary in ability, the two foundational data sources we use are the ordinal property of the items and the students’ ability levels. We order the items according to their difficulty—percentages of students getting each item correct— and order students according to their level of ability to respond correctly—percentages of items the student answers correctly.

6. John. Wait a minute. Back to the 10 multiplication problems. So we order them one through ten in difficulty based on percentages of kids who get each one correct. That doesn’t mean the difference in difficulty is the same between the hardest vs. #2 hardest and 6th hardest vs. 7th hardest. What if 64% got the hardest correct, 63% got #2 hardest correct; and 84% got #6 hardest correct, but 90% got #7 hardest correct. Now what do we do? {{I come back to this below.}}

You are correct. "Hardest" is item difficulty. It comes from a Rasch analysis that starts with ordering the items in difficulty according to the percentages of students who get each item correct. That ordering process would mean that the 64% item of #2 put it second in difficulty, the 63% on #2 put it the hardest item, the 84% on #6 put it third in difficulty, and the 90% on #7 made it the easiest item. You stated 64% got the hardest item correct. It’s not the hardest. #2 is. Either you made a mistake here or you are somehow attributing a difficulty to the item separate from its rank among the percentage of students answering items correctly. We don’t get to attribute the difficulty level to the item. The item’s result determines its rank in difficulty.

But I read you to mean a more important point that the intervals between the numbers of correct items are not equal, and you are right. Remember, we have to derive the scale using the Rasch analysis before we have a measure of the items’ difficulties. Then, the item difficulties show up along the scale, and we can see what their spacing is. For your items, we will find a large distance separating some items on the scale and a small distance along the scale between other items. But until the underlying scale is derived, we don’t know what the spacing is since all we have is the rank ordering of the items. The problem is how to get from these ordinal data to the abstraction of an equal interval scale.

To do that, we do two things. We use both the order of the items and the order of the students to determine the probability of each response being correct. For example, a hard item and a low ability student combine to create a very low probability that the item will be answered correctly. If it is, of course, something is not working according to our model. This partially solves the problem you raise that things are not equal among the scores. The second thing we do is convert these probabilities to a log scale. Then the log probabilities form an equal interval scale. This scale’s units are now logits, not counts of correct answers.

Both the individual items and individual students can be measured along this logit scale. It will show us how large the distance is between any two items and let us compare these distances at any points along the scale. That equal interval quality of this scale that is independent of either the students or the items means we would have a powerful tool for measuring student growth or curriculum topics we want students to learn. We could measure how difficult it is for students to learn 7 x 6 versus 2 x 4. Or we could measure either the current level of each student or the growth of a student across time and have it be directly comparable to other grade levels and curriculum