Sunday, June 12, 2005

Rasch Response 4

4. Dick: But for a high ability student who gets most items correct, getting one more correct, a very difficult items, means much more in terms of the distance along the scale. So everyone agrees that the raw scores are not equal intervals along some theoretic scale.

4. John: This is where I start getting confused. In multiplication, I can understand that 10 x 10 is easier than 57 x 79. But if we use 10 problems, and Smart kid gets 9 right and Smarter kid gets 10 right, we don’t know which one Smart got wrong. Maybe he got a medium-difficulty problem wrong, not the toughest one.

If so, let’s say the RIT scale has just a 2-point differential between getting 5 right and getting 6 right, then why should it have a 6 point differential between getting 9 right and getting 10 right, where the one wrong answer might not have been the toughest problem? Why doesn’t it make more sense to give different numbers of points for problems depending on their difficulty [10 x 10 is worth 2 points; 57 x 79 is worth 6 points]?




4. Your interpretation of scores based on counting correct responses isn’t quite correct. For a Rasch calibrated test, the student’s score doesn’t depend on the number correct but where on the scale he falls as a result of his responses to the items. If he got a lot of hard item correct and no easy ones, he would score higher than another student who answered more items correct but didn’t get as many difficult items correct (although the test results would show a serious misfit for this student meaning something was wrong with the way the student responded).

It doesn’t matter how many items but which items he answers correctly. We could give a short version of the test and still get the same Rasch score for a student because the items only serve to locate where on the scale the student is according to which items he answers correctly. The scale has nothing to do with the number correct. Because the items don’t fall on the same place on the scale, if a student answers a moderately hard one correctly and a very hard one incorrectly we know his ability level on the scale falls between the difficulty level on the scale of these two items. In theory, if our test was perfect we would only need the two items that exactly bracketed the student to measure the student.

Once we have the scale, then the scale doesn’t depend on the items as long as we have the difficulty level of each item. New items can be measured with the scale to see where they fall in their difficulty. Then we can use these as items for measuring students’ ability levels. The items, in effect, stand for points on the scale.

The same is exactly true for students. Once we have the scale, we can see where students fall in their ability level. Then we can use them to stand for certain points on the scale just as we can locate items at their point on the scale. The results don’t depend on what items or what students we measure. The scale is permanently fixed.

If we gave total points, then, even if items were weighted as you suggest, we would not longer have an independent scale. We would have a count of responses. Rasch scales don’t represent a simple counting of correct responses. It depends on the difficulty level of the items the student answers correctly. But of course we can only use the scale to measure items or students if we already have the scale developed. So the use of the scale to measure items or students is different than the procedure for abstracting the underlying scale in the first place. With a norm referenced test based on counting correct answers, we don’t have any underlying scale except for the counts of correct answers by a certain group of students and for a certain set of items. That counting procedure is not very objective.