In response to Rob Kremer’s critique of the CIM/CAM and the state tests, I pointed out that the tests are technically brilliant but conceptually bankrupt. I can say this since no validity studies on them have been conducted that I know of. I can also say they are technically very well constructed since the multiple choice tests make use of some of the most advanced Rasch techniques for items calibration and scaling. So we need be careful in our criticism, and that means we need to understand this distinction between validity of the tests and their technical properties.
I don’t think the testing issue matters much because the larger principle is that the testing itself doesn’t cause change and improvement. It simply isn’t a strategy that will produce significant change and improvement in the system. So improving the tests isn’t a reform I would worry much about. None-the-less, the tests do matter and huge amounts of time and attention are devoted to the state’s tests. So we need to understand them. We need to ask questions and insist on responsesresponses like this reader responding to my comment at Rob Kremer:
I do not get the Rasch/RIT scale. I have somewhere read that the RIT scale is alleged to be an interval scale, which I understand to mean that a 10-point difference should be equal whether it is a difference between 190 and 200 or between 240 and 250, just as a one-foot difference in the shot put is the same. I understand you can create equal intervals when measuring how far someone can put the shot, but I have a harder time accepting that when it comes to test results.
The Oregon department of education for the fifth grade sample math test, 2001-02, provided a table converting ‘number correct’ to the RIT score. Total 25 questions. One correct equals 169.0. Two correct jumps 9.8 points to 178.8. After that, for each additional correct answer, the RIT gains trend from over six points to barely over two points at the middle range of 12-13 correct answers, then trend back up [24 correct gets a RIT of 250.8; 25 correct gets a RIT of 260.6, an increase of 9.8 points].
So how can every single one-RIT-point interval be deemed equal?
If a kid gets 24 correct and scores 250.8, there is no way to know which one of the 25 questions he got wrong; but the kid who gets 25 correct gets 9.8 points more. This treats every one of the 25 questions as if it is worth 9.8 points. But that makes no sense. It might sort of make sense if total correct answers was tied to a normal curve. But then the relative difficulty of individual questions does not relate.
Can you explain?
I don’t know how well I can explain it in simple terms but let me try.
When we measure something, we need a measuring instrument that compares an attribute of an object to a scale. The bathroom scale is the tool you use to compare your weight to a weight scale. The scale has a standard unit and an anchor at 0. Thermometers measure temperature of an object against a Fahrenheit scale of standard degree units and an anchor at boiling and freezing. We take one measurement.
A multiple choice test item is a measurement tool but it only measures whether a student exceeds a single point on the scale. It can’t measure a range of values, only a single value or answer. So we use many items hopefully at many points along a scale. To create an anchor for the scale we group all the scores from a group of students and find their average. To create the scale we use their distribution along a count of correct responses for each item and covert these to different scales, raw scores, percentage correct, percentiles of students, standard deviations, etc. There’s another way to create a scale and measuring tool. It starts with two kinds of data that are different from norms based on a group of students and a particular set of test items.
Items vary in difficulty. Students vary in ability to respond correctly. We might be using many items of average difficulty with few that are hard or easy. So we can imagine that getting a few more or less correct for the average student, even though the count of correct answers changes, doesn’t mean much in terms of movement along some theoretic scale. The student may still be almost average. But for a high ability student who gets most items correct, getting one more correct, a very difficult items, means much more in terms of the distance along the scale. So everyone agrees that the raw scores are not equal intervals along some theoretic scale. How then do we get to a theoretic scale if not by using the average and distribution of a group of students on a test of selected items we choose to use in measuring the students?
Since the items vary in difficulty and students vary in ability, the two foundational data sources we use are the ordinal property or the items and the students’ ability levels. We order the items according to their diffriculty—percentages of students getting each item correct— and order students according to their level of ability to respond correctly—percentages of items the student answers correctly. To use these ordered series, we can then put every response into a matrix ordered by the difficulty level of the items and the ability level of students. As a result, we should see a Guttman scaling in which students tend to get all items correct that are below their ability level and no items correct that are above their ability level. From the matrix rows and columns, we also have the percentage of items correct for each student and percentage of students answering correctly for each item.
The percentages don’t form a linear scale because to get a hard item correct depends both on an increase in item difficulty and an increase in student ability. The odds of getting the extreme items correct changes exponentially; so we convert these percentages into a logarithmic expression and the resulting log scale is linear. Then we can figure the probability for each cell according to the percentages of a correct response in each cell according to the difficulty level of the item and ability of the student. We express this probability in terms of our logarithmic scale. So the odds of getting an item correct right at the students ability level is 50% while expressed as log odds on the logarithmic equal interval scale, its probability is 0. Getting a harder items correct has lower odds and getting easier items correct. An item in a cell at the 55%tile probability has a log odds of +0.20. One at the 95% has a log odds or +2.94. So many cells for a low ability student will have low odds of her producing correct responses in those cells while the corresponding cells for a high ability will have much higher odds of getting those same items correct. The log odds intervals are called logits or Rasch units or the scale can be transformed to more convenient units and given another name.
But to compute the difficulty level for an item, we need to know the ability level of students. To know the difficulty level of students we need to know the items’ difficulty levels. We can’t figure one without first knowing the other however there is an easy way to do it with a computer. What we do is estimate one and compute the other. Then we reverse the process. Then we keep iterating the process and each time watch our results as the computed levels for each item and each student keep refining themselves until their change is negligible. Now these log odds creating our scale to measure either the item difficulties or student ability levels don’t depend on anything further; for example, we haven’t computed averages or distributions. All we did we express the probabilities of being correct for each response as a function of student ability and item difficulty. It doesn’t matter what kind of distribution we have of students or items. We end up with the difficulty levels for every item and for every student.
With an independent scale, we can add more items or more students and the scale won’t change since it doesn’t depend on any sample of students or any the set of items. The scale doesn’t depend on what test we use or what group of students we test, and that student and item independence is very useful. We have the ability to measure the difficulty level of each current or any future test items. We could extend the test upwards or downwards. We can omit items or add them since the only thing that is used is the difficulty level of each item and ability level of each student independent of all other items and students. We could tie or versions tests together since we can anchor them to the permanent logit scale. The logit scale expresses probabilities, and items and students might appear anywhere on the scale. They might be clustered or not. The scale exists independently of the individual log odds of each response to items by each student. So we have an interval scale that isn’t affected by items. This means we can add items at the difficulty level we set for a cutoff so we have a very fine grained tool at this point on the scale. Getting another item correct moves along the scale only a tiny bit. It gives us great accuracy.
Now this Rasch model—based on the idea that correct responses are a function of item difficulty and ability level—has powerful use in test development. Since the response patterns form an ordered series they should form a Guttman scale either for the responses of a single student or all students’ responses for a single item, Responses should be all positives until they turn negative in which case they should continue negative. If either the items or the students don’t’ follow this pattern, something is wrong either with the item or with the testing situation of a student. We say the item or the student doesn’t fit the model.
The misfit the model. We can take each of these misfits that are out of place responses and figure how much error variance they produce. An out of order item close to the students’ ability level doesn’t produce much error variance but a response farther removed from the transition point of correct and incorrect responses for a student does. These error variances are called "residuals," and they are used to compute a "fit" score for each student and each item.
When I develop a test, I must throw out or change items that have poor fit scores. Likewise, we cannot use the students test results if his results have a poor fit score. These fit scores tell us whether the statistical results are matching our theoretic model for generating test items. The verify the soundness of the test. We also have error scores and other statistics describing the quality of the test items that are used for calibrating tests.
The ODE, Cathy Brown, didn’t, and wouldn’t, use these tools when the math problem solving tests were developed, so the quality of the test were unknown. They were not calibrated. As a result, the abject failure of the tests didn’t show up until after they were implemented. It makes us wonder whether officials at the ODE understand the Rasch model. The Rasch model is much more powerful than norm referencing but it has nothing to do with the validity of test items. Its statistics can be used to verify certain theoretic ideas we attempt to express in our test items have but it doesn’t generate valid tests. We should encourage the use and disclosure of the Rasch model and its results but our real concern should be the validity of the state tests. More importantly we should question whether standards and testing are even a reform strategy or if they are only indicators of the level of performance.
