Tuesday, May 24, 2005
Can We Make Overachieving a Normal Expectation?
The chart atEduwonk.com shows Roxbury Prep's students and statewide scores by subgroup. So we have an unusual school. Can we expect all schools to be unusual merely because one is an outlier on the distribution curve? Is the message: this school performs well so all schools should be able to do as well? This hope strategy in ed reform won't work. We've got to get the fundamentals right.
Labels:
intelligence,
public schooling
Sunday, May 22, 2005
Rasch Versus Norm Referenced Tests
In response to Rob Kremer’s critique of the CIM/CAM and the state tests, I pointed out that the tests are technically brilliant but conceptually bankrupt. I can say this since no validity studies on them have been conducted that I know of. I can also say they are technically very well constructed since the multiple choice tests make use of some of the most advanced Rasch techniques for items calibration and scaling. So we need be careful in our criticism, and that means we need to understand this distinction between validity of the tests and their technical properties.
I don’t think the testing issue matters much because the larger principle is that the testing itself doesn’t cause change and improvement. It simply isn’t a strategy that will produce significant change and improvement in the system. So improving the tests isn’t a reform I would worry much about. None-the-less, the tests do matter and huge amounts of time and attention are devoted to the state’s tests. So we need to understand them. We need to ask questions and insist on responsesresponses like this reader responding to my comment at Rob Kremer:
I do not get the Rasch/RIT scale. I have somewhere read that the RIT scale is alleged to be an interval scale, which I understand to mean that a 10-point difference should be equal whether it is a difference between 190 and 200 or between 240 and 250, just as a one-foot difference in the shot put is the same. I understand you can create equal intervals when measuring how far someone can put the shot, but I have a harder time accepting that when it comes to test results.
The Oregon department of education for the fifth grade sample math test, 2001-02, provided a table converting ‘number correct’ to the RIT score. Total 25 questions. One correct equals 169.0. Two correct jumps 9.8 points to 178.8. After that, for each additional correct answer, the RIT gains trend from over six points to barely over two points at the middle range of 12-13 correct answers, then trend back up [24 correct gets a RIT of 250.8; 25 correct gets a RIT of 260.6, an increase of 9.8 points].
So how can every single one-RIT-point interval be deemed equal?
If a kid gets 24 correct and scores 250.8, there is no way to know which one of the 25 questions he got wrong; but the kid who gets 25 correct gets 9.8 points more. This treats every one of the 25 questions as if it is worth 9.8 points. But that makes no sense. It might sort of make sense if total correct answers was tied to a normal curve. But then the relative difficulty of individual questions does not relate.
Can you explain?
I don’t know how well I can explain it in simple terms but let me try.
When we measure something, we need a measuring instrument that compares an attribute of an object to a scale. The bathroom scale is the tool you use to compare your weight to a weight scale. The scale has a standard unit and an anchor at 0. Thermometers measure temperature of an object against a Fahrenheit scale of standard degree units and an anchor at boiling and freezing. We take one measurement.
A multiple choice test item is a measurement tool but it only measures whether a student exceeds a single point on the scale. It can’t measure a range of values, only a single value or answer. So we use many items hopefully at many points along a scale. To create an anchor for the scale we group all the scores from a group of students and find their average. To create the scale we use their distribution along a count of correct responses for each item and covert these to different scales, raw scores, percentage correct, percentiles of students, standard deviations, etc. There’s another way to create a scale and measuring tool. It starts with two kinds of data that are different from norms based on a group of students and a particular set of test items.
Items vary in difficulty. Students vary in ability to respond correctly. We might be using many items of average difficulty with few that are hard or easy. So we can imagine that getting a few more or less correct for the average student, even though the count of correct answers changes, doesn’t mean much in terms of movement along some theoretic scale. The student may still be almost average. But for a high ability student who gets most items correct, getting one more correct, a very difficult items, means much more in terms of the distance along the scale. So everyone agrees that the raw scores are not equal intervals along some theoretic scale. How then do we get to a theoretic scale if not by using the average and distribution of a group of students on a test of selected items we choose to use in measuring the students?
Since the items vary in difficulty and students vary in ability, the two foundational data sources we use are the ordinal property or the items and the students’ ability levels. We order the items according to their diffriculty—percentages of students getting each item correct— and order students according to their level of ability to respond correctly—percentages of items the student answers correctly. To use these ordered series, we can then put every response into a matrix ordered by the difficulty level of the items and the ability level of students. As a result, we should see a Guttman scaling in which students tend to get all items correct that are below their ability level and no items correct that are above their ability level. From the matrix rows and columns, we also have the percentage of items correct for each student and percentage of students answering correctly for each item.
The percentages don’t form a linear scale because to get a hard item correct depends both on an increase in item difficulty and an increase in student ability. The odds of getting the extreme items correct changes exponentially; so we convert these percentages into a logarithmic expression and the resulting log scale is linear. Then we can figure the probability for each cell according to the percentages of a correct response in each cell according to the difficulty level of the item and ability of the student. We express this probability in terms of our logarithmic scale. So the odds of getting an item correct right at the students ability level is 50% while expressed as log odds on the logarithmic equal interval scale, its probability is 0. Getting a harder items correct has lower odds and getting easier items correct. An item in a cell at the 55%tile probability has a log odds of +0.20. One at the 95% has a log odds or +2.94. So many cells for a low ability student will have low odds of her producing correct responses in those cells while the corresponding cells for a high ability will have much higher odds of getting those same items correct. The log odds intervals are called logits or Rasch units or the scale can be transformed to more convenient units and given another name.
But to compute the difficulty level for an item, we need to know the ability level of students. To know the difficulty level of students we need to know the items’ difficulty levels. We can’t figure one without first knowing the other however there is an easy way to do it with a computer. What we do is estimate one and compute the other. Then we reverse the process. Then we keep iterating the process and each time watch our results as the computed levels for each item and each student keep refining themselves until their change is negligible. Now these log odds creating our scale to measure either the item difficulties or student ability levels don’t depend on anything further; for example, we haven’t computed averages or distributions. All we did we express the probabilities of being correct for each response as a function of student ability and item difficulty. It doesn’t matter what kind of distribution we have of students or items. We end up with the difficulty levels for every item and for every student.
With an independent scale, we can add more items or more students and the scale won’t change since it doesn’t depend on any sample of students or any the set of items. The scale doesn’t depend on what test we use or what group of students we test, and that student and item independence is very useful. We have the ability to measure the difficulty level of each current or any future test items. We could extend the test upwards or downwards. We can omit items or add them since the only thing that is used is the difficulty level of each item and ability level of each student independent of all other items and students. We could tie or versions tests together since we can anchor them to the permanent logit scale. The logit scale expresses probabilities, and items and students might appear anywhere on the scale. They might be clustered or not. The scale exists independently of the individual log odds of each response to items by each student. So we have an interval scale that isn’t affected by items. This means we can add items at the difficulty level we set for a cutoff so we have a very fine grained tool at this point on the scale. Getting another item correct moves along the scale only a tiny bit. It gives us great accuracy.
Now this Rasch model—based on the idea that correct responses are a function of item difficulty and ability level—has powerful use in test development. Since the response patterns form an ordered series they should form a Guttman scale either for the responses of a single student or all students’ responses for a single item, Responses should be all positives until they turn negative in which case they should continue negative. If either the items or the students don’t’ follow this pattern, something is wrong either with the item or with the testing situation of a student. We say the item or the student doesn’t fit the model.
The misfit the model. We can take each of these misfits that are out of place responses and figure how much error variance they produce. An out of order item close to the students’ ability level doesn’t produce much error variance but a response farther removed from the transition point of correct and incorrect responses for a student does. These error variances are called "residuals," and they are used to compute a "fit" score for each student and each item.
When I develop a test, I must throw out or change items that have poor fit scores. Likewise, we cannot use the students test results if his results have a poor fit score. These fit scores tell us whether the statistical results are matching our theoretic model for generating test items. The verify the soundness of the test. We also have error scores and other statistics describing the quality of the test items that are used for calibrating tests.
The ODE, Cathy Brown, didn’t, and wouldn’t, use these tools when the math problem solving tests were developed, so the quality of the test were unknown. They were not calibrated. As a result, the abject failure of the tests didn’t show up until after they were implemented. It makes us wonder whether officials at the ODE understand the Rasch model. The Rasch model is much more powerful than norm referencing but it has nothing to do with the validity of test items. Its statistics can be used to verify certain theoretic ideas we attempt to express in our test items have but it doesn’t generate valid tests. We should encourage the use and disclosure of the Rasch model and its results but our real concern should be the validity of the state tests. More importantly we should question whether standards and testing are even a reform strategy or if they are only indicators of the level of performance.
Labels:
research methods
Wednesday, May 18, 2005
Will HB3162 Improve Education?
I doubt it. Rob Kremer thinks so and he tackles the CIM/CAM and pushes HB 3162 as the cure for the failure of ed reform. We do need a discussion about the CIM/CAM and the tests as to whether they can be the means to reform of our public schools. But the discussion needs to focus on whether any standardizing and testing can served as the road to reform. Why would changing the tests matter? The strategy hasn't changed. HB3162 doesn't change or improve the existing strategy of reform. It retains the false hope that testing somehow matters. HB3162 only replaces "their" tests with "his" tests. It doesn't change the system.
Standards and testing aren't a strategy for reform. They only describe the performance of the system. And they reveal the massive failure of the system. Of the 75% of our kids who graduate, only 1/3 achieve the CIM. That means the system has a sucess rate of only 25%. That dismal performance should be intolerable and cause for dramatic responses. We have known about this dismal performance for the last 20 years but the legislature is still avoiding the problem. It has no response, no theory of action for how it will bring about rapid, significant change and improvement.
So my criticism of CIM/CAM does not concern the quality of the tests. It lies with the absurd notion that we can make schools change by putting standards and testing in place. Testing more, or differently, or uniformly, is not a strategy for change. It's only measurement. The fundamental inertia and resistance to change that is built into the system by its very design will continue to give us the stable,dismal performance we have known about for the last 20 years.
The tests are irrelevant to change. They are not the problem or the cure. It is the system design that is bad. It is the means of delivery. Until the legislature honestly admits that trying to use tests, any tests, to force the schools to improve is not a strategy for change, we will not have a strategy that produces significant change and improvement in our schools. We must understand and change the design of the system, the means of delivery of education.
Standards and tests, whether flawed or not, are only the messenger of the bad news. The means we are trying to use to deliver public education services is a dynsfunctional system, an old public utility model of a government bureucratic organization with a protected exclusive franchise over all public education services and money that we hope will somehow respond to the need for dramatic change merely because the tests once again show it to be increasingly costly and consistent in its poor performance.
We need to focus on the system design, the fact that the legislature has created a design that guarantees the public schools their revenues, jobs, benefits whether the kids learn or not. The kids' failure to thrive is irrelefvant to the adults' success. The kids are hurt. The adults are rewarded, and HB3162 changes that immoral state of affairs not one whit.
Labels:
public schooling
Cultural Competency
Rob Kremer tackles the cultural competency bill in the legislature and the letters to the Oregonian. I looked up PSU's mission statement for their graduate teacher education. Here it is:
Guiding Principles:
1. We create and sustain educational environments that serve all students and address diverse needs.
2. We encourage and model exemplary programs and practices across the life span.
3. We build our programs on the human and cultural richness of the University's urban setting.
4. We develop collaborative efforts that foster our mission.
5. We challenge assumptions about our practice and accept the risks inherent in following our convictions.
6. We develop our programs to promote social justice, especially for groups that have been historically disenfranchised.
7. We strive to understand the relationships among culture, curriculum, and practice and the long-term implications for ecological sustainability.
8. We model thoughtful inquiry as a basis for sound decision-making.
I would ask PSU to put some value into being guided by and pursuing scientific research. By making the above their guiding principles, they appear to be subordinating the proper role of the university to a secondary social agenda.
Guiding Principles:
1. We create and sustain educational environments that serve all students and address diverse needs.
2. We encourage and model exemplary programs and practices across the life span.
3. We build our programs on the human and cultural richness of the University's urban setting.
4. We develop collaborative efforts that foster our mission.
5. We challenge assumptions about our practice and accept the risks inherent in following our convictions.
6. We develop our programs to promote social justice, especially for groups that have been historically disenfranchised.
7. We strive to understand the relationships among culture, curriculum, and practice and the long-term implications for ecological sustainability.
8. We model thoughtful inquiry as a basis for sound decision-making.
I would ask PSU to put some value into being guided by and pursuing scientific research. By making the above their guiding principles, they appear to be subordinating the proper role of the university to a secondary social agenda.
Labels:
public schooling
Subscribe to:
Posts (Atom)
