John, you posed some ideas in response to my earlier attempt to explain Rasch test calibration. You put some care and time in your reponses so I want to give each of them careful thought. You helpfully numbered them, so here’s some of my statements from my earlier post, your responses, and my new response:
1. One of the principles of measuring things is that we don’t combine several attributes or properties into a single measurement result. If the decathlon measures a single underlying trait of athletic ability, then we could create a valid scale from the results. However, the abilities underlying events are probably diverse such that a person could be of high ability for one skill, say distance running, and low for another, say high jumping. We wouldn’t measure objects by lumping together length, weight, color and then say we have a test of something called physical properties even though there might be a correlation among the items.
The same is true of achievement tests. When they measure many different abilities so as to mix together underlying attributes into some global notion of "achievement" they are not uni-dimensional. Achievement tests often struggle to get items that follow an order of development since there are several developments at work. With norm referencing, the problem doesn’t reveal itself since we don’t look at whether items fall into a consistent order of difficulty for all students.
2. I’m going to make a suggestion as to why you and I see things differently. I’m leaping off the deep end, here, and this long post may not be useful but I’d like to try to capture something I think important. It may be that differences over Rasch may be due to some rather profound differences in what we each think we are measuring. Those differences are what I’m going to try to explain.
Mathematics as a Body of Knowledge You seem to see the individual math items as the bits or pieces of mathematical knowledge that students should accumulate and therefore the bits that we should test to see if they are present in the student. I read you as saying that if we test for some items doesn’t tell us about whether the other items are present. So the Rasch idea of a scale doesn’t make sense. I’m not certain whether I captured your view properly or not but if someone did think of the math knowledge as made up of items that are learned and accumulated, then math achievement would, in essence, be the total amount of aggregation of the things that students’ learn. That is, from this point of view, which seems to be yours, is that math achievement is math knowledge made up of the accumulation of an aggregate or collection of items that is added by learning each new fact. That is, math knowledge is a body of knowledge built up by adding pieces.
This idea of math knowledge as a body or corpus forms an analogy to a physical body or object in which we can add pieces to the physical object and make it bigger. A physical piece and the object itself are static entities. They can be moved and added together to make a bigger whole. But this analogy of math achievement to the accumulation of pieces to a growing aggregate of things isn’t the way the mind works. If the growth of math knowledge does not work in the same way that physical pieces are added together, we will need a different explanation, theory, metaphor, analogy to describe the growth of mathematical knowledge.
Mathematics as a Living System of Thought The alternative is to see math knowledge more like a living system of mental activity, not a physical body. Math knowledge is more like the development of a biological or mental ability in which things are not received and stored but is a capability that is developed, reinforced, extended, related and integrated with other capabilities. It would mean that a fact like 4 + 3 involves a mental activity that reconstructs the idea of 4 and 3 and the combining of them into 7, something that occurs slowly through a laborious series of activities counting out objects or something that has achieved an almost instantaneous reproduction. But the point is that the fact isn’t a static piece of knowledge but rather a capability or ability of mental thought that has become organized so as to reconstruct the fact.
Physical Versus Biological So we find two views of mathematics, or actually a range of views between two poles. On the one end, we have math knowledge envisioned as math-as-a-body-made-up-of-pieces-learned-bit-by-bit; let’s call this the physicalist view. On the other end, we have math knowledge as math-as-a-living-system-of mental-activities-reproducing-itself; let’s call this the biological view. Physical bodies have static parts, and these parts can be added together to form larger wholes. Living bodies don’t add pieces to make larger wholes. They "grow" by developing their capabilities. So we have two models of mathematical knowledge.
Which of the views we hold of mathematics, whether it’s best understood using a physical metaphor or biological one, obviously determines how and what we measure. For the physicalist, we can’t assume the student has received and retained one fact merely because he has possession of another. So you say if the smart kid gets 9 of 10 right we don’t know which one he got wrong.
Behavior Versus Thought For the biologist, we can assume there is an underlying ability to learn and remember facts, and that means that since facts differ in the difficulty of the relationships they embody, then remembering math facts will differ according to their difficulty level and the developing math ability of the student. That is people with even a somewhat biological view of math achievement assume there is an underlying dimension to math knowledge that is reflected in which items students answer correctly.
For them, it isn’t the items themselves that are the essence of what needs to be counted, and so it isn’t them that are being measured. They are only proxies. They only stand for something deeper. They serve as a way to get at the deeper, underlying dimension.
So from the non-physicalist perspective (my perspective) trying to use careful methods to get closer to an objective measurement of math achievement is what is important. It means attempting to use the math test items to get at the underlying dimension of some kind of developing math knowledge. It means trying to determine a scale of intervals of this underlying dimension. We don’t want to measure actual observed responses of students, their overt behavior. We want to get past behavior to something deeper within student thought itself. In a nutshell, we want to measure the development of thought, not behavior.
Since there is something more to knowledge than just overt responses indicating which pieces of math knowledge the student has learned and accumulated, we want to use the students’ performances on test items to make inferences about the underlying dimension. But the quality of inference depends on the method of inference. If an inference is just a guess, than it is subjective and an inference of little use as to the thoughts of students. If it is based on carefully developed logico-mathematical methods that let us more accurately model and represent what really is going on in the students’ thoughts, then it is a more objective inference as to the nature of student thought.
In short, what we are doing is trying to go deeper into students’ actual test performance on items to get to the underlying math ability that is producing the responses. The underlying ability is more stable and permanent than the facts which it sometimes reproduce incorrectly or which it may use incorrectly due to perhaps a too quick and misinterpretation of the situation posed by a test item. So this underlying developing ability to reproduce math facts and math knowledge in general would appear to be more important in the long run and therefore what should be measured.
