Tuesday, June 14, 2005

Plugging the Book on Rasch



For anyone really serious about understanding at least the basics of the Rasch model, Trevor Bond and Christine Fox wrote the best book, Applying the Rasch Model, I know of so far for educators. I say "so far" because the book is actually written for those scholars who are conducting Rasch analyses and so it is geared to how to use it. It means it is a bit long and technical for an educator who just wants to understand general principles. None-the-less, it has some readable chapters on the basics that are perhaps the best source so far for a simple introduction to the Rasch model.

Trevor Bond is a remarkable teacher educator and researcher from James Cook University in Australia with a world wide reputation and a huge vita of work in developmental theory, Rasch modeling, and leadership in all kinds of organizations, scholarly journals, and conferences. He is currently on sabbatical as Professor at Hong Kong Institute of Education in the New Territories where he provides academic leadership to the Dept of Educational Psychology, Counseling and Learning Needs. He is a popular speaker at conventions, and several times in the past, he has graciously agreed to come to Portland for conferences we organized. One of our conferences was a joint conference with the NWREL focused on the Rasch model.

All this to say, Rasch is an extremely important research tool that every grad student in education should learn. Since it is replacing norm-referencing as a method of test development and calibration, it is very important for educators to understand. Every educator should understand the basics of measurement even if the technical details are a bit difficult to grasp quickly.

For example, understanding the power of the Rasch model makes clear some of the disappoints we have with the ODE such as in the attempt by the math specialist to develop math problem solving tests. Another example. For teachers, the use of a simple matrix of items by students showing the pattern of correct responses that Trevor describes in his book can easily be done by any teacher for any test they give, and should be. For educators in general, they should understand how to interpret test scores, use a fit statistic for an item or a student, and interpret the error size and its placement on the logit scale. Sadly, these essential aspects of professional knowledge of assessment are lacking in most if not all teachers and administrators. The simplistic views of testing seems to restrict reform efforts to inside-the-box notions of testing as demanding and counting the right responses we want from students. John’s exemplary effort to understand the Rasch model far surpasses the level of apathy in understanding measurement I find among educators.

Perhaps someday we might entice Dr. Bond to return to Portland to help us advance our woeful lack of understanding of measurement. We need some serious discussions of the issues, theories, and ideologies surrounding state testing. Wouldn’t that be interesting?

My Rasch Error


John pointed out a mistake in my assertion that the number of items correct doesn’t matter in the score. He is correct.

In your most recent response for item 4, you wrote:

"For a Rasch calibrated test, the student’s score doesn’t depend on the number correct but where on the scale he falls as a result of his responses to the items.

"If he got a lot of hard item correct and no easy ones, he would score higher than another student who answered more items correct but didn’t get as many difficult items correct (although the test results would show a serious misfit for this student meaning something was wrong with the way the student responded).

"It doesn’t matter how many items but which items he answers correctly. We could give a short version of the test and still get the same Rasch score for a student because the items only serve to locate where on the scale the student is according to which items he answers correctly. The scale has nothing to do with the number correct."

Below is something I copied into a Word file from the ODE web page a couple years ago and I did not save the link. It seems to me that what you wrote above flatly contradicts the table below. If not, I am really confused. John

Sample Test, Benchmark 2, Grade 5
2001-03

CONVERTING TO A RIT SCORE—

No. Correct RIT
1 169.0
2 178.8
3 184.9
4 189.5
5 193.3
6 196.6
7 199.6
8 202.3
9 204.8
10 207.2
11 209.5
12 211.8
13 214.0
14 218.4 o
15 220.7
16 222.9
17 225.3
18 227.8
19 230.5 oo
20 233.3
21 236.6
22 240.3
23 244.8
24 250.8
25 260.6
Recommendations for Level Test Placement:
o Likely to meet Benchmark II Standards
oo Likely to exceed Benchmark II Standards

I was flat out wrong when I said the number of items correct doesn't directly convert to a score. In the examples I gave, the two students with the same number correct would have the same score. The one student would have a large misfit statistic showing his answer pattern didn't match the difficulty sequence of the items. Sorry.

I put the question to Trevor Bond who wrote the book Applying the Rasch Model: Fundamental Measurement in the Human Sciences. He replied:

N - the number right - is the sufficient statistic; all convert to the same logit value
Fit statistics reflect the pattern
good fit means the pattern is close enough for Rasch that the N=logit conversion holds

So, John, you are correct. This new way of thinking about test construction is not easy for me, either.

Monday, June 13, 2005

Rasch Response 9 and 10



Here again are more comments by John and my responses.

9. Dick: The odds of getting the extreme items correct changes exponentially; so we convert these percentages into a logarithmic expression and the resulting log scale is linear.

9. John: Whoa Nellie! I’m lost in a logarithmic fog. I don’t understand this well enough even to ask a question but let me try this. Are you saying something like this: we have 10,000 kids take my 10-question exam. Ignoring the success rates for the first 8 easiest problems, let’s jump to the two toughest. Let’s say on the # 2 hardest, only 100 kids get it right. The # 1 hardest is so much tougher that only 10 kids get it right. Is this sort of where you are headed on the log stuff?

Not really since the Rasch analysis figures out the difficult level of each item and the ability level of each student. Perhaps I over simplified the idea of a logarithmic scale. If we know something increases by a set amount each interval, like the distance we travel at a given rate of speed, then we know this is best thought of as an arithmetic function or graph. The graph will show this arithmetic relationship between distance and time as a straight line. But we know some things change differently over time so we need a different model to understand them. If something doubles each unit of time, then we have an exponential rate of change. We have a rate of change that we can model on the exponential series of 2. So it proceed 2 to the first power, 2 to the second power, and on. We know that fruit flies multiply exponentially so that the graph of number of flies per unit of time isn’t flat, it curves upward. We immediately see that we need to imagine an exponential rate of change and so use an exponential growth curve to model the actual rate of multiplication.

Charter schools in the nation do not increase arithmetically but also follow an algebraic curve, like an exponential curve. So we have phenomena that follow different kinds of mathematical relationships or models. The normal distribution has a different rate of change, rather complicated, but one which statisticians see as following certain mathematical patterns which they can formulate.

So to understand various phenomena, we attribute a mathematical relationships to them. Lots of phenomena follow exponential relationships particularly where growth is involved, and so if we want to show these exponential relationships on a graph, we could use an arithmetic scale for each axis but since the rate of change is exponential, why not convert the scale to a logarithmic scale since a logarithm is simply an exponent? So now the intervals on our scale shows arithmetic intervals that are the increase in the logarithm. Place values positions represent exponential values. Each one goes up by one factor or exponent of ten. Since using a log scale matches the exponential rate of change of a growth phenomenon, the graph appears as a flat line and we can better see changes in its basic growth rate. The log graph shows us much more clearly whether the rate of change, whether the exponential rate of change is constant or not. So since we changed our scale, our graph matches the mathematical relationship we are using to model the change we find in reality, and we have an interval scale for an exponential type of relationship.

Lots of things in the social sciences follow logarithmic relationships. In modeling testing data, algebraic relationships, not arithmetic ones, predominate. Anytime there is more than a simple addition of things, then we need mathematical models to match the algebraic relationships of multiple factors. The probabilities of correct responses to test items also is an algebraic relationship. The curve of the chances from everyone getting the item correct to no one getting it correct varies from 0 up to a midpoint and then back to 0. It is not a straight line going up to a peak and then abruptly angling back down. So we need an algebraic model for probability distributions. We transform the probability by representing the probabilities in terms of a logarithm. Now the arithmetic increase in the logarithm expresses the logarithmic change in the probabilities. The scale is transformed into a simple interval scale that accurately models the rate of change.

Fundamental to the Rasch model is this transformation of the probabilities of each item into its logarithmic representation in order to create an interval scale. It doesn’t mean each item’s value is an equal interval away from the next. It is not counting items we are doing. It means we create an underlying logarithmic scale and use it to represent the logarithmic value of the probability of each item. We create a log scale, or a scale of units called logits which we can then convert to something convenient like the state’s logit scale which they call a RIT scale.
10. Dick: Then we can figure the probability for each cell according to the percentages of a correct response in each cell according to the difficulty level of the item and ability of the student. We express this probability in terms of our logarithmic scale.

10. John: It would really help to use examples for this. Maybe you could play off my math exam illustration and show how it might work at this step.

To conclude, you might try thinking of each item as reprenting a certain difficulty level. A test is like using many individual single measure comparisons to the student so as to see if the student is above or below each item. We essentially have an ordered series of test items placed at various points along the interval scale, and we are attempting to find out where in that series the student fits. So he gets the easy ones below his ability level correct and the hard one above his level wrong. We know where he belongs on the scale from the items. The test items are like stick of different known lengths arranged in a series (although the spacing between them will vary depending on their actual lengths) and one by one we compare an object to each stick to see where in the series the student belongs. If the student is above a stick 4 1/2 feet and below a stick 4 7/8, then we know his length is 4 11/16 with a large error width since the student actually could be anywhere between 4 1/2 and 4 7/8.

Rasch modeled measurements are the only method that can currently be used to formulate tests that can provide an interval scale. They mark a huge advance over norm-referenced tests.

Rasch Response 8



More…

8. Dick: The percentages don’t form a linear scale because to get a hard item correct depends both on an increase in item difficulty and an increase in student ability.

8. John: I think your point here is that where the test is designed correctly, you won’t have a situation like I suggested in paragraph 6, where the difference in difficulty is so small between some questions that it doesn’t really take greater ability to get the slightly harder one correct. In other words, if your norming process showed that 64% got the hardest problem correct, and 63% got #2 hardest correct, you have not selected the two hardest problems correctly—you need a more difficult "hardest" problem that substantially fewer than 63% get right, so that there is a significant ability difference or gap between the smartest kids who get the hardest one right vs the next-smartest who get nine correct but get the hardest one wrong.

It doesn’t matter whether two items are close together in difficulty. The analysis will show their difficulty level as close together. For students, since the items are about the same difficulty, the chances of answering either item correctly won’t be different. Getting the slightly easier one wrong and the harder one right, might occur often with students if the probabilities of answering each one correctly overlap so that there isn’t a significant difference between them. In other words, since they are close to each other on the scale, we expect some minor inconsistency.

But we can have many items very close together at one point on the scale if we want increased accuracy. If we have big gaps between items, then we can measure students accurately. So the state actually does pack in many more items at each of its cut scores so as to reduce the error of measurement at that point.

Remember, we aren’t using a norming process. We are using a different model, one based on determining the probabilities of correct responses for each item and converting those to a scale. There are no norms unless we then use our test and scale to see how many students score at various points, or what the average might be for a group of students, or whatever. But these data have nothing to do with the test calibration and design. By freeing the test from dependence on norms, we can create a measurement that is independent of the students or the items we use to develop the test. It simply doesn’t matter how difficult or how easy or how many of each kind of items we have. They simply show up as points on the scale. Of course, in general, we probably want a nice even distribution of the items along the scale so that the scale is equally precise all along it. But we can add or remove items as we like to get the kind of precision we want.

Sunday, June 12, 2005

Rasch Response 7

More on Rasch and again, here’s my original statement, then John’s, then my latest response…

7 Dick: To use these ordered series, we can then put every response into a matrix ordered by the difficulty level of the items and the ability level of students. As a result, we should see a Guttman scaling in which students tend to get all items correct that are below their ability level and no items correct that are above their ability level. From the matrix rows and columns, we also have the percentage of items correct for each student and percentage of students answering correctly for each item.

7. John: I have not seen a Guttman scale, but I think I basically understand this paragraph. It seems to me that the premise of this method is contradictory to my statement in paragraph 4 above—

"But if we use 10 problems, and Smart kid gets 9 right and Smarter kid gets 10 right, we don’t know which one Smart got wrong. Maybe he got a medium-difficulty problem wrong, not the toughest one."

The premise of your methodology here is that most likely we will in fact know which one problem Smart kid got wrong, and it will most likely be the #1 hardest problem.

Remember to distinguish between the calibration process of the test items with the derivation of the scale, and the subsequent use of the items and scale for assessing students after we have the scale and items analyzed. When we start out and first give students some unknown items we think might measure some underlying dimension of student growth, we don’t have any way of knowing anything about whether an underlying scale actually exists or what the difficulty level of the items is or their position and separation along the scale we think we might find.

When we use the Rasch model to analyze test items and students, we must use exact knowledge of which items each student answers correctly and which students answered an item correctly. The Rasch analysis depends on having this information. In other words, it is not a count of correct items we use, it is a huge matrix of each item in their order of correct responses and each student in their order or correct responses that we use. The cells in this (ordered item by ordered student) matrix shows whether each item was answered correctly or not. If the response data show a consistent pattern from no correct to suddenly all correct when we move along either a column or a row, then both the students and the items are following an order of difficulty or ability and we know we can find the underlying scale. It is from that matrix that we can derive probabilities for each cell and the conversion of the probabilities to a logarithmic scale.

Rasch Response 6



Here's more in the Rasch discussion…

6. Dick: Since the items vary in difficulty and students vary in ability, the two foundational data sources we use are the ordinal property of the items and the students’ ability levels. We order the items according to their difficulty—percentages of students getting each item correct— and order students according to their level of ability to respond correctly—percentages of items the student answers correctly.

6. John. Wait a minute. Back to the 10 multiplication problems. So we order them one through ten in difficulty based on percentages of kids who get each one correct. That doesn’t mean the difference in difficulty is the same between the hardest vs. #2 hardest and 6th hardest vs. 7th hardest. What if 64% got the hardest correct, 63% got #2 hardest correct; and 84% got #6 hardest correct, but 90% got #7 hardest correct. Now what do we do? {{I come back to this below.}}

You are correct. "Hardest" is item difficulty. It comes from a Rasch analysis that starts with ordering the items in difficulty according to the percentages of students who get each item correct. That ordering process would mean that the 64% item of #2 put it second in difficulty, the 63% on #2 put it the hardest item, the 84% on #6 put it third in difficulty, and the 90% on #7 made it the easiest item. You stated 64% got the hardest item correct. It’s not the hardest. #2 is. Either you made a mistake here or you are somehow attributing a difficulty to the item separate from its rank among the percentage of students answering items correctly. We don’t get to attribute the difficulty level to the item. The item’s result determines its rank in difficulty.

But I read you to mean a more important point that the intervals between the numbers of correct items are not equal, and you are right. Remember, we have to derive the scale using the Rasch analysis before we have a measure of the items’ difficulties. Then, the item difficulties show up along the scale, and we can see what their spacing is. For your items, we will find a large distance separating some items on the scale and a small distance along the scale between other items. But until the underlying scale is derived, we don’t know what the spacing is since all we have is the rank ordering of the items. The problem is how to get from these ordinal data to the abstraction of an equal interval scale.

To do that, we do two things. We use both the order of the items and the order of the students to determine the probability of each response being correct. For example, a hard item and a low ability student combine to create a very low probability that the item will be answered correctly. If it is, of course, something is not working according to our model. This partially solves the problem you raise that things are not equal among the scores. The second thing we do is convert these probabilities to a log scale. Then the log probabilities form an equal interval scale. This scale’s units are now logits, not counts of correct answers.

Both the individual items and individual students can be measured along this logit scale. It will show us how large the distance is between any two items and let us compare these distances at any points along the scale. That equal interval quality of this scale that is independent of either the students or the items means we would have a powerful tool for measuring student growth or curriculum topics we want students to learn. We could measure how difficult it is for students to learn 7 x 6 versus 2 x 4. Or we could measure either the current level of each student or the growth of a student across time and have it be directly comparable to other grade levels and curriculum

Rasch Response 5


More comments from John…

5. Dick: How then do we get to a theoretic scale if not by using the average and distribution of a group of students on a test of selected items we choose to use in measuring the students?

5. John: I don’t really understand what the theoretic scale is, and I really don’t understand the "equal intervals" idea, unless we are truly measuring weight, distance, etc., rather than giving points for a test that measures a wide range of knowledge and skills. In the decathlon, for individual events we can measure things in equal intervals of distance or time, but frankly I have no idea how the scores attributable to each distance or time are calculated. But the overall points for each event and for the whole thing are certainly not intended to be equal intervals [e.g. the difference of one point between 4,905 and 4,906 vs. 8,905 and 8,906].

We are not "giving points" as you put it for right answers. We are using test items like a one pound item that is a certain weight on a weight scale. We can compare another object to the one pounder to see whether it is heavier or lighter. If we have several different benchmark items, we can find out where an object is by finding out for which items it is heavier and for which it is lighter.

Your decathlon example has two measures. One is of distance which has nothing to do with the decathlon. The other is for a decathlon type of athletic ability. The measure of this ability depends on both how difficult the specific test is, say a 4 minute mile, and the ability level of the other athletes. As far as I know, the underlying scale has never been abstracted using the Rasch model although it could be. Instead, a count of points is used to represent some level of the athlete. And you are right, the counts do not form equal intervals. We can abstract equal intervals using the Rasch model but they do not correspond to counts or the number correct but on the underlying scale.

Rasch Response 4

4. Dick: But for a high ability student who gets most items correct, getting one more correct, a very difficult items, means much more in terms of the distance along the scale. So everyone agrees that the raw scores are not equal intervals along some theoretic scale.

4. John: This is where I start getting confused. In multiplication, I can understand that 10 x 10 is easier than 57 x 79. But if we use 10 problems, and Smart kid gets 9 right and Smarter kid gets 10 right, we don’t know which one Smart got wrong. Maybe he got a medium-difficulty problem wrong, not the toughest one.

If so, let’s say the RIT scale has just a 2-point differential between getting 5 right and getting 6 right, then why should it have a 6 point differential between getting 9 right and getting 10 right, where the one wrong answer might not have been the toughest problem? Why doesn’t it make more sense to give different numbers of points for problems depending on their difficulty [10 x 10 is worth 2 points; 57 x 79 is worth 6 points]?




4. Your interpretation of scores based on counting correct responses isn’t quite correct. For a Rasch calibrated test, the student’s score doesn’t depend on the number correct but where on the scale he falls as a result of his responses to the items. If he got a lot of hard item correct and no easy ones, he would score higher than another student who answered more items correct but didn’t get as many difficult items correct (although the test results would show a serious misfit for this student meaning something was wrong with the way the student responded).

It doesn’t matter how many items but which items he answers correctly. We could give a short version of the test and still get the same Rasch score for a student because the items only serve to locate where on the scale the student is according to which items he answers correctly. The scale has nothing to do with the number correct. Because the items don’t fall on the same place on the scale, if a student answers a moderately hard one correctly and a very hard one incorrectly we know his ability level on the scale falls between the difficulty level on the scale of these two items. In theory, if our test was perfect we would only need the two items that exactly bracketed the student to measure the student.

Once we have the scale, then the scale doesn’t depend on the items as long as we have the difficulty level of each item. New items can be measured with the scale to see where they fall in their difficulty. Then we can use these as items for measuring students’ ability levels. The items, in effect, stand for points on the scale.

The same is exactly true for students. Once we have the scale, we can see where students fall in their ability level. Then we can use them to stand for certain points on the scale just as we can locate items at their point on the scale. The results don’t depend on what items or what students we measure. The scale is permanently fixed.

If we gave total points, then, even if items were weighted as you suggest, we would not longer have an independent scale. We would have a count of responses. Rasch scales don’t represent a simple counting of correct responses. It depends on the difficulty level of the items the student answers correctly. But of course we can only use the scale to measure items or students if we already have the scale developed. So the use of the scale to measure items or students is different than the procedure for abstracting the underlying scale in the first place. With a norm referenced test based on counting correct answers, we don’t have any underlying scale except for the counts of correct answers by a certain group of students and for a certain set of items. That counting procedure is not very objective.

Rasch Response 3


To continue…

3 Dick: To create an anchor for the scale we group all the scores from a group of students and find their average. To create the scale we use their distribution along a count of correct responses for each item and covert these to different scales, raw scores, percentage correct, percentiles of students, standard deviations, etc.

3. John: I like my multiplication example here because in explaining things we can suggest, let’s say, creating a selection of ten two 2-digit by two 2-digit problems. Then we can do what you just explained for each of the ten problems and for an overall score. Am I understanding correctly?


3. I think so. I am not describing anything new here. I was just describing the traditional method of test calibration in formulating an anchor for a scale, not the Rasch model. I’m not sure I was clear that this was traditional test theory.

The point I was making is that for norm referenced tests the anchor for the scale depends on the group of students tested, i.e., their "norm." That is not a particularly good idea. We must have a scale with an anchor that is free from any group. In fact, we don’t want anything about the scale to depend on the particular group of students or any particular items we use. If the scale is to be objective, it must have its own defined, objective properties that exist independently without being an outgrowth of the performance properties of a certain group of students on a certain group of items.

I’m not sure whether I answered your question since I’m not sure I understood what you were getting at.

Rasch Response 1 & 2


John, you posed some ideas in response to my earlier attempt to explain Rasch test calibration. You put some care and time in your reponses so I want to give each of them careful thought. You helpfully numbered them, so here’s some of my statements from my earlier post, your responses, and my new response:

1. Dick: When we measure something, we need a measuring instrument that compares an attribute of an object to a scale. The bathroom scale is the tool you use to compare your weight to a weight scale. The scale has a standard unit and an anchor at 0. Thermometers measure temperature of an object against a Fahrenheit scale of standard degree units and an anchor at boiling and freezing. We take one measurement.

1. John: How is this for an analogy? The decathlon. If we are trying to measure "track and field athletic ability" the decathlon is one way to measure. It seems to me the decathlon is like trying to measure math knowledge and skills, which has lots of ‘events’ that must be measured.

1. One of the principles of measuring things is that we don’t combine several attributes or properties into a single measurement result. If the decathlon measures a single underlying trait of athletic ability, then we could create a valid scale from the results. However, the abilities underlying events are probably diverse such that a person could be of high ability for one skill, say distance running, and low for another, say high jumping. We wouldn’t measure objects by lumping together length, weight, color and then say we have a test of something called physical properties even though there might be a correlation among the items.

The same is true of achievement tests. When they measure many different abilities so as to mix together underlying attributes into some global notion of "achievement" they are not uni-dimensional. Achievement tests often struggle to get items that follow an order of development since there are several developments at work. With norm referencing, the problem doesn’t reveal itself since we don’t look at whether items fall into a consistent order of difficulty for all students.

2. Dick: A multiple choice test item is a measurement tool but it only measures whether a student exceeds a single point on the scale. It can’t measure a range of values, only a single value or answer. So we use many items hopefully at many points along a scale.

2. John: Let’s say we want to create a math test, and one skill we want to measure is multiplying two 2-digit numbers [from 10 to 99]. There are 8,100 possible problems [90 x 90]. We could test kids on all 8,100, but that has obvious drawbacks. We could test just one item, say 10 x 10. But that’s your "one point on the scale." Using just that one item, we know only whether the kids can do that one problem. But whichever one problem we pick, it might range from very easy to very difficult. So getting data from just one problem may not give us a good idea of every kid’s overall ability to do any other 8,099 problems. [A similar analogy for geography would be knowing the 50 state capitals].


2. I’m going to make a suggestion as to why you and I see things differently. I’m leaping off the deep end, here, and this long post may not be useful but I’d like to try to capture something I think important. It may be that differences over Rasch may be due to some rather profound differences in what we each think we are measuring. Those differences are what I’m going to try to explain.

Mathematics as a Body of Knowledge You seem to see the individual math items as the bits or pieces of mathematical knowledge that students should accumulate and therefore the bits that we should test to see if they are present in the student. I read you as saying that if we test for some items doesn’t tell us about whether the other items are present. So the Rasch idea of a scale doesn’t make sense. I’m not certain whether I captured your view properly or not but if someone did think of the math knowledge as made up of items that are learned and accumulated, then math achievement would, in essence, be the total amount of aggregation of the things that students’ learn. That is, from this point of view, which seems to be yours, is that math achievement is math knowledge made up of the accumulation of an aggregate or collection of items that is added by learning each new fact. That is, math knowledge is a body of knowledge built up by adding pieces.

This idea of math knowledge as a body or corpus forms an analogy to a physical body or object in which we can add pieces to the physical object and make it bigger. A physical piece and the object itself are static entities. They can be moved and added together to make a bigger whole. But this analogy of math achievement to the accumulation of pieces to a growing aggregate of things isn’t the way the mind works. If the growth of math knowledge does not work in the same way that physical pieces are added together, we will need a different explanation, theory, metaphor, analogy to describe the growth of mathematical knowledge.

Mathematics as a Living System of Thought The alternative is to see math knowledge more like a living system of mental activity, not a physical body. Math knowledge is more like the development of a biological or mental ability in which things are not received and stored but is a capability that is developed, reinforced, extended, related and integrated with other capabilities. It would mean that a fact like 4 + 3 involves a mental activity that reconstructs the idea of 4 and 3 and the combining of them into 7, something that occurs slowly through a laborious series of activities counting out objects or something that has achieved an almost instantaneous reproduction. But the point is that the fact isn’t a static piece of knowledge but rather a capability or ability of mental thought that has become organized so as to reconstruct the fact.

Physical Versus Biological So we find two views of mathematics, or actually a range of views between two poles. On the one end, we have math knowledge envisioned as math-as-a-body-made-up-of-pieces-learned-bit-by-bit; let’s call this the physicalist view. On the other end, we have math knowledge as math-as-a-living-system-of mental-activities-reproducing-itself; let’s call this the biological view. Physical bodies have static parts, and these parts can be added together to form larger wholes. Living bodies don’t add pieces to make larger wholes. They "grow" by developing their capabilities. So we have two models of mathematical knowledge.

Which of the views we hold of mathematics, whether it’s best understood using a physical metaphor or biological one, obviously determines how and what we measure. For the physicalist, we can’t assume the student has received and retained one fact merely because he has possession of another. So you say if the smart kid gets 9 of 10 right we don’t know which one he got wrong.

Behavior Versus Thought For the biologist, we can assume there is an underlying ability to learn and remember facts, and that means that since facts differ in the difficulty of the relationships they embody, then remembering math facts will differ according to their difficulty level and the developing math ability of the student. That is people with even a somewhat biological view of math achievement assume there is an underlying dimension to math knowledge that is reflected in which items students answer correctly.

For them, it isn’t the items themselves that are the essence of what needs to be counted, and so it isn’t them that are being measured. They are only proxies. They only stand for something deeper. They serve as a way to get at the deeper, underlying dimension.

So from the non-physicalist perspective (my perspective) trying to use careful methods to get closer to an objective measurement of math achievement is what is important. It means attempting to use the math test items to get at the underlying dimension of some kind of developing math knowledge. It means trying to determine a scale of intervals of this underlying dimension. We don’t want to measure actual observed responses of students, their overt behavior. We want to get past behavior to something deeper within student thought itself. In a nutshell, we want to measure the development of thought, not behavior.

Since there is something more to knowledge than just overt responses indicating which pieces of math knowledge the student has learned and accumulated, we want to use the students’ performances on test items to make inferences about the underlying dimension. But the quality of inference depends on the method of inference. If an inference is just a guess, than it is subjective and an inference of little use as to the thoughts of students. If it is based on carefully developed logico-mathematical methods that let us more accurately model and represent what really is going on in the students’ thoughts, then it is a more objective inference as to the nature of student thought.

In short, what we are doing is trying to go deeper into students’ actual test performance on items to get to the underlying math ability that is producing the responses. The underlying ability is more stable and permanent than the facts which it sometimes reproduce incorrectly or which it may use incorrectly due to perhaps a too quick and misinterpretation of the situation posed by a test item. So this underlying developing ability to reproduce math facts and math knowledge in general would appear to be more important in the long run and therefore what should be measured.