Wednesday, December 21, 2005

Analogous Situations with Educational Systems


If our understanding of systems theory is correct, we ought to be able to analyze the misconceptions in economics as structurally correspondent assimilations to particular kinds of understandings of living systems. Educational, psychological, economic, social, biological systems are all subclasses of the general class of living systems; they therefore should have some general properties in common as well as more specific properties that distinguish then one from another.

For example, the exchanges between the biological organism and its environment are material exchange whereas the cognitive exchanges between the subject and the subject's object of thought are purely functional with no material "content." Both are systems of exchange so the correspondence is between the organism/environment and the subject/object. In social systems, the environment or object is itself another living system so the exchanges are not between an active agent and a passive object but between two active agents. In economics, the buyer seller relationship that constitutes an economic exchange is the correspondent structure. Thus, one of the most basic relationships we might examine in people’s economic thinking is whether both sides of the exchange are understood as mutually constituent aspects of a single interactive exchange or whether exchanges tend to be understood in terms of one-way relationships. Don Boudreaux mentions a myth in which people think "8. Prices and wages are arbitrarily set by businesses.". Do not consumers have a role? Apparently not. So, we might investigate the hypothesis that a person who expresses this view is seeing the dynamics of only one side of the relationship, namely the choices of sellers in the setting the price without seeing the reciprocal role of buyers' choices in what they are willing to pay. In fact, the generic form of the buyer/seller relationship is that both are exchanging something of lessor value for something of greater value. Whether the values are embodied in products or money or time or opportunity costs can all be generalized to exchanges of values. If this understanding is not general as the form of economic understanding, then other misperceptions may arise as well. For example, in the case in which people think if tolls are charged as congestion pricing, road usage is no longer free. In this case, they don’t see the value of drivers time as a cost. Hence, congested roads are free of any exchange of value on their part.

The task then is to investigate these understandings and explain, from a systems point of view, the exact nature of these assimilations and how they develop such that they produce the misconceptions and the more advanced forms of understanding of economists.

This kind of systems analysis seems particularly challenging because the tendency is to stop with the expression of a misconception and simply dismiss it as wrong while then proceeding to explain the correct interpretation. We do that all the time with students. We essentially ignore their system of understanding. We only correct their mistakes, we do not try to understand them. Some mistakes are, of course, simple mistakes but in education, misconceptions are generally cognitive systems that are inadequate according to an adult standard of validity but these systems of understanding are natural and useful, even adequate, methods of understanding from the students point of view.

Now if we really do intend to promote a change in thinking, we must begin with the system as it is currently operating in the learner and work with it to promote its change, not with simply attempting to impose our completed form of understanding on the learner. There is a caveat to this principle. When the learner's system of understanding is close to developing the form of understanding we want them to have, our explanations can have great effect as Inhelder et al found in their learning studies. It is when there is a large lag between the learner's system of understanding and the one we are teaching that verbal methods are insufficient and ineffective, perhaps even negative in their effects. In that case, the emphasis on verbal presentations seems to produce memorized effects in the learner (when the learner is so motivated) rather than change in understanding. Since the learner can "get the right answer", the problem appears to be solved and their may be no cognitive discomfort as in the case when a response doesn't seem to work.

We found this verbal learning when we investigated students' understanding of place value. Verbal instructions were often memorized as procedures to follow such as in borrowing operations during subtraction. In other words, the importance of collecting economic misperceptions and misconceptions is not to correct them but to see that they form the raw data for beginning research into the system of thought that produces the notions so as to understand how the system operates the way it does and thereby to appear to distort the facts.

Is African Poverty Caused By Their Policies or By Our Stinginess?




We often hear people advocating more aid for poor African nations (Bono and now Blair). It seems compassionat but the assumption seems to be that the cause of their poverty can be cured or at least ameliorated by more foreign aid. Don Boudreau, economist at George Mason, points out what is commonly known—that we have given African nations almost a half trillion dollars in aide over the last forty years but these impoverished nations are no better off. Why then do people focus on us giving more aid as if we are stingy and at fault for the plight of poor nations?

Aid clearly does not alleviate, poverty. Its real effect seems to be to support repressive government. The alternative policy would be to put pressures on corrupt governments to mend their ways. But advocating change in nations’ repressive economic and political policies seems heartless, like blaming the victim. Is that is why we are so reticent in discussing the root causes of their poverty or is it simply misunderstanding of the principles of economic development?

Sunday, December 04, 2005

What To Do Regarding Misperceptions About the Poor



A Federal Researve Bank of Dallas annual report "By Our Own Bootstraps," 1995 Annual Report. FRB, Dallas lists some of the gloomy and common misperceptions folks commonly hold regarding the poor. See if these aren’t familiar:

o The rich are getting richer, and the poor are getting poorer. Most of us are getting nowhere.
o Upward mobility is the privilege of a select few, those lucky enough to win at life's lottery.
o The middle class is vanishing.
o Today's 20-something job-seekers face meager prospects and may be the first Americans in history not to live as well as their parents.

The report discusses the factual evidence that shows these perceptions to be largely false; they are misperceptions of the situation of the poor. Here’s why:

This picture of the income distribution would be useful if America were a caste society with rigid class lines keeping those in the bottom today there tomorrow. But if ours is not a caste society, such statistics tell us virtually nothing—particularly about opportunity. By nature, opportunity is personal, an assessment of how well-off you can be tomorrow relative to today. Even the most sophisticated income distribution studies fail to tell us what we really want to know: are most Americans losing their birthright—a chance at upward mobility?

Tracking individuals' incomes over time gives a startlingly different view of the forces shaping America's income distribution. Let's begin with the people who were in the bottom fifth of income earners in 1975. The conventional view leads us to think they were worse off in the 1990s. Nothing could be further from the truth. In the University of Michigan sample, only 5 percent of those in the bottom quintile in 1975 were still there in 1991.


This account doesn’t come from a newspaper columnist or politician. I chose this report because it was written by the vice president of the Dallas Federal Reserve Bank W. Michael Cox, and another economist, Richard Alm. Mr. Cox conducted the research, and in his position, he obviously must produce carefully documented findings and conclusions. He must be objective, factual, research based in his conclusions since holding his highly respected position requires that he be an economic scholar and leader for everyone he serves no matter their ideology or politics; his job depends on how well he develops objective conclusions derived from honest facts. He dare not distort things for ideological purposes, and he has no incentive to do so. He can be considered a solid source if we want the truth of the matter from an objective scholar with advanced knowledge of economics. So, having established his credibility, here’s more of what they say:

Even more important, a majority of these people had made it to the top 60 percent of the income distribution—middle class or better—over that 16-year span. Almost 29 percent of them rose to the top quintile. This is a far cry from the popular vision of a society in which the poor are getting poorer. In fact, the evidence suggests that low income is largely a transitory experience for those willing to work, a place where people may visit but rarely choose to live…(I omitted his charts.)

There's further evidence that being in the low-income bracket isn't, for a large majority of people, permanent. Less than 0.5 percent of the sample showed up in the bottom quintile every year from 1975 to 1991.[3] Nearly a quarter of those in the bottom tier in 1975 moved up the next year and never again returned. More than three-quarters of the lowest 20 percent in 1975 made it into the top 40 percent of income earners for at least one year by 1991. In fact, the poor made the most dramatic gains in the income distribution. Those who started in the bottom quintile in 1975 had a $25,322 average gain in real income by 1991. In the top quintile, the increase was $3,974. In other words, the rich have gotten a little richer, but the poor have gotten much richer. (I omitted another chart.)

The patterns are similar in other quintiles. Among the second poorest quintile in 1975, more than 70 percent had moved to a higher bracket by 1991—with 26 percent going all the way to the top tier. From the middle grouping, almost half of the income earners managed to make themselves better off. A third of the people in the second highest quintile made it to the highest fifth during these 17 years. All through the University of Michigan data, there's a consistent, powerful thrust…salary income was primarily responsible for pushing people upward in the distribution, indicating that work, not luck, is the widest path to opportunity. Ours is not a
Wheel of Fortune economy.


Now these are old data and conditions may have radically changed. There have been follow-on studies showing the mobility not as great. But I strongly doubt that the factors of economic mobility have somehow stopped functioning altogether to now make the misperception suddenly become true but we need research to know for sure. The factor of mobility is apparently real and measurable. But besides more research into income mobility, there is something else from a research point of view that would be quite interesting and useful. I would find research into the prevalence and stability of the misperceptions regarding the income levels and mobility not only interesting but then leading us to investigate why and how these misperceptions arise and present stable phenomena. It would be useful research because the misperceptions are probably strongly correlated with political ideology. So if that is true, what seems to me to be of even great importance would be some kind of research into the sources and origins of these misperceptions.

This kind of misperception research has been conducted in many other areas of human thought. Researchers particularly have found stable, pervasive misperceptions by naïve people in the area of science, geometry and math. Usually research into misperceptions uses children at various ages as a source of data. Children are shown a situation often involving a change in some aspect of the situation. The researcher poses the change as a problem to be explained and interviews the children about what they see and how they explain the changes they see.

For example, yesterday on NPR there was a discussion about the gender differences that relied on research into perceptions of the tilting jar. The tilting jar task has a well established research history. One way to trigger this phenomenon is to cover, or even to leave uncovered, a jar partially filled with water. Tip the jar. Put it aside and then ask students what they saw regarding the change in the water level. Students are shown or asked to draw pictures of what they saw regarding the change in the water level. Some people have a great deal of difficulty seeing that the water level remains parallel to the horizon; for others, they see it as horizontal no matter how the jar is tipped. There are quit remarkable differences in the ability to see the water level objectively as always horizontal. And differences appear between male and female. So, what is the source of these differences and what should be done about them as a matter of public ed policy? Hence the NPR radio show.

Now the researchers don’t simple measure differences in perceptions, some being correct and others more or less incorrect (naïve people do not draw the water line as horizontal when the jar is tipped). Researchers are interested in the underlying conception that forms the ability to objectively see and record the water level. It is the underlying ability that is key and that ability can be more precisely identified as the particular concept that is at work. That conceptual ability is that of the system of horizontal and vertical spatial axes that appears in human thought around the late elementary grades. In the developing child, the sense of up and down appears very early if not a birth but the representation in thought of up and down requires the development of some cognitive machinery that puts together the mental system of axes so that it can then be used to interpret situations factually, see the facts objectively consciously in the mind’s eye, and re-present the facts in thought when the objects are no longer present.

The basic principle here is that concept and percept go together. What we see depends on what we see with. So the challenge for the researcher is to study and eventually explain mentally of exactly what the system of spatial axes consists once it is fully formed and how it arises and develops in the mental activity of human thought. This kind of study of the origins of various kinds of scientific knowledge, that is, how knowledge as a phenomenon of the human species arises in the form of human thought about the world, forms the science of epistemology (as opposed to the philosophic version of epistemology).

Now this same kind of epistemological research into the origins of scientific thought has not yet been applied to economic conceptions. When we find economic misperceptions that are common and stable in spite of the facts, these percepts could give rise to research into the form of the concepts they depend upon. In the case of the income of the poor, we should look at the concepts people implicitly or explicitly use to arrive at their view of the situation. We can take the experet economist views and compare them to naïve views. We could investigate how what people see regarding income of the poor is shaped by what they see with. Their percepts are controlled by their concepts. So an investigation into misperceptions actually investigates what misconceptions are producing the misperceptions. Misperceptions are only the outward indicators of a specific conceptual ability at work so it is the conceptual abilities that will prove important.

In the income misperceptions we may have some clues what these differences in conceptual ability might consist such that they could lead to hypotheses we might investigate. The conceptual differences between those holding the misperceptions and the economists’ views may lie in whether the income groups are understood as descriptive of a fixed group of people that comprise an income level or whether the income levels represent varying membership. Are the poor as a group represented by an income level or as a group with more or less income mobility.

There seems to be an essential difference between whether income is seen properly as a descriptions of a fixed group of people or whether the income of people should be understood as is a variable measuring change or mobility across income levels. If income of the various levels is a fixed property of a single group of people, then people in the low income group aren’t progressing because the low income level hasn’t changed much. But if income is a variable of people such that they move around the income levels, then mobility rather than the income levels becomes the essential factor describing the poor since an income level at different points in time represent a different groups of people. From this point of view of income mobility, the poor are much better off because they aren’t conceived as a group of people represented by the bottom income level.

The misperceptions about income may well depend on the notion that each of the income levels represents a single group of the same people at various points in time. The "over time" is important because it is essential to the idea of income mobility. It is necessary if we want to shift from a static description of the states of people at various income levels to measuring changes in people’s income over time, income mobility as a factor. When scientific description advances beyond static descriptions of properties to reconceptualize descriptive concepts as changes in the concepts, that is, as variables, then these changes over time can be organized into larger explanations, that is, into causal explanations that transcend simple description. They become part of theories. For example, we can teach young children to understand mass, energy, and speed as descriptions of objects but helping the to organize them as variables into an equation for energy and its equivalent to mass times the speed of light squared is conceptually much more advanced and difficult. They same relative difficulty levels may exist in economics in understanding the income of people versus understanding changes in income level over time particularly as to its causes.

From the public policy perspective if we want to do something about poverty, this kind of research could be important for it makes all the difference in the world whether we base our policies on a view of the poor as a fixed group of people described by the lowest income level or whether we base our policies on the variable of income mobility of the poor. The first view of fixed position probably leads to entitlements policies, job guarantees, and the welfare state while the second view may lead to expanding opportunities by reducing barriers like minimum wage, licensure requirements, etc. Research could be important in understanding these policy differences both of which share a common goal but which have radically different means.

Since we find glaring differences among people in their economic perceptions of the very same situations, we might suppose that differences in conceptual machinery exists. Unless there is no objective reality and we deny the fact that human scientific activity makes progress over time in more objectively understanding reality, we have to postulate that some perceptions are obviously closer to the objective truth of a situation than are others. Why and how do these subjective misperceptions develop into more objective perceptions as they appear differently in people’s thought such that they produce differences regarding the same reality being addressed? The psychological fact that these misperceptions might exist in the face of contrary economic facts indicates something very powerful is at work shaping and controlling people’s perceptions of what they believe to be fact. By extending misconception research from other areas into the area of economics, we can formulate hypotheses that the differences in perceptions have to do with the adequacy of the conceptual machinery people use to see and record their perceptions. We can study the development of the conceptions that are necessary in objectively understanding an economic situation.

Since something other than facts shape people’s perceptions, namely concepts, changing their economic thought means promoting their concept development, not correcting their facts, a quite different approach. Without conceptual change, people holding misperceptions not only will but must continue to believe in their misperceptions regardless of more facts simple because their facts depend on their concepts. Even though it is most likely that most people have never been given the facts about poverty, the widespread nature of the misperceptions suggests that the problem lies with their concepts, not their perceptions of facts. As result, we should not expect their convictions to be easily changed by facts. Their views may exist in spite of the objective facts, not because of the facts. But w cannot know if this is true in economics as it is in science until research is conducted.

Most ideological disputes, most contentious public policy issues, most political rhetoric apparently arises over very different but very stable conceptions and perceptions. Over time these may change with intellectual growth but we don’t know yet what exactly causes change over time. For example, recently there was an Israel/Palestine debate hosted by Harvard between Chomsky and Dershowitz. It presented such a challenge from the point of view of differences in perceptions that one student did some fact checking, and his results were reported thanks to Power Line: Webber fact-checks Chomsky. These kind of debates usually remain at the level of throwing out assertions of perceptions of the situation. It’s not really possible to dig into the conceptual machinery of the participants, and the complexity of human thought and the situation would make it impossible to do so in a debate, but how fascinating it would be if a researcher could isolate one difference in perception among people and investigate the conceptual differences that account for the differences in perceptions of the facts. Here in the Harvard debate are two scholars well along in their careers with radically different perceptions. How can that be? What exactly about their differences in conceptual machinery is it that leads to their vast differences in perception?

Less spectacular are the misperceptions the economists regularly note regarding free trade, minimum wage, licensure, drug safety, welfare, and poverty. (I find Don Boudreaux one of the best at identifying and setting misperceptions aright.) Rather than arguing these issues themselves, each of these issues could raise interesting research possibilities into how misperceptions arise in people's thought first as misperceptions and then with more advanced thought as more objective perceptions. We surely can ask whether "the rich are getting richer, and the poor are getting poorer" and "Most of us are getting nowhere" are true but once the economists have conducted the research and produce objective results, then a new type of research into the question of why these myths persist becomes important. It is the only way I know of that holds the possibility of resolving contentious economic issues.

Saturday, December 03, 2005

Iraq Perceptions and Research




The issue of how to deal with terrorism and evil regimes in the world seems to provoke strong polarized opinions. With these often come misperceptions of the actual state of affairs. People seem to selectively use and distort information to support their position. We see this when we ask people whether Iraqi citizens are better or worse off than they were before the iraqi war; the split in opinion seems clearly along ideological lines. Some perceive, often strongly believe, the Iraqi’s are worse off. Their living conditions are worse, they have poorer life because of the destruction of their infrastructure, they have no economy or income, and the war against the terrorism is going badly. Of course on the other side, people paint a picture of improving living conditions. Their perceptions of conditions in Iraq are exactly the opposite. With this difference of opinion about conditions, what seems interesting and important might be the source of people’s information, whether they seem to have actual facts that support their judgments. Sorting this all out is of course a highly charged task, so not only do we have the challenge of locating the facts of the situation but then finding facts and explanations for why such large difference of opinion might exist, some in spite of the actual facts and other because of them.

Now we don’t make progress in reaching agreement by throwing assertions at each other but research using objective methods provides real progress. We should be able to conduct research and find facts that we can all agree upon and then do the same into the secondary problem of why those how misperceive the facts hold those views. For example to establish the facts of the condition of Iraq, IBD reports that before 2003 there were no independent media outlets in Iraq. Today there are 44 commercial TV stations, 72 radio stations, and over 100 newspapers. Are these objective facts reflecting the Iraqi condition? What kinds of indicators should we choose to investigate? Max Boot reports per capita income has doubled since 2003 and the Iraqi economy is projected to grow at 16.8% next year. There are 5 times more cars than under Saddam, 5 times more telephone subscribers, 32 times more internet users, and 14 of the 18 provinces are violence free. And of course we probably can agree that the genocide, the rape rooms, the torture and killing that occurred under Saddam has stopped.

These facts, if they are facts, don’t seem to be determined by one’s ideological predisposition. We could also count terrorist attacks. I don’t have good data here but the facts I have seen report there were 347 attacks on Americans and Iraqis before the January elections and 13 attacks before the later October elections. If we have objective methods for investigating Iraq’s conditions, these should be reasonably objective facts that we can easily agree on but they do seem in scarce supply in the rather opinionated columns and news stories. Typically columnist throw out generalities without substantiating them with citations of research and specific findings.

But regardless of these difficulties in finding objective facts reflecting the truth of the matter, since the average citizen doesn’t have the time or need to conduct a personal investigation into the facts, many may adopt second hand opinions such as those they perceive from the mainstream media. So it is also important to have some sense of the public’s sources for their opinions particularly the rather notorious reputation polling shows the public sometimes has for stupid ideas and misperceptions. This ignorance should not be considered unusual since it is rational on the part of the citizens not to waste their time conducting factual research into things such as foreign policy that are not an important part of their daily lives. People do seem to usually have a rather surprisingly profound and comprehensive knowledge surrounding their jobs but on topics tangential to their daily concerns they often express a degree of ignorance that we should expect. So we might expect citizens to generally adopt views and perceptions regarding foreign policy from easily and widely available media sources. So a search for the source of public misperceptions (if they are misperceptions) could focus on the media and their possible biases.

But again this secondary matter of whether there is media bias seems itself to be in dispute. To assert that public opinion and perceptions of American foreign policy in Iraq are shaped by the media itself depends on an ideologically loaded perception. How would one transcend ideology to factually examine the issue of media bias? Does it exist or is this another case of a fact that depends on one’s point of view?

Again, research based on objective methods should lead the way. One way would be to conduct research that attempts to actually measure bias. For example, the Media Research Center reviewed every Iraq story on the evening news programs of ABC, CBS, and NBC from January through September and found 61% of the stories were negative or pessimistic while only 15% were positive or optimistic, a 4 to 1 ratio of negative to positive. It found 79 of the stories focused primarily on allegations of wrongdoing by American forces in Iraq including Abu Ghraib rehashes while 8 focused on heroism or good works.

Now if we had good research into these questions, the problem then becomes determining whether citizens’ ideological predispositions lead them to read and adopt views consonant with their points of view or whether they are rather passive consumers and adopters of what is presented to them. Can biased media significantly alter public opinion or is opinion rather stable and tied to ideological position?

So it seems there are interesting research questions here made important by the great credence given in our Democratic society to public opinion. Who would do this? I'm not sure. It seems our research function is largely filled by the universities but I wonder whether they could or would do it. Think tanks might be another possibility. Certain non-profits have established reputations for objective research and often draw in researchers from across the ideologic spectrum. But research into the source and origins of perceptions and conceptions that form the basis for public policy seems almost essential if we are to make progress in overcoming fundamental differences of opinion.

Friday, December 02, 2005

Are We Still In a Recession?



People worry about the economy and whether we are making it out of the recession. Are we in a recession? Apparently 43% of Americans think the economy is in recession while only 28% thought so in July (as reported by IBD 12/1/05).

The facts are that America has doubled its total wealth over the last decade, to $50 trillion. The GDP is growing at a 4.3% rate, unemployment is down to 5%, core inflation is at 2%, energy prices are dropping (gas in Sandy is $2.04 per gallon), predictions from many sources are for continued growth next year in the 3-4% range, the dollar hasn’t declined as predicted, and interest rates still remain relatively low.

But the most remarkable fact is that if the economy continues to grow, as is predicted, this will be the first time in history that the economy will have gone through a significant Fed tightening (it raised rates incrementally 12 quarter points) without a slowdown or recession. In a speech Alan Greenspan remarked that the most surprising thing to him during his service as Fed Chairman was the resilience of the U.S. economy. The economy has withstood several disasters without faltering that in the past always dragged it down. He found that a new phenomenon and truly remarkable.

By no economic indicators are we in a recession; just the opposite has been true in recent years. Now why, in the face of positive economic facts and good news about the growth of the economy do so many people think we are in a recession?

Wednesday, November 30, 2005

What Causes the Economic Development of Nations?


What causes the economic development of nations in their rise out of poverty? What prevents the development of poor nations?

The poverty of African nations has seemed intractable but advances in our economic understanding of development is encouraging. What the development research shows is that Africa’s poverty has remained constant in spite of World Bank efforts and foreign aid and that there is another, more important factor that causes development. The Meltzer Commission in 2000 found that the World Bank’s aid projects failed 55% to 60% of the time reports Investor’s Business Daily.

Research into what causes economic development shows that simplistic solutions like more foreign aid fail because most of the money, about 80%, is diverted, stolen, by corrupt governments to maintain repressive regimes. As a result, the per capita GDP of Africa declined by nearly 0.6% over the past 25 year period in spite of $450 billion provided by rich nations, and the per capita income in Africa is 11% lower than it was in 1960. And comparing Africa to South Asia’s progress from 1975 to 200 shows that Africa’s lack of progress occurred in spite of significant foreign aid while South Asia grew. South Asia grew at an annual per capita average of nearly 3% while getting only about 20% of the aid Africa receive on a per capita basis.

This Asian growth is good news because it means that economic development can be triggered without huge development costs; and we now know conclusively how to raise a nation out of poverty. It’s relatively simple (although politically challenging). The difference between economically healthy nations and those remaining poor is the degree to which they possess the fundamentals of economic freedom. The key to lifting a nation out of poverty lies in the establishment of freedom whose essentials are personal choice rather than collective choice, voluntary exchange through free markets rather than through political processes, free entry and activity in markets, and the protection of persons and their property from aggression. In short, freedom that comes from protection from repressive, excessive governments and from crime.

These fundamentals of economic health in societies do not require massive infusions of money, only people’s freedom to engage in economic exchange. Massive foreign aid apparently is even counter productive. Giving money to nations with corrupt governments largely tends to be diverted to supporting the repression, not freeing the people for constructive economic activity.

Economic aid is not necessary for economic development; the only necessary and sufficient factor in producing a nation’s development is freedom. The massive data collected for the Economic Freedom Index by the Fraser Institute clearly verify that the essential factor of development and the way out of poverty is the economic freedom of a nation. The ranking of nations according to their economic freedom directly correlates with their economic development and per capita income. In those cases where nations—for example Hong Kong, Singapore, New Zealand, etc.—have changed their governmental policies from a repressive, high taxation central government to one allowing more freedom of economic exchanges, their development has soared. These are now both free and rich nations with very high per capita incomes. But the African nations with repressive governments are at the bottom of the freedom rankings, and they cannot produce per capita income above the poverty level nor can they even sustain healthy growth. In fact, many are actually economically regressing.

The "foreign aid" that can actually help poor nations is to give people freedom, not money. And we can promote freedom not only at no cost to rich nations but at a net gain to all nations both rich and poor. Free trade never hurts anyone and always helps everyone. We can eliminate poverty through free trade with these nations. Without it, poor nations cannot develop.

And there is evidence that free trade puts pressure on repressive governments to reform. For example, in China a change in government policy is happening where its own economic activity, not a concern for human rights, has caused the Chinese government to expand property rights. If we truly want to aid poor nations not only can we open trade with them, we can also put direct pressure on their governments to reform, and when there is a moral imperative, to encourage the overthrow of a repressive regime.

The path out of poverty by establishing economic freedom and trade is now clearly established, freedom works, and enlightened leaders no longer need to rely on the misconception that more monetary aid will somehow help. We can now see how it instead tends to reinforce the poverty of nations. What causes the economic development out of poverty of nations is freedom, not foreign aid. The only foreign aid that works is to give people freedom, not money, if we want to help them grow and develop.

Tuesday, June 14, 2005

Plugging the Book on Rasch



For anyone really serious about understanding at least the basics of the Rasch model, Trevor Bond and Christine Fox wrote the best book, Applying the Rasch Model, I know of so far for educators. I say "so far" because the book is actually written for those scholars who are conducting Rasch analyses and so it is geared to how to use it. It means it is a bit long and technical for an educator who just wants to understand general principles. None-the-less, it has some readable chapters on the basics that are perhaps the best source so far for a simple introduction to the Rasch model.

Trevor Bond is a remarkable teacher educator and researcher from James Cook University in Australia with a world wide reputation and a huge vita of work in developmental theory, Rasch modeling, and leadership in all kinds of organizations, scholarly journals, and conferences. He is currently on sabbatical as Professor at Hong Kong Institute of Education in the New Territories where he provides academic leadership to the Dept of Educational Psychology, Counseling and Learning Needs. He is a popular speaker at conventions, and several times in the past, he has graciously agreed to come to Portland for conferences we organized. One of our conferences was a joint conference with the NWREL focused on the Rasch model.

All this to say, Rasch is an extremely important research tool that every grad student in education should learn. Since it is replacing norm-referencing as a method of test development and calibration, it is very important for educators to understand. Every educator should understand the basics of measurement even if the technical details are a bit difficult to grasp quickly.

For example, understanding the power of the Rasch model makes clear some of the disappoints we have with the ODE such as in the attempt by the math specialist to develop math problem solving tests. Another example. For teachers, the use of a simple matrix of items by students showing the pattern of correct responses that Trevor describes in his book can easily be done by any teacher for any test they give, and should be. For educators in general, they should understand how to interpret test scores, use a fit statistic for an item or a student, and interpret the error size and its placement on the logit scale. Sadly, these essential aspects of professional knowledge of assessment are lacking in most if not all teachers and administrators. The simplistic views of testing seems to restrict reform efforts to inside-the-box notions of testing as demanding and counting the right responses we want from students. John’s exemplary effort to understand the Rasch model far surpasses the level of apathy in understanding measurement I find among educators.

Perhaps someday we might entice Dr. Bond to return to Portland to help us advance our woeful lack of understanding of measurement. We need some serious discussions of the issues, theories, and ideologies surrounding state testing. Wouldn’t that be interesting?

My Rasch Error


John pointed out a mistake in my assertion that the number of items correct doesn’t matter in the score. He is correct.

In your most recent response for item 4, you wrote:

"For a Rasch calibrated test, the student’s score doesn’t depend on the number correct but where on the scale he falls as a result of his responses to the items.

"If he got a lot of hard item correct and no easy ones, he would score higher than another student who answered more items correct but didn’t get as many difficult items correct (although the test results would show a serious misfit for this student meaning something was wrong with the way the student responded).

"It doesn’t matter how many items but which items he answers correctly. We could give a short version of the test and still get the same Rasch score for a student because the items only serve to locate where on the scale the student is according to which items he answers correctly. The scale has nothing to do with the number correct."

Below is something I copied into a Word file from the ODE web page a couple years ago and I did not save the link. It seems to me that what you wrote above flatly contradicts the table below. If not, I am really confused. John

Sample Test, Benchmark 2, Grade 5
2001-03

CONVERTING TO A RIT SCORE—

No. Correct RIT
1 169.0
2 178.8
3 184.9
4 189.5
5 193.3
6 196.6
7 199.6
8 202.3
9 204.8
10 207.2
11 209.5
12 211.8
13 214.0
14 218.4 o
15 220.7
16 222.9
17 225.3
18 227.8
19 230.5 oo
20 233.3
21 236.6
22 240.3
23 244.8
24 250.8
25 260.6
Recommendations for Level Test Placement:
o Likely to meet Benchmark II Standards
oo Likely to exceed Benchmark II Standards

I was flat out wrong when I said the number of items correct doesn't directly convert to a score. In the examples I gave, the two students with the same number correct would have the same score. The one student would have a large misfit statistic showing his answer pattern didn't match the difficulty sequence of the items. Sorry.

I put the question to Trevor Bond who wrote the book Applying the Rasch Model: Fundamental Measurement in the Human Sciences. He replied:

N - the number right - is the sufficient statistic; all convert to the same logit value
Fit statistics reflect the pattern
good fit means the pattern is close enough for Rasch that the N=logit conversion holds

So, John, you are correct. This new way of thinking about test construction is not easy for me, either.

Monday, June 13, 2005

Rasch Response 9 and 10



Here again are more comments by John and my responses.

9. Dick: The odds of getting the extreme items correct changes exponentially; so we convert these percentages into a logarithmic expression and the resulting log scale is linear.

9. John: Whoa Nellie! I’m lost in a logarithmic fog. I don’t understand this well enough even to ask a question but let me try this. Are you saying something like this: we have 10,000 kids take my 10-question exam. Ignoring the success rates for the first 8 easiest problems, let’s jump to the two toughest. Let’s say on the # 2 hardest, only 100 kids get it right. The # 1 hardest is so much tougher that only 10 kids get it right. Is this sort of where you are headed on the log stuff?

Not really since the Rasch analysis figures out the difficult level of each item and the ability level of each student. Perhaps I over simplified the idea of a logarithmic scale. If we know something increases by a set amount each interval, like the distance we travel at a given rate of speed, then we know this is best thought of as an arithmetic function or graph. The graph will show this arithmetic relationship between distance and time as a straight line. But we know some things change differently over time so we need a different model to understand them. If something doubles each unit of time, then we have an exponential rate of change. We have a rate of change that we can model on the exponential series of 2. So it proceed 2 to the first power, 2 to the second power, and on. We know that fruit flies multiply exponentially so that the graph of number of flies per unit of time isn’t flat, it curves upward. We immediately see that we need to imagine an exponential rate of change and so use an exponential growth curve to model the actual rate of multiplication.

Charter schools in the nation do not increase arithmetically but also follow an algebraic curve, like an exponential curve. So we have phenomena that follow different kinds of mathematical relationships or models. The normal distribution has a different rate of change, rather complicated, but one which statisticians see as following certain mathematical patterns which they can formulate.

So to understand various phenomena, we attribute a mathematical relationships to them. Lots of phenomena follow exponential relationships particularly where growth is involved, and so if we want to show these exponential relationships on a graph, we could use an arithmetic scale for each axis but since the rate of change is exponential, why not convert the scale to a logarithmic scale since a logarithm is simply an exponent? So now the intervals on our scale shows arithmetic intervals that are the increase in the logarithm. Place values positions represent exponential values. Each one goes up by one factor or exponent of ten. Since using a log scale matches the exponential rate of change of a growth phenomenon, the graph appears as a flat line and we can better see changes in its basic growth rate. The log graph shows us much more clearly whether the rate of change, whether the exponential rate of change is constant or not. So since we changed our scale, our graph matches the mathematical relationship we are using to model the change we find in reality, and we have an interval scale for an exponential type of relationship.

Lots of things in the social sciences follow logarithmic relationships. In modeling testing data, algebraic relationships, not arithmetic ones, predominate. Anytime there is more than a simple addition of things, then we need mathematical models to match the algebraic relationships of multiple factors. The probabilities of correct responses to test items also is an algebraic relationship. The curve of the chances from everyone getting the item correct to no one getting it correct varies from 0 up to a midpoint and then back to 0. It is not a straight line going up to a peak and then abruptly angling back down. So we need an algebraic model for probability distributions. We transform the probability by representing the probabilities in terms of a logarithm. Now the arithmetic increase in the logarithm expresses the logarithmic change in the probabilities. The scale is transformed into a simple interval scale that accurately models the rate of change.

Fundamental to the Rasch model is this transformation of the probabilities of each item into its logarithmic representation in order to create an interval scale. It doesn’t mean each item’s value is an equal interval away from the next. It is not counting items we are doing. It means we create an underlying logarithmic scale and use it to represent the logarithmic value of the probability of each item. We create a log scale, or a scale of units called logits which we can then convert to something convenient like the state’s logit scale which they call a RIT scale.
10. Dick: Then we can figure the probability for each cell according to the percentages of a correct response in each cell according to the difficulty level of the item and ability of the student. We express this probability in terms of our logarithmic scale.

10. John: It would really help to use examples for this. Maybe you could play off my math exam illustration and show how it might work at this step.

To conclude, you might try thinking of each item as reprenting a certain difficulty level. A test is like using many individual single measure comparisons to the student so as to see if the student is above or below each item. We essentially have an ordered series of test items placed at various points along the interval scale, and we are attempting to find out where in that series the student fits. So he gets the easy ones below his ability level correct and the hard one above his level wrong. We know where he belongs on the scale from the items. The test items are like stick of different known lengths arranged in a series (although the spacing between them will vary depending on their actual lengths) and one by one we compare an object to each stick to see where in the series the student belongs. If the student is above a stick 4 1/2 feet and below a stick 4 7/8, then we know his length is 4 11/16 with a large error width since the student actually could be anywhere between 4 1/2 and 4 7/8.

Rasch modeled measurements are the only method that can currently be used to formulate tests that can provide an interval scale. They mark a huge advance over norm-referenced tests.

Rasch Response 8



More…

8. Dick: The percentages don’t form a linear scale because to get a hard item correct depends both on an increase in item difficulty and an increase in student ability.

8. John: I think your point here is that where the test is designed correctly, you won’t have a situation like I suggested in paragraph 6, where the difference in difficulty is so small between some questions that it doesn’t really take greater ability to get the slightly harder one correct. In other words, if your norming process showed that 64% got the hardest problem correct, and 63% got #2 hardest correct, you have not selected the two hardest problems correctly—you need a more difficult "hardest" problem that substantially fewer than 63% get right, so that there is a significant ability difference or gap between the smartest kids who get the hardest one right vs the next-smartest who get nine correct but get the hardest one wrong.

It doesn’t matter whether two items are close together in difficulty. The analysis will show their difficulty level as close together. For students, since the items are about the same difficulty, the chances of answering either item correctly won’t be different. Getting the slightly easier one wrong and the harder one right, might occur often with students if the probabilities of answering each one correctly overlap so that there isn’t a significant difference between them. In other words, since they are close to each other on the scale, we expect some minor inconsistency.

But we can have many items very close together at one point on the scale if we want increased accuracy. If we have big gaps between items, then we can measure students accurately. So the state actually does pack in many more items at each of its cut scores so as to reduce the error of measurement at that point.

Remember, we aren’t using a norming process. We are using a different model, one based on determining the probabilities of correct responses for each item and converting those to a scale. There are no norms unless we then use our test and scale to see how many students score at various points, or what the average might be for a group of students, or whatever. But these data have nothing to do with the test calibration and design. By freeing the test from dependence on norms, we can create a measurement that is independent of the students or the items we use to develop the test. It simply doesn’t matter how difficult or how easy or how many of each kind of items we have. They simply show up as points on the scale. Of course, in general, we probably want a nice even distribution of the items along the scale so that the scale is equally precise all along it. But we can add or remove items as we like to get the kind of precision we want.

Sunday, June 12, 2005

Rasch Response 7

More on Rasch and again, here’s my original statement, then John’s, then my latest response…

7 Dick: To use these ordered series, we can then put every response into a matrix ordered by the difficulty level of the items and the ability level of students. As a result, we should see a Guttman scaling in which students tend to get all items correct that are below their ability level and no items correct that are above their ability level. From the matrix rows and columns, we also have the percentage of items correct for each student and percentage of students answering correctly for each item.

7. John: I have not seen a Guttman scale, but I think I basically understand this paragraph. It seems to me that the premise of this method is contradictory to my statement in paragraph 4 above—

"But if we use 10 problems, and Smart kid gets 9 right and Smarter kid gets 10 right, we don’t know which one Smart got wrong. Maybe he got a medium-difficulty problem wrong, not the toughest one."

The premise of your methodology here is that most likely we will in fact know which one problem Smart kid got wrong, and it will most likely be the #1 hardest problem.

Remember to distinguish between the calibration process of the test items with the derivation of the scale, and the subsequent use of the items and scale for assessing students after we have the scale and items analyzed. When we start out and first give students some unknown items we think might measure some underlying dimension of student growth, we don’t have any way of knowing anything about whether an underlying scale actually exists or what the difficulty level of the items is or their position and separation along the scale we think we might find.

When we use the Rasch model to analyze test items and students, we must use exact knowledge of which items each student answers correctly and which students answered an item correctly. The Rasch analysis depends on having this information. In other words, it is not a count of correct items we use, it is a huge matrix of each item in their order of correct responses and each student in their order or correct responses that we use. The cells in this (ordered item by ordered student) matrix shows whether each item was answered correctly or not. If the response data show a consistent pattern from no correct to suddenly all correct when we move along either a column or a row, then both the students and the items are following an order of difficulty or ability and we know we can find the underlying scale. It is from that matrix that we can derive probabilities for each cell and the conversion of the probabilities to a logarithmic scale.

Rasch Response 6



Here's more in the Rasch discussion…

6. Dick: Since the items vary in difficulty and students vary in ability, the two foundational data sources we use are the ordinal property of the items and the students’ ability levels. We order the items according to their difficulty—percentages of students getting each item correct— and order students according to their level of ability to respond correctly—percentages of items the student answers correctly.

6. John. Wait a minute. Back to the 10 multiplication problems. So we order them one through ten in difficulty based on percentages of kids who get each one correct. That doesn’t mean the difference in difficulty is the same between the hardest vs. #2 hardest and 6th hardest vs. 7th hardest. What if 64% got the hardest correct, 63% got #2 hardest correct; and 84% got #6 hardest correct, but 90% got #7 hardest correct. Now what do we do? {{I come back to this below.}}

You are correct. "Hardest" is item difficulty. It comes from a Rasch analysis that starts with ordering the items in difficulty according to the percentages of students who get each item correct. That ordering process would mean that the 64% item of #2 put it second in difficulty, the 63% on #2 put it the hardest item, the 84% on #6 put it third in difficulty, and the 90% on #7 made it the easiest item. You stated 64% got the hardest item correct. It’s not the hardest. #2 is. Either you made a mistake here or you are somehow attributing a difficulty to the item separate from its rank among the percentage of students answering items correctly. We don’t get to attribute the difficulty level to the item. The item’s result determines its rank in difficulty.

But I read you to mean a more important point that the intervals between the numbers of correct items are not equal, and you are right. Remember, we have to derive the scale using the Rasch analysis before we have a measure of the items’ difficulties. Then, the item difficulties show up along the scale, and we can see what their spacing is. For your items, we will find a large distance separating some items on the scale and a small distance along the scale between other items. But until the underlying scale is derived, we don’t know what the spacing is since all we have is the rank ordering of the items. The problem is how to get from these ordinal data to the abstraction of an equal interval scale.

To do that, we do two things. We use both the order of the items and the order of the students to determine the probability of each response being correct. For example, a hard item and a low ability student combine to create a very low probability that the item will be answered correctly. If it is, of course, something is not working according to our model. This partially solves the problem you raise that things are not equal among the scores. The second thing we do is convert these probabilities to a log scale. Then the log probabilities form an equal interval scale. This scale’s units are now logits, not counts of correct answers.

Both the individual items and individual students can be measured along this logit scale. It will show us how large the distance is between any two items and let us compare these distances at any points along the scale. That equal interval quality of this scale that is independent of either the students or the items means we would have a powerful tool for measuring student growth or curriculum topics we want students to learn. We could measure how difficult it is for students to learn 7 x 6 versus 2 x 4. Or we could measure either the current level of each student or the growth of a student across time and have it be directly comparable to other grade levels and curriculum

Rasch Response 5


More comments from John…

5. Dick: How then do we get to a theoretic scale if not by using the average and distribution of a group of students on a test of selected items we choose to use in measuring the students?

5. John: I don’t really understand what the theoretic scale is, and I really don’t understand the "equal intervals" idea, unless we are truly measuring weight, distance, etc., rather than giving points for a test that measures a wide range of knowledge and skills. In the decathlon, for individual events we can measure things in equal intervals of distance or time, but frankly I have no idea how the scores attributable to each distance or time are calculated. But the overall points for each event and for the whole thing are certainly not intended to be equal intervals [e.g. the difference of one point between 4,905 and 4,906 vs. 8,905 and 8,906].

We are not "giving points" as you put it for right answers. We are using test items like a one pound item that is a certain weight on a weight scale. We can compare another object to the one pounder to see whether it is heavier or lighter. If we have several different benchmark items, we can find out where an object is by finding out for which items it is heavier and for which it is lighter.

Your decathlon example has two measures. One is of distance which has nothing to do with the decathlon. The other is for a decathlon type of athletic ability. The measure of this ability depends on both how difficult the specific test is, say a 4 minute mile, and the ability level of the other athletes. As far as I know, the underlying scale has never been abstracted using the Rasch model although it could be. Instead, a count of points is used to represent some level of the athlete. And you are right, the counts do not form equal intervals. We can abstract equal intervals using the Rasch model but they do not correspond to counts or the number correct but on the underlying scale.

Rasch Response 4

4. Dick: But for a high ability student who gets most items correct, getting one more correct, a very difficult items, means much more in terms of the distance along the scale. So everyone agrees that the raw scores are not equal intervals along some theoretic scale.

4. John: This is where I start getting confused. In multiplication, I can understand that 10 x 10 is easier than 57 x 79. But if we use 10 problems, and Smart kid gets 9 right and Smarter kid gets 10 right, we don’t know which one Smart got wrong. Maybe he got a medium-difficulty problem wrong, not the toughest one.

If so, let’s say the RIT scale has just a 2-point differential between getting 5 right and getting 6 right, then why should it have a 6 point differential between getting 9 right and getting 10 right, where the one wrong answer might not have been the toughest problem? Why doesn’t it make more sense to give different numbers of points for problems depending on their difficulty [10 x 10 is worth 2 points; 57 x 79 is worth 6 points]?




4. Your interpretation of scores based on counting correct responses isn’t quite correct. For a Rasch calibrated test, the student’s score doesn’t depend on the number correct but where on the scale he falls as a result of his responses to the items. If he got a lot of hard item correct and no easy ones, he would score higher than another student who answered more items correct but didn’t get as many difficult items correct (although the test results would show a serious misfit for this student meaning something was wrong with the way the student responded).

It doesn’t matter how many items but which items he answers correctly. We could give a short version of the test and still get the same Rasch score for a student because the items only serve to locate where on the scale the student is according to which items he answers correctly. The scale has nothing to do with the number correct. Because the items don’t fall on the same place on the scale, if a student answers a moderately hard one correctly and a very hard one incorrectly we know his ability level on the scale falls between the difficulty level on the scale of these two items. In theory, if our test was perfect we would only need the two items that exactly bracketed the student to measure the student.

Once we have the scale, then the scale doesn’t depend on the items as long as we have the difficulty level of each item. New items can be measured with the scale to see where they fall in their difficulty. Then we can use these as items for measuring students’ ability levels. The items, in effect, stand for points on the scale.

The same is exactly true for students. Once we have the scale, we can see where students fall in their ability level. Then we can use them to stand for certain points on the scale just as we can locate items at their point on the scale. The results don’t depend on what items or what students we measure. The scale is permanently fixed.

If we gave total points, then, even if items were weighted as you suggest, we would not longer have an independent scale. We would have a count of responses. Rasch scales don’t represent a simple counting of correct responses. It depends on the difficulty level of the items the student answers correctly. But of course we can only use the scale to measure items or students if we already have the scale developed. So the use of the scale to measure items or students is different than the procedure for abstracting the underlying scale in the first place. With a norm referenced test based on counting correct answers, we don’t have any underlying scale except for the counts of correct answers by a certain group of students and for a certain set of items. That counting procedure is not very objective.

Rasch Response 3


To continue…

3 Dick: To create an anchor for the scale we group all the scores from a group of students and find their average. To create the scale we use their distribution along a count of correct responses for each item and covert these to different scales, raw scores, percentage correct, percentiles of students, standard deviations, etc.

3. John: I like my multiplication example here because in explaining things we can suggest, let’s say, creating a selection of ten two 2-digit by two 2-digit problems. Then we can do what you just explained for each of the ten problems and for an overall score. Am I understanding correctly?


3. I think so. I am not describing anything new here. I was just describing the traditional method of test calibration in formulating an anchor for a scale, not the Rasch model. I’m not sure I was clear that this was traditional test theory.

The point I was making is that for norm referenced tests the anchor for the scale depends on the group of students tested, i.e., their "norm." That is not a particularly good idea. We must have a scale with an anchor that is free from any group. In fact, we don’t want anything about the scale to depend on the particular group of students or any particular items we use. If the scale is to be objective, it must have its own defined, objective properties that exist independently without being an outgrowth of the performance properties of a certain group of students on a certain group of items.

I’m not sure whether I answered your question since I’m not sure I understood what you were getting at.

Rasch Response 1 & 2


John, you posed some ideas in response to my earlier attempt to explain Rasch test calibration. You put some care and time in your reponses so I want to give each of them careful thought. You helpfully numbered them, so here’s some of my statements from my earlier post, your responses, and my new response:

1. Dick: When we measure something, we need a measuring instrument that compares an attribute of an object to a scale. The bathroom scale is the tool you use to compare your weight to a weight scale. The scale has a standard unit and an anchor at 0. Thermometers measure temperature of an object against a Fahrenheit scale of standard degree units and an anchor at boiling and freezing. We take one measurement.

1. John: How is this for an analogy? The decathlon. If we are trying to measure "track and field athletic ability" the decathlon is one way to measure. It seems to me the decathlon is like trying to measure math knowledge and skills, which has lots of ‘events’ that must be measured.

1. One of the principles of measuring things is that we don’t combine several attributes or properties into a single measurement result. If the decathlon measures a single underlying trait of athletic ability, then we could create a valid scale from the results. However, the abilities underlying events are probably diverse such that a person could be of high ability for one skill, say distance running, and low for another, say high jumping. We wouldn’t measure objects by lumping together length, weight, color and then say we have a test of something called physical properties even though there might be a correlation among the items.

The same is true of achievement tests. When they measure many different abilities so as to mix together underlying attributes into some global notion of "achievement" they are not uni-dimensional. Achievement tests often struggle to get items that follow an order of development since there are several developments at work. With norm referencing, the problem doesn’t reveal itself since we don’t look at whether items fall into a consistent order of difficulty for all students.

2. Dick: A multiple choice test item is a measurement tool but it only measures whether a student exceeds a single point on the scale. It can’t measure a range of values, only a single value or answer. So we use many items hopefully at many points along a scale.

2. John: Let’s say we want to create a math test, and one skill we want to measure is multiplying two 2-digit numbers [from 10 to 99]. There are 8,100 possible problems [90 x 90]. We could test kids on all 8,100, but that has obvious drawbacks. We could test just one item, say 10 x 10. But that’s your "one point on the scale." Using just that one item, we know only whether the kids can do that one problem. But whichever one problem we pick, it might range from very easy to very difficult. So getting data from just one problem may not give us a good idea of every kid’s overall ability to do any other 8,099 problems. [A similar analogy for geography would be knowing the 50 state capitals].


2. I’m going to make a suggestion as to why you and I see things differently. I’m leaping off the deep end, here, and this long post may not be useful but I’d like to try to capture something I think important. It may be that differences over Rasch may be due to some rather profound differences in what we each think we are measuring. Those differences are what I’m going to try to explain.

Mathematics as a Body of Knowledge You seem to see the individual math items as the bits or pieces of mathematical knowledge that students should accumulate and therefore the bits that we should test to see if they are present in the student. I read you as saying that if we test for some items doesn’t tell us about whether the other items are present. So the Rasch idea of a scale doesn’t make sense. I’m not certain whether I captured your view properly or not but if someone did think of the math knowledge as made up of items that are learned and accumulated, then math achievement would, in essence, be the total amount of aggregation of the things that students’ learn. That is, from this point of view, which seems to be yours, is that math achievement is math knowledge made up of the accumulation of an aggregate or collection of items that is added by learning each new fact. That is, math knowledge is a body of knowledge built up by adding pieces.

This idea of math knowledge as a body or corpus forms an analogy to a physical body or object in which we can add pieces to the physical object and make it bigger. A physical piece and the object itself are static entities. They can be moved and added together to make a bigger whole. But this analogy of math achievement to the accumulation of pieces to a growing aggregate of things isn’t the way the mind works. If the growth of math knowledge does not work in the same way that physical pieces are added together, we will need a different explanation, theory, metaphor, analogy to describe the growth of mathematical knowledge.

Mathematics as a Living System of Thought The alternative is to see math knowledge more like a living system of mental activity, not a physical body. Math knowledge is more like the development of a biological or mental ability in which things are not received and stored but is a capability that is developed, reinforced, extended, related and integrated with other capabilities. It would mean that a fact like 4 + 3 involves a mental activity that reconstructs the idea of 4 and 3 and the combining of them into 7, something that occurs slowly through a laborious series of activities counting out objects or something that has achieved an almost instantaneous reproduction. But the point is that the fact isn’t a static piece of knowledge but rather a capability or ability of mental thought that has become organized so as to reconstruct the fact.

Physical Versus Biological So we find two views of mathematics, or actually a range of views between two poles. On the one end, we have math knowledge envisioned as math-as-a-body-made-up-of-pieces-learned-bit-by-bit; let’s call this the physicalist view. On the other end, we have math knowledge as math-as-a-living-system-of mental-activities-reproducing-itself; let’s call this the biological view. Physical bodies have static parts, and these parts can be added together to form larger wholes. Living bodies don’t add pieces to make larger wholes. They "grow" by developing their capabilities. So we have two models of mathematical knowledge.

Which of the views we hold of mathematics, whether it’s best understood using a physical metaphor or biological one, obviously determines how and what we measure. For the physicalist, we can’t assume the student has received and retained one fact merely because he has possession of another. So you say if the smart kid gets 9 of 10 right we don’t know which one he got wrong.

Behavior Versus Thought For the biologist, we can assume there is an underlying ability to learn and remember facts, and that means that since facts differ in the difficulty of the relationships they embody, then remembering math facts will differ according to their difficulty level and the developing math ability of the student. That is people with even a somewhat biological view of math achievement assume there is an underlying dimension to math knowledge that is reflected in which items students answer correctly.

For them, it isn’t the items themselves that are the essence of what needs to be counted, and so it isn’t them that are being measured. They are only proxies. They only stand for something deeper. They serve as a way to get at the deeper, underlying dimension.

So from the non-physicalist perspective (my perspective) trying to use careful methods to get closer to an objective measurement of math achievement is what is important. It means attempting to use the math test items to get at the underlying dimension of some kind of developing math knowledge. It means trying to determine a scale of intervals of this underlying dimension. We don’t want to measure actual observed responses of students, their overt behavior. We want to get past behavior to something deeper within student thought itself. In a nutshell, we want to measure the development of thought, not behavior.

Since there is something more to knowledge than just overt responses indicating which pieces of math knowledge the student has learned and accumulated, we want to use the students’ performances on test items to make inferences about the underlying dimension. But the quality of inference depends on the method of inference. If an inference is just a guess, than it is subjective and an inference of little use as to the thoughts of students. If it is based on carefully developed logico-mathematical methods that let us more accurately model and represent what really is going on in the students’ thoughts, then it is a more objective inference as to the nature of student thought.

In short, what we are doing is trying to go deeper into students’ actual test performance on items to get to the underlying math ability that is producing the responses. The underlying ability is more stable and permanent than the facts which it sometimes reproduce incorrectly or which it may use incorrectly due to perhaps a too quick and misinterpretation of the situation posed by a test item. So this underlying developing ability to reproduce math facts and math knowledge in general would appear to be more important in the long run and therefore what should be measured.

Tuesday, May 24, 2005

Can We Make Overachieving a Normal Expectation?


The chart atEduwonk.com shows Roxbury Prep's students and statewide scores by subgroup. So we have an unusual school. Can we expect all schools to be unusual merely because one is an outlier on the distribution curve? Is the message: this school performs well so all schools should be able to do as well? This hope strategy in ed reform won't work. We've got to get the fundamentals right.

Sunday, May 22, 2005

Rasch Versus Norm Referenced Tests



In response to Rob Kremer’s critique of the CIM/CAM and the state tests, I pointed out that the tests are technically brilliant but conceptually bankrupt. I can say this since no validity studies on them have been conducted that I know of. I can also say they are technically very well constructed since the multiple choice tests make use of some of the most advanced Rasch techniques for items calibration and scaling. So we need be careful in our criticism, and that means we need to understand this distinction between validity of the tests and their technical properties.

I don’t think the testing issue matters much because the larger principle is that the testing itself doesn’t cause change and improvement. It simply isn’t a strategy that will produce significant change and improvement in the system. So improving the tests isn’t a reform I would worry much about. None-the-less, the tests do matter and huge amounts of time and attention are devoted to the state’s tests. So we need to understand them. We need to ask questions and insist on responsesresponses like this reader responding to my comment at Rob Kremer:
I do not get the Rasch/RIT scale. I have somewhere read that the RIT scale is alleged to be an interval scale, which I understand to mean that a 10-point difference should be equal whether it is a difference between 190 and 200 or between 240 and 250, just as a one-foot difference in the shot put is the same. I understand you can create equal intervals when measuring how far someone can put the shot, but I have a harder time accepting that when it comes to test results.

The Oregon department of education for the fifth grade sample math test, 2001-02, provided a table converting ‘number correct’ to the RIT score. Total 25 questions. One correct equals 169.0. Two correct jumps 9.8 points to 178.8. After that, for each additional correct answer, the RIT gains trend from over six points to barely over two points at the middle range of 12-13 correct answers, then trend back up [24 correct gets a RIT of 250.8; 25 correct gets a RIT of 260.6, an increase of 9.8 points].

So how can every single one-RIT-point interval be deemed equal?

If a kid gets 24 correct and scores 250.8, there is no way to know which one of the 25 questions he got wrong; but the kid who gets 25 correct gets 9.8 points more. This treats every one of the 25 questions as if it is worth 9.8 points. But that makes no sense. It might sort of make sense if total correct answers was tied to a normal curve. But then the relative difficulty of individual questions does not relate.

Can you explain?


I don’t know how well I can explain it in simple terms but let me try.

When we measure something, we need a measuring instrument that compares an attribute of an object to a scale. The bathroom scale is the tool you use to compare your weight to a weight scale. The scale has a standard unit and an anchor at 0. Thermometers measure temperature of an object against a Fahrenheit scale of standard degree units and an anchor at boiling and freezing. We take one measurement.

A multiple choice test item is a measurement tool but it only measures whether a student exceeds a single point on the scale. It can’t measure a range of values, only a single value or answer. So we use many items hopefully at many points along a scale. To create an anchor for the scale we group all the scores from a group of students and find their average. To create the scale we use their distribution along a count of correct responses for each item and covert these to different scales, raw scores, percentage correct, percentiles of students, standard deviations, etc. There’s another way to create a scale and measuring tool. It starts with two kinds of data that are different from norms based on a group of students and a particular set of test items.

Items vary in difficulty. Students vary in ability to respond correctly. We might be using many items of average difficulty with few that are hard or easy. So we can imagine that getting a few more or less correct for the average student, even though the count of correct answers changes, doesn’t mean much in terms of movement along some theoretic scale. The student may still be almost average. But for a high ability student who gets most items correct, getting one more correct, a very difficult items, means much more in terms of the distance along the scale. So everyone agrees that the raw scores are not equal intervals along some theoretic scale. How then do we get to a theoretic scale if not by using the average and distribution of a group of students on a test of selected items we choose to use in measuring the students?

Since the items vary in difficulty and students vary in ability, the two foundational data sources we use are the ordinal property or the items and the students’ ability levels. We order the items according to their diffriculty—percentages of students getting each item correct— and order students according to their level of ability to respond correctly—percentages of items the student answers correctly. To use these ordered series, we can then put every response into a matrix ordered by the difficulty level of the items and the ability level of students. As a result, we should see a Guttman scaling in which students tend to get all items correct that are below their ability level and no items correct that are above their ability level. From the matrix rows and columns, we also have the percentage of items correct for each student and percentage of students answering correctly for each item.

The percentages don’t form a linear scale because to get a hard item correct depends both on an increase in item difficulty and an increase in student ability. The odds of getting the extreme items correct changes exponentially; so we convert these percentages into a logarithmic expression and the resulting log scale is linear. Then we can figure the probability for each cell according to the percentages of a correct response in each cell according to the difficulty level of the item and ability of the student. We express this probability in terms of our logarithmic scale. So the odds of getting an item correct right at the students ability level is 50% while expressed as log odds on the logarithmic equal interval scale, its probability is 0. Getting a harder items correct has lower odds and getting easier items correct. An item in a cell at the 55%tile probability has a log odds of +0.20. One at the 95% has a log odds or +2.94. So many cells for a low ability student will have low odds of her producing correct responses in those cells while the corresponding cells for a high ability will have much higher odds of getting those same items correct. The log odds intervals are called logits or Rasch units or the scale can be transformed to more convenient units and given another name.

But to compute the difficulty level for an item, we need to know the ability level of students. To know the difficulty level of students we need to know the items’ difficulty levels. We can’t figure one without first knowing the other however there is an easy way to do it with a computer. What we do is estimate one and compute the other. Then we reverse the process. Then we keep iterating the process and each time watch our results as the computed levels for each item and each student keep refining themselves until their change is negligible. Now these log odds creating our scale to measure either the item difficulties or student ability levels don’t depend on anything further; for example, we haven’t computed averages or distributions. All we did we express the probabilities of being correct for each response as a function of student ability and item difficulty. It doesn’t matter what kind of distribution we have of students or items. We end up with the difficulty levels for every item and for every student.

With an independent scale, we can add more items or more students and the scale won’t change since it doesn’t depend on any sample of students or any the set of items. The scale doesn’t depend on what test we use or what group of students we test, and that student and item independence is very useful. We have the ability to measure the difficulty level of each current or any future test items. We could extend the test upwards or downwards. We can omit items or add them since the only thing that is used is the difficulty level of each item and ability level of each student independent of all other items and students. We could tie or versions tests together since we can anchor them to the permanent logit scale. The logit scale expresses probabilities, and items and students might appear anywhere on the scale. They might be clustered or not. The scale exists independently of the individual log odds of each response to items by each student. So we have an interval scale that isn’t affected by items. This means we can add items at the difficulty level we set for a cutoff so we have a very fine grained tool at this point on the scale. Getting another item correct moves along the scale only a tiny bit. It gives us great accuracy.

Now this Rasch model—based on the idea that correct responses are a function of item difficulty and ability level—has powerful use in test development. Since the response patterns form an ordered series they should form a Guttman scale either for the responses of a single student or all students’ responses for a single item, Responses should be all positives until they turn negative in which case they should continue negative. If either the items or the students don’t’ follow this pattern, something is wrong either with the item or with the testing situation of a student. We say the item or the student doesn’t fit the model.
The misfit the model. We can take each of these misfits that are out of place responses and figure how much error variance they produce. An out of order item close to the students’ ability level doesn’t produce much error variance but a response farther removed from the transition point of correct and incorrect responses for a student does. These error variances are called "residuals," and they are used to compute a "fit" score for each student and each item.

When I develop a test, I must throw out or change items that have poor fit scores. Likewise, we cannot use the students test results if his results have a poor fit score. These fit scores tell us whether the statistical results are matching our theoretic model for generating test items. The verify the soundness of the test. We also have error scores and other statistics describing the quality of the test items that are used for calibrating tests.

The ODE, Cathy Brown, didn’t, and wouldn’t, use these tools when the math problem solving tests were developed, so the quality of the test were unknown. They were not calibrated. As a result, the abject failure of the tests didn’t show up until after they were implemented. It makes us wonder whether officials at the ODE understand the Rasch model. The Rasch model is much more powerful than norm referencing but it has nothing to do with the validity of test items. Its statistics can be used to verify certain theoretic ideas we attempt to express in our test items have but it doesn’t generate valid tests. We should encourage the use and disclosure of the Rasch model and its results but our real concern should be the validity of the state tests. More importantly we should question whether standards and testing are even a reform strategy or if they are only indicators of the level of performance.

Wednesday, May 18, 2005

Will HB3162 Improve Education?



I doubt it. Rob Kremer thinks so and he tackles the CIM/CAM and pushes HB 3162 as the cure for the failure of ed reform. We do need a discussion about the CIM/CAM and the tests as to whether they can be the means to reform of our public schools. But the discussion needs to focus on whether any standardizing and testing can served as the road to reform. Why would changing the tests matter? The strategy hasn't changed. HB3162 doesn't change or improve the existing strategy of reform. It retains the false hope that testing somehow matters. HB3162 only replaces "their" tests with "his" tests. It doesn't change the system.

Standards and testing aren't a strategy for reform. They only describe the performance of the system. And they reveal the massive failure of the system. Of the 75% of our kids who graduate, only 1/3 achieve the CIM. That means the system has a sucess rate of only 25%. That dismal performance should be intolerable and cause for dramatic responses. We have known about this dismal performance for the last 20 years but the legislature is still avoiding the problem. It has no response, no theory of action for how it will bring about rapid, significant change and improvement.

So my criticism of CIM/CAM does not concern the quality of the tests. It lies with the absurd notion that we can make schools change by putting standards and testing in place. Testing more, or differently, or uniformly, is not a strategy for change. It's only measurement. The fundamental inertia and resistance to change that is built into the system by its very design will continue to give us the stable,dismal performance we have known about for the last 20 years.

The tests are irrelevant to change. They are not the problem or the cure. It is the system design that is bad. It is the means of delivery. Until the legislature honestly admits that trying to use tests, any tests, to force the schools to improve is not a strategy for change, we will not have a strategy that produces significant change and improvement in our schools. We must understand and change the design of the system, the means of delivery of education.

Standards and tests, whether flawed or not, are only the messenger of the bad news. The means we are trying to use to deliver public education services is a dynsfunctional system, an old public utility model of a government bureucratic organization with a protected exclusive franchise over all public education services and money that we hope will somehow respond to the need for dramatic change merely because the tests once again show it to be increasingly costly and consistent in its poor performance.

We need to focus on the system design, the fact that the legislature has created a design that guarantees the public schools their revenues, jobs, benefits whether the kids learn or not. The kids' failure to thrive is irrelefvant to the adults' success. The kids are hurt. The adults are rewarded, and HB3162 changes that immoral state of affairs not one whit.

Cultural Competency

Rob Kremer tackles the cultural competency bill in the legislature and the letters to the Oregonian. I looked up PSU's mission statement for their graduate teacher education. Here it is:

Guiding Principles:
1. We create and sustain educational environments that serve all students and address diverse needs.
2. We encourage and model exemplary programs and practices across the life span.
3. We build our programs on the human and cultural richness of the University's urban setting.
4. We develop collaborative efforts that foster our mission.
5. We challenge assumptions about our practice and accept the risks inherent in following our convictions.
6. We develop our programs to promote social justice, especially for groups that have been historically disenfranchised.
7. We strive to understand the relationships among culture, curriculum, and practice and the long-term implications for ecological sustainability.
8. We model thoughtful inquiry as a basis for sound decision-making.


I would ask PSU to put some value into being guided by and pursuing scientific research. By making the above their guiding principles, they appear to be subordinating the proper role of the university to a secondary social agenda.

Tuesday, March 01, 2005

RoguePundit: Blogging as Public Service

Here's an example of public service in promoting public policy discussion by a blogger; he is pushing the civic discussion deeper into issues than we typically get from the the mainstream media. The Rogue Pundit takes a hard look at cities who are Feeling the Pinch in their attempts to at least break on their convention centers.
Yes, convention centers create jobs and draw attendees who spend money in the area, both of which boost the economy.  But, beware the estimates of those trying to sell conference centers.  As I noted a couple days ago, the consultants who develop the estimates for such benefits typically exaggerate, as is being shown in the many under-utilized, money-pit convention centers across the country.

He also lists about 30 "good Oregon blogs." Somebody please classify and rate these blogs. But we can use more.

We need to convince our best minds to get into blogging and share their thoughts. Cascade Policy Institute, where are you when we need you? It would be helpful if you would step forward and get your academic advisors on line. If we are going to have serious public discussions about the important issues we face in Oregon, we need our best and brightest sharing their thoughts in the blogosphere.

Sunday, February 27, 2005

Business Management Books are Misleading for Education Administrators



Some college education courses like to assign popular management books from the business sector for reading and discussion. It's another instance of ITBT at work since reading these books and thinking about good management misses the fundamental assumption that the books must make. It is an assumption that makes these books largely irrelevant to educational administrators. Businesses operate within markets. School districts operate within a government system. Managing a government bureaucracy operating inside a geographical district as the single, protected organization that has the exclusive power to deliver public education is fundamentally different from business management that must develop products and offer them in a competitive market to consumers. Consumers determine the value of the products and therefore the relative success among businesses. For districts, the students and families have no role since revenues do not depend on student learning or parent satisfaction. Districts don’t determine their services; they implement democratically or bureaucratically determined services, and they must comply with standards and processes passed down from legislature, the ODE, the local board, and the feds, not consumers standards or processes of choosing.

The differences between these two sectors create the fundamental differences in the management tasks, values, capabilities, determinants of success and so on that exist between these two sectors. It may be enjoyable for educators to dream about having a district organization that could be managed the way the business books describe organizational management but it is only a dream until education is opened to entrepreneurship and multiple providers. The arrangements in the public school sector for controlling delivery of services using the old public utility model must be changed by the legislatures. Until then, administrative success in education depends on being good at methods of top down management while remaining popular and therefore employable. If a superintendent actually led the way the business management books say to lead, the superintendent’s board and the big ed organizations would work to get rid of him. The public education sector designs static districts. Market sectors produce dynamic organizations.

By law, public education is locked into an old, top-down system that is sustained by the popular myth of democratic governance and the confusion over protecting the institution of public schools districts with protecting public education. This district structure of the institution insures what Ted Kolderie and Virginia Postrel call stasis within a sector. The institutional church needs reformation to preserve the faith in public education. And it is not within the power of the superintendent to change that institutional system. The superintendent’s job is to make it appear to work. And, to lead the charge for more money. The most important goal of school districts (according to what often heads the list of "reforms" educators say is needed in education) is almost always more money to run the organization, not R&D devoted to improving the organizations product, student learning.

In the private sector, business freely evolves because it is a dynamic system that allows new providers to start up or to consolidate or to go out of business. There is a constant effort to develop better products that will earn consumers’ approval and purchase. The goal and key of business success is its product development and marketing by making it available for consumer choice. Profits are not the goal, only the economic indicator of this success. Profits are what happen if the fundamentals are done right. They are an outcome that results from good products. To maintain a successful business organization means to marshal the people and resources in the efforts to develop quality products attractive to consumers. Profits follow. The necessary focus on organizational innovation of new products means business writers have long ago understood the importance of the evolution of business management away from top down management to a networking of cooperative relationships. . Michael Barone writing in National Journal (02/25/2005) sees it happening now in politics, another dynamic sector. In business, he says we are evolving

from an industrial command-and-control America to post-industrial, Information Age, network-connected America. In 2004, our politics followed.


Both political parties, by shifting strategies toward networking and grass roots activism, increased turnout (although Kerry only increased Democratic turnout by 7% versus Bush’s 16% increase in Republican turnout). The product and marketing of the Republicans attracted more consumers’ choice on the ballot. In a dynamic sector, management makes a difference. In a sector such as public schooling, which by legal structure is designed to be static, success takes on a different meaning.

Educational administrators must look like they are leading but they must not alter fundamental arrangements or benefits of those with a vested in the system. It means administrative success is essentially about political acceptability, about making the existing dysfunction system appear to work. And to increase "profits," the administrator mujst attempt to trade promises for increase performance for more money. To market this strategy, the superintendent helps the board blame the legislature for lack of revenues. This strategy deflects attention away from the internal design of the organization that is preventing it from controlling its costs.

For school administrators in the public school sector, taking business management books seriously could be dangerous to their careers but college profs never point this out nor do they admit that business management methods really don’t apply to the protected government districts. They foist a delusion on administrators when they should be helping administrators understand these sector differences and the resulting differences in management methods.

So why don’t these profs help administrators understand the sector in which administrators must operate? Are they evil or just stupid for foisting this dream on education administrators? As Sandy says, never underestimate stupidity.