Showing posts sorted by relevance for query complexity. Sort by date Show all posts
Showing posts sorted by relevance for query complexity. Sort by date Show all posts

Friday, May 15, 2009

AP Exams and Complexity

The College Board offers AP tests, which are generally accepted as college credit, in certain subjects. High school courses are offered as preparation for the tests--a very nice arrangement for the College Board. The relationship between grades in the AP courses versus scores on the exams were an issue in this article in the Jacksonville News, viz:

Duval [County public school] students passed 80 percent of their AP courses last year with a "C" or better. But only 23 percent of the national AP exams, taken near the end of those courses, were passed.

The national exam pass rate for public schools was 56 percent.

In other words, students successfully complete a course that is essentially a preparation for a standardized test, and then fail said test. The College Board, which reviewed the results, blames the effect on under-prepared students taking the courses combined with inexperienced teachers. When I read this, I thought it was perhaps a good opportunity to look for evidence of a phenomenon I've said should exist: that standardized tests are more valid for low-complexity subjects of study. Here, complexity means in the computational sense (search for the word in my blog for lots more on the subject). If we assumed all things equal (dubious, but I have no choice) then preparation courses for lower complexity subjects ought to be more effective than for higher complexity subjects. This would manifest itself in the correlation between passing tests and passing AP exams. This is all possible, because the statistics for the school district in question are posted online.

In pure complexity terms, math is low complexity and languages are higher complexity. This is easy to see--math is a foreign language with little new vocabulary and a few rules. Spoken languages have massive vocabulary and many, often arbitrary-seeming, rules of grammar. So if my theory is any good, it ought to be the case that math courses can prepare a student better than language courses for a standardized test, all else being equal. Also, the overlap in students between the two courses is probably pretty good since college-bound seniors will be taking both a foreign language and math. Of course, even learning a foreign language is mostly committing deductive processes to memory, and hence of not the highest complexity (inductive processes would be). So this is a contest between low complexity and, shall we say, medium complexity.

Here are the results:

I debated whether or not to include both calculus courses (clearly, more advanced students are in the BC section). I also assumed that the three languages are equally complex, although in practice Spanish dominates. If a student passed a calculus course, he or she had a 48% chance of passing the AP subject test. For languages (the more complex subject), only 39% of those who passed the course went on to succeed on the exam.

Does this prove anything? Not really--there are too many uncontrolled variables. But it's still fun to push this as far as it can go. If the complexity to difficulty relationship holds, we would expect that the subject with the worst test/course pass ratio would be the most complex subject. Of course, sample size plays a role, so let's agree (before I look at the numbers) that there had to be at least N=50 to qualify. For all tests combined, the average test/course ratio was 29%. Anything lower than that would be lower than average preparation for the test (and higher complexity maybe). The least effective (or most complex) course was a three-way tie, with a 16% conditional probability of passing the AP test, given that the course had been passed. The three subjects were World History, Human Geography, and Micro Economics. These each had enrollments in the hundreds and thousands.

Are these subjects the hardest to test because of complexity? It's easy to guess that the first two might be, cluttered with endless facts and fuzzy theories. Micro Economics is much more like chemistry or physics, one would think. Chemistry scored a low 22%, but Physics B was 58%.

This was fun, but it's hard to really make the case that complexity is the driving force here. It does give me some ideas, however, about comparing difficulty (course pass rate) versus complexity (course to test pass ratio). Meanwhile, I still have the placement test project to try out...

[Update: here's an interesting article about the effectiveness of the calculus AP test]

Wednesday, May 06, 2009

Difficultie, er, Difficulty vs C0mp13xi7y

Sorry for the l337 garbage in the title. I've been absorbing the ideas from Dr. Ed Nuhfer (see previous post) about knowledge surveys, and in particular the idea that complexity and difficulty are orthogonal. That is to say, different dimensions. We could think of an abstraction of learning that incorporates these into a graph. Mathematicians love graphs. Well, applied mathematicians love graphs anyway.
I have used my considerable expertise with Paint to produce the glorious figure shown above. It's supposed to lay bare an abstraction of the learning process if we only consider difficulty (expressed here inversely as probability of success) and complexity. I think it's fair to say that we sometimes think of learning happening this way--that no matter what the task, probability of success is increased by training, and that more complex tasks are inherently more difficult than less complex ones. This may, of course, be utterly wrong.

We encountered Moravec's Paradox in a previous episode here. The idea is that some things that are very complex seem easy. For example, judging another person's character, or determining if another beer is worth four bucks at the bar. So, it may be that the vertical dimension of the graph isn't conveying anything of interest. But I have a way to wiggle out of that problem.

If we restrict ourselves to thinking about purely deductive reasoning tasks, then success depends on the learner's ability to execute an algorithm. Multiplication tables or solving standard differential equations--it's all step-by-step reasoning. In this case, it seems reasonable to assume that increased complexity (number of steps involved) reduces chance of success. In fact, if we abstract out a constant probability of success for each step, then we'd expect an exponentially decaying probability of success as complexity increases because we're multiplying probabilities together (if they're independent anyway).

We could test this with math placement test results. An analysis of the problems on a standard college placement test should give us pretty quickly the approximate number of steps required to solve any given problem (with suitable definitions and standards). The test results will give us a statistical bound on probability of success. Looking at the two together should be interesting. If we assume that Pr[success on item of complexity c] = exp( d*c), where d is some (negative) constant to be determined, and c is the complexity measured in solution-steps, then we could analyze which problems are unexpectedly easy or difficult. That is, we could analyze the preparation of students taking the placement not just by looking at the number of problems they got correct, but the level of complexity they can deal with.

I've written a lot about complexity and analytical thinking, which you can sieve from the blog via the search box if you're interested. I won't recapitulate here.

I'll see if I can find some placement test data to try this out on. It should make for an interesting afternoon some day this summer. Stay tuned.

Update: It occurs to me upon reflection that if the placement test is multiple-choice, this may not work. Many kinds of deterministic processes are easier to verify backwards than they are to run forwards. For example, if a problem asks for the roots of a quadratic equation, it's likely easier to plug in the candidates for the answers and see if they work than it is to use the quadratic equation or complete the square (i.e. actually solve the problem directly). This would make the complexity weights dubious.

Friday, September 18, 2009

Complexity and Reboots

One of two stacks of books on my nightstand. The thin volume on top is on simplicity. Under it is a thick textbook on complexity theory.

Complexity is one of those really profound ideas that is surprisingly simple to access. The basic idea is that the complexity of some set of data is the size of the smallest description. So a one with a million zeros is not very complex because you can write it as 10^1000000. Or 10^(10^6), which shortens it by one character. This idea wouldn't work if all methods of description weren't somehow interchangeable. We're really talking about computer-code-like descriptions, and the argument goes like this: Any computer language that has certain basic functionality can be used to create an interpreter for any other kind of language. So we can switch between one and the other by paying the cost of this translation. It is in this sense that descriptions are language-invariant, and so there is something like an absolute measure of complexity in this formal sense.

It's a powerful idea, even in mundane affairs. I'm a simplicity hawk when it comes to creating new bureaucracy, for example, or creating a web interface. Most things end up with more whistles and bells than are good for them. And the more complex a system is, the more unexpected its behavior can be. This is probably why evolution hasn't invested a lot of "effort" in making organisms live forever.

Ever wonder about that? Why do our bodies suffer senescence? Why didn't we evolve with robust repair mechanisms so that we could go on reproducing ad infinitum? Think of all the effort it requires biologically to start all over again with a baby and create a new specimen capable of reproduction. There must be some good reason.

It's like a reboot. Running computers, unless they're running perfectly stable software, tend to accumulate problems in the machine's state. As I'm sure you've experienced, this can cause the thing to blue-screen and freeze. Rebooting eliminates the accumulated complexity. Same thing with biological reproduction, I assume. I imagine there's a limiting factor that works like this: any self-repair mechanism incurs a complexity cost. That is, it's easier to build a system that can't self-repair than one that can. So you have to add complexity in order to reduce it. At some point, it's self-defeating to try to add more repair ability because it isn't capable of recovering its own cost. I have no idea if this is true, but it's an amusing theory.

What does this have to do with higher education? More generally, it has a lot to do with self-governance. Faculty senates, state and federal legislative bodies, and so forth are run by systems of rules, at least in theory. Over time these accumulate complexity as exceptions are created and new conditions included: think of the tax code. I would hypothesize that it would be a good thing to build a reboot procedure into such things.

Consider a faculty senate, which in a moment of singular passion, changes its bylaws so that all future motions can only be carried with every voting member present and voting for the motion. This is tantamount to self-destruction. On the other hand, the rules and committees could proliferate to such a degree that processes slowed to a crawl. I almost wrote "glacial speeds" but the ice mountains are moving faster these days than most committees I observe.

Saturday, December 10, 2011

Randomness and Prediction

I saw a question at stackoverflow.com asking why computer programs can't produce true random numbers. I can't locate the exact page now, but here's a similar one. The question spooked around in my head all day, despite my head-down work to catch up on paperwork after being away at the SACSCOC meeting (see the tweets). After coming home, I finally gave in and wrote some notes down on the topic. It has applications to assessment, believe it or not.

According to complexity theory, "random" means infinitely complex. Complexity is the size of a perfect description of, for example, a list of numbers. If we are given an infinitely long list of numbers like
1, 1, 1, 1, 1, ....
it's easy to see that it has very low complexity. We can describe the sequence as "all ones."  Similarly, the powers of two makes a low complexity sequence, or any other simple arithmetic sequence. We could create a more complicated computer program that tries to produce numbers that are as "mixed-up" as possible--this is what pseudo-random number generators do, but if we have access to the program (i.e. the description of the sequence), we could perfectly predict the numbers in the sequence. It's hard to call that random.

Truly random numbers (as far as we know) come from real-world phenomena like radioactive decay. You can have a certain amount of this so-called "entropy" for free from internet sources like Hotbits. I use such services for my artificial life experiments (insert maniacal laugh here). Real randomness is a valuable commodity, and I'm constantly running over my limit for what I can get for free from these sites. Here's a description from their site of where the numbers come from:
HotBits is an Internet resource that brings genuine random numbers, generated by a process fundamentally governed by the inherent uncertainty in the quantum mechanical laws of nature, directly to your computer in a variety of forms. HotBits are generated by timing successive pairs of radioactive decays detected by a Geiger-Müller tube interfaced to a computer.
What would be involved if you wanted to predict a sequence of such numbers (which will come in binary as ones and zeros)? As far as we know, radioactive decay is not predictable from knowing the physical state of the system (see Bell's Theorem for more on such things).

Even in a mechanical system such as a spinning basket of ping-pong balls, like those used for selecting the winning numbers in lotteries, a complete description of the system that is sufficient to allow you to predict which balls will emerge to declare the winner would be a very long set of formulas and data. In other words, even if it's not infinitely complex, it's very, very complex.

But what if we want partial credit? This was the big idea that had half my brain working all day. What if we are content to predict some fraction of the sequence, and not every single output? (Like "lossy" compression of image files versus exact compressors for executables.) For example, if I flip a coin over and over, and I confidently "predict" for each flip that it will come up heads, I will be right about half the time (in fact, any predictor I use is going to be right half the time). So even with the simplest possible predictor, I can get 50% accuracy.

Imagine that we have an infinitely complex binary sequence S-INF that comes from some real source like radioactive decay. We write a program to do the following:
For the first 999,999 bits, we output a 1 each time. For the one-millionth bit we output the first S-INF bit, which is random. Then we repeat this process forever, with very long strings of 1s followed by a random one or zero. Call this sequence S-1
It should be clear that a perfect predictor of S-1 is impossible with a finite program. It's still infinitely complex because of the interjection of bits from S-INF. But on the other hand, we can accurately predict the sequence a large portion of the time. We'll be wrong once in every two million bits, on average, if we just predict that the output will be 1 every time.

There is a big difference between randomness and predictability. If we take this another step, we could imagine making a picture of prediction-difficulty for a given sequence. An example from history may make this clearer.
Before Galileo, many people presumably thought that heavy objects fell faster than lighter objects. This is a simple predictor of a physical system. It works fine with feathers and rocks, but it gives the wrong answer for rocks of different weights. Galileo and then Newton added more description (formulas, units, measurements) that allowed much better predictions. These turned out to be insufficient for very large scale cases, and Einstein added even more description (more math) to create a more complex way of predicting gravitational effects. We know that even relativity is incomplete, however, because it doesn't work on very small scales, so theorists are trying ideas like string theory to find an even more complex predictor that will work in more instances. This process of scientific discovery increases the complexity of the description and increases the accuracy of predictions as it does. 
Perfect prediction in the real world can be very, very complex, or even infinitely complex. That means that there isn't enough time or space to do the job perfectly. As someone else has noted (Arthur C. Clark?), some systems are so complex that the only way to "predict" what happens is to watch the system itself. But even very complex systems may be predictable with high probability, as we have seen. What is the relationship between the complexity of our best predictor and the probability of a correct prediction? This will depend on the system we are predicting--there are many possibilities. Below are some graphs to illustrate predictors of a binary sequence. Recall that a constant "predictor" can't do worse that 50%.

The optimistic case is pictured below. This is embodied in the philosophy of progress--as long as we keep working hard, creating more elaborate (and accurate) formulas, the results will come in the form of better and better predictors.

 The worst case is shown below. No matter how hard we try, we can't do better than guessing. This is the case with radioactive decay (as far as anyone knows).
The graph below is more like the actual progress of science as a "punctuated equilibrium." There are increasingly large complexity deserts, where no improvement is seen. Compare the relatively few scientists that led to Newton's revolution or the efforts of Einstein and his collaborators to the massive undertaking that is string theory (and its competition, like loop quantum gravity).
Note that merely increasing the complexity of a predictor is easy. The hard part is figuring out how to increase prediction rates. You can always make a formula or description more complex, but by doing so it doesn't guarantee that the predictors are any better. Generally speaking, there is no computable (that is, systematic or deterministic) method for automatically finding optimal predictors for a given complexity level. You might think that you could just try every single program of a given complexity level and proceed by exhausting the possibilities, but you run into the Halting Problem. There are practical ways to tackle the problem though. This is a topic from the part of computer science called machine learning. A new tool that appeared this year from Cornell University is Eureqa, a program for finding formulas to fit patterns in data sets using an evolutionary approach.

Next time I will apply this idea to testing and outcomes assessment. It's very cool.

Wednesday, August 05, 2009

Open Courses

A technology article in The Chronicle yesterday describes the Obama administration's plan to develop open courses and give them away. This is a fascinating idea--the sort of thing that I mused about in the post "Unplanned Obsolescence" a few weeks ago. Carnegie Mellon's Open Learning Initiative is given as the model for this plan. The delivery of the course is potentially fully online, but example cited is a hybrid course with content developed by a group of experts and leveraging online capabilities for what I would call the low-complexity part. This is a natural way to optimize the delivery of instruction, and you already see it with course packs delivered from text-book publishers. Rote learning, or learning low-complexity modes of thought can be easily and tirelessly done through computer training. It's potentially more fun too. The Rosetta Stone language software is like this, using instant feedback, images and sound to train vocabulary (my wife used it to learn some Arabic). Where computers fail is at answering open-ended questions. If you've ever called your bank with an odd problem and had to navigate the automated phone tree (press one if you're still breathing...) you know how frustrating it is to try to solve high complexity problems with low complexity tools. So the need for actual instructors still exists, but their time could be applied more usefully to high-complexity tasks. This effect is noted in the article:
Carnegie's materials have already changed how Logan Stark's professor at California Polytechnic State University approaches her widely feared biochemistry-for-nonmajors class. Anya L. Goodman used to work from a prepared lecture, starting with the basics so she didn't lose anyone. Now she puts the burden on students to learn the basics online. She focuses class time on clearing up misconceptions, applying the materials to real life, and working in small groups.
The idea has merit, but there are certainly some big problems to overcome. If the courseware project is run and funded by the government, it may be open, but it surely won't be cheap. How, then do the course materials get updated? This isn't too big of a problem with, say, Euclidean Geometry, but for something like Finance, I imagine the books get re-written all the time. The article supposes that this might continue to be funded by the government, but this doesn't sound like a wonderful idea to me. It would be far better, methinks, if a culture evolved similar to the open source software movement. Rather than a set of "perfect" open courses designed and maintained centrally, a whole ecology of work collocated in some natural place--analogous to sourceforge.net--could grow and evolve, tagged with comments and other metadata. This depends on willing practitioners doing all the work. Faculty members taking the time to update old materials, probably. It seems unlikely on the face of it, but somehow it works for software. It works for Wikipedia.

A really interesting question is how course assessment ties into the courseware. Would it be developed and delivered in parallel, integrated with the materials? Or will assessment remain a second-thought tack-on for another decade? But if it is to be integrated, then the learning objectives need to be clear. There seems to be an opportunity for the assessment profession to get involved with this train before it leaves the station. Polish up your resume.

In a recent report I asked for, a group of twenty-plus institutions like ours had an average total expenditure on instruction-related items of 37% of total budget. We might suppose then that the asymptotic limit for reducing administration (assuming that all academic support is free somehow, libraries and such) is a 63% reduction in the cost of delivering programs. The cost of instruction could be reduced too, if the low-complexity components are off-loaded to technology. Moreover, competition in the fluid digital domain would tend to force prices down. I don't think it's unrealistic to estimate that a bachelor's degree could be delivered for about 25% of what it costs at a traditional college now.

Marshal Smith, senior counselor to the Secretary of Education seems to be the the guiding light behind this open ed plan. You can read his ideas if you have a subscription to Science in his article "Opening Education." You can also browse MIT's version of open courseware here.

Wednesday, January 25, 2012

Closed and Open Thinking

Most readers will know William of Occam's principle about not multiplying eventualities unnecessarily. It's commonly thought of as "the simplest explanation is the best explanation." I learned about a countervailing principle in Arora and Barak's Computational Complexity: A Modern Approach. It's even older than the venerable Mr. Occam, dating back to the Epicureans, and it states that we should not abandon any explanation that is consistent with the facts. I have mentioned this before, but I had an interesting thought at lunch today: what if this tension between efficiency and open-mindedness is at the heart of the Dunning-Kruger effect? In case you've missed that bit of news, here's the introduction from the Wikipedia entry:
The Dunning–Kruger effect is a cognitive bias in which unskilled people make poor decisions and reach erroneous conclusions, but their incompetence denies them the metacognitive ability to recognize their mistakes.[1] The unskilled therefore suffer from illusory superiority, rating their ability as above average, much higher than it actually is, while the highly skilled underrate their own abilities, suffering from illusory inferiority. 
Actual competence may weaken self-confidence, as competent individuals may falsely assume that others have an equivalent understanding. As Kruger and Dunning conclude, "the miscalibration of the incompetent stems from an error about the self, whereas the miscalibration of the highly competent stems from an error about others" (p. 1127).[2] The effect is about paradoxical defects in cognitive ability, both in oneself and as one compares oneself to others.
This just puts some research behind what Bertrand Russell is quoted as having said:
The trouble with the world is that the stupid are cocksure and the intelligent are full of doubt.
So what we have is two epistemologies, and we shouldn't be hasty to choose one as better than the other, despite the obvious bias of the quotes above.

Method 1 (Closed). Obtain a small amount of evidence, and create the most restrictive explanation that fits the facts. Subsequent facts that come to surface do not affect the conclusion.

William of Occam would probably sue me for defamation if he were around to read this. I have intentionally restated his principle in a very narrow sense in order to contrast it with:

Method 2 (Open). Continually gather information and create increasingly complex explanations that account for all the observations. Although the current explanation may be the simplest one that fits the facts, no explanation is ever final--all the others that are consistent with facts are kept in reserve.

I have given the methods intuitive names for convenience (closed vs open), not as prejudgments. The closed method will be the better one in situations where observations can be explained simply. This may be because the underlying cause and effect relationship is of low complexity, or perhaps that the variance in observed characteristics is small.  "All dogs have four legs" would be an example of the latter. "Stuff falls when you drop it" applies to the former.

The most basic structure of language is a verb applied to a noun, which is a model for the closed epistemology. "Birds fly," "Fire burns," and so on, are summaries of real world observations that can be arrived at accurately from just a few examples and without much error. It's an easy conjecture that these simple relationships became so integral to understanding that exceptions were met with challenge. Such as: "If an ostrich doesn't fly, then it can't be a bird." This is what school children encounter when they learn that a whale isn't a fish. The language we use rather gracelessly allows these exceptions in the form of conjunctive appendices, but this is clearly a hack. I will suggest below that a formal language is required to overcome that difficulty (for example, expressions of formal logic, which defines a consistent way of using "or" and "and," and allows unlimited nesting of exceptions, so that any true/false relationship can be expressed unambiguously).

Quickly assembling a set of closed rules for a new environment seems like a good idea. It's a fast best-guess approach to finding useful cause and effect relationships.

Of course, the closed method is not suitable to doing science. Khun's The Structure of Scientific Revolutions suggests that closed outlooks solidify at any level of complexity, and require some bashing to break up. An example would be the certainty (due to Aristotle) that celestial bodies move in perfect circles. This is like Gould's idea of "punctuated equilibrium" in biological evolution. I graphed the associated relationship between predictability and complexity recently in "Randomness and Prediction."

The question is when to use the open versus closed approach.  Historically, I think the closed approach may have had a blanket "explanation" in the form of mystical associations of cause and effect, which provides a putative low-complexity relationship. "Joe got struck by lightning because he displeased the weather god" has the appearance of an explanation, except that it's not actually predictive. It takes a dedicated effort to discover that fact, however. For that we need an open method.

The disadvantages of the open method make a long list. First, it's energy intensive--you have to continually be making observations, comparing what you see to what you think you should see (e.g. three-legged cat), and updating the every-growing explanation.  It also takes more energy to use or communicate the current explanation, and as soon as you do, it's out of date again.

These are not fatal flaws, but ones to be considered. For some phenomena, this is probably how we naturally reason, if in a limited way. For example, our memory and minds do something like Bayesian reasoning (updating the probability of an event based on how frequently we encounter it), although our on-board system has been shown to be deeply flawed (see Daniel Kahneman's recent book, for this and a lot more).

Perhaps the open process needs a kind of empirical 'clean-up' to be really useful. Elegant explanations generally only work with clean data. That is, if you want to discover Newtonian mechanics, it's unlikely that you can do this with just your eyes and ears. When Galileo began measuring the "drop" times on an inclined plane, he was onto something.

In addition to a solid empirical methodology, an open method also needs a way to reduce the size of an explanation while retaining its predictive power. In my graphs in "Randomness and Prediction," I plotted predictability versus complexity, not size. It works like this.

Suppose I have an observed relationship that I have cataloged like this: (1,2), (2,4), (3,8), (4,16), where this might be thought of as a cause and effect. A one 'causes' a two, and so on. Because my empirical methods are sound, I trust that there's not too much error in the observed values. As the list grows by using the open method, I have a better and better 'explanation' of past events and a better and better predictor of future ones (fine print about the inductive hypothesis goes here...). But the list will become too unwieldy to remember, communicate, or use effectively, as the observations accumulate. What I need is a kind of data compression to reduce the list to a manageable size. If I do this correctly, the explanation doesn't change, nor does the complexity, but the size does. I can reduce it to effect = 2^cause if I have the idea of an exponential function. We might call this data reduction the creation of a formal theory.

Conclusions
I started by wondering if people who don't know things, and further don't know that they don't know them, could be attributed to one of the two epistemologies mentioned at the beginning. I think the argument above shows that it's possible that the two barriers of empiricism and abstract thinking needed to effectively use an open method are too formidable for a lot of people. For one thing, it's not hard to get by using closed systems, and it may require formal education in scientific method and meta-cognition to effectively use open systems.

One final note appropriate to the calendar in the US: it's a lot easier to communicate closed explanations than open ones. Even with data compression, "things fall" is less complex than Newton's laws. So in a debate made with sound bites from political candidates, the closed epistemology wins. It's easier, it's comfortable to the listener--the whole construct of English is build to 'hack' a closed way of thinking by adding a few contingencies ("Cats have four legs, but I once saw one with three.")--and the explanations take up less time to say. You have to expand "Drill!" into "Drill, baby drill!" to make it bigger because the basic message can be summed up in one word, and that may seem too short for some audiences as a serious thought.

This is just another reason why we should be deliberate about teaching science and meta-cognition in school, not as alien ways of thinking that only people in white coats use at work, but as the mode of thinking that differentiates us from the other mammals, and might allow us someday to collectively make good decisions.

Thursday, April 02, 2009

Why Documenting Learning Outcomes is Hard, Part One of Infinity

Pat Williams recently asked if anyone is closing the proverbial loop on her blog. Quote:

For almost all of us the question, "What are you going to do about it?" has proven the toughest to answer.
If you've ever written assessment reports yourself, or had to judge ones others have written, you'll probably find it hard to disagree with this statement. When I first confronted the task of documenting those for our SACS report, it became obvious that there were some institutional disconnects at the accreditation level. For example, if you look back at the old SACS general education requirement (3.5.1), it was written as a minimum standards policy (you must set standards and show that students meet them), but in practice, every single person I talked to at the SACS conferences used a continuing improvement philosophy to actually judge compliance! This included IR people who'd been on many campus visits as well as senior vice presidents of the Association speaking in formal presentations. I found this bizarre and almost Kafkaesque. I pointed this out in emails when SACS invited comment. I don't know how my suggestions were received, but the standard was changed.

The point here isn't really about SACS. It's about the confusion surrounding standards-based vs. continuing improvement models. Public schools are primarily standards-based in their approach. Results are determined by standardized tests. If the learning objectives are simple enough (in the formal sense of complexity, which I've written about here many times), this approach should work well enough. This is because the items on the test can correspond pretty precisely to the material being taught. For example, if we want to know if third graders know their multiplication tables, this is relatively straightforward.

Things go wrong when the assessment doesn't align with the actual goals. This is a question of validity of the test, but it often goes unnoticed. By believing too much in the wrong assessment, we can create what I call degenerate assessment loops (more here). You can read a nice parable about this in The Well Intentioned Commissar.

Continuous improvement is orthogonal to standards. If one mixes up standards-based assessment and continuous improvement, confusion quickly sets in. Even in the SACS 3.3.1 standard of old, there was ambiguous language (I complained about that here, and it changed too, but it's probably coincidence). The problem is that many people assume that you should be able to show continuous improvement in some particular metric. This is virtually impossible to do once learning outcomes have reached a certain level of complexity.

We think of many assessment situations as falling into one of these categories:
  • Summative and transparent An example of this is low-complexity tasks like learning multiplication tables. It's easy to judge how things are going, and relatively easy to effect change (that's the transparent part). One can apply standards here effectively.
  • Summative and Opaque Here it's easy to see where you are, but hard to know how to get where you want to go. An example of this is the stock market. In learning outcomes assessment, this category applies to standardized tests of complex behaviours, which necessarily reduce the complexity (and hence validity). It's easy to see if the numbers go up or down, but hard to know how to affect them (without, say, directly teaching students how to take the test). Standards can be applied here, but it's very hard to use them for accountability, since there's no guaranteed way to reach new heights of achievement.
  • Formative and transparent In this case, subjectivity is valued more than objectivity, and standardization gives way to things like portfolio review, interviews, and other "fuzzy" types of assessment. This is most appropriate for complex outcomes like analytical thinking or creative writing. Transparency, or the ability to turn results into effective action, depends much on the specificity of the outcome. In other words, reducing the complexity to something manageable. For example, a speech instructor giving advice to a student who has just delivered his first performance. Rubrics can help identify simplifications, but shouldn't be thought of as a panacea. It's not always true that the whole is the sum of the parts. For example, numerically summing up subjective ratings on a rubric to combine to some whole is really not defensible as science or epistemology. Continuous improvement models work in this situation; standards-based are less effective.
  • Formative and opaque These are the hardest nuts to crack. I'd say assessing 'critical thinking' falls into this category. I recommend rethinking the problem, and redefining one's mission if you find yourself unable to come to grips with a problem even subjectively. Neither standards-based nor continuous improvement are going to help you out here. Abandon all hope.

Stay tuned for part two. [update: Part two is here]

Wednesday, March 04, 2009

Data Compressing the Curriculum

I've written about vodcasting lectures for out-of-class consumption and about the idea of information theory in the past. A recent article in the Chronicle puts these together. The piece highlights a practice of condensing a lecture or topic into a one to three minute video that encapsulates the essence of the material to be taught. This would be followed by an exercise to reinforce the material.

You can see one here:




I can see that this technique would work better with some ideas than others. In particular, this is the kind of delivery that would have the best chance of being successful with low-complexity learning: facts, simple pattern recognition, or simple examples of deductive reasoning. In fact, it occurs to me that one could scientifically study the complexity of a task by reducing the length of the mini-lecture until students don't understand anymore. This would be an upper bound on complexity. So perhaps solving a 2x2 system of linear differential equations has 25-minute complexity, whereas learning the basic facts about the Norman Invasion has 2.4 minute complexity. If this worked, you could actually create a taxonomy of knowledge and skills for a discipline that had some (minimal) scientific grounding.

Thursday, June 03, 2010

Assessment Crosstalk

A recent InsideHigherEd article "The Faculty Role in Assessment" has stayed on the site's Most Popular list for several days.  In "Fixing Assessment" I posted my solution: differentiating clearly between strategic and tactical goals and pursuing those at an appropriate level.  It is evident from the comments on the IHE article that there is deep interest in the issue, and in this post I will take a stab at categorizing those comments.

There's no proof assessment does any good is the argument made by RSP.  This turns the tables on the psychometric-speak employed by the academic arm of the assessment profession (up to and including standardized test vendors).  The argument is natural: if you can measure learning, show me how much students have learned and how much assessment matters, qualitatively and rigorously. The community of assessment professionals is aware of this contradiction, but don't seem to take it seriously.  The obvious solution to that is to be more modest about claims. 

Assessment is big business is the ad hominum comment by skeptical_provost.  It is ironically book-ended by John B. later on, who is selling the product. Now an ad hominum may not be a bad argument.  If I learned you invested with Mr. Madoff, I may have reason to be worried for your finances without further information.  In this case, the industrialization of testing has produced undesirable side effects.  Marketing overstates what the tests can do, this and politicization seem to have cultivated the idea even at the Dept of Ed that we can measure minds like weighing potatoes.  This underlines the previous point of don't believe the advertising.  Do some critical thinking, right?

Oversimplification of assessments  is a problem first raised by Jonathan.  I've argued this point many times in this blog: the complexity of an assessment has to match the complexity of what you're assessing.  Or else you should be very suspicious of the results.  This relates to the previous comments in at least two ways.  First, the industrialization of assessment requires convenience (for economic reasons) and reliability (in order to make the argument that measurement is being done).  Both of these serve to reduce the complexity of assessments to standardized instruments and standardized scoring.  As a prosaic example, imagine hiring a babysitter.  Would you rely solely on some standardized test to tell you how good a candidate is?  Or would you want to interview the person, see if you can suss out how good his/her judgment is, how well the person communicates, how bright they seem?  You might want to do both to cover your bases, but the second (high complexity) step tells you things the low-complexity test cannot.  This problem raises a particular question for the more aggressive claims of assessment: how do you account for complexity?  Here's another example.
A limerick is a well-defined form, which you can read about here.  We can fairly easily create a rubric and standardized assessment for limericks.  Imagine you've done that, and then assess the following one (borrowed from John D. Barrow's Impossibility):

There was a young man of Milan
Whose rhymes they never would scan;
When asked why it was,
He said, 'It's because
I always try to cram as many words into the last line as ever I possibly can.'
Rated honestly, you'd have to flunk it, because it's not technically a limerick at all.  It fails the most basic requirements.  However, it could be considered a meta-limerick, which is arguably more interesting.  It's unlikely your rubric can handle this.  It might explode.

It does make a difference, responds Jeremy Penn (RSP's comment).  But there is a subtle shift here.  Jeremy references what happens in the classroom and instruction techniques, NOT assessment writ large.  This mismatch of meaning and vocabulary is at the heart of much misunderstanding.  Let me call these points of view:
Assessment as Pedagogy (AP): use of subjective and some standardized techniques to improve the delivery of instruction at the course and program level, tied very closely to the faculty who teach the classes and the material that's being taught.  It doesn't need heavy bureaucracy or assumptions about reductionism.

Assessment as Science (AS): the idea that we can actually measure learning precisely enough to determine things like the "value added" by a general education curriculum.  The hallmarks of AS are large claims tied to narrow standardized instruments, theoretical models that pin parameter estimates on test statistics, and a positivist/reductionist approach.
So Jeremy is talking about AP (if I understand the comment), and RSP's comment is probably addressed at AS. 

Accountability and efficiency are mentioned by Luisa, who points out that bureaucracies are not known for either.  One fragment:
[I] become better at what I do through meaningful dialogue and narrative, rather than through inaccurate/invented data and endless bullet lists.

[...] What goes on in my classroom on a daily basis does not 'count.' What 'counts' is 'documented' learning, ie, the product-as-educational-widget.
This contrasts AP to AS.  Luisa also makes the distinction between outcome and process, which I wrote about here. 

Grades are assessment says Cal.  It is by now canon in the assessment world that grades are not assessments.  This does not stop researchers and test-makers from using GPA correlations as evidence of validity, but never mind.  But if we look at grades in the lens of the AP/AS definitions, it's clear that grades aren't really either one. 

Allow me to pull together an observation combining Luisa and Cal's comments:
Assessment as AP can allow us to overcome bureaucratic limitations natural to higher education.
Good assessment (as AP) can allow us to do things that wouldn't otherwise be thunk of.  Grades are statistical goo that don't tell us enough about what a student actually did. There are easy ways to get richer information (I mean really easy, and free too).  I recently helped our nascent performing arts programs work through their assessment plans. One of their goals was to develop creativity in students.  This is something that takes time to mature, and can best be observed outside the bureaucracy of class blocks.  It was a great discussion, which led them to create a feedback form for students, to be used across the curriculum.  A draft is shown below.

The idea is that students get feedback in an organized way throughout the whole curriculum.  This can't be done with grades.  It also doesn't require a big bureaucracy or expensive software to pull off.  Note also, that this form is very much AP.  Students will get feedback in the form of written comments telling them why their presentation (for example) is at the Freshman/Sophomore level of expectation rather than Graduate (meaning ready to graduate).  It is not reductionist (as AS would be), and does not attempt to measure anything--it's a pure, subjective, professional assessment by faculty.  There is a rubric (not shown) that communicates basic expectations for each level, but it remains a very flexible tool for tracking progress.   I've used this technique in math too--it's a good exercise to have students rate themselves and then compare to your own ratings.  In the arts, critiques are important, so you can do peer ratings also.

By way of contrast, the AS approach is to reduce learning into narrow bins (to increase reliability), assign ratings, and then aggregate to create dimension-challenged statistical goo.  This quickly becomes so disconnected to what's happening in the classroom, that it's a head-scratcher to figure out what actions might be implied by the numbers (see this, for example).

Emphasis on input vs output is one of several points made by commenter G. Starkey.  This may be another way to refer to process versus outcomes.  Note that AP is (or can be) deeply involved with process, and its messy subjectivity.  AS tries to stand aloof from that by looking only at outcomes.  Starkey mentions "the hierarchical nature of assessment," which sounds like a reference to AS.  This is underlined by subsequent comments, so the criticism seems to me to be "AS is imposed from the top for accountability," which is an "inherently ill-defined problem."  Although AS produces simple, standardized results, the inputs (students, processes) are anything but standardized.

A charge of incompetence is leveled by Charles Bittner, who sees failure in public schools and links that to "failed methods and techniques" of education departments.  I presume this to refer to AS-type activities, since massive low-complexity standardized tests are in place in public education.

Fuzziness of rubrics and metrics is not to be taken seriously, says Gratefully uninvolved in an gratuitously elitist comment (referring to those who have never taught at "real universities").

Assessment makes a difference argues the aptly named Ms. Assessment, who supports this with anecdotes about curriculum maps and rubrics and feedback.  Although there are not enough details to be sure, it seems to me that she is referring to AP approaches that have direct influence in the classroom and across classes, not a grand AS approach with standard deviations to be moved.  But I could be wrong.

AP and AS are different, is the point of BDL, when translated into my definitions above.  BDL identifies the "meaning gap" better than I did, talking about the question "what are our students learning, and how do we know this?":
This is no simple bean-counting question for summative accountability. Instead it is an ongoing question of self-reflection that ought to be integral in everything from our curriculum planning, to our classroom practice, and whatever occurs between.
BDL also points at bureaucracy as being unhelpful:
And we continue to measure "student success" in opaque grades, with transcripts that like so many 19th century factories account, mostly, for how many hours students have "spent" in class, over how many years. What they've learned, we can only assert as a matter of faith and belief.
If this idea resonates with you, you might be interested in "Getting Rid of Grades."

AS is symbolic, AP is useful, is my translation of Charles McClintock:
While assessment clearly has a symbolic use to external audiences to signal that resources are being used wisely, it can have its greatest instrumental value for students.
He gives examples of AP being successful.  For example "The greatest value of these learning outcomes, in my view, is that they communicate to students what is expected of them in their graduate work."

AS is difficult or impossible, says Raoul Ohio.

Taking issue with Ms. Assessment's comment is Peter C. Herman, who says anonymous anecdotes are not evidence.  Here he assumes that she is talking about AS, whereas I assume she's talking about AP, in which case the question of evidence is moot.  He underlines this by talking about the problems with an AS program.  For example: "'Assessment' assumes that students are inert objects, to be made or marred by their professors, and that is just not a valid assumption."  Commenter Bear says "me too" to this idea.  Both seem to assume that Assessment = AS, and ignore AP.  This is not perhaps a conscious decision, because the context at their respective institutions may be a solely AS-oriented approach. 

The next comment, by midwest prof, also criticizes AS:
Standardized measures throughout became the benchmark of effective assessment, with resistance depicted as evidence that faculty either did not know how to teach or simply did not want to be held accountable.
midwest prof advises to "find value in what faculty are already doing and find a way to organize it in a way that those who feel assessment is important can utilize."  I don't know if this is cynical (tell the administration some good news and get on with things) or constructive use of crowd-sourcing  (see Evolutionary Thinking).  Another thread to this comment is lack of preparation in graduate school for assessment, and recommends mentoring.

It takes money, says Unfunded Mandate, including faculty time.  It's not clear what kind of assessment program is the subject here, but something time-consuming and expensive is a good guess.  This applies more to AS than AP, I think, although both require investment in staff positions and professional development. The next comment, by Wendell Motter, puts some numbers to time spent doing classwork-based (probably AP) assessment, and makes that case that time is proportional to students, and therefore there is a limit beyond which it is not reasonable to expect good results. 

Decoupling is the result of bureaucratization argues sk: that (presumably) institutional top-down AS isolates process from result.  I interpolated a lot there.

Administration detracts from teaching, asserts Professor C, and argues against "endless rubrics and data formats."  This seems like a criticism from the Assessment = AS perspective.


Rubrics are not a panacea, is the theme from an Alfie Kohn article linked by Professor D.  The crux of the criticism is perhaps in this paragraph:
Consistent and uniform standards are admirable, and maybe even workable, when we’re talking about, say, the manufacture of DVD players.  The process of trying to gauge children’s understanding of ideas is a very different matter, however. It necessarily entails the exercise of human judgment, which is an imprecise, subjective affair.   
Again, this is AS contrasted to AP.  Rubrics can be used for either--it depends on the approach and use of results.  Rubrics as scientific data-measuring are AS, rubrics as communications tools and guides are AP. The whole article is worth reading--it contrasts the two approaches to assessment in different language than I have used here, but it's the same issue.

SA isn't proven to work, is how I interpret John W. Powell's comment, who writes:
Rubrics, simplistic embedded assessments constructed with an eye to efficient end-of-term review, and the resulting reports which overwhelm chairs and assessment coordinators, all betray a severely stunted vision of the multifarious answers to the question, what is education for?
John goes further, challenging the quality of debate: "[...Outcomes assessment] looks more and more like a fundamentalist religious faith." Given the nature of the argument (e.g. "its justifications abjure the science we would ordinarily require"), the critique is best addressed to an AS approach.

Inappropriate measures are a problem, says humanities prof, comparing actual learning outcomes:
I try to teach my students to systematically break down and evaluate complex, abstract and subjective questions.
to the official assessmen:
 A four to five question multiple choice pre- and post-test consisting of basic terms and concepts.
Humanities prof blames this on the need to quantify and standardize, which is a charge that would only apply to AS.


Crosstalk is noted by Betsy Vane, who says that advocates are arguing about the reasons for doing assessment and opponents are criticizing the means.  Betsy suggests a rapprochement, but I don't see that as possible without deciding first to what extent assessment is going to be AS or AP.  I don't see faculty ever buying in to the pure form of AS, for example.  Administrations out there spending kilo-dollars to implement data-tracking systems based on the AS idea might want to think hard about this.

Stating the AS agenda is Brian Buerke, who advises that "The key for assessing higher-order skills is to break them down into fundamental, well-defined elements and then assessing each element separately."  So that: "[S]tudent progress on specific elements can be charted. Aggregating the results over multiple elements allows student learning to be assessed in the more complex areas."  I put this in the AS category.  Depending on how seriously you took the numbers generated, it's possible this could be AP.  But I doubt it. One of the problems I have with the sort of aggregation I imagine is meant here is that you end up averaging different dimensions and then calling the result some other (new) dimension.  For example, if you sum up the rubric scores for "correctness", "style",  and "audience" and call the result "writing", it's not much different from adding up height, weight, and hat size of a person and calling it "bigness."  Maybe you can abstract something useful out of bigness once in a while, but it's become statistical goo and suspect as information.

The idea that you can take something complex like "critical thinking" and create linear combinations of other stuff to add up to it is suspect to me.  I've written a lot about that elsewhere.  I think part of the problem is that we get the idea of complexity backwards sometimes because of Moravec's Paradox.  For example, solving a linear system of differential equations is probably far lower complexity than judging someone's emotion by looking at their face, even though the first may be hard for us and the second easy.

Arguing against the validity of AS is FK, who writes
Many of the processes that permeate our teaching, in the humanities and in the natural and social sciences, cannot be encapsulated in how we measure students’ learning of facts or methods of analysis through tests or papers, exams or presentations, etc.
I interpret this as complexity argument, which is a theme here.  FK describes the practical problem this poses.
If part of the “assessment” is describing what our learning outcomes are, and there are no definitive ways of measuring these outcomes after a course or within the 4-years of the students’ academic career, do we de-value these outcomes and only focus on what is measurable—even when certain courses, especially in the humanities, involve a lot more intangible outcomes than tangible ones?

Note:  So you know what my stake is in this, I'll describe my own bias.  I'm responsible for making sure that we do assessment and the rest of institutional effectiveness at a small liberal arts school.  All my direct experience is with such schools.  I've tried and failed to make AS work both as a faculty member and now as admin. I can see the usefulness of an AS approach in very limited circumstances, but I find the claims (for example) of the usefulness of standardized tests of critical thinking to be dubious.  On the other hand, I've had good success with AP methods in my own classroom and in working with other programs.  So I have a bias in favor of AP because of this.  Given the comments I've tried to synthesize, it seems that I'm not alone.

Thursday, April 23, 2009

Rules, Damned Rules, and Policy

My daughter has this thing about wanting pets. Because we aren't well situated to play host family to them (the gerbils were a disaster on our first attempt), I try various ruses to change the subject. When she was younger I latched onto the idea of using a laser pointer as a dog substitute (named Spot, inevitably, although Red was a contender). We took it for a lot of walks, always after dark, and watched proudly as it visited all the trees on the block.

Yesterday it was plants. Pet plants need minimal care, I figure. So we went to Home Despot to look. I particularly wanted rosemary, to replace the nice bush we had at the old house. We bought some 'pet' vegetables, but there was no rosemary, so we walked down to Kmort to look. Their outdoor section looked much like a concentration camp, and I'm sure if I spoke plant-ese I'd have been moved to tears by their pleas. My daughter wanted to ask the salespeople if she was allowed to water the plants, but there was no one to ask.

None of this has anything to do with rules, in case you're wondering. That connection came next, when I noticed a gaming store next to Kmort. In graduate school I co-authored a board game of sorts with a friend, so I wanted to peek in and wallow in the ambiance for a bit. My almost-teen daughter was properly horrified, which added to the attraction.

The place was filled with gamers at tables piled with miniatures from a dizzy variety of genres. There were sci-fi tableaux, with someone asking about how far pulse rifles could shoot, fantasy sorts of things I couldn't recognize, and historical battles with box-like formations of hand-painted troops. Around the walls were stocked the complicated rules books I remembered.

You've probably figured out by this point that I played a lot of geeky games as a teen, with rule books that resemble the 1040 tax instruction booklet. The rules got increasingly more complex as time went on, until half the games seemed to consist of searching for the right sub-section with the table on the chance of successfully napkin-folding or whatever. For me, there reached a point where I didn't find it enjoyable anymore. That's when I started coding up rules in Applesoft, to use the computer to keep track of the complicated bits, and truly descended into full-fledged geekdom, from which I never really emerged.

This all by way of introduction to "rules and the academy". (I hope this one worked better than sock muffins did.) There are places where rules are absolutely essential. You can identify them by the lack of thinking that's required by the tasked staffer. Storing backup tapes somewhere safe, following procedure with regard to transcripts, keeping track of financial accounts properly, and so on, are good examples. The lower the complexity, the more suitable for rules-making a process is. At the other end of the spectrum lies general responsibilities like "being president," which is too fuzzy to be described in a President's Operating Manual, or something.

If you like rules, stop by human resources. Here, as in other areas like IT, rules can be used to simply block things that staff don't want to do and generally accumulate power and influence. I read a (possibly apocryphal, but plausible) account of a man interviewing for a job, who was asked by the HR interviewer for the phone number of his previous employer, so as to verify the information on his resume. I can't find the original now, but it went something like this:
"I need to call your previous employer, Mr. Snark."
"Well, I was self-employed, so that would be me."
"Fine. What's the phone number?"
Mr. Snark, bemused, gives the number and watches the numbers being dialed. He pulls out his cell phone and answers on the first ring.
"Mr. Snark?"
"Yes, that's me."
"I need to verify some information about a previous employee."
In this Kafka-esque drama, the interview plays out in full, after which Mr. Snark is given the explanation that rules, after all, must be followed.

It's debated whether or not evolution produces more complexity in living things. I think it's probably true that complexity is more valuable in some situations than others. Like a string in a drawer, systems seem to bow to entropy almost immediately and become more complicated without much effort. Straightening them out is hard. As Machiavelli put it:
It must be considered that there is nothing more difficult to carry out nor more doubtful of success nor more dangerous to handle than to initiate a new order of things; for the reformer has enemies in all those who profit by the old order, and only lukewarm defenders in all those who would profit by the new order; this lukewarmness arising partly from the incredulity of mankind who does not truly believe in anything new until they actually have experience of it.
Whereas complexity naturally emerges in the form of additional rules, in order to properly simplify the resultant mess, the reformer has to overcome the natural resistance of those who benefit from the complications. Think of the US tax code.

So, like biological bodies renewing themselves through reproduction after entropy has corrupted them, it's a healthy process to change administrations and shake things up once in a while. It is perhaps the case that the most insidious rules are not affected by this, however. They may be invisible.

At least too-complex rules can be seen. There are many quasi-rules that are merely implied. Dress codes often are, and are enforced through social conventions, although explicit ones aren't uncommon. More dangerous, I think, to the mission of the academy are the implied rules that pertain to learning. Here are a few. You can add to the list.
  • Teaching only takes place in formal sessions
  • Learning is not as important as ratings like grades
  • Education proceeds by check marks on a sheet
  • Education is a service one pays for, just like having your car washed
  • Learning experiences can be made uniform, like an assembly line
  • Student collaboration, unless explicitly allowed, is cheating
The bureaucracy of higher education is necessary to organizing the massive endeavor, no doubt. But leaving unexamined the implications of these practices blinds us to some pernicious effects. Do we really have a right to complain if students see courses as milestones to be passed on a linear journey--points of momentary interest that can be forgotten? Doesn't the very structure of the process from advising through transcripts encourage that point of view?

In order to gain perspective, a complete rethink is in order. I was impressed recently with an idea by Gary Brown at Washington State University about a way to redefine the hoary old idea of a gradebook. No, I don't mean moving it to Excel. You can read more about this "harvesting gradebook" idea on the blog Center for Teaching, Learning & Technology. The authors of this project seem to be questioning what exactly grading is--a review of the associated implicit rules, as it were. It will be interesting to see where it leads. This is an example of the creative disruption that is called for in order to reach more than a superficial review of the stew of formal and informal complexity that comprise the academy.

Update: In the pursuit of sensible database policies this morning, I found myself wandering through the wilds of the FERPA rules, and discovered this gem [pdf].

Under FERPA a school may not disclose a student’s grades to another student without the prior written consent of the parent or eligible student. “Peer-grading” is a common educational practice in which teachers require students to exchange homework assignments, tests, and other papers, grade one another’s work, and then either call out the grade or turn in the work to the teacher for recordation. Even though peer-grading results in students finding out each other’s grades, the U.S. Supreme Court in 2002 issued a narrow holding in Owasso that this practice does not violate FERPA because grades on students’ papers are not “maintained” under the definition of “education records” and, therefore, would not be covered under FERPA at least until the teacher has collected and recorded them in the teacher’s grade book, a decision consistent with the Department’s longstanding position on peer-grading. The Court rejected assertions that students were “parties acting for” an institution when they scored each other’s work and that the student papers were, at that stage, “maintained” within the meaning of FERPA. Among other considerations, the Court expressed doubt that Congress intended to intervene in such a drastic fashion with traditional State functions or that the “federal power would exercise minute control over specific teaching methods and instructional dynamics in classrooms throughout the country.” The final regulations create a new exception to the definition of education records” that excludes grades on peer-graded papers before they are collected and recorded by a teacher. This change clarifies that peer-grading does not violate FERPA.

Exceptions are the hallmark of complexity. If trivial exceptions can't be dealt with by simple common-sense methods, you're stuck with arguing trivialities at the highest, most formal level of adjudication. This is a recipe for entropy-induced "heat death," as it's called when one speaks of the end of time.

Friday, August 07, 2009

Creating Meeting Discipline

Anybody reading this blog has probably sat in generous number meetings. I'm guessing you haven't had a lot of those meetings I read about where everyone stands rather than sitting, in order abbreviate the proceedings. I've blogged here before about various aspects of meetings, which you can find here.

I like meetings where you walk out with a sense of accomplishment. There's this particular feeling that comes with successful group-decision-making that must be a pale shadow of what it's like to be a node in a hive mind, if such a thing really exists. One of the striking realizations I had in reading My Stroke of Insight was that our brain hemispheres are like two very closely cooperating minds. When the right people and right habits of mind combine something magical occurs. I think I remember first reading about it in Michael Herr's Dispatches, but I can't be sure that's the book. After searching on google books, I couldn't locate the quote. [Edit: it's A Rumor of War, by Philip Caputo, and the quote is here.] Like the rest of my generation, Vietnam was the "last war" and held a certain fascination that led me to read a lot of books about it. In any event, the scene I remember is a platoon leader recollecting the experience of directing his troops in a sweep, and his description of an exalted feeling of the extension of his own body, almost a proprioception, into the men following his direction. This, of course, isn't exactly the same thing as colleagues deciding general education around the table, but sometimes it seems like it.

Sci-fi has lots to say about the idea of group minds (as opposed to group-think, which may be its opposite), and naturally pushes the envelope. See the excellent novel A Darkness in the Sky by Vernor Vinge, for example. Science itself does too, if you will tolerate one more digression before pulling the chair fully up to the table of "what to do about bad meetings."

The term "complexity" became more confusing at some point because it came to mean, in addition to the extant meanings, the study of how systems emerge out of goop. Of course, this is not the technical definition, which you can find here. I've made a lot of hay in this blog with another kind of complexity: computational complexity, but the two are different. John Holland, who works at the Sante Fe Institute now (think Manhattan Project) developed some cool ideas about intelligence. Well, artificial intelligence, anyway, but who's counting. The idea that sticks in my mind is that of competing algorithms (like voices) that sound an alarm when they think they can contribute something useful. If their input is actually valuable, it's rewarded. Otherwise they may be ignored the next time, like the proverbial boy who cried "lupus!" or some other auto-immune disease, I forget. This is very like the members of a team or committee, who have to individually decide when their input is valuable, and slowly accrue or leak social capital with their reputation for effective contributions. With humans, of course, there are many other complexities, like how to speak, what words to use, how to interact with others, and so on. There are lots of things that can go wrong. I think mostly they do, in fact, because the meetings that are exceptionally productive seem few and far betwixt. This certainly cannot be a limitation of the people involved, but one of method, I present to you. But what method? By what cryptic scheme can a meeting be set in order?

I don't know. But, practicing what I preach--viz., that complex problems can be approached through an evolutionary method--I herein propose a starting point. I think the genesis for this was something I read long ago in the C User's Journal or Dr. Dobbs. Or not. Anyway I read about a protocol for conducting a meeting. And I don't mean Robert's Rules of Odor. Nothing is more annoying than the pedantry with which meeting minutes get presented and approved, after which all rules disappear. Nothing against Robert, 'natch. I just don't think the answer is complicated bureaucracy. Douglas Hofstadter created a game out of rules of order whereby one tries to bring the whole system to illogic, a frozen halting state. That's what I think of when I think of rules of order: a computational engine guaranteed to lock up.

No, what I have in mind is more like a game. We may as well call it the Committee Game. Here are the first draft rules. They aren't meant to be comprehensive.
  1. A referee is assigned to loosely enforce the following rules, with a liberal dose of common sense. It would be best if this were the committee chair, at least to start with.
  2. The meeting agenda needs to spell out the level of detail an item will be addressed in. Typically, "tactical" or "strategic" suffice to do this, but you may want "administrative" to talk about office functions or "meta" to talk about the functioning of the committee as a whole. Roll your own.
  3. Someone--probably not a committee member--is tasked to take timings. This consists of watching a second hand and noting how long each speaker talks. Doesn't need to be perfect. Simple statistics are generated from this for feedback later.
  4. At the end of the meeting, if anyone spoke longer than 30 seconds continuously, the longest speaker is fined a buck.
  5. If someone seems to go off topic, the referee should note this immediately by holding up some symbolic object, like a stuffed animal. Once the referee has the floor, he or she asks the group if they really want to take up that topic. If not, the offending member gets the off-topic symbol to 'own' until the next offense.
  6. A particular type of "off topic" is when someone begins to talk about tactical considerations during a strategic discussion or vice-versa. The procedure in #5 should be applied, and the scope (tactical, strategic, meta, whatever) explicitly noted, so that there will be a greater awareness of the level of discussion the committee is engaged in.
  7. Responsibility for being referee should rotate, so as to build a culture where time and effectiveness are valued.
  8. Defer to the chair of the committee when extraordinary measures are required, such as changing the agenda during the meeting, or suspending the rules. The chair is ultimately responsible for setting and executing the agenda, but defers actual meeting discipline to the referee.
  9. Periodically, the committee reviews its performance, considering how well agendas have been executed, the timing statistics, and other general considerations that apply. Improvements to the rules are made as deemed reasonable. Sample questions are:
  • Rate the overall effectiveness of the committee.
  • Are contributors getting to the point quickly enough?
  • Do you feel that decisions are being reached quickly, but with due consideration?
  • Do committee members feel fairly treated?
  • Are deadlines being met?
  • What could be improved and how?
  • Do you enjoy coming to meetings?
Even this level of formality may not be necessary. In practice, I've noticed a marked improvement in the deliberations of one of my standing committees simply by the group acknowledgment that meeting time is valuable. We agreed not to tolerate digressions, for example. It happened anyway, but I noticed that there was an awareness of it: "I know this is a digression...". At the end of the meeting we informally evaluated our own performance. We had accomplished a rather complex task in record time, leaving an hour for a sub-committee to polish the proposal. I think there was a general good feeling about the effort to make our deliberations more efficient and then seeing the outcome of that effort. It's evolution in action, and a pretty thing to watch.

Thursday, April 24, 2008

Why I don't understand the assessment results...

...and what I can do about it.

This post is for comments related to my session on Friday, April 25 in Cary, NC, at the NCSU Undergraduate Assessment Symposium. The description is:
In pursuit of generating general education assessment results, we may enthusiastically adopt methods that are convenient but artificial. We explore ways to ‘find’ authentic assessments and present summary reports that your grandma can understand. The benefits are a better integration of assessment with the curriculum, and a much easier time convincing faculty that you know what you’re doing. With examples, colorful charts, and a manageable dose of epistemology.
The PowerPoint slides will be posted on the college's assessment site. There you can also find information about Coker College's general education assessment program, the open source software we developed to manage accreditation documents, our compliance report, and other stuff.

UPDATE: thanks to
Dr. Pamela Steinke for noting some great questions during the talk. Here are my answers.

Q: To what degree does your institutional mission statement reflect your general education goals?

A: From the discussion, it seems as though Coker is a bit unusual in having liberal arts goals explicitly stated in the mission. The middle paragraph of our mission lists analytical and creative thinking, effective speaking and writing as educational goals. This makes it easier to rally the faculty around these.

Q: How often do the results of assessment data and assessment results get shared and acted upon at your institution?

A: For our general education assessment, we have a faculty meeting at the beginning of the year where a report is given on the big picture. Then departments get reports at the program and individual student level. Additionally, committees like the institutional effectiveness committee will use assessment data sporadically throughout the year.

Q: When trying to assess complex processes, the greater the reliability of your measure the less the validity. Do you have examples of this?

A: Remember that validity is in the eye of the beholder to a large extent. As reasoning creatures, we model phenomenon with simple relationships (like linear ones, for example), and the word "complex" in common useage can mean “hard to predict”. Complexity has a technical definition that is quite useful (see Kolmogorov Complexity on the web), but requires a fuller explanation than I can give here. It’s easy to create examples of high complexity skills that can be reliably tested. For example, a single question can test if someone can pilot a 747: “Can you pilot a 747?”. This would very likely be quite reliable in the sense that subsequent repetitions would give the same response. Is it valid? I wouldn’t trust my life to it! On the other hand, impressions from a first date can be assumed to be meaningful (valid), even though the reliability is zero—you can’t have two first dates with the same person, so the whole concept of reliability is meaningless in this instance. In daily life, most things we do fall in the second category.

A point I made in the session is that complex phenomena manifest themselves in more ways than simple ones. Testing for knowledge of simple multiplication, for example, consists of verifying that the subject can solve a few problem types. If we wanted to test for knowledge of the US tax code, on the other hand, how would we have complete confidence in our results? We'd literally have to test every part of the tax code--an impossible task for any individual I imagine. Thus we tend to reduce complex outcomes to some subset. This can cause a merelogical fallacy--substituting the part for the whole, like in the 747 example. This reduces validity if you care about these details.

Q: How important is it that general education assessments be done in context of the discipline? Can general education skills be assessed with validity outside of disciplinary context?

A: We assess liberal arts skills across the curriculum. I'm not quite sure which discipline the question refers to, but we get enough inter-rater reliability with our method to be happy about it (exact matches over half the time with different raters on the same student).

Q: Means should only be calculated across the same unit. Do you have examples of inappropriate use of means?

A: Three pounds of grapes plus four pounds of nuts is seven pounds, but not of grapenuts. Aggregation works great with quantities, but not so well with qualities. (Averaging is just aggregation with a division afterwards.) So averaging math assessments with writing assessments is probably questionable. But this is exactly how GPAs are computed, of course! The average is pretty meaningless as a number. What a high GPA tells you, for example, is that a student did well in most of his or her classes. A proportion would do a better job. If you set B grades as your threshold, you could look at the percentage of course grades that meet that threshold and get, say Johnny has 34% and Mary 95%, which would more directly tell you what you’re interested in. A student who has a GPA of 2.0 could have had a 4.0 for a year, and then had a really rotten year because of some personal problem. Or it could be a solid C student. Are these the same thing? No. I realize that this sound heretical, since grade point averages are the currency of the registrar's office, but it's the kind of figure you should take with a grain of salt. It tells you something, but maybe not the most important thing you care about. As in any kind of data compression, you run the risk of losing something important. There's a joke about engineers who want to account for the presence of humans in a building in order to anticipate their impact on the wireless network (water interferes with it). Their first assumption is "assume a human is a one meter diameter sphere of water." This data compression might work well in one context, but not another.

Q: What tools have you found especially helpful with data analysis and reporting. Pivot tables, logistic regression….

A: The combination of database + pivot table is extremely powerful. Logistic regression is a special need kind of thing, but is designed to predict binary events like attrition. A very useful trick with pivot tables is to redefine scalar quantities (decimal numbers or integers) into a 0 or 1. For example, you could classify students with 3.0 GPA or above a 1 and the others as 0. When you use this field in the pivot table as data, tell it to average, and display as a percentage. It will then read off the percentage of the group (e.g. demographic) that falls in that classification. When I get a chance, I’ll put more detailed instructions with screenshots on this blog.

Q: Share ways you get more useful results by analyzing data categorically rather than looking at means.

A: You can more easily see extremes. How many students perform very well or very poorly? This can get averaged out easily, and hence lost on average reports. But it’s individual students we deal with, not average students, so these extremes matter. Another example is to look at assessment data by gender, ethnicity, or classroom performance (grades).

Q: What important details do the averages hide? Do you have any examples in which reporting the minimum and maximum would be more appropriate?

A: Averages hide the composition of the results, obliterating the distribution. Would you rather teach a class with half brilliant students and half remedial, or one with nearly uniform preparation and ability? The average ability (if such a thing were to exist) is the same in both. A prominent figure in the testing industry visited our campus a couple of years ago, and I had the opportunity to review our results from an instrument he’d designed. I brought the sheets with averages on them. He said “I don’t know why we even publish this stuff—where are the distributions?” Unfortunately, it took me another two years to figure out what he meant. Instead of averages, consider defining a cut-off for acceptable/unacceptable. This creates a meaningful reduction in data that can be used simply with pivot tables, for example. Maxes and mins or the whole distribution can be given fairly easily. These are quite informative—usually more so than the average.

Q: How can I present the data in a way that is most useful for faculty for improvement?

A: This may be a question to pose to the faculty, but certainly I would avoid the “lines go up” global graph of data, unless it’s just to provide context. If the report doesn’t connect their conceptual model of what they can change to the assessment results, they’ll simply be puzzled about what to do with it. The best case is if they are the ones generating the data to begin with—then they’ll have a better idea than anyone of what it means in functional terms. You can read how we try to accomplish that on our assessment website. We call it Assessing the Elephant.