Showing posts sorted by relevance for query noncognitive. Sort by date Show all posts
Showing posts sorted by relevance for query noncognitive. Sort by date Show all posts

Monday, June 08, 2009

Noncognitives as Indicators of Success

As Institutional Research Director at Coker College some years ago, I was involved in creating a predicted GPA matrix, which was used to set merit awards in our financial aid leveraging process. I noticed that the amount of variance explained by the traditional measures of high school GPA and SAT (or ACT equivalent) was not great--perhaps 30% in the linear model using both variables. I tried using other data at hand, and found that these two variables were about the best usable ones at our disposal. Unusable ones included sex and income. This last was troubling because it was clear that students from a low socio-economic group got lower grades and standardized test scores on average, and hence received less merit aid. I began to wonder if it could really be true that these students were less able to do college work. Clearly some of them could do quite well--even at the very bottom of our acceptance pool, rated by predicted GPA, about half the students would do quite well academically. This was the genesis of the idea to try to find ways to identify the actual potential of these poorly-rated students. There are two advantages to this idea. The first is that it's fairer to students--rather than the usual (implicit) practice of adding merit aid on top of high family incomes, a closer inspection can even out the aid (and accepted applications, for that matter). Secondly, it's beneficial to the institution. This is because high GPA, high SAT students are competed for among many institutions--they are visibly "high potential" applicants, despite the 70% of variance in actual performance that is unexplained by these predictors. By developing better predictors, it would be possible to partially avoid the bidding-war that drives up discount rates and puts the cost of running a university disproportionally on low-predicted students (who, as we've noted are on average of lower socio-economic status to begin with).

It's somewhat curious that with this win-win advantage to better predictors that there is not wider acknowledgment that the usual cognitive predictors (GPA, SAT, class rank, etc.) are not sufficient. It's doubly odd since the College Board itself makes no secret of the limitations of the SAT, for example. This applies in general as well as to particular demographics. From the 2008 College Board Research Report Differential Validity and Prediction of the SAT: "The results for race/ethnicity show that for the individual SAT sections, the SAT is most predictive for white students, with correlations ranging from 0.46 to 0.51. The SAT appears less predictive for underrepresented groups, in general, with correlations ranging from 0.40 to 0.46." Those correlations are really only useful as an adjunct to high school grades (and perhaps rank), so one has to take them with a grain of salt. Even so, squaring these to get r2 yields 20% or so of the variance in first year college grades. This is hardly enough to justify the confidence that the general public and even higher education administrators seem to have in the test. I imagine that college ratings systems like U.S. News accentuate this distortion.

We have therefore established that there is a great potential for indicators other than the usual cognitive ones in order to chip away at the 70% or so unknown variance in predictive power. Why noncognitives?

Recently results from an unscientific but interesting faculty poll were announced in a staff meeting at my university. The faculty members had been asked what qualities they would like to see in students. As the deans read their lists, I categorized them for my own amusement and found that two thirds of the traits listed were things like "works hard," "is interested in the subject," and "takes studies seriously." In other words, descriptions of noncognitives--attitudes and behaviors other than grades and test scores. It had been long clear to me by that point that we needed better insight into those qualities of our students, but I was surprised to see how much research had been done on it already.

Even the big standardized test companies are becoming interested in noncognivites. In 2008 ETS announced that it would begin testing the use of noncognitives for augmenting the GRE with a "Personal Potential Index". In this case, this means using standardized letters of recommendation from advisors and professors. The College Board has funded a project (GRANT) at Michigan State University to study noncognitives.

Below are a few selected sources from the literature on noncognitives.

The University of Chicago Chronicle, Jan. 8, 2004, Vol. 23 No 7:

Heckman, the Henry Schultz Distinguished Service Professor in Economics and one of the world’s leading figures in the study of human capital policy, has found that programs that encourage non-cognitive skills effectively promote long-term success for participants.

Although current policies, particularly those related to school reform, put heavy emphasis on test scores, practical experience and academic research show that non-cognitive skills also lead to achievement.
"Numerous instances can be cited of people with high IQs who fail to achieve success in life because they lacked self-discipline and of people with low IQs who succeeded by virtue of persistence, reliability and self-discipline," Heckman writes in the forthcoming book, Inequality in America: What Role for Human Capital Policies? (Massachusetts Institute of Technology Press), which he co-authored with Alan Krueger.
Noncognitive predictors of academic performance: Going beyond the traditional measures, by Susan DeAngelis in Journal of Allied Health, Spring 2003.
The purpose of this study was to provide an initial investigation into the potential use of the PSI as a noncognitive predictive measure of academic success. As programs continue to experience demands by allied health professions for their graduates, admissions committees should employ highly predictive and valid criteria to select the most qualified applicants. Although it is impossible to select only candidates who are ultimately successful, admissions decisions must be based on the best data available to reduce the risk of attrition and to increase the number of newly licensed graduates entering the profession. The preliminary findings indicate that the PSI moderately enhanced the predictive capacity of the traditional cognitive measures of entering GPA and ACT score.
Probably the best source for the practical use of noncognitives in higher education, and a rich source of literature on the topic, is William E. Sedlacek's Beyond the Big Test: Noncognitive Assessment in Higher Education. The book describes itself thus:

William E. Sedlacek--one of the nation's leading authorities on the topic of noncognitive assessment--challenges the use of the SAT and other standardized tests as the sole assessment tool for college and university admissions. In Beyond the Big Test, Sedlacek presents a noncognitive assessment method that can be used in concert with the standardized tests. This assessment measures what students know by evaluating what they can do and how they deal with a wide range of problems in different contexts. Beyond the Big Test is filled with examples of assessment tools and illustrative case studies that clearly show how educators have used this innovative method to:

• Select a class diverse on dimensions of race, gender, and culture in a practical, legal, and ethical way

• Teach a diverse class employing techniques that reach all students

• Counsel and advise students in ways that consider their culture, race, and gender

• Award financial aid to students with potential who do not necessarily have the highest grades and test scores

• Assess the readiness of an institution to educate and provide services for a diverse student body
Sedlacek identifies eight dimensions of interest:

1. Positive self-concept
2. Realistic self-appraisal
3. Successfully handling the system
4. Preference for long-term goals
5. Availability of strong support person
6. Leadership experience
7. Community involvement
8. Knowledge acquired in a field

He also gives research findings and processes for reviewing applications, for example, to identify and rate these traits.

Research like the sources cited above show promise in the use of noncognitive predictors. And there is clearly a need for better predictors of student success. Nevertheless, we must be careful not to underestimate the difficulties or imagine that "noncognitives" are a panacea. Dr. Sedlacek's work and others indicates to modest gains in predictive power based on surveys and other methods of gather information about attitudes and behaviors. In his paper with Julie R. Ancis "Predicting the Academic Achievement of Female Students Using the SAT and Noncognitive Variables," the noncognitive instrument only predicted a handful of percentage points in the variance of GPA, a result that has been indicated by others as well.

Additional complications are introduced if the methods of assessing noncognitives are standardized and published. Even tests of cognitive skills are gamed, as with SAT preparation tests that teach not just content, but also test-taking strategies. It's easy to imagine how much easier it would be for an applicant to a competitive institution to "game" a formalized noncognitive survey. This imposes constraints on how we can approach the problem.

Within these limitations, there is still opportunity to develop the use of noncognitives in useful ways. Obviously, applying such indicators in an admissions setting must be done carefully. But the usefulness of knowledge about what traits lead to success is not limited to admissions. Student intervention, orientation, and even the curriculum itself can benefit from consideration of skills that are not traditionally academic. The AAC&U's LEAP initiative, for example, lists noncognitives in their recommendations for general education (teamwork and civic engagement, for example).

In summary: there is motivation to find new predictors of academic performance, and there are indications that noncognitives are promising. There are, however, obstacles that can be best met with a collaborative, wide-ranging effort that focuses on practical effect.

Monday, September 28, 2009

Predicting Success

Nature abhors a vacuum it's said. My daughter showed me the other day how her science class used this "principle" to determine the amount of oxygen in the air by using a candle and a test tube in water to measure before and after. Of course, it's isn't really abhorrence but air pressure that makes vacuum a chore to create down here where we live. But maybe economics really does abhor a passed-over opportunity. As in the old joke where one economist says "hey, there's a hundred dollar bill laying there on the ground," to which the other replies "can't be so--someone would have picked it up."

I've argued for a while that the low predictive validity of GPA + SAT creates market opportunities for those willing to experiment with other demonstrations of achievement. Here's a list of previous posts related to the topic:
In Malcolm Gladwell's Outliers, he has the following observation about the international math and science test called TIMSS . He notes that the test is accompanied by a 120-question survey, which many students don't complete. In his words:
Now, here's the interesting part. As it turns out, the average number of items answered on that questionnaire varies from country to country. It is possible, in fact, to rank all the participating countries according to haw many items their students answer on the questionnaire. Now, what do you think happens if you compare the questionnaire rankings with the math rankings on the TIMSS? They are exactly the same. (pg. 247)
He concludes a page later that "We should be able to predict which countries are best at math simply by looking at which national cultures place the highest emphasis on effort and hard work."

One could certainly ask for better analysis--why not correlate by student the number of survey items completed against the math score, rather than aggregating by country? But the sentiment is certainly the same expressed in noncognitive literature: there's more to success than the ability to do mental gymnastics.

In InsideHigherEd today there was an article "Next Stages in Testing Debate" talking about institutions that de-emphasize SAT in admissions decisions:
[A] common idea was that decreasing reliance on the SAT does not mean any loss of academic rigor and can in fact lead to the creation of classes that do better academically (and are more diverse).
This may require a rethink and additional training of admissions staff:
That report said that for too many admissions officers, the only training they receive on the use of testing may come from the technical training provided by testing companies, entities that have a vested interest in the continued use of testing.
Some of the other types of accomplishments sought by admissions officers are evidence of creativity, practical skills, wisdom about how to promote the common good (Tufts), and essays (George Mason). This idea is something that seems to be blooming. As Mr. Gladwell might say, it's blinking toward an outlying tipping point.

Evidence of this arrived in my in-box the other day: a forwarded email from the Law School Admisssion Council (LSAC) with the following news:
LSAC has funded research on noncognitive skills for some time. A study funded by LSAC—Identification, Development, and Validation of Predictors for Successful Lawyering by Marjorie Shultz and Sheldon Zedeck—identified 26 noncognitive factors that make for successful lawyering. The study included suggestions for assessments that might measure those factors prior to admission to law school.
I hadn't heard of Shultz and Zedeck, so I scurried off to the tubes that comprise the Internets to find out more. You can find the whole 100-page report and more here. The executive summary has an imposing, lawyerly warning on the front page:
NOT TO BE USED FOR COMMERCIAL PURPOSES NOT TO BE DISTRIBUTED, COPIED, OR QUOTED WITHOUT PERMISSION OF AUTHORS
I suppose this implies an argument that fair use somehow doesn't apply to this work. In any event, since I've obviously already violated the terms with the quote above, I may as well proceed.

Rather than focus on something easy like just predicting law school GPA, the researchers actually assessed job performance and got LSAT and law school performance data. Then they threw a bunch of tests at the problem, like Hogan Personality Inventory, Hogan Development Survey, Motives, Values, Preferences Inventory, and Self Monitoring Scale, a Situational Judgment Test, and Biographical Information Inventory. This on a sample of more than 1100 subjects. I think the word I'm looking for is "wow."

The results found successful indicators of effectiveness (per their definition) that seemed to be assessing independent characteristics, adding dimensionality to the standard predictors. In fact, the LSAT didn't seem to predict success at all. Go look at the executive summary for details--it's not very long.

All in all, this seems to be a solid study that shows that noncognitives are important--perhaps better--predictors of professional effectiveness than the traditional cognitive ones.

Other noncog stories in the news:
The College Board is even getting into the act. A 2004 publication talks about "individualized review" that includes factors beyond GPA and test. I am led to understand that they have a big project underway now on noncogs, but I can't find the website.

Okay, remember the vacuum? What if we managed to identify these noncognitive variables and start to use them? The LSAC folks outline what can happen next:
A major concern about developing an assessment for noncognitive factors is the possibility that the test would be so coachable that its results would be unreliable in the high-stakes environment of law school admissions.
I argued here (Zog's lemma), here, and here that any imperfect predictor invites error inflation for economic gain, but it doesn't take a genius to see that if checking "I'm a hard worker" gets me more financial aid, I'll be more inclined to overestimate my industriousness. This is a game theory problem with no solution that's likely to be mass marketed. Imagine if the assessment of prospective students had to be done laboriously by hand by highly trained admissions staff, instead of relying on a convenient test that cranks out a one-dimensional predictor. I'm not sure that's a bad thing.

Saturday, March 07, 2009

SAT Validity

The College Board has a nice page on Data and Reports, with a report on SAT Validity showing correlations between SAT and high school grades (HSGPA) and first year college grades (FYGPA). The table below from the report shows these statistics.

There is also a report on demographic breakdowns called Differential Validity and Prediction of SAT, which shows differences between race and gender groups. On the whole, predictive validity doesn't vary much, especially considering that what we're interested in usually is R-squared.

Since I've been interested in using noncognitive indicators to predict achievement, I looked for some statistics to guide me. All I've found so far is in the paper "Predicting the Academic Achievement of Female Students Using the SAT and Noncognitive Variables" by Julie R. Ancis and William E. Sedlacek. In the study used in the paper, two of the noncognitive dimensions seemed to be the best predictors (my interpretation): realistic self-appraisal and community service. There wasn't information about the residual when GPA and SAT are also taken into account, but putting these together, one can create an approximate "best case" scenario based on the statistics. This is shown in the chart below.



The contribution of the noncognitives was to me disappointingly low, and it is likely even smaller because of the correlation with SAT and GPA, which is unknown here. That is, the green slice might actually overlap with the red or blue ones.

It's interesting to note that that "real-life" correlations like the one in the Ancis and Sedlacek study are lower between SAT and FYGPA than in the College Board's research reports. This has been my experience too, and I can't explain the difference. I came across a rather impassioned argument for dropping the SAT as a predictor in a 1992 proposal from Jonathan Baron at University of Pennsylvania. His correlations are even lower than mine for SAT and FYGPA, and he argues that there's not enough value added by the SAT (after taking into account other predictive variables the university uses) to justify its continued use. He makes an interesting point about the emphasis on SAT in the admission process:
College admissions criteria have major effects on high-school
education. College-bound high-school students do what they think
will help them get into a good college. Students now spend a
considerable amount of time preparing for the SAT. (When my son
was taught how to take multiple-choice exams in Kindergarten,
when he was in a group of children who could already read, I
complained that this was an inappropriate activity, and I was
told that it's never too early to start preparing for the SAT!)
If they were told that the SAT was not important but their grades
and their achievement test scores WERE important, they might
spend more time trying to learn something and less time trying to
learn how to appear to be intelligent on a test. This might be
reason enough to drop the SAT, even if it were somewhat useful
for prediction.
The trend toward standardized testing as a measure of minds is not limited to SAT. Now we have the whole No Child Left Behind apparatus and an increasing appetite at the Department of Education to infect higher education with this philosophy using the likes of the CLA. I hope that instruments using noncognitive assessment can get more attention and be developed into something useful.

Monday, February 09, 2009

Beyond the Big Test

This weekend I got my copy of William E. Sedlacek's text with that title, which is also subtitled Noncognitive Assessment in Higher Education. Anyone who's read my blog in the last year will know I've scratched my head over the evident value that admissions processes typically put on ACT/SAT, and to a lesser extend high school grades (see this article, for example). Our work with surveys like the CIRP shows that there are behavioral and attitudinal traits that make good predictors of, for example attrition. These are generally ignored in the admissions process. Or it might be better to say they're not looked for.

Well, silly me. Apparently this subject has a long history in the literature and is a quite well developed concept. Some major schools like North Carolina State University have used these methods successfully. So my research program to find variables for use in predicting academic success can accellerate considerably--we just have to customize the work others have done.

I have not finished the book, but can tell already that it's a wonderful resource. The background and history of noncognitive assessment is given, as well as solid research findings, actual survey instruments, and examples of how to coach staff in looking for these traits in interviews, on existing application materials, essays, etc.

The focus of the book is toward evening the playing field for what the author calls non-traditional students. In his usage, this means anyone who isn't a white male. My purpose is more targeted, but the material is no less useful for it.

The specific noncognitive traits identified, with some help from factor analysis in the process of ascertaining construct validity of surveys, is as follows (pg. 7):
  1. Positive self-concept
  2. Realistic self-appraisal
  3. Successfully handling the system
  4. Preference for long-term goals
  5. Availability of strong support person
  6. Leadership experience
  7. Community involvement
  8. Knowledge acquired in a field
These are described in detail, of course, along with methods of assessing them.

I'll be recommending that our university proceed full speed ahead with this project, to catch what we can of the current cycle. This will undoubtably make Mr. Sedlacek happy, as it will entail buying many more copies of his book for distribution.

Sunday, October 11, 2009

Michigan State Noncognitive Study

A "Report of the First-Year Follow-up of College Applicants at Twelve Universities" was forwarded to me by a colleague. According to him, the report is public now, but the appendices with survey instruments are not. I haven't seen the report itself online, and because of copyright don't feel like I can post it. But herein are some highlights.

The authors are Neal Schmitt, Abigail Billington, Juliya Golubovich, Jessica Keeney, Timothy Pleskac, Matthew Reeder, Ruchi Sinha, and Mark Zorzie, and the work was supported by the College Board, which is interesting.

The basics:
  • Participants: Earlham College, Furman University, Johnson & Wales University at Providence, Kenyon College, Lafayette College, Meredith College, Michigan State University, Ohio State University, Purdue University, University of North Carolina at Chapel Hill, University of Southern California, and University of Washington
  • 844 students provided enough data for analysis, initially through a College Board web survey of applicants on background, interests, and judgment, then through a follow-up for those who enrolled. The overall yield rate (enroll/app) was 26%. Of those who enrolled, 42% responded to the follow-up survey (there was a $20 incentive).
  • The kinds of data considered were biodata (background and life history, not blood pressure), a situational judgment test (SJT) representing behaviors, demographics, self-reported performance (BARS), citizenship behaviors (positive or negative), academic satisfaction, social satisfaction, grades and standardized test scores, an inventory of "shock" events a student may have experienced, the big five personality traits, substance abuse, use of time, and self-reported withdrawal tendency.
The report is chock-full of tables of data with correlations and crosstabs, but I'll skip to the regression results. Here are some key findings, quoted from the executive summary (pg. 3):
  • First-year college GPA is predicted significantly by several biodata scales, most notably Knowledge, Ethics, and Perseverance, but HSGPA and SAT/ACT scores are much more predictive of college GPA than are biodata and SJT.

  • Self ratings of performance (BARS), Organizational Citizenship Behavior (OCB), and student self-reports of Deviance were especially well predicted by the biodata measures and SJT while HSGPA and SAT/ACT were relatively uncorrelated with these outcomes.
It's disappointing that first year GPA isn't better predicted by this gob of noncognitive variables. But completion is a better goal, and the study hasn't had time to mature to that point. For example, any actionable information about first-year retention would be worth its weight in undergrads. Stay tuned for the next report.

Update: Dr. Neal Schmitt gave me a link to a publications page, which will soon include the body of the report I cited.

Friday, August 21, 2009

The Use of Grades

There's been an interesting discussion on the ASSESS listserv about assessments vs. grades, which led me to think about the relative uses of each. Part of the problem in addressing that difference is that grades come in all sorts of flavors that may or may not resemble an outcomes assessment. For example, a carefully designed calculus final exam may pass quite nicely as a summative assessment. By contrast, a student who increases his final grade by attending an extracurricular event (say in an orientation course) has little claim that the grade is an reflection of his performance in some cognitive skill. The summing of individual grades into a cloudy average makes this worse (see: Statistical Goo). The same can be done for assessments, of course.

One fundamental difference between grades and assessments is that the former is used to motivate students. In fact, we've created a whole industry that depends on this kind of motivation, from accreditation downwards to the classroom: a kind of "do this or else" mentality. Surely this can't be ideal. I appreciate the requirements of state education to more or less force kids out of their beds in the morning, onto busses, and subject themselves to ideas that hurt to absorb. But higher education? Especially in the liberal arts, we talk about general goals like creating life-long learners and such. How does that square with our methods of delivery?

That discussion comes back to the topic of noncognitive traits in learners: how motivated is Tatiana to learn computer programming? For a motivated learner, assessments are better than grades because they get straight to the point of "how well am I doing?" without the coercive baggage that's inherited from elementary school. This is, by the way, an argument for detailed reports in assessments as well. The more gooey they become, the more they resemble grades and the less useful they are. Compare:
Tatiana sees her score of 79% on the C++ test and concludes that she is doing okay, but not excelling.

Tatiana reads her C++ assessment and sees that while basic control structures are second nature to her, she really doesn't understand pointers.
Only the second of these is actionable--Tatiana can increase her skills by practicing with pointers, getting tutoring, reading what others have to say about the subject. Learning is about details, and assessments should be too. With the ubiquity of modern information systems, keeping track of details isn't a problem--we just have bad habits left over from the grade school mentality of reducing a semester's work to a single letter. It's absurd, if you think about it.

Can we find another way to motivate students? I don't know, but I don't think we're really trying. I can imagine a culture that fosters a more inquisitive approach to self-improvement, but don't know how that might be engineered. And yet, we don't really want to produce graduates who only perform in order to get the pellet in a Skinner box, do we? I think the first step is to start to assess certain noncognitive traits, and bring them into the curriculum (and not just in orientation course). I'm not saying that all students lack self-motivation, of course. We have and value the go-getters in class who drive classroom discussions, ask for new things to read, and always come for help when they don't understand. How do we harness that energy to help pull along students who aren't so energetic in their learning practice?

Until such a motivation exists, it's probably best to keep grades to push students along and keep the political pressure off assessments. It's very convenient for the Assessment Director to not have to worry about the kind of scrutiny that the registrar requires. But it is a capitulation of sorts.

Tuesday, September 15, 2009

Motivation, Outcomes, and Bandwidth

I suppose every generation creates its share of hideous neologisms and circumlocutions. For me "on a regular basis" (instead of "regularly"), is one of the worst. I made the mistake of telling my daughter how much I hate it, and now she uses it regularly just to annoy me. But "incentivize" must also rank up there the list of 20th century abominations. The idea is simple, and is probably linked to the recently debunked myth of markets that are perfectly efficient and people who always act in their own best economic interests. Paul Krugman describes this empirical enlightenment of economists vividly in his recent New York Times piece "How Did Economists Get It So Wrong?".

It reminds me of something I read a long time ago about an engineer doing research on the effect of a crowd on wireless transmission, beginning with: assume that a person is a one-meter diameter sphere of water... Assumptions and approximations have to be watched carefully.

I was reminded of "incentivize" a couple of days ago when I came across a TED talk by Dan Pink on the science of motivation. It's a 17 minute video you can see here. I will summarize some of his points here, but you might find it more interesting to watch the video first. The topic of the talk was motivation: what types work in what circumstances. This is an interesting topic in higher education because of various obvious reasons, including one you may not think of right off. More on that later.

Motivation isn't as simple as it seems, it seems. Consider the crazy things we do in the name of motivation. A car cuts us off in traffic and we may get angry, wave interesting gestures and honk at the driver. We might even call the police and report it if the behavior is egregious enough. Why are we doing that? At the bottom of it, I would posit that we are unconsciously trying to incentivize the driver not to do such things again through negative reinforcement. But in most cases, the odds that we will ever encounter this particular driver in that situation again are probably remote. (Consider how you might behave differently if it were your neighbor rather than a stranger, to see some of the complexities here.) So we are wasting our time at best, and possibly even acting against our own best interests. But such inclinations run deep. It's worth bringing their effects to the light of day.

We seem to have an instinct to incentivize (I'm going to get that word out of my system), even when it isn't likely to do any good. As pointed out above, our actions could make the situation worse. I assume that there are evolutionary reasons for these tendencies: a million years of living in social groups must have had some effect on our programming. There seems to be a general societal purpose to such actions (see this article, for example).

So now to Dan Pink's talk (spoilers ahead). He makes the same point forcefully, applied to the workplace: managers don't understand incentives and are mostly doing the wrong thing to motivate employees. Here are some quotes from the transcript.

Sam Glucksberg did a Candle Problem experiment to learn about the effect of incentives:
He gathered his participants. And he said, "I'm going to time you. How quickly you can solve this problem?" To one group he said, I'm going to time you to establish norms, averages for how long it typically takes someone to solve this sort of problem.

To the second group he offered rewards. He said, "If you're in the top 25 percent of the fastest times you get five dollars. If you're the fastest of everyone we're testing here today you get 20 dollars."
It took the second group three and a half minutes longer on average to solve the problem. This is counter-intuitive. As Dan Pink puts it:
You've got an incentive designed to sharpen thinking and accelerate creativity. And it does just the opposite. It dulls thinking and blocks creativity.
And most amazing is the claim:
This has been replicated over and over and over again, for nearly 40 years. These contingent motivators, if you do this, then you get that, work in some circumstances. But for a lot of tasks, they actually either don't work or, often, they do harm. This is one of the most robust findings in social science. And also one of the most ignored.
Further research, again using the Candle Problem, showed that incentives can positively affect performance, but only when the problem was simplified to a rote (I would say low-complexity analytical) task. The result reinforces the idea that the prospect of immediate reward (or punishment) causes us to reduce creative, conceptual approaches to problems in favor of direct obvious connections.

Since this is September, it's easy to make the leap to 9/11 as an example. In the early 1900s, the czarist Russian security organ had an imaginative idea: what if the terrorists (which were blooming everywhere) got it into their heads to crash an airplane into a building? I read about this in Orlando Figes' A People's Tragedy. Compare that creative idea to the actual response to the horrible actuality: a very narrow focus on preventing exactly the same thing from happening again. This isn't criticism--it's very natural and sensible--the point is that a big whomping motivation narrowed attention to what exactly the problem was seen to be in an obvious sense, not what it could be in the larger sense. If you read my last post about generalizing through recursion, you'll see what I mean. Here's a list of Bad Things That Can Happen, which have nothing to do with taking off your shoes before you board a plane:
  1. A near-earth asteroid could hit us
  2. The caldera at Yellowstone could blow
  3. Gene hacking of biological viruses becomes as common as computer-virus hacking, and some 18 year-old sets off a catastrophe (you do remember the 90s?)
  4. Nanotech goo takes over the world
  5. The oceans turn to acid as the climate heats up
  6. Environmental toxins are having epigenetic effects that will last generations
  7. A housing bubble could blow up the banking industry.
These are the sort of broad-spectrum dangers you're not likely to think about while your house is on fire. I think it's fair to summarize the findings Dan Pink describes as "incentives narrow focus."

If these findings are valid, much of the way management is done in business, including higher education, is wrong. In the video, Mr. Pink talks about some alternatives.

Relating this to education. If you've been in the classroom, you've probably been as frustrated as I have been by the question "is this going to be on the test?" To the mind of a "lifelong learner" this is entirely the wrong attitude. But it's easy to see from the perspective of motivational cause and effect that the incentives we apply would lead directly to that question. To use that awful word again, we incentivize students with grades. Why should we find it surprising that they tend to narrowly fixate on grades?

Dan Pink has created a consulting business out of this idea, where he pitches:
And to my mind, that new operating system for our businesses revolves around three elements: autonomy, mastery and purpose. Autonomy, the urge to direct our own lives. Mastery, the desire to get better and better at something that matters. Purpose, the yearning to do what we do in the service of something larger than ourselves.
Whether or not this is a formula that works, or is Utopian dream is unknown, but the ideas are certainly worth considering. Notice that of autonomy, mastery, and purpose, only the second one is a cognitive skill. The other two are affective, or noncognitive, or in jargon-free plain English: emotions.
I think there is untapped opportunity to try to engineer ways to motivate students more sensibly--to model and inculcate autonomy and purpose, and illuminate the role of zeal in creating mastery. That's all good, and I think there are opportunities especially for small liberal arts schools here. It's particularly ironic that incentivizing may actively hinder the teaching of critical thinking, which is supposed to be what grades-bound liberal arts colleges are suppose to be good at. But there's a bigger question.

"A Virtual Revolution is Brewing for Colleges" from The Washington Post is the latest article I've seen predicting the doom of traditional higher education. I blogged about this topic recently here. The argument is that for many subjects, education can happen conveniently and cheaply over the Internet, and that competition will drive the bricks and mortarboard model to ruin (except for the elite institutions that rely on deep pockets or have massive self-sustaining endowments).

What, aside from inertia, stands in the way of this transformation? I think one of the biggest problems for online education is the noncognitive load it places on the consumer: they have to be motivated to log in and do the work. They have to minimize the distractions of Facebook and a million other things while working on the computer. I don't have any statistics to back this up, so I may be completely wrong. But it seems to me that one current advantage of a residential school is the pervasive culture that comes with it. As imperfect as it is, there is social pull to come to class and not humiliate oneself by flunking every test. In the nearly anonymous hyperspace of online classes, I imagine that this is less so. In-person interactions are naturally more engaging that online ones. Is that really true? If so, how long will it remain true?

For the moment, I can't conceive that online teaching can approach the richness of a good professor's interactions with students in class, the dining hall, and in the office--the social engagement that includes mentoring and a kind of tribe-like kinship that comes from the circle of mutual acquaintances and shared experiences that play out in full-color, real-time, three-D.

In short, the bandwidth for real life ("rl" in cyberspeak, contrasting with, say, "vr" for virtual reality, or specifically "sl" for the online world Second Life) is still far superior to anything modems can deliver. But rather than leveraging this advantage, rl institutions waste most of the bandwidth. We're generally not engaging students on autonomy and purpose in the pursuit of mastery. We care about grades and bureaucracy and grants from the government.

Tentative conclusions. There may be a niche for second-tier institutions in the new education landscape to provide premium rl education, but only if they seriously address the engagement problem. George Kuh of the NSSE comes at this from another angle, and he actually does have some statistics. The point isn't just the survival of the traditional model, it's to provide a service that is superior to online education because of bandwidth and proximity: a million years of evolution has programmed us to live reasonably well together in social groups, and that should be taken advantage of.

De-emphasizing traditional grades is one step in that direction. Read my post about Western Governor's University to see a model for how that is already happening (online). But that's only part of it. A portfolio that a student can carry with them (and accumulate as a life-long resume) could contain evidence of not just subject mastery but also explicitly address noncognitive traits like purpose. Higher education has been allergic to "purpose" since it became largely secular. It's time to reconnect with the big "why" questions outside of a hermetically sealed philosophy course.

Online education isn't going to stand still, of course. Already there are very motivated people working on the problem of creating learning communities online. These visionaries think big, and for the most part, think "cheap" or "free" (search open education on this blog). Bandwidth will increase, rl will become more conflated with vr, and the next generations will perhaps feel at home in a warm LCD glow as they do in rl. For institutions frittering away their bandwidth advantage now, remember what happened to CDs. MP3s are generally inferior to CDs because the latter are compressed. But MP3 rule because bandwidth loses to convenience.

We might think of online education as a low-pass filter that employs only the deep end of the spectrum, like the telephone company only transmits a small range of frequencies when you talk in order to save money. What value is the high-frequency stuff? Can it be used to engage students in ways that vr can't approach? I don't know, but I think for the time being the answer is yes. I'll now take off my sackcloth, shave my beard, abandon the giant urn, and give up this prophetic conceit (I can't go to work like this) after one final prognostication:

It's a good time to be an energetic new college or university president with a creative, entrepreneurial spirit. There are opportunities. It's a bad time to be locked into the traditional model of private higher education that demands a high price and then blows the bandwidth.

Thursday, August 06, 2009

Grit, Grades, and Graduation

It's funny how categories affect thinking. Since becoming interested in the noncognitive traits of students, their assessment and consequence predictive powers, I've begun to see noncognitives everywhere. The August 2nd article "The Truth about Grit" in the Boston Globe is an example.

The point of the article is that intelligence doesn't guarantee success; success also requires perseverance in the face of obstacles, or "grit."
The hope among scientists is that a better understanding of grit will allow educators to teach the skill in schools and lead to a generation of grittier children.

[...]

The new focus on grit is part of a larger scientific attempt to study the personality traits that best predict achievement in the real world.
The US Army has supported the research with some interesting findings. The example of the US Army's military academy West Point is compelling:
The Army has long searched for the variables that best predict whether or not cadets will graduate, using everything from SAT scores to physical fitness. But none of those variables were particularly useful.
What did work was a survey that assessed perseverance. The article notes that this echoes Francis Galton's 1869 research findings that concluded that a prerequisite for higher-order achievement was “ability combined with zeal and the capacity for hard labour.”

The whole article is worth reading. There is a link given (indirectly) to a grit survey developed by A.L. Duckworth at the University of Pennsylvania. The project is applicable to higher education:
Duckworth has recently begun analyzing student resumes submitted during the college application process, as she attempts to measure grit based on the diversity of listed interests. While parents and teachers have long emphasized the importance of being well-rounded - this is why most colleges require students to take courses in all the major disciplines, from history to math - success in the real world may depend more on the development of narrow passions.
This should be terribly interesting for liberal arts schools particularly. If perseverance is tied to singular passions, then how does this interact with the "broadening of the mind" mission of the institution? Is the implicit goal of success after graduation secondary, or should we try to pull off both? This has direct implications to learning outcomes assessment, as noted in the article:
In recent decades, the American educational system has had a single-minded focus on raising student test scores on everything from the IQ to the MCAS. The problem with this approach, researchers say, is that these academic scores are often of limited real world relevance.
Supposing we took the idea of grit seriously. That would start with using instruments like the one Dr. Duckworth has developed to estimate it. Can we then teach it? Should we?

Dr. Dweck at Standford University is quoted referring to a "growth mindset" versus a "fixed mindset," the difference being what we believe about our abilities. Can we grow them, or are we fixed with them? Fixed mindset learners are more likely to give up when encountering obstacles, assuming they're just not talented enough. Dweck's research, according to the article, demonstrates that the growth mindset can be taught effectively.

Ironically, praising children for their intelligence may make them less likely to succeed because it reinforces the fixed mindset. A better strategy is to reward effort and hard work. I've blogged about this subject before, and Malcolm Gladwell's Outliers is in much the same vein, I believe (I have only read reviews of it to date).

Should we assess these traits? One of the results of reporting assessment results for learning outcomes is that it becomes evident that students can earn decent grades and graduate without scoring high on assessments. If you've been in the classroom any length of time, you've probably encountered this student, whom we'll call Joe. Joe works very hard in your Finite Math class, forming study groups outside of class, turning in all the homework, coming to office hours. But the material just doesn't click with him, and he has a terrible time of coming to grips with the key concepts. Nevertheless, he gives each exam his full efforts and manages to pass with a B. Everyone is happy about this achievement, no? After grades are posted, Joe announces he wants to major in math. You consult with your colleagues and worry together for a bit. Although Joe has worked very hard and earned his B in everyone's eyes, he clearly lacks some spark of imagination that makes doing math rewarding. You fear he's setting himself up for failure.

The point is two-fold. One is that we assess outcomes (subjectively and informally in Joe's case) and we assign grades, but they mean different things. Tied up in grading is the notion of perseverance, of meeting each of the many assignments head-on and getting through them. But in the gestalt there may be something missing; the pieces do not always make the whole. The second point is related: we attempt to estimate the cognitive outcomes with our formal assessments, but generally do not assess the noncognitives. Shouldn't we doing so? If the science quoted is valid, then grit has as much to do with success as intelligence.

If you've browsed my Assessing the Elephant piece, or read enough of this blog (e.g. here), you can guess where I'm going. Why not assess grit along with thinking and communication skills across the curriculum, in a minimally-intrusive survey? It works for the cognitive skills, so there's a chance it will work for the noncognitives. Wouldn't it be fascinating to be able to compare cognitive skills, grit, grades, and graduation rates?

Please note that this gives institutions a way to include students themselves in their assessment reports. My last post was about publicly reporting learning outcomes and results. For example, Capella University's website has a page for Learning Outcomes. This is the future--gaze well upon it, ye assessment directors:
As this practice becomes common, the question will be asked "why doesn't every student get the highest rating?" Is this a fault with our education? At present, we could only shrug our shoulders and perhaps mutter about SAT scores of incoming freshmen, or if accreditation is imminent the glorious plans for improvement we've printed in reports. But if we actually assessed noncognitive traits like grit, we could report out richer details. We might note that the students who failed to graduate also were rated low in perseverance. We could isolate those with the most assessed grit and see how this relates to cognitive development. We could include noncognitives in the curriculum itself and begin to take it seriously, just like effective writing. For liberal arts institutions, the tricky question of balancing broadness with singleness of purpose could be explored with at least some minimal data.

It's an experiment worth doing.

Tuesday, August 25, 2009

Zog's Lemma: Assessment and the Amplification of Error

The concept of outcomes assessment is like a Swiss Army spatula: it's used in all kinds of ways. As opposed to the proverbial Russian Army hardware, pictured below (ubiquitous on the Internets):
To some, outcomes assessment is a touchy-feely endeavor of encouraging the practitioners of higher education to do the right thing, close the right loop, bring a glowing smile to the visiting team. All that. I'm comfortable with that.

To others, it's a more serious matter, more scientific in approach, rather like making tick-marks on the door's threshold on a child's birthday to signify evidence of growth. We know that is serious business: with shoes or without, and where exactly is the top of the head? Does hair count or not? It grows too, after all.

The scientific approach requires belief in theory: that assessments are valid to the purposes we employ them for. In reality we usually have no real way to know that based on a solid physical theory. The belief has to be defended by what statistics can be summoned to make a case. (In truth, beliefs don't need to be defended at all; they only need to be believed.) What comprises this validity? I'd like to address that question sideways. The discussion of what constitutes validity is a groove cut deeply in the literature of psychometrics, and I'd rather ask a more important question: of what use is it to believe in the validity of an assessment?

It's easy to make hay from the fact that the definition of validity itself isn't settled, but this is unfair, I think. It's a difficult philosophical nut to crack. For my purposes, predictive validity is the most important aspect of assessment. If an assessment doesn't tell us anything about what's going to happen in the future, I don't see much use in it. The Wiki article on predictive validity contains an interesting observation that is the crux of the value of an assessment:
[T]he utility (that is the benefit obtained by making decisions using the test) provided by a test with a correlation of .35 can be quite substantial.
Utility is a concept from economics that allows for a weighting of outcomes to balance outcomes that would otherwise be numerically indistinguishable. A classic example is the fact that $1000 means more to a poor person than it does to a rich person, even though it won't buy any more for the former. The relative worth to the penniless is more. The point is that even imperfect predictive assessments are useful.

An example of this is the college admissions process. Let's suppose that with the data gathered on the application, the institution can estimate the probability of "success" of student. Success could be retention to second year, or GPA > 2.5, or graduation, or whatever you deem important. The statistics can be generated with a logistic regression, which will yield a model with a certain amount of predictive power. No model is perfect, and even in retrospect (feeding the original data back into the model) it will not correctly classify applicants as "predict success" or "predict failure" with 100% accuracy. There will be some proportion of false positives and false negatives. This can be visualized on a Receiver Operating Characteristic (ROC) curve. You can see a bunch of them on google images. The usefulness of the ROC curve is that it lets you visually explore the decision of where to set a threshold for decision-making. If you set it too high, you get too many false negatives (reject too many qualified candidates). Too low and you admit too many false positives.

For an institution, such a tool is obviously useful to believe in: one can work backwards from the desired number size of the entering class to see where the threshold should be set. You could even estimate the number of false positives. Any power to discriminate between successful and unsuccessful students is better than none. Of course there are other factors, such as ability to pay, that make this more complicated. To keep things simple, I won't consider these distractions further.

One can imagine a utopia springing from this arrangement, where the assessments continually get better and the applicants learn to distinguish themselves by giving signals that the assessments can recognize. The first doesn't seem to be happening, and the second has an unfortunate twist. An analogy from biology serves us well here.

The story of the peacock, according to evolutionary biologists, is that a showy mating display is worth the biological cost of making all those pretty feathers because of the payoff in reproduction. The similarity is that when mates choose each other they have limited information to go on--an assessment we could characterize with a ROC curve if we had all the facts. This produces a distortion that favors any apparent advantage in a potential mate, or in the case of the peacock, creates over time a completely artificial means of assessment. The peacock is saying "look, I'm so healthy I can drag around all this useless plumage and still escape predators."

So, if the value of partial assessments is clear from the institutional vantage, it's a different picture altogether from the applicant's point of view. It's interesting that InsideHigherEd has an article this morning on this very topic. According to the article, in response to increasing competitiveness at "top institutions":
[H]igh school students could respond to the pressure by taking more rigorous courses and studying more -- or they could focus their attentions on gaming the system and trying to impress.
A study by John Bound, Brad Hershbein, and Bridget Terry Long shows that while this perhaps motivates students to take a more rigorous curriculum, it also prompts them to spend more time in test preparation, or in games like trying to engineer more time to take the test. Peacock plumage? It's hard to say without seeing actual success rates. It could be that having the willingness to spend all that extra effort is itself a noncognitive predictor of success (that is, not related to the scores themselves, but to the personality traits of the applicant). As Rich Karlgaard put it in Forbes magazine, a degree from an elite institution is valuable because:
The degree simply puts an official stamp on the fact that the student was intelligent, hardworking and competitive enough to get into Harvard or Yale in the first place.
What is clear is that those with the means to game the system are better off than those who do not. So we see things like test-prep for kindergartners at $450/hr in the upper crust of society, but probably not so much in housing projects. This likely produces more false positives among the select group: it's as simple as money buying better access. [Update: see this InsideHigherEd article for some dramatic numbers to that effect. Year over year, SAT scores increased 8-9 points for $200,000+ families, and 0/1 points for the poorest group.]

The economic demand for false positives has become an industry under No Child Left Behind, and that doesn't seem to be changing. The New York Times has published letters from teachers on this topic. Some quotes:
  • [T]he use of test data for purposes of evaluating and compensating teachers will work against the education of the most vulnerable children. It is a mistake to conceptualize education as a “Race to the Top” (as federal grants to schools are titled) — for children or schools. (Julie Diamond)
  • Linking teacher evaluations to faulty standardized tests ignores the socioeconomic impact on a nation that is both rich and poor. Can a teacher confronting the poverty of some children in Bedford-Stuyvesant be made to compete with a teacher instructing affluent children in Scarsdale?(Maurice R. Berube)
  • My job went from teaching children to teaching test preparation in very little time. Many of our nation’s teachers have left their profession because the focus on testing leaves little room for passion, creativity or intellect. (Darcy Hicks)
  • Standardized tests are, by their nature, predictable. Most administrators and teachers, fearing failure and loss of position and/or bonuses, de-emphasize or delete those parts of the curriculum least likely to be tested. The students sense this and neglect serious studying because they know that they will be prepped for the big exams. (Martin Rudolph)
The economics of false positives is clear: test prep is an industry, teachers and administrator and schools are rated by how many they generate. Of course the object is not to create false positives, that's just the result of so much emphasis on an imperfect assessment.

Less attention is paid to false negatives. Research points to noncognitive traits like grit, planning for the future, and self-assessment as being important to actual success, but these are not directly accounted for in standardized assessments. To be sure, college admissions officers look at extra-curricular activities to try to add value to SAT, GPA, and curriculum, but I think it's safe to say that the cognitive assessments are primary. How badly do we underestimate actual performance?

That question comes up when we evaluate or create a predictive model for applicants. The resulting predicted (first year) GPA can be used for admissions decisions and financial aid awards. In the Noel-Levitz leveraging schema applicants are sorted into "low-ability" to "high-ability" bins for individual attention. The question of false negatives is the same as the question "how accurate is the description 'low-ability', based on the predictors?" Not very good, as it turns out.

Every time I ran the statistics I got the same answer: at the very bottom end of our admit pool--those Presidential and provisional admits who were supposed to have the hardest time--about half of them performed well (I usually use GPA > 2.5 for that distinction). Since we reject students below that line, we should assume that about half of those students just below the cutoff would have done as well too. If this is typical, there are a LOT of false negatives.

Remember, this line of thought applies to all outcomes assessments to one degree or another. Let's take a concrete example. Suppose we want to assess vocabulary knowledge of German language students with an exam. Memorizing a single word (including declensions for nouns and conjugations for verbs) is low-complexity. The total complexity is that for one word times, let's say 10,000 total items, less any compressibility of this data. All told, this is a lot of complexity (measured in bits) for a human. So it's a suitable subject for assessing, and the low complexity per item means that each of them can be assessed with confidence.

Supposing that we do not have the resources to test our learners on all 10,000 items of vocabulary, we'll have to sample randomly and hope that the ratios are representative (or else spend a lot of time checking correlations and such). Maybe our assessment is only on 100 items, chosen at random from the 10,000. If Tatiana actually only knows K of the items, then there is a chance of K/10,000 that she will know an individual word, and theoretically score K/10,000 on average (on any sized test). Given this arrangement, what is the trade-off between false positives and false negatives?

We will assume we can use the normal distribution to estimate these Bernoulli trials. This only requires that K not be too large or too small--in those cases the chance of an error diminishes anyway. We already know the mean is p=K/10,000, and the standard deviation is given by SQRT(p(1-p)). Two standard errors is about 3% for large and small values of p, ranging to about 5% for those in the middle. These would decrease for a number of items larger than 100, and increase for a smaller test.

The conclusion is that if we set our cutoff for passing in a usual spot, say 70%, we should expect at least half of the scores in the range 65-75% to be either false positives or false negatives. For example, if Tatiana actually knows 70% of the vocabulary items, she has a 50% chance of passing the test because the distribution of her scores over all possible tests is (almost) symmetrical and centered on 70%.

The actual number of false positives and negatives depends on where the skill ranges of the test-takers lie in relation to the cutoff value. The more there are close to the cutoff, the more errors there will be. I actually witnessed a multiple-choice placement test being graded one time, and inquired about the cutoff to find that the grade to pass was actually less than the average result expected by chance! This certainly reduces false negatives, but I'm not sure about the overall result.

The vocabulary example is a best case, where the items themselves are of low complexity and (we hope) not subject to a lot of other kinds of error. When the predictive validity is actually quite low (as in the example where we explain a small percentage of the variance in the outcome), the proportion of errors in both directions is far worse. What does this mean, then when we "believe" in a test like the SAT and adopt it as a de facto industry standard? Even together with high school GPA, these can typically explain only about 33% of the variance of first year college grades.

First, it is of benefit to the institution to be able to imperfectly sort students by desirability. It is costly if the institution has to bid for the apparent best applicants with institutional aid, and there is incentive to look at noncognitives and better predictors, but I can't see a lot of progress in that direction. So the false positives get bid up with the rest. Meanwhile, the false negatives--those applicants who would succeed but don't show it on the predictor--get passed over for admission or merit aid.

A tentative conclusion is that a weak predictor gets amplified in a competitive environment. This is probably true in lots of domains, like the evolutionary biology example of the peacock. Probably some economist has his/her name attached to it: Zog's Lemma or something. I'll have to ask around. (I just googled it--apparently there is no Zog's Lemma.)

For the assessment types, there are two lessons. First, there are going to be errors in classification. Second, those errors get amplified as the assessment gains importance. This is an argument against high-stakes assessments, I suppose. I've always gotten good results by keeping assessments free from political pressures like instructor or program review. "Accountability" creates problems as it tries to solve them. Acknowledging that would be a fine thing.

PS, if you're interested in other kinds of amplifiers in science, read this.

UPDATE: see Amplification Amplification

Sunday, September 13, 2009

Recursive Critical Thinking

I was going to entitle this piece "critical thinking squared" as a cute way to imply critical thinking about critical thinking, but the imprecision bothered me. Squared means multiplication by self, and that's not the same as applying a process to itself. What multiplication means in this context isn't precise either, but you can possibly make a sense of it by considering a combinatorial factorization into dimensions like critical thinking = (analysis, creativity, communication). If we abbreviate critical thinking = CT, then CT2 might look like a matrix:


AnalysisCreativityCommunication
AnalysisAnalysisCreativity * AnalysisCommunication * Analysis
CreativityAnalysis * CreativityCreativity Communication * Creativity
CommunicationAnalysis * CommunicationCreativity * CommunicationCommunication

This assumes that each dimension is idempotent (meaning S*S = S), and that "*" is some way of combining the two dimensions. You still have to figure out what the Creativity * Analysis combination means, but at least you have a way to produce detail from the squaring operation. But this is all rather silly, and the reason I don't like the "CT squared" idea.

Here's a better way to think about it. If you talk about talking, that isn't (talking)2, but rather talking(talking), expressed here as a function that takes itself as an input.

This is even practical. For example, you could write a function get_loc(...) in the C programming language to take a function and return its address in memory. Then you could ask for get_loc(get_loc) to retrieve its own location. This sort of thing is called recursion, and it's a big deal in computer science. In common parlance, we might stick the prefix meta- in front of the concept to show that it's recursive, as in metacognition, which in the right context we justifiably call a noncognitive trait: the reflective practice of thinking about one's own thinking process. More on the relationship between CT and noncogs later. First, let's take a closer look at the role of recursion in problem solving.

In a few stolen moments this morning I was sipping an iced latte, enjoying a cool breeze, and trying to make some progress on a research project regarding survival in a certain abstract sense. You can read the actual paper here, but in a nutshell you can imagine an environment that poses survival challenges to organisms (all abstracted into computer language, which you can learn more about and download a simulator here). There are two different questions regarding the complexity of a given environment:
  1. What's the simplest thing that can survive the given conditions?
  2. What's the simplest recursive process that can find the solution to #1?
Here, recursion means that some process can be tried over and over again, feeding the output of the last iteration into the input of the next. Like natural selection, for example, acting recursively on the gene pool to blindly hone the fitness of the survivors.

Let me give a more down-to-earth example, as a simple "critical thinking" problem. Suppose Tatiana works all day in retail, and part of her job is to calculate sales reductions for coupons, sale prices, and so on. In addition she has to add sales tax to total amounts. For reasons known only to management, they skimped on her point of sale (cash register) and these functions are not included. So all day long she has to do stuff like:
  1. Find actual cost of an item by reducing for sale price, coupon, etc.
  2. Sum adjusted prices
  3. Calculate tax
  4. Add to get total
This is pretty tedious and prone to error, so there's an advantage to having the best possible way of doing this. In the context of my framing questions, we should ask:
  1. What's the best way of doing her job?
  2. How do we find it?
In practice, we have to address the second before the first. The second we might call a critical thinking exercise, requiring analytical and creative thought.

Tatiana need not be reflective. She probably already has a solution, and may not care that it's not optimal. I see this all the time in real stores. A clerk wants to reduce an item by 15%, say. Most often they multiply the price by 15% (.15) on a calculator, write down this number, and then subtract from the original price. Sometimes I tell them a quicker way to do it: just multiply the original price by .85 and you're done.

In a complex environment, you're never really finished with the "how do we find a better solution?" step. The question itself is recursive: "how do we find better ways of finding better solutions?" Mathematics is full of this sort of thing. You can follow the chain of meta-thought all the way up to something called category theory, where logic itself can be generalized (logic about logic).

So without really thinking about the definition, the idea of "critical thinking" can get you rapidly into the deep part of the pool. For me, it's very important to keep straight the difference between knowing a good solution to a problem and finding a good solution to a problem. I don't think this is as appreciated as it should be. The first is analytical/deductive and the second is creative/inductive, and they require very different preparations.

Think of an old-timey telephone switchboard operator, plugging and unplugging wires all day.
(photo courtesy of Wikipedia). There's a vast difference between knowing how to operate the switchboard and knowing how to design one or improve existing designs. The link between these two questions, as with the ones above is the "why" operator. Here's a possible chain of why-iterations for thinking about telephones:
  1. Q: Why are you moving those wires and plugs around?
    A: I'm operating a switchboard according to the procedures I've been trained in
  2. Q: Why is there a switchboard?
    A: To facilitate telephone calls.
  3. Q: Why are telephone calls useful?
    A: So people can communicate across distances.
  4. Q: Why do people need to communicate across distances?
    A: So they can lead better lives.
  5. Q: Why do people need to live better lives?
Each one of these increasingly general domains has its own problems and solutions. Solving the general ones can make the specific ones go away. If we keep asking why (see this related article), we end up with very general questions like "what problem does my existence solve?" and "why does anything exist?" I've tried to portray this recursion graphically below.
I'm assuming that the creation of knowledge is scientific (what art does is something different from what I mean here). I've quoted Bertrand Russell before on this point (here):
All definite knowledge--so I should contend--belongs to science; all dogma as to what surpasses definite knowledge belongs to theology. But between theology and science there is a No Man's Land, exposed to attack from both sides; this No Man's Land is philosophy.
Philosophy, theology, and other avenues of inquiry that remain immune to the scientific method are lumped together at the bottom of my graph. Wouldn't it be nice if we showed our students of critical thinking how this works? The unveiling of the breadth of meta-thought ought to be a stunning moment of realization for an undergraduate. Consider the following question and meta-question:
Why is the sky blue? (proposed answer here)
Why ask why?
From a scientific question we arguably leap directly over all of science to a philosophical one. Not only that, to my eyes it seems like a fixed point under meta-recursion. That is:
Why ask "why ask why?"? is the same as Why ask why?
Which would make the question the most profound one possible, I suppose. This could be a great starting point for a course on critical thinking. Note that I kind of cheated in my one-step leap to "why ask why?" Figuring out how is your meta-cognition homework. :-)

How does all this fit with existing literature on critical thinking? A colleague recently pointed me to the 1988 publication "The Delphi Report" on critical thinking, which arguably kicked off recent interest in the teaching and assessment of said skill. In the executive summary, which is linked to the title, a consensus statement reads:
We understand critical thinking to be purposeful, self-regulatory judgment which results in interpretation, analysis, evaluation, and inference, as well as explanation of the evidential, conceptual, methodological, criteriological, or contextual considerations upon which that judgment is based.
This is quite different from the line I've taken above, isn't it? Maybe it's not even useful to use the term "critical thinking" for both. In the above definition, which is the kind usually used, it's described as an activity with a particular type of result. The activity itself is not described here other than purposeful, self-regulatory judgment. These are noncognitive descriptors, please note. In fact, the definition elaborates on this point with a vivid description of the thinker:
The ideal critical thinker is habitually inquisitive, well-informed, trustful of reason, open-minded, flexible, fairminded in evaluation, honest in facing personal biases, prudent in making judgments, willing to reconsider, clear about issues, orderly in complex matters, diligent in seeking relevant information, reasonable in the selection of criteria, focused in inquiry, and persistent in seeking results which are as precise as the subject and the circumstances of inquiry permit.
As far as I can tell, in practice most programs in critical thinking don't actually pay much attention to the "affective" or noncogitives listed. But that's not too unexpected--academics in general seems to be allergic to modeling and teaching personal attributes. I find it increasingly odd that this is so.

Into the meat of the executive summary we do find particular cognitive skills. You'll see these or similar ones in rubrics and learning taxidermy.
  1. interpretation
  2. analysis
  3. evaluation
  4. inference
  5. explanation
  6. self-regulation
I don't see how self-regulation is cognitive, but maybe it is in a metacognition sort of way (self-reflection). One of the findings is that evaluating one's own thinking is a way to improve it.

Although they don't get around to saying it this way, the authors note the importance of analytical skills:
Although the identification and analysis of CT skills transcend, in significant ways, specific subjects or disciplines, learning and applying these skills in many contexts requires domain-specific knowledge.
Knowing how to solve an urgent problem while sailing is different from solving one while flying a plane.

This debate is important. Lots of institutions put "critical thinking" on their to-do list. Good definitions should lead to good implementations and good assessments.

Although I appreciate the value of the work that's been done in traditional meta-critical thinking, I don't much like the result--those lists of vague terms like interpretation and evaluation. I know they can be used successfully, and they can probably produce a good curriculum and assessment. But to me they're just a disjoint collection of loosely-defined techniques that a committee came up with. There's no underlying structure or theory. No way to make sense of it all by asking the meta-question: why is critical thinking the way it's described in "The Delphi Report?" You can only answer that the experts agreed that this is what it should be. In Russell's description, this makes it dogmatic. And if critical thinking has any value at all, it's to question dogma, no? That makes it ironic, but we can't judge too harshly on this account. Even Karl Popper freely admitted that his system could not be proven to be self-consistent (i.e. prove that nothing is every really proven). You have to start somewhere.

It may just be my bias coming from a computer science/math background, where a more natural schema is the study of algorithms and complexity, but I'd like more than the opinion of a panel of experts. I want to keep asking why until there is a self-consistent answer, if possible. In the meantime, here's my recipe for teaching CT:
  1. Analytical thinking in various domains
  2. Creative thinking in domains where analytical thinking is well-developed by the student
  3. Training in recursive thought [ including CT(CT) ]
  4. Communications skills (otherwise, what's the point?)
  5. Noncognitives like intellectual honesty and open-mindedness and willing use of the above skills.
I know the first two can be effectively assessed; I've never tried to do the third because it just occurred to me today. The last one ought to be on everyone's to-do list.

Thursday, February 26, 2009

Ignorance => Meta-Ignorance

In the last article here, I speculated about "unknown knowns," those bits of institutional knowledge that may be locked away by silos and rigid processes. I suggested that it might be in the institution's best interests to shake those out. It's a natural effect of a new administration taking over, or probably should be. Someone passed along the following advice about new administrations: keep the best one third of the current leadership, bring in one third new from the outside, and promote one third from within. It seems to me that this combinatorical shuffle would have the effect of breaking up old processes and modes of thought and allowing a temporary meritocracy of ideas to prevail. If only we could do that with the tax code!

It all surely comes down to the continued development of professional expertise of everyone on the job, I think. Encouraging subordinates to challenge our ideas may slow things down a bit occasionally, but in my experience is a good way to improve decisions. Isn't that what academia is all about anyway? You can't create new knowledge without challenging an existing mode of thought or 'best practice.' (The label 'best practice' makes me grit my teeth--surely any practice can be improved, no? It sounds like an admission of failure. 'Accepted practice' is more honest.)

Ignorance is meta is the conclusion of an article in the New York Times' science section from January 18, 2000 called "Among the Inept, Researchers Discover, Ignorance Is Bliss." The article suggests that there is a double-whammy to being uninformed. The ignorant don't know, and they don't know they don't know. That is, they are confident in their knowledge, even when they have little. Author Erica Goode explains:
One reason that the ignorant also tend to be the blissfully self-assured, the researchers believe, is that the skills required for competence often are the same skills necessary to recognize competence.
Cornell Psychology professors Dunning and Kruger, who researched this idea, make some interesting points, as quoted in the article:
  • Not only do they reach erroneous conclusions and make unfortunate choices, but their incompetence robs them of the ability to realize it.
  • This deficiency in "self-monitoring skills," the researchers said, helps explain the tendency of the humor-impaired to persist in telling jokes that are not funny.
  • Some college students, Dr. Dunning said, evince a similar blindness: after doing badly on a test, they spend hours in his office, explaining why the answers he suggests for the test questions are wrong.
If you have followed this blog on the topic of noncognitive assessment, you may recall that realistic self-appraisal is one of the predictors of success. On the other hand, the most able subjects in the study conducted by the researchers were the most likely to underestimate their own abilities.
The researchers attributed this to the fact that, in the absence of information about how others were doing, highly competent subjects assumed that others were performing as well as they were -- a phenomenon psychologists term the "false consensus effect."
There is some hope: Kruger and Dunning were able to 'train in' more realistic self-appraisal skills for those lacking them. The problem, they suggest, is lack of feedback. If you're doing a lousy job and no one tells you, how will you learn otherwise? A certain amount of humility is a good thing.

Of course, there has to be a balance. Paralysis through analysis is no good either. Being too timid to act on a new idea because there is no way to find out if it's good or bad prevents real leadership. After all, if all decisions are obvious, why are they paying you that fat administrative salary? Unfortunately, the Total Quality Management model that accreditors are fond of these days assumes that with enough information, good decisions can be made. That isn't always the case--just look at the stock market. A lot of very smart people with a lot of very good information get it wrong about half the time.

Therein lies the key to good leadership: entertaining new ideas on the one hand, but in spite of little information to go on, intuiting which of them are disastrous. I think this is a very rare trait. As Niccolo Machiavelli wrote in The Prince:
There is nothing more difficult to take in hand, more perilous to conduct, or more uncertain in its success than to take the lead in the introduction of a new order of things.
Readers of this column will not be too surprised that I have an anecdote to supply on this general topic. It was first published in a small college literary magazine, probably a decade ago. I will not insult your intelligence by highlighting my own examples of meta-ignorance in the story. You'll find them easily enough.

Geniosity

I’m a genius. Well, perhaps I should modify that statement just a tad. I was a genius. In fact, on two separate occasions during my life, I have had pure Eureka! moments that lifted me from the mundane to the ethereal. The descent was just as sudden, but at least I got a glimpse of what it must be like to be a real genius—you know, the kind that wakes up and goes to bed still in the bug-eyed goddamn I’m smart state. I suppose some people must find it addictive, having instant blinding flashes of that leave them gasping. I wouldn’t want it to happen all the time, though, or I might drive into a tree just as I’d solved the global deforestation problem, for example. All in all, I found it to be quite pleasant, although I noticed right away about how other people aren’t very interested in moments of clarity, unless it’s their own, in which case it’s hard to get them to shut up afterwards. So I got a bumper sticker that reads My dog had its day that I put right beside the Towers will be violatedand Don’t void where prohibited stickers. It’s a little obscure, but I figure that’s okay because obscurity is hard to tell from profundity sometimes.

It happened while I was doing dishes. The sink in my kitchen is divided into two stainless steel basins, with a faucet arm that can be swiveled to either side. Both sides drain down the same pipe. The problem is that every time you turn on the garbage disposer, which is attached to the right side, it backwashes filthy water up into the left basin. Since that’s usually where I place the dishes to dry, it’s a less than perfect situation. Before my instant of genius, I resorted to turning on the disposer in short bursts so as not to give it time to spew much water back up the other side. I had done this for years. But last week, I had finished stacking the last plate into the rack in the left basin, and was contemplating the pool of foaming dirty water in the right basin waiting to be drained when I had my geniosity (one of the perks of geniushood, even if only a part-timer, is the permission to create new words). I realized that if I ran some clean water into the left basin before turning on the garbage disposer, at the very worst only clean water would come back up! It worked beautifully, and I have switched entirely to my new method of draining the sink.

I was beginning to wonder if I’d lost the touch, because my only previous geniosity had occurred when I was in kindergarten, some thirty years before. There was the possibility that that earlier one had been a fluke, but now I’m convinced that if I wait another thirty years something equally profound will occur to me. I’m thinking of starting a newsletter. Anyway, back to kindergarten: it was one of those special days when something extraordinary happens. In this case, we had a magician coming to perform for us in the auditorium. I was hoping it would be the good kind—magicians that do magic tricks, instead of the bad kind—magicians that just play music. It was some time before I realized that musician is a whole different word. We were led into the auditorium in single file, and row-by-row filled up the folding chairs set up on the floor. They started with the back row, and I ended up in the second row, a prime spot for watching the tricks, if they were to materialize. As I planted myself into the child-sized folding chair, I noticed the kids who were being led into the row in front of me. The child about to sit directly before me was Bobbie. This was before last names were invented. Bobbie was a troubled child. He had announced one day on the playground that his real name was Robert, for which the rest of us laughed him to scorn. Really! We might have only been five, but we weren’t stupid enough to believe that you’d call a thing something other than what it was. A Bobbie was a Bobbie, and a Robert was something quite different.

As Bobbie prepared to sit, I had my geniosity. If I were to pull his chair back, I thought, he would miss it and end up on the ground! No sooner had inspiration struck than did I put it into action, and to my amazement it worked! Bobbie plopped right on to the floor, and then looked around with the most bewildered expression, which could be interpreted as how did I miss a twelve-inch wide chair with a six-inch wide butt? His universe had changed forever, as had mine. He had discovered The Unexplained, and myself a profound moral question:

Are some geniosities best left unimplemented?

Sadly, this question hasn’t gotten the attention it deserves, despite my well-received article (cf. “Gravitational effects of translation while sitting,” Kids Today 1969, v102, pp 87-89) which raised the issue. It is even more important today that it was then. Do you think those guys at Los Alamos, mucking around in the desert, really thought they could build an atomic bomb? Ironically, the typical defense used when a geniosity is misapplied is the stupidity claim. I’m ashamed to say that’s exactly what I used to explain Bobbie’s unexpected contact with the floor. The resulting Q&A with my teacher Miss Birdbalm is instructive. My comments are in brackets.

Q: Did you pull Robert’s chair out from under him? [Direct question, a tough nut.]

A: That’s not Robert it’s Bobbie! [First attempt—misdirection, obfuscation]

Q: Did you? [Miss Birdbalm was not easily distracted.]

A: Yes, but I didn’t know that Bobbie would sit on the ground. [The stupidity defense.]

Q: Why did you pull the chair back? [Her first mistake, questioning my intentions.]

A: I thought it would help. [Who can argue with good intentions?]

Q: How would pulling Bobbie’s chair back help him? [She’s a gonner now.]

A: I noticed that he walks with a limp, so I calculated the moment of inertia about his ankles and concluded that his head would be whiplashed against the back of the chair upon sitdown. Unfortunately I overcompensated and pulled the chair out too far. Believe me, it would have been worse if I’d done nothing at all. I’ve got all the bugs worked out now... [You get the idea.]

The problem with geniosities is that they are too precious not to be implemented, so that Whoa, I could do THIS!, is inevitably followed in short order by Whoa, I did THAT! Maybe this is the march of progress, but I’d like to propose a moratorium on geniosities until we get this sorted out. Except for mine, of course. I only have helpful ideas now.