Showing posts sorted by date for query noncognitive. Sort by relevance Show all posts
Showing posts sorted by date for query noncognitive. Sort by relevance Show all posts

Monday, November 28, 2011

Link Salad

A Monday's worth of interesting education-related links:

On non-cognitives, we have two articles from the Boston Globe. The first is "How College Prep is Killing High School":
A number of economists, including Nobel economist James Heckman, have documented the need for noncognitive or so-called soft skills in the labor market, such as motivation, perseverance, risk aversion, self-esteem, and self-control.
The second is "How Willpower Works":
In dozens of studies conducted over the past 25 years, Baumeister has found that taking on specific habits - like brushing your teeth with the opposite hand you’d normally use - can increase levels of self-control. In a phone interview, he likened willpower to a muscle: “If you exercise it, you can make it stronger. There’s nothing magical about it.’’
Then there is the less optimistic offering from the New York Times "The Dwindling Power of a College Degree," which contains a warning for all of us:
A general guideline these days is that people are rewarded when they can do things that take trained judgment and skill — things, in other words, that can’t be done by computers or lower-wage workers in other countries.
The Wall Street Journal has a scorecard of career salaries by degree, in case you're keeping score. The highest 75th percentile salary goes to math and computer science combined. Compare it to math education:

A partial listing of the WSJ salary/major list found here.
The quote in the New York Times article about computers replacing us is especially interesting when juxtaposed to the ambitious research plan described in "Mining the Language of Science," from Phyorg.com:
Scientists are developing a computer that can read vast amounts of scientific literature, make connections between facts and develop hypotheses.
Stanford University is offering a free online course on machine learning if you want to learn how to make a computer smarter than yourself (true story).

 To round out that topic, here are two articles on the limits of human understanding. First from Physorg.com again is "People are Biased against Creative Ideas, Studies Find," including these findings:
  • Creative ideas are by definition novel, and novelty can trigger feelings of uncertainty that make most people uncomfortable. 
  •  People dismiss creative ideas in favor of ideas that are purely practical -- tried and true. 
  •  Objective evidence shoring up the validity of a creative proposal does not motivate people to accept it. 
  • Anti-creativity bias is so subtle that people are unaware of it, which can interfere with their ability to recognize a creative idea.
The second article, from SciGuru, is "Ignorance is bliss when it comes to challenging social issues."
The less people know about important complex issues such as the economy, energy consumption and the environment, the more they want to avoid becoming well-informed, according to new research published by the American Psychological Association. And the more urgent the issue, the more people want to remain unaware [...]
This illustrates the mechanism I described in "Self-limiting Intelligence."  You can test yourself on these last two points. Here's a creative idea from Business Insider, and a challenging social issue from The Economist. Good luck!


Friday, October 01, 2010

Course Evaluations and Learning Outcomes

I've posted recently about some of our course evaluation statistics, and the effect of going from paper to electronic. A while back I also showed a summary of our Faculty Assessment of Core Skills learning assessment. I'm trying to put the two together by re-engineering the course evaluation to focus squarely on learning. The old version was a standardized one with fifty-five items, only five of which addressed learning at all, and these not very well. Here's my first draft of a new version, with comments afterward. The scale is indicated after each item. The exact wording is still in development.

  1. What was the quality of instruction in this course as it contributed to your learning? (try to set aside your feelings about the course content)
    --(ineffective to very effective)
  2. How much effort did you put into this course
    --(minimal to maximum)
  3. How much did you know about the course content before taking the course?
    --(nothing to a lot)
  4. How much do you know about the course content now?
    --(nothing to a lot)
  5. How much your skills in analytical/deductive thinking (knowing facts, following rules and formulas, learning standard methods) increase in this course?
    --(none to a lot)
  6. How much did your skills in creative/inductive thinking (trial-and-error, development of ideas, taking chances) increase in this course?
    --(none to a lot)
  7. How much did your ability to speak effectively increase in this course?
    --(none to a lot)
  8. How much did your ability to write effectively increase in this course?
    --(none to a lot)
  9. How much did this course help you understand yourself?
    --(none to a lot)
  10. How much did this course spark your interest in the content?
    --(none to a lot)
  11. Was the course enjoyable?
    --(not at all to very much)
  12. How much course content (the subject area, like chemistry or psychology) do you think you learned in this course?
    --(none to a lot)
  13. What overall rating would you give this course as a learning experience?
    --(poor to excellent)

Comments.

This is a radical departure from what we do now. The first question is what we use now on the evaluation form, and is the only one used for evaluation. Question 13 is a validity check on it because the answers should be very much the same.

The questions all focus on learning, except numbers 2 and 11. In the old evaluation, almost all of the questions were about the process of teaching, which makes a lot of assumptions about the value of those processes, and doesn’t transfer well to styles like online learning or hybrid courses.

The learning questions are split between the content area and general liberal-arts skills. This gives us a natural complement to the Faculty Assessment of Core Skills (FACS), which we launched very successfully last spring. Taken together, the teacher view and the student view will give us excellent insight into gen ed outcomes across the whole curriculum.

Question 2 is included because it matches the one on the FACS. The noncognitive “effort” is very important to performance. Here’s the graph from the spring FACS, with GPA in red and credits earned in blue numbers. More effort means better grades and better chance of advancing.

Questions 3 and 4 get at how much content was learned by asking in terms of before/after. This is checked for reliability with question 12.

Questions 5-9 are about general learning outcomes. No course would be expected to get max scores in all of these—it’s an environmental scan to help us understanding where students feel what kind of learning is happening where. It complements the NSSE, the QEP, and the FACS, and will be a gold mine of information.

Question 9 is from the temple of Apollo at Delphi: “know thyself.”

Question 11 will raise some hackles, but it’s there as a control. We know from research that students who rate courses as enjoyable also rate everything else higher. This allows us to investigate that phenomenon locally. If we get to the point where we can administer electronically, we can do these studies ourselves by comparing to course grade. With an anonymous paper survey, we’ll have less ability to do that, but can still do intra-response correlations. We could be more direct and just ask “how happy are you right now?” but that would turn off some students.

There are two free-response questions we'll carry over from the old survey. This will let students write on topics they care about most.

The survey is short for two reasons. First, we’ll get better reliability because students won’t get survey fatigue. Second, this leaves room for other surveys customized by a program, department, or college, to be administered in parallel. For example, the Lit folks could ask detailed content-related questions if they wanted, or conversely ask all about processes (office hours, syllabus, etc.).

Sunday, October 11, 2009

Michigan State Noncognitive Study

A "Report of the First-Year Follow-up of College Applicants at Twelve Universities" was forwarded to me by a colleague. According to him, the report is public now, but the appendices with survey instruments are not. I haven't seen the report itself online, and because of copyright don't feel like I can post it. But herein are some highlights.

The authors are Neal Schmitt, Abigail Billington, Juliya Golubovich, Jessica Keeney, Timothy Pleskac, Matthew Reeder, Ruchi Sinha, and Mark Zorzie, and the work was supported by the College Board, which is interesting.

The basics:
  • Participants: Earlham College, Furman University, Johnson & Wales University at Providence, Kenyon College, Lafayette College, Meredith College, Michigan State University, Ohio State University, Purdue University, University of North Carolina at Chapel Hill, University of Southern California, and University of Washington
  • 844 students provided enough data for analysis, initially through a College Board web survey of applicants on background, interests, and judgment, then through a follow-up for those who enrolled. The overall yield rate (enroll/app) was 26%. Of those who enrolled, 42% responded to the follow-up survey (there was a $20 incentive).
  • The kinds of data considered were biodata (background and life history, not blood pressure), a situational judgment test (SJT) representing behaviors, demographics, self-reported performance (BARS), citizenship behaviors (positive or negative), academic satisfaction, social satisfaction, grades and standardized test scores, an inventory of "shock" events a student may have experienced, the big five personality traits, substance abuse, use of time, and self-reported withdrawal tendency.
The report is chock-full of tables of data with correlations and crosstabs, but I'll skip to the regression results. Here are some key findings, quoted from the executive summary (pg. 3):
  • First-year college GPA is predicted significantly by several biodata scales, most notably Knowledge, Ethics, and Perseverance, but HSGPA and SAT/ACT scores are much more predictive of college GPA than are biodata and SJT.

  • Self ratings of performance (BARS), Organizational Citizenship Behavior (OCB), and student self-reports of Deviance were especially well predicted by the biodata measures and SJT while HSGPA and SAT/ACT were relatively uncorrelated with these outcomes.
It's disappointing that first year GPA isn't better predicted by this gob of noncognitive variables. But completion is a better goal, and the study hasn't had time to mature to that point. For example, any actionable information about first-year retention would be worth its weight in undergrads. Stay tuned for the next report.

Update: Dr. Neal Schmitt gave me a link to a publications page, which will soon include the body of the report I cited.

Monday, September 28, 2009

Predicting Success

Nature abhors a vacuum it's said. My daughter showed me the other day how her science class used this "principle" to determine the amount of oxygen in the air by using a candle and a test tube in water to measure before and after. Of course, it's isn't really abhorrence but air pressure that makes vacuum a chore to create down here where we live. But maybe economics really does abhor a passed-over opportunity. As in the old joke where one economist says "hey, there's a hundred dollar bill laying there on the ground," to which the other replies "can't be so--someone would have picked it up."

I've argued for a while that the low predictive validity of GPA + SAT creates market opportunities for those willing to experiment with other demonstrations of achievement. Here's a list of previous posts related to the topic:
In Malcolm Gladwell's Outliers, he has the following observation about the international math and science test called TIMSS . He notes that the test is accompanied by a 120-question survey, which many students don't complete. In his words:
Now, here's the interesting part. As it turns out, the average number of items answered on that questionnaire varies from country to country. It is possible, in fact, to rank all the participating countries according to haw many items their students answer on the questionnaire. Now, what do you think happens if you compare the questionnaire rankings with the math rankings on the TIMSS? They are exactly the same. (pg. 247)
He concludes a page later that "We should be able to predict which countries are best at math simply by looking at which national cultures place the highest emphasis on effort and hard work."

One could certainly ask for better analysis--why not correlate by student the number of survey items completed against the math score, rather than aggregating by country? But the sentiment is certainly the same expressed in noncognitive literature: there's more to success than the ability to do mental gymnastics.

In InsideHigherEd today there was an article "Next Stages in Testing Debate" talking about institutions that de-emphasize SAT in admissions decisions:
[A] common idea was that decreasing reliance on the SAT does not mean any loss of academic rigor and can in fact lead to the creation of classes that do better academically (and are more diverse).
This may require a rethink and additional training of admissions staff:
That report said that for too many admissions officers, the only training they receive on the use of testing may come from the technical training provided by testing companies, entities that have a vested interest in the continued use of testing.
Some of the other types of accomplishments sought by admissions officers are evidence of creativity, practical skills, wisdom about how to promote the common good (Tufts), and essays (George Mason). This idea is something that seems to be blooming. As Mr. Gladwell might say, it's blinking toward an outlying tipping point.

Evidence of this arrived in my in-box the other day: a forwarded email from the Law School Admisssion Council (LSAC) with the following news:
LSAC has funded research on noncognitive skills for some time. A study funded by LSAC—Identification, Development, and Validation of Predictors for Successful Lawyering by Marjorie Shultz and Sheldon Zedeck—identified 26 noncognitive factors that make for successful lawyering. The study included suggestions for assessments that might measure those factors prior to admission to law school.
I hadn't heard of Shultz and Zedeck, so I scurried off to the tubes that comprise the Internets to find out more. You can find the whole 100-page report and more here. The executive summary has an imposing, lawyerly warning on the front page:
NOT TO BE USED FOR COMMERCIAL PURPOSES NOT TO BE DISTRIBUTED, COPIED, OR QUOTED WITHOUT PERMISSION OF AUTHORS
I suppose this implies an argument that fair use somehow doesn't apply to this work. In any event, since I've obviously already violated the terms with the quote above, I may as well proceed.

Rather than focus on something easy like just predicting law school GPA, the researchers actually assessed job performance and got LSAT and law school performance data. Then they threw a bunch of tests at the problem, like Hogan Personality Inventory, Hogan Development Survey, Motives, Values, Preferences Inventory, and Self Monitoring Scale, a Situational Judgment Test, and Biographical Information Inventory. This on a sample of more than 1100 subjects. I think the word I'm looking for is "wow."

The results found successful indicators of effectiveness (per their definition) that seemed to be assessing independent characteristics, adding dimensionality to the standard predictors. In fact, the LSAT didn't seem to predict success at all. Go look at the executive summary for details--it's not very long.

All in all, this seems to be a solid study that shows that noncognitives are important--perhaps better--predictors of professional effectiveness than the traditional cognitive ones.

Other noncog stories in the news:
The College Board is even getting into the act. A 2004 publication talks about "individualized review" that includes factors beyond GPA and test. I am led to understand that they have a big project underway now on noncogs, but I can't find the website.

Okay, remember the vacuum? What if we managed to identify these noncognitive variables and start to use them? The LSAC folks outline what can happen next:
A major concern about developing an assessment for noncognitive factors is the possibility that the test would be so coachable that its results would be unreliable in the high-stakes environment of law school admissions.
I argued here (Zog's lemma), here, and here that any imperfect predictor invites error inflation for economic gain, but it doesn't take a genius to see that if checking "I'm a hard worker" gets me more financial aid, I'll be more inclined to overestimate my industriousness. This is a game theory problem with no solution that's likely to be mass marketed. Imagine if the assessment of prospective students had to be done laboriously by hand by highly trained admissions staff, instead of relying on a convenient test that cranks out a one-dimensional predictor. I'm not sure that's a bad thing.

Tuesday, September 15, 2009

Motivation, Outcomes, and Bandwidth

I suppose every generation creates its share of hideous neologisms and circumlocutions. For me "on a regular basis" (instead of "regularly"), is one of the worst. I made the mistake of telling my daughter how much I hate it, and now she uses it regularly just to annoy me. But "incentivize" must also rank up there the list of 20th century abominations. The idea is simple, and is probably linked to the recently debunked myth of markets that are perfectly efficient and people who always act in their own best economic interests. Paul Krugman describes this empirical enlightenment of economists vividly in his recent New York Times piece "How Did Economists Get It So Wrong?".

It reminds me of something I read a long time ago about an engineer doing research on the effect of a crowd on wireless transmission, beginning with: assume that a person is a one-meter diameter sphere of water... Assumptions and approximations have to be watched carefully.

I was reminded of "incentivize" a couple of days ago when I came across a TED talk by Dan Pink on the science of motivation. It's a 17 minute video you can see here. I will summarize some of his points here, but you might find it more interesting to watch the video first. The topic of the talk was motivation: what types work in what circumstances. This is an interesting topic in higher education because of various obvious reasons, including one you may not think of right off. More on that later.

Motivation isn't as simple as it seems, it seems. Consider the crazy things we do in the name of motivation. A car cuts us off in traffic and we may get angry, wave interesting gestures and honk at the driver. We might even call the police and report it if the behavior is egregious enough. Why are we doing that? At the bottom of it, I would posit that we are unconsciously trying to incentivize the driver not to do such things again through negative reinforcement. But in most cases, the odds that we will ever encounter this particular driver in that situation again are probably remote. (Consider how you might behave differently if it were your neighbor rather than a stranger, to see some of the complexities here.) So we are wasting our time at best, and possibly even acting against our own best interests. But such inclinations run deep. It's worth bringing their effects to the light of day.

We seem to have an instinct to incentivize (I'm going to get that word out of my system), even when it isn't likely to do any good. As pointed out above, our actions could make the situation worse. I assume that there are evolutionary reasons for these tendencies: a million years of living in social groups must have had some effect on our programming. There seems to be a general societal purpose to such actions (see this article, for example).

So now to Dan Pink's talk (spoilers ahead). He makes the same point forcefully, applied to the workplace: managers don't understand incentives and are mostly doing the wrong thing to motivate employees. Here are some quotes from the transcript.

Sam Glucksberg did a Candle Problem experiment to learn about the effect of incentives:
He gathered his participants. And he said, "I'm going to time you. How quickly you can solve this problem?" To one group he said, I'm going to time you to establish norms, averages for how long it typically takes someone to solve this sort of problem.

To the second group he offered rewards. He said, "If you're in the top 25 percent of the fastest times you get five dollars. If you're the fastest of everyone we're testing here today you get 20 dollars."
It took the second group three and a half minutes longer on average to solve the problem. This is counter-intuitive. As Dan Pink puts it:
You've got an incentive designed to sharpen thinking and accelerate creativity. And it does just the opposite. It dulls thinking and blocks creativity.
And most amazing is the claim:
This has been replicated over and over and over again, for nearly 40 years. These contingent motivators, if you do this, then you get that, work in some circumstances. But for a lot of tasks, they actually either don't work or, often, they do harm. This is one of the most robust findings in social science. And also one of the most ignored.
Further research, again using the Candle Problem, showed that incentives can positively affect performance, but only when the problem was simplified to a rote (I would say low-complexity analytical) task. The result reinforces the idea that the prospect of immediate reward (or punishment) causes us to reduce creative, conceptual approaches to problems in favor of direct obvious connections.

Since this is September, it's easy to make the leap to 9/11 as an example. In the early 1900s, the czarist Russian security organ had an imaginative idea: what if the terrorists (which were blooming everywhere) got it into their heads to crash an airplane into a building? I read about this in Orlando Figes' A People's Tragedy. Compare that creative idea to the actual response to the horrible actuality: a very narrow focus on preventing exactly the same thing from happening again. This isn't criticism--it's very natural and sensible--the point is that a big whomping motivation narrowed attention to what exactly the problem was seen to be in an obvious sense, not what it could be in the larger sense. If you read my last post about generalizing through recursion, you'll see what I mean. Here's a list of Bad Things That Can Happen, which have nothing to do with taking off your shoes before you board a plane:
  1. A near-earth asteroid could hit us
  2. The caldera at Yellowstone could blow
  3. Gene hacking of biological viruses becomes as common as computer-virus hacking, and some 18 year-old sets off a catastrophe (you do remember the 90s?)
  4. Nanotech goo takes over the world
  5. The oceans turn to acid as the climate heats up
  6. Environmental toxins are having epigenetic effects that will last generations
  7. A housing bubble could blow up the banking industry.
These are the sort of broad-spectrum dangers you're not likely to think about while your house is on fire. I think it's fair to summarize the findings Dan Pink describes as "incentives narrow focus."

If these findings are valid, much of the way management is done in business, including higher education, is wrong. In the video, Mr. Pink talks about some alternatives.

Relating this to education. If you've been in the classroom, you've probably been as frustrated as I have been by the question "is this going to be on the test?" To the mind of a "lifelong learner" this is entirely the wrong attitude. But it's easy to see from the perspective of motivational cause and effect that the incentives we apply would lead directly to that question. To use that awful word again, we incentivize students with grades. Why should we find it surprising that they tend to narrowly fixate on grades?

Dan Pink has created a consulting business out of this idea, where he pitches:
And to my mind, that new operating system for our businesses revolves around three elements: autonomy, mastery and purpose. Autonomy, the urge to direct our own lives. Mastery, the desire to get better and better at something that matters. Purpose, the yearning to do what we do in the service of something larger than ourselves.
Whether or not this is a formula that works, or is Utopian dream is unknown, but the ideas are certainly worth considering. Notice that of autonomy, mastery, and purpose, only the second one is a cognitive skill. The other two are affective, or noncognitive, or in jargon-free plain English: emotions.
I think there is untapped opportunity to try to engineer ways to motivate students more sensibly--to model and inculcate autonomy and purpose, and illuminate the role of zeal in creating mastery. That's all good, and I think there are opportunities especially for small liberal arts schools here. It's particularly ironic that incentivizing may actively hinder the teaching of critical thinking, which is supposed to be what grades-bound liberal arts colleges are suppose to be good at. But there's a bigger question.

"A Virtual Revolution is Brewing for Colleges" from The Washington Post is the latest article I've seen predicting the doom of traditional higher education. I blogged about this topic recently here. The argument is that for many subjects, education can happen conveniently and cheaply over the Internet, and that competition will drive the bricks and mortarboard model to ruin (except for the elite institutions that rely on deep pockets or have massive self-sustaining endowments).

What, aside from inertia, stands in the way of this transformation? I think one of the biggest problems for online education is the noncognitive load it places on the consumer: they have to be motivated to log in and do the work. They have to minimize the distractions of Facebook and a million other things while working on the computer. I don't have any statistics to back this up, so I may be completely wrong. But it seems to me that one current advantage of a residential school is the pervasive culture that comes with it. As imperfect as it is, there is social pull to come to class and not humiliate oneself by flunking every test. In the nearly anonymous hyperspace of online classes, I imagine that this is less so. In-person interactions are naturally more engaging that online ones. Is that really true? If so, how long will it remain true?

For the moment, I can't conceive that online teaching can approach the richness of a good professor's interactions with students in class, the dining hall, and in the office--the social engagement that includes mentoring and a kind of tribe-like kinship that comes from the circle of mutual acquaintances and shared experiences that play out in full-color, real-time, three-D.

In short, the bandwidth for real life ("rl" in cyberspeak, contrasting with, say, "vr" for virtual reality, or specifically "sl" for the online world Second Life) is still far superior to anything modems can deliver. But rather than leveraging this advantage, rl institutions waste most of the bandwidth. We're generally not engaging students on autonomy and purpose in the pursuit of mastery. We care about grades and bureaucracy and grants from the government.

Tentative conclusions. There may be a niche for second-tier institutions in the new education landscape to provide premium rl education, but only if they seriously address the engagement problem. George Kuh of the NSSE comes at this from another angle, and he actually does have some statistics. The point isn't just the survival of the traditional model, it's to provide a service that is superior to online education because of bandwidth and proximity: a million years of evolution has programmed us to live reasonably well together in social groups, and that should be taken advantage of.

De-emphasizing traditional grades is one step in that direction. Read my post about Western Governor's University to see a model for how that is already happening (online). But that's only part of it. A portfolio that a student can carry with them (and accumulate as a life-long resume) could contain evidence of not just subject mastery but also explicitly address noncognitive traits like purpose. Higher education has been allergic to "purpose" since it became largely secular. It's time to reconnect with the big "why" questions outside of a hermetically sealed philosophy course.

Online education isn't going to stand still, of course. Already there are very motivated people working on the problem of creating learning communities online. These visionaries think big, and for the most part, think "cheap" or "free" (search open education on this blog). Bandwidth will increase, rl will become more conflated with vr, and the next generations will perhaps feel at home in a warm LCD glow as they do in rl. For institutions frittering away their bandwidth advantage now, remember what happened to CDs. MP3s are generally inferior to CDs because the latter are compressed. But MP3 rule because bandwidth loses to convenience.

We might think of online education as a low-pass filter that employs only the deep end of the spectrum, like the telephone company only transmits a small range of frequencies when you talk in order to save money. What value is the high-frequency stuff? Can it be used to engage students in ways that vr can't approach? I don't know, but I think for the time being the answer is yes. I'll now take off my sackcloth, shave my beard, abandon the giant urn, and give up this prophetic conceit (I can't go to work like this) after one final prognostication:

It's a good time to be an energetic new college or university president with a creative, entrepreneurial spirit. There are opportunities. It's a bad time to be locked into the traditional model of private higher education that demands a high price and then blows the bandwidth.

Sunday, September 13, 2009

Recursive Critical Thinking

I was going to entitle this piece "critical thinking squared" as a cute way to imply critical thinking about critical thinking, but the imprecision bothered me. Squared means multiplication by self, and that's not the same as applying a process to itself. What multiplication means in this context isn't precise either, but you can possibly make a sense of it by considering a combinatorial factorization into dimensions like critical thinking = (analysis, creativity, communication). If we abbreviate critical thinking = CT, then CT2 might look like a matrix:


AnalysisCreativityCommunication
AnalysisAnalysisCreativity * AnalysisCommunication * Analysis
CreativityAnalysis * CreativityCreativity Communication * Creativity
CommunicationAnalysis * CommunicationCreativity * CommunicationCommunication

This assumes that each dimension is idempotent (meaning S*S = S), and that "*" is some way of combining the two dimensions. You still have to figure out what the Creativity * Analysis combination means, but at least you have a way to produce detail from the squaring operation. But this is all rather silly, and the reason I don't like the "CT squared" idea.

Here's a better way to think about it. If you talk about talking, that isn't (talking)2, but rather talking(talking), expressed here as a function that takes itself as an input.

This is even practical. For example, you could write a function get_loc(...) in the C programming language to take a function and return its address in memory. Then you could ask for get_loc(get_loc) to retrieve its own location. This sort of thing is called recursion, and it's a big deal in computer science. In common parlance, we might stick the prefix meta- in front of the concept to show that it's recursive, as in metacognition, which in the right context we justifiably call a noncognitive trait: the reflective practice of thinking about one's own thinking process. More on the relationship between CT and noncogs later. First, let's take a closer look at the role of recursion in problem solving.

In a few stolen moments this morning I was sipping an iced latte, enjoying a cool breeze, and trying to make some progress on a research project regarding survival in a certain abstract sense. You can read the actual paper here, but in a nutshell you can imagine an environment that poses survival challenges to organisms (all abstracted into computer language, which you can learn more about and download a simulator here). There are two different questions regarding the complexity of a given environment:
  1. What's the simplest thing that can survive the given conditions?
  2. What's the simplest recursive process that can find the solution to #1?
Here, recursion means that some process can be tried over and over again, feeding the output of the last iteration into the input of the next. Like natural selection, for example, acting recursively on the gene pool to blindly hone the fitness of the survivors.

Let me give a more down-to-earth example, as a simple "critical thinking" problem. Suppose Tatiana works all day in retail, and part of her job is to calculate sales reductions for coupons, sale prices, and so on. In addition she has to add sales tax to total amounts. For reasons known only to management, they skimped on her point of sale (cash register) and these functions are not included. So all day long she has to do stuff like:
  1. Find actual cost of an item by reducing for sale price, coupon, etc.
  2. Sum adjusted prices
  3. Calculate tax
  4. Add to get total
This is pretty tedious and prone to error, so there's an advantage to having the best possible way of doing this. In the context of my framing questions, we should ask:
  1. What's the best way of doing her job?
  2. How do we find it?
In practice, we have to address the second before the first. The second we might call a critical thinking exercise, requiring analytical and creative thought.

Tatiana need not be reflective. She probably already has a solution, and may not care that it's not optimal. I see this all the time in real stores. A clerk wants to reduce an item by 15%, say. Most often they multiply the price by 15% (.15) on a calculator, write down this number, and then subtract from the original price. Sometimes I tell them a quicker way to do it: just multiply the original price by .85 and you're done.

In a complex environment, you're never really finished with the "how do we find a better solution?" step. The question itself is recursive: "how do we find better ways of finding better solutions?" Mathematics is full of this sort of thing. You can follow the chain of meta-thought all the way up to something called category theory, where logic itself can be generalized (logic about logic).

So without really thinking about the definition, the idea of "critical thinking" can get you rapidly into the deep part of the pool. For me, it's very important to keep straight the difference between knowing a good solution to a problem and finding a good solution to a problem. I don't think this is as appreciated as it should be. The first is analytical/deductive and the second is creative/inductive, and they require very different preparations.

Think of an old-timey telephone switchboard operator, plugging and unplugging wires all day.
(photo courtesy of Wikipedia). There's a vast difference between knowing how to operate the switchboard and knowing how to design one or improve existing designs. The link between these two questions, as with the ones above is the "why" operator. Here's a possible chain of why-iterations for thinking about telephones:
  1. Q: Why are you moving those wires and plugs around?
    A: I'm operating a switchboard according to the procedures I've been trained in
  2. Q: Why is there a switchboard?
    A: To facilitate telephone calls.
  3. Q: Why are telephone calls useful?
    A: So people can communicate across distances.
  4. Q: Why do people need to communicate across distances?
    A: So they can lead better lives.
  5. Q: Why do people need to live better lives?
Each one of these increasingly general domains has its own problems and solutions. Solving the general ones can make the specific ones go away. If we keep asking why (see this related article), we end up with very general questions like "what problem does my existence solve?" and "why does anything exist?" I've tried to portray this recursion graphically below.
I'm assuming that the creation of knowledge is scientific (what art does is something different from what I mean here). I've quoted Bertrand Russell before on this point (here):
All definite knowledge--so I should contend--belongs to science; all dogma as to what surpasses definite knowledge belongs to theology. But between theology and science there is a No Man's Land, exposed to attack from both sides; this No Man's Land is philosophy.
Philosophy, theology, and other avenues of inquiry that remain immune to the scientific method are lumped together at the bottom of my graph. Wouldn't it be nice if we showed our students of critical thinking how this works? The unveiling of the breadth of meta-thought ought to be a stunning moment of realization for an undergraduate. Consider the following question and meta-question:
Why is the sky blue? (proposed answer here)
Why ask why?
From a scientific question we arguably leap directly over all of science to a philosophical one. Not only that, to my eyes it seems like a fixed point under meta-recursion. That is:
Why ask "why ask why?"? is the same as Why ask why?
Which would make the question the most profound one possible, I suppose. This could be a great starting point for a course on critical thinking. Note that I kind of cheated in my one-step leap to "why ask why?" Figuring out how is your meta-cognition homework. :-)

How does all this fit with existing literature on critical thinking? A colleague recently pointed me to the 1988 publication "The Delphi Report" on critical thinking, which arguably kicked off recent interest in the teaching and assessment of said skill. In the executive summary, which is linked to the title, a consensus statement reads:
We understand critical thinking to be purposeful, self-regulatory judgment which results in interpretation, analysis, evaluation, and inference, as well as explanation of the evidential, conceptual, methodological, criteriological, or contextual considerations upon which that judgment is based.
This is quite different from the line I've taken above, isn't it? Maybe it's not even useful to use the term "critical thinking" for both. In the above definition, which is the kind usually used, it's described as an activity with a particular type of result. The activity itself is not described here other than purposeful, self-regulatory judgment. These are noncognitive descriptors, please note. In fact, the definition elaborates on this point with a vivid description of the thinker:
The ideal critical thinker is habitually inquisitive, well-informed, trustful of reason, open-minded, flexible, fairminded in evaluation, honest in facing personal biases, prudent in making judgments, willing to reconsider, clear about issues, orderly in complex matters, diligent in seeking relevant information, reasonable in the selection of criteria, focused in inquiry, and persistent in seeking results which are as precise as the subject and the circumstances of inquiry permit.
As far as I can tell, in practice most programs in critical thinking don't actually pay much attention to the "affective" or noncogitives listed. But that's not too unexpected--academics in general seems to be allergic to modeling and teaching personal attributes. I find it increasingly odd that this is so.

Into the meat of the executive summary we do find particular cognitive skills. You'll see these or similar ones in rubrics and learning taxidermy.
  1. interpretation
  2. analysis
  3. evaluation
  4. inference
  5. explanation
  6. self-regulation
I don't see how self-regulation is cognitive, but maybe it is in a metacognition sort of way (self-reflection). One of the findings is that evaluating one's own thinking is a way to improve it.

Although they don't get around to saying it this way, the authors note the importance of analytical skills:
Although the identification and analysis of CT skills transcend, in significant ways, specific subjects or disciplines, learning and applying these skills in many contexts requires domain-specific knowledge.
Knowing how to solve an urgent problem while sailing is different from solving one while flying a plane.

This debate is important. Lots of institutions put "critical thinking" on their to-do list. Good definitions should lead to good implementations and good assessments.

Although I appreciate the value of the work that's been done in traditional meta-critical thinking, I don't much like the result--those lists of vague terms like interpretation and evaluation. I know they can be used successfully, and they can probably produce a good curriculum and assessment. But to me they're just a disjoint collection of loosely-defined techniques that a committee came up with. There's no underlying structure or theory. No way to make sense of it all by asking the meta-question: why is critical thinking the way it's described in "The Delphi Report?" You can only answer that the experts agreed that this is what it should be. In Russell's description, this makes it dogmatic. And if critical thinking has any value at all, it's to question dogma, no? That makes it ironic, but we can't judge too harshly on this account. Even Karl Popper freely admitted that his system could not be proven to be self-consistent (i.e. prove that nothing is every really proven). You have to start somewhere.

It may just be my bias coming from a computer science/math background, where a more natural schema is the study of algorithms and complexity, but I'd like more than the opinion of a panel of experts. I want to keep asking why until there is a self-consistent answer, if possible. In the meantime, here's my recipe for teaching CT:
  1. Analytical thinking in various domains
  2. Creative thinking in domains where analytical thinking is well-developed by the student
  3. Training in recursive thought [ including CT(CT) ]
  4. Communications skills (otherwise, what's the point?)
  5. Noncognitives like intellectual honesty and open-mindedness and willing use of the above skills.
I know the first two can be effectively assessed; I've never tried to do the third because it just occurred to me today. The last one ought to be on everyone's to-do list.

Monday, August 31, 2009

Money, Genes, and College

The New York Times has a recent article on SAT related to family income here. It shows that incomes and scores have high positive correlation for 2009. I've reproduced the graph from the article below.
The article doesn't say that income causes higher scores (as I recently speculated), but the article is nevertheless criticised by an economics professor here. In his blog post, Prof. Mankiw suggests that a significant part of the slope is due to genes, with an argument along the following lines, which I've made more explicit here:
  1. IQ correlates positively with income
  2. The ability to perform well on an IQ test is influenced by heredity
  3. IQ correlates positively with SAT
Therefore, students of wealthy parents should have higher SAT scores merely because they are smarter. This is presented as a bias for explaining part of the slope of the curve evident above (which is greatly exaggerated by the scale used, please note). How much of the SAT bonus is due to IQ is not spelled out, other than the concluding note in Dr. Mankiw's article:
It would be interesting to see the above graph reproduced for adopted children only. I bet that the curve would be a lot flatter.
I interpret "a lot flatter" to mean that the IQ contribution accounts for a significant part of the slope. This is all reminiscent of the The Bell Curve and the controversy of genes vs. environment in the creation of intelligence.

Some analysis is in order. The inheritability of intelligence is a very political topic. The left would like to assume that all people really are created equal, and that the "blank slate" is there to be written upon. This legitimizes interventions that affect socio-economic status (SES). The right would like to believe that interventions are counter-productive because intelligence is fixed at birth. Both argue from the conclusions back to reasons for believing them, which is probably some kind of tragedy of the commons in the public realm: in order to stay in power, parties have to act sometimes in a way that is contrary to the common good. Since I don't have to be elected I can freely wish a pox on both their houses. Where is the science on the matter?

1. Income and IQ

There seems to be good evidence that IQ and income correlate positively. It's important to remember that IQ and intelligence are not the same thing. The first is a monological definition based on a particularly kind of cognitive test, and the second is a vocabulary item in common usage. In The Bell Curve and subsequent articles, the authors try to make the case that intelligence stratifies the employment landscape with the Very Dull at the bottom, hardly educated and hardly employable, and the Very Bright at the top. (Those words annoy me because dull is not the opposite of bright. Dim is.) In reading those arguments, it conjures up for me a kind of overarching social history that's implied--something on the order of what Marx sparked. In any case, it's a very strong conclusion they try to reach. You can spend a lot of time reading the debate about that book and related concepts.

It's noteworthy that physical attractiveness seems to also be linked to higher income.

2. Genes and Intelligence

The problem with an approach like The Bell Curve is that it's the wrong discipline. If you want to make conclusions about genetics, you need to actually look at some genes. This is especially true if you want to make big conclusions. The authors try very hard to control for environmental factors, but that doesn't substitute for identifying DNA that is causally linked to intelligence. That kind of work is being done by biologists, however. Here are some examples you can find on ScienceDaily.com:
Of course, there is science pointing to environmental influence over intelligence as well:
Largely ignored in the grand debate over environment vs. genes are new findings about epigenetics, which you can survey here and here.

I think a reasonable person has to conclude that genes, epigenetics, and environment all play a role. Because genes are discrete, it ought to be possible to identify consequent effects with more precision than the other categories. The third article linked in the list above pins a particular gene to an R^2 of 3% of IQ.

One of the puzzling things about IQ is that if it really describes a primarily genetic effect, how can we explain the dramatic rise in scores in the last decades? This is the so-called Flynn Effect, summarized in Wikepedia as the change over a 30-year period where:
  1. the mean IQ had increased by 9.7 points (the Flynn effect),
  2. the gains were concentrated in the lower half of the distribution and negligible in the top half, and
  3. the gains gradually decreased from low to high IQ.[reference]
What would a genetic-based explanation look like? I think you'd have to assume that the genes for Dullness became expressed less often, perhaps because they are going extinct. This would suggest a harsher survival environment for said genes. It seems to me, however, that's it's becoming easier to survive without developing intelligence, but that's just my take on it. Others have commented on this at great length. See a rather critical piece here.

I think it's also fair to conclude from the evidence we have that smarter parents on average will produce smarter offspring, all other things being equal, but that this is not fully deterministic. The question of how much intelligence is inheritable is still open.

3. IQ and SAT

The link between IQ and SAT is also controversial, at least from test-maker's point of view. See an overview here. There seems to be a strong correlation between the two, which you can read about in this research article.

SAT isn't very good at predicting first year college grades, but that's what it's designed for. It's even less good at predicting success beyond the first year. This isn't perhaps surprising: the content of the test resembles high school and college freshmen academic work. There are many other factors that influence success. I was interested to read here that:
Bates College, which dropped all pre-admission testing requirements in 1990, first conducted several studies to determine the most powerful variables for predicting success at the college. One study showed that students' self-evaluation of their "energy and initiative" added more to the ability to predict performance at Bates than did either Math or Verbal SAT scores.
Is IQ similarly limited in predicting college success? I couldn't find anything definitive about that in the time I had at my disposal, so I'll leave the question open.

Conclusion.

If we can reach any conclusion, it's a very weak one. Almost certainly income is indirectly linked to SAT scores through the associations advanced. However, how much that bonus is remains unclear. My blog post last time pointed out that it's not only that higher incomes get higher SATs, but that the year-to-year differential is also higher. If this is not some statistical artifact, it doesn't jibe with a purely genetic explanation, as it would require the gene pool to be evolving at a very high rate, implying extraordinary pressure from a fitness gradient.

I looked for more historical data on this increase as a function of wealth, but unfortunately it's only in the last two years that the SAT included the salary range from zero all the way to $200,000+. The old version only went to $100,000, and the most interesting part of the curve is above that. Having said that, I did not find any evidence for my theory about the economic value of error with the data that is available. We'll have to wait for next year's results.

There is, however, a direct causal explanation that is difficult to imagine away. It contrasts with the somewhat specious illustration Prof. Mankiw gives:
Suppose we were to graph average SAT scores by the number of bathrooms a student has in his or her family home. That curve would also likely slope upward.
Correlation and causation are different things. But consider another scenario. Suppose we were to graph SAT scores by the number of prep-tests a student attended, or the number of times they took the test (guaranteed on average to raise the maximum score), or the quality of the high school they attended. All of those have plausible causal connections to how well a student performs on a test of high-school cognitive material, no? And which students have access to the best high schools? Can afford to take the test multiple times? Can afford prep-tests that are advertised to raise scores 100 points?

The ironic thing is that the economic value of raising the SAT is real because of the usual policies in awarding financial aid. So the wealthiest are in the best position to reap that aid for both reasons of intelligence and the ability to buy error (over-prediction due to coaching or taking multiple tests). This trend has been documented for a long time here.

This line of thought gave me an interesting idea some time ago, which I hope to turn into a project. At my current institution we have embarked on a shift toward recruiting better students (students who have a better chance of success). But the conversation with the board, as well as public perception, includes the median SAT scores of the students we admit. My preference is to use better predictors (including noncognitive ones) to find the best students, but this is largely independent of their SATs, which creates a tension between actual goals and perceptions.

The idea is to create a grant-funded summer camp for the low-SAT that are nevertheless predicted to do well. In this camp they would receive SAT test-taking preparation and then retake the test immediately afterwards. This would raise median SATs without affecting our ability to get the students we want.

Update: here are some references:

Tuesday, August 25, 2009

Zog's Lemma: Assessment and the Amplification of Error

The concept of outcomes assessment is like a Swiss Army spatula: it's used in all kinds of ways. As opposed to the proverbial Russian Army hardware, pictured below (ubiquitous on the Internets):
To some, outcomes assessment is a touchy-feely endeavor of encouraging the practitioners of higher education to do the right thing, close the right loop, bring a glowing smile to the visiting team. All that. I'm comfortable with that.

To others, it's a more serious matter, more scientific in approach, rather like making tick-marks on the door's threshold on a child's birthday to signify evidence of growth. We know that is serious business: with shoes or without, and where exactly is the top of the head? Does hair count or not? It grows too, after all.

The scientific approach requires belief in theory: that assessments are valid to the purposes we employ them for. In reality we usually have no real way to know that based on a solid physical theory. The belief has to be defended by what statistics can be summoned to make a case. (In truth, beliefs don't need to be defended at all; they only need to be believed.) What comprises this validity? I'd like to address that question sideways. The discussion of what constitutes validity is a groove cut deeply in the literature of psychometrics, and I'd rather ask a more important question: of what use is it to believe in the validity of an assessment?

It's easy to make hay from the fact that the definition of validity itself isn't settled, but this is unfair, I think. It's a difficult philosophical nut to crack. For my purposes, predictive validity is the most important aspect of assessment. If an assessment doesn't tell us anything about what's going to happen in the future, I don't see much use in it. The Wiki article on predictive validity contains an interesting observation that is the crux of the value of an assessment:
[T]he utility (that is the benefit obtained by making decisions using the test) provided by a test with a correlation of .35 can be quite substantial.
Utility is a concept from economics that allows for a weighting of outcomes to balance outcomes that would otherwise be numerically indistinguishable. A classic example is the fact that $1000 means more to a poor person than it does to a rich person, even though it won't buy any more for the former. The relative worth to the penniless is more. The point is that even imperfect predictive assessments are useful.

An example of this is the college admissions process. Let's suppose that with the data gathered on the application, the institution can estimate the probability of "success" of student. Success could be retention to second year, or GPA > 2.5, or graduation, or whatever you deem important. The statistics can be generated with a logistic regression, which will yield a model with a certain amount of predictive power. No model is perfect, and even in retrospect (feeding the original data back into the model) it will not correctly classify applicants as "predict success" or "predict failure" with 100% accuracy. There will be some proportion of false positives and false negatives. This can be visualized on a Receiver Operating Characteristic (ROC) curve. You can see a bunch of them on google images. The usefulness of the ROC curve is that it lets you visually explore the decision of where to set a threshold for decision-making. If you set it too high, you get too many false negatives (reject too many qualified candidates). Too low and you admit too many false positives.

For an institution, such a tool is obviously useful to believe in: one can work backwards from the desired number size of the entering class to see where the threshold should be set. You could even estimate the number of false positives. Any power to discriminate between successful and unsuccessful students is better than none. Of course there are other factors, such as ability to pay, that make this more complicated. To keep things simple, I won't consider these distractions further.

One can imagine a utopia springing from this arrangement, where the assessments continually get better and the applicants learn to distinguish themselves by giving signals that the assessments can recognize. The first doesn't seem to be happening, and the second has an unfortunate twist. An analogy from biology serves us well here.

The story of the peacock, according to evolutionary biologists, is that a showy mating display is worth the biological cost of making all those pretty feathers because of the payoff in reproduction. The similarity is that when mates choose each other they have limited information to go on--an assessment we could characterize with a ROC curve if we had all the facts. This produces a distortion that favors any apparent advantage in a potential mate, or in the case of the peacock, creates over time a completely artificial means of assessment. The peacock is saying "look, I'm so healthy I can drag around all this useless plumage and still escape predators."

So, if the value of partial assessments is clear from the institutional vantage, it's a different picture altogether from the applicant's point of view. It's interesting that InsideHigherEd has an article this morning on this very topic. According to the article, in response to increasing competitiveness at "top institutions":
[H]igh school students could respond to the pressure by taking more rigorous courses and studying more -- or they could focus their attentions on gaming the system and trying to impress.
A study by John Bound, Brad Hershbein, and Bridget Terry Long shows that while this perhaps motivates students to take a more rigorous curriculum, it also prompts them to spend more time in test preparation, or in games like trying to engineer more time to take the test. Peacock plumage? It's hard to say without seeing actual success rates. It could be that having the willingness to spend all that extra effort is itself a noncognitive predictor of success (that is, not related to the scores themselves, but to the personality traits of the applicant). As Rich Karlgaard put it in Forbes magazine, a degree from an elite institution is valuable because:
The degree simply puts an official stamp on the fact that the student was intelligent, hardworking and competitive enough to get into Harvard or Yale in the first place.
What is clear is that those with the means to game the system are better off than those who do not. So we see things like test-prep for kindergartners at $450/hr in the upper crust of society, but probably not so much in housing projects. This likely produces more false positives among the select group: it's as simple as money buying better access. [Update: see this InsideHigherEd article for some dramatic numbers to that effect. Year over year, SAT scores increased 8-9 points for $200,000+ families, and 0/1 points for the poorest group.]

The economic demand for false positives has become an industry under No Child Left Behind, and that doesn't seem to be changing. The New York Times has published letters from teachers on this topic. Some quotes:
  • [T]he use of test data for purposes of evaluating and compensating teachers will work against the education of the most vulnerable children. It is a mistake to conceptualize education as a “Race to the Top” (as federal grants to schools are titled) — for children or schools. (Julie Diamond)
  • Linking teacher evaluations to faulty standardized tests ignores the socioeconomic impact on a nation that is both rich and poor. Can a teacher confronting the poverty of some children in Bedford-Stuyvesant be made to compete with a teacher instructing affluent children in Scarsdale?(Maurice R. Berube)
  • My job went from teaching children to teaching test preparation in very little time. Many of our nation’s teachers have left their profession because the focus on testing leaves little room for passion, creativity or intellect. (Darcy Hicks)
  • Standardized tests are, by their nature, predictable. Most administrators and teachers, fearing failure and loss of position and/or bonuses, de-emphasize or delete those parts of the curriculum least likely to be tested. The students sense this and neglect serious studying because they know that they will be prepped for the big exams. (Martin Rudolph)
The economics of false positives is clear: test prep is an industry, teachers and administrator and schools are rated by how many they generate. Of course the object is not to create false positives, that's just the result of so much emphasis on an imperfect assessment.

Less attention is paid to false negatives. Research points to noncognitive traits like grit, planning for the future, and self-assessment as being important to actual success, but these are not directly accounted for in standardized assessments. To be sure, college admissions officers look at extra-curricular activities to try to add value to SAT, GPA, and curriculum, but I think it's safe to say that the cognitive assessments are primary. How badly do we underestimate actual performance?

That question comes up when we evaluate or create a predictive model for applicants. The resulting predicted (first year) GPA can be used for admissions decisions and financial aid awards. In the Noel-Levitz leveraging schema applicants are sorted into "low-ability" to "high-ability" bins for individual attention. The question of false negatives is the same as the question "how accurate is the description 'low-ability', based on the predictors?" Not very good, as it turns out.

Every time I ran the statistics I got the same answer: at the very bottom end of our admit pool--those Presidential and provisional admits who were supposed to have the hardest time--about half of them performed well (I usually use GPA > 2.5 for that distinction). Since we reject students below that line, we should assume that about half of those students just below the cutoff would have done as well too. If this is typical, there are a LOT of false negatives.

Remember, this line of thought applies to all outcomes assessments to one degree or another. Let's take a concrete example. Suppose we want to assess vocabulary knowledge of German language students with an exam. Memorizing a single word (including declensions for nouns and conjugations for verbs) is low-complexity. The total complexity is that for one word times, let's say 10,000 total items, less any compressibility of this data. All told, this is a lot of complexity (measured in bits) for a human. So it's a suitable subject for assessing, and the low complexity per item means that each of them can be assessed with confidence.

Supposing that we do not have the resources to test our learners on all 10,000 items of vocabulary, we'll have to sample randomly and hope that the ratios are representative (or else spend a lot of time checking correlations and such). Maybe our assessment is only on 100 items, chosen at random from the 10,000. If Tatiana actually only knows K of the items, then there is a chance of K/10,000 that she will know an individual word, and theoretically score K/10,000 on average (on any sized test). Given this arrangement, what is the trade-off between false positives and false negatives?

We will assume we can use the normal distribution to estimate these Bernoulli trials. This only requires that K not be too large or too small--in those cases the chance of an error diminishes anyway. We already know the mean is p=K/10,000, and the standard deviation is given by SQRT(p(1-p)). Two standard errors is about 3% for large and small values of p, ranging to about 5% for those in the middle. These would decrease for a number of items larger than 100, and increase for a smaller test.

The conclusion is that if we set our cutoff for passing in a usual spot, say 70%, we should expect at least half of the scores in the range 65-75% to be either false positives or false negatives. For example, if Tatiana actually knows 70% of the vocabulary items, she has a 50% chance of passing the test because the distribution of her scores over all possible tests is (almost) symmetrical and centered on 70%.

The actual number of false positives and negatives depends on where the skill ranges of the test-takers lie in relation to the cutoff value. The more there are close to the cutoff, the more errors there will be. I actually witnessed a multiple-choice placement test being graded one time, and inquired about the cutoff to find that the grade to pass was actually less than the average result expected by chance! This certainly reduces false negatives, but I'm not sure about the overall result.

The vocabulary example is a best case, where the items themselves are of low complexity and (we hope) not subject to a lot of other kinds of error. When the predictive validity is actually quite low (as in the example where we explain a small percentage of the variance in the outcome), the proportion of errors in both directions is far worse. What does this mean, then when we "believe" in a test like the SAT and adopt it as a de facto industry standard? Even together with high school GPA, these can typically explain only about 33% of the variance of first year college grades.

First, it is of benefit to the institution to be able to imperfectly sort students by desirability. It is costly if the institution has to bid for the apparent best applicants with institutional aid, and there is incentive to look at noncognitives and better predictors, but I can't see a lot of progress in that direction. So the false positives get bid up with the rest. Meanwhile, the false negatives--those applicants who would succeed but don't show it on the predictor--get passed over for admission or merit aid.

A tentative conclusion is that a weak predictor gets amplified in a competitive environment. This is probably true in lots of domains, like the evolutionary biology example of the peacock. Probably some economist has his/her name attached to it: Zog's Lemma or something. I'll have to ask around. (I just googled it--apparently there is no Zog's Lemma.)

For the assessment types, there are two lessons. First, there are going to be errors in classification. Second, those errors get amplified as the assessment gains importance. This is an argument against high-stakes assessments, I suppose. I've always gotten good results by keeping assessments free from political pressures like instructor or program review. "Accountability" creates problems as it tries to solve them. Acknowledging that would be a fine thing.

PS, if you're interested in other kinds of amplifiers in science, read this.

UPDATE: see Amplification Amplification

Friday, August 21, 2009

The Use of Grades

There's been an interesting discussion on the ASSESS listserv about assessments vs. grades, which led me to think about the relative uses of each. Part of the problem in addressing that difference is that grades come in all sorts of flavors that may or may not resemble an outcomes assessment. For example, a carefully designed calculus final exam may pass quite nicely as a summative assessment. By contrast, a student who increases his final grade by attending an extracurricular event (say in an orientation course) has little claim that the grade is an reflection of his performance in some cognitive skill. The summing of individual grades into a cloudy average makes this worse (see: Statistical Goo). The same can be done for assessments, of course.

One fundamental difference between grades and assessments is that the former is used to motivate students. In fact, we've created a whole industry that depends on this kind of motivation, from accreditation downwards to the classroom: a kind of "do this or else" mentality. Surely this can't be ideal. I appreciate the requirements of state education to more or less force kids out of their beds in the morning, onto busses, and subject themselves to ideas that hurt to absorb. But higher education? Especially in the liberal arts, we talk about general goals like creating life-long learners and such. How does that square with our methods of delivery?

That discussion comes back to the topic of noncognitive traits in learners: how motivated is Tatiana to learn computer programming? For a motivated learner, assessments are better than grades because they get straight to the point of "how well am I doing?" without the coercive baggage that's inherited from elementary school. This is, by the way, an argument for detailed reports in assessments as well. The more gooey they become, the more they resemble grades and the less useful they are. Compare:
Tatiana sees her score of 79% on the C++ test and concludes that she is doing okay, but not excelling.

Tatiana reads her C++ assessment and sees that while basic control structures are second nature to her, she really doesn't understand pointers.
Only the second of these is actionable--Tatiana can increase her skills by practicing with pointers, getting tutoring, reading what others have to say about the subject. Learning is about details, and assessments should be too. With the ubiquity of modern information systems, keeping track of details isn't a problem--we just have bad habits left over from the grade school mentality of reducing a semester's work to a single letter. It's absurd, if you think about it.

Can we find another way to motivate students? I don't know, but I don't think we're really trying. I can imagine a culture that fosters a more inquisitive approach to self-improvement, but don't know how that might be engineered. And yet, we don't really want to produce graduates who only perform in order to get the pellet in a Skinner box, do we? I think the first step is to start to assess certain noncognitive traits, and bring them into the curriculum (and not just in orientation course). I'm not saying that all students lack self-motivation, of course. We have and value the go-getters in class who drive classroom discussions, ask for new things to read, and always come for help when they don't understand. How do we harness that energy to help pull along students who aren't so energetic in their learning practice?

Until such a motivation exists, it's probably best to keep grades to push students along and keep the political pressure off assessments. It's very convenient for the Assessment Director to not have to worry about the kind of scrutiny that the registrar requires. But it is a capitulation of sorts.

Thursday, August 06, 2009

Grit, Grades, and Graduation

It's funny how categories affect thinking. Since becoming interested in the noncognitive traits of students, their assessment and consequence predictive powers, I've begun to see noncognitives everywhere. The August 2nd article "The Truth about Grit" in the Boston Globe is an example.

The point of the article is that intelligence doesn't guarantee success; success also requires perseverance in the face of obstacles, or "grit."
The hope among scientists is that a better understanding of grit will allow educators to teach the skill in schools and lead to a generation of grittier children.

[...]

The new focus on grit is part of a larger scientific attempt to study the personality traits that best predict achievement in the real world.
The US Army has supported the research with some interesting findings. The example of the US Army's military academy West Point is compelling:
The Army has long searched for the variables that best predict whether or not cadets will graduate, using everything from SAT scores to physical fitness. But none of those variables were particularly useful.
What did work was a survey that assessed perseverance. The article notes that this echoes Francis Galton's 1869 research findings that concluded that a prerequisite for higher-order achievement was “ability combined with zeal and the capacity for hard labour.”

The whole article is worth reading. There is a link given (indirectly) to a grit survey developed by A.L. Duckworth at the University of Pennsylvania. The project is applicable to higher education:
Duckworth has recently begun analyzing student resumes submitted during the college application process, as she attempts to measure grit based on the diversity of listed interests. While parents and teachers have long emphasized the importance of being well-rounded - this is why most colleges require students to take courses in all the major disciplines, from history to math - success in the real world may depend more on the development of narrow passions.
This should be terribly interesting for liberal arts schools particularly. If perseverance is tied to singular passions, then how does this interact with the "broadening of the mind" mission of the institution? Is the implicit goal of success after graduation secondary, or should we try to pull off both? This has direct implications to learning outcomes assessment, as noted in the article:
In recent decades, the American educational system has had a single-minded focus on raising student test scores on everything from the IQ to the MCAS. The problem with this approach, researchers say, is that these academic scores are often of limited real world relevance.
Supposing we took the idea of grit seriously. That would start with using instruments like the one Dr. Duckworth has developed to estimate it. Can we then teach it? Should we?

Dr. Dweck at Standford University is quoted referring to a "growth mindset" versus a "fixed mindset," the difference being what we believe about our abilities. Can we grow them, or are we fixed with them? Fixed mindset learners are more likely to give up when encountering obstacles, assuming they're just not talented enough. Dweck's research, according to the article, demonstrates that the growth mindset can be taught effectively.

Ironically, praising children for their intelligence may make them less likely to succeed because it reinforces the fixed mindset. A better strategy is to reward effort and hard work. I've blogged about this subject before, and Malcolm Gladwell's Outliers is in much the same vein, I believe (I have only read reviews of it to date).

Should we assess these traits? One of the results of reporting assessment results for learning outcomes is that it becomes evident that students can earn decent grades and graduate without scoring high on assessments. If you've been in the classroom any length of time, you've probably encountered this student, whom we'll call Joe. Joe works very hard in your Finite Math class, forming study groups outside of class, turning in all the homework, coming to office hours. But the material just doesn't click with him, and he has a terrible time of coming to grips with the key concepts. Nevertheless, he gives each exam his full efforts and manages to pass with a B. Everyone is happy about this achievement, no? After grades are posted, Joe announces he wants to major in math. You consult with your colleagues and worry together for a bit. Although Joe has worked very hard and earned his B in everyone's eyes, he clearly lacks some spark of imagination that makes doing math rewarding. You fear he's setting himself up for failure.

The point is two-fold. One is that we assess outcomes (subjectively and informally in Joe's case) and we assign grades, but they mean different things. Tied up in grading is the notion of perseverance, of meeting each of the many assignments head-on and getting through them. But in the gestalt there may be something missing; the pieces do not always make the whole. The second point is related: we attempt to estimate the cognitive outcomes with our formal assessments, but generally do not assess the noncognitives. Shouldn't we doing so? If the science quoted is valid, then grit has as much to do with success as intelligence.

If you've browsed my Assessing the Elephant piece, or read enough of this blog (e.g. here), you can guess where I'm going. Why not assess grit along with thinking and communication skills across the curriculum, in a minimally-intrusive survey? It works for the cognitive skills, so there's a chance it will work for the noncognitives. Wouldn't it be fascinating to be able to compare cognitive skills, grit, grades, and graduation rates?

Please note that this gives institutions a way to include students themselves in their assessment reports. My last post was about publicly reporting learning outcomes and results. For example, Capella University's website has a page for Learning Outcomes. This is the future--gaze well upon it, ye assessment directors:
As this practice becomes common, the question will be asked "why doesn't every student get the highest rating?" Is this a fault with our education? At present, we could only shrug our shoulders and perhaps mutter about SAT scores of incoming freshmen, or if accreditation is imminent the glorious plans for improvement we've printed in reports. But if we actually assessed noncognitive traits like grit, we could report out richer details. We might note that the students who failed to graduate also were rated low in perseverance. We could isolate those with the most assessed grit and see how this relates to cognitive development. We could include noncognitives in the curriculum itself and begin to take it seriously, just like effective writing. For liberal arts institutions, the tricky question of balancing broadness with singleness of purpose could be explored with at least some minimal data.

It's an experiment worth doing.