Showing posts with label SAT. Show all posts
Showing posts with label SAT. Show all posts

Saturday, November 06, 2010

SAT-W: Does Size Matter?

This ABC News story tells of 14-year-old Milo Beckman's conclusion that longer SAT essays lead to higher scores. Interestingly, the College Board's response dances around the issue but doesn't deny it:
Our own research involving the test scores of more than 150,000 students admitted to more than 100 colleges and universities shows that, of all the sections of the SAT, the writing section is the most predictive of college success, and we encourage all students to work on their writing skills throughout their high school careers.
It would have been easy to say that with such-and-such alpha level, controlling for blah and blah, the length of the essay is not a significant predictor of the total score. But they didn't. It is interesting that the writing score is supposed to be the most predictive. It's not hard to find the validity report on the College Board SAT page. Here are the gross averages from the report.
The average high school GPA is much higher than I would have expected. Notice too that the first year college GPA is .63 less. The obvious question here is whether this difference has increased over time due to grade inflation. An ACT report "Are High School Grades Inflated?" compares 1991 to 2003 high school grades and answers in the affirmative:
Due to grade inflation and other subjective factors, postsecondary institutions cannot be certain that high school grades always accurately depict the abilities of their applicants and entering first-year students. Because of this, they may find it difficult to make admissions decisions or course placement decisions with a sufficient level of confidence based on high school GPA alone.
This is somewhat self-serving. Colleges don't need to really predict the abilities of their applicants--that's not how admissions works. What we do is try to get the best ones we can for the net revenue we need. That is, a rank is sufficient, and even with grade inflation, high school grades still do that. In the end, SAT and ACT are used to rank students for admissions decisions too.

The grade inflation is remarkable. Let's take a look at this graph from the ACT report.

Two things are obvious--first, the relationship between grades and the standardized test are almost linear. Especially in the 1991 graph, there's very little average information added by the test (by which I mean the deviation from a straight line is very small). Second, there's an inflation-induced compression effect as the high achievers get crammed up against the high end of the grading scale, reducing its power to discern.

Recall from above that the average SAT-taker's high school GPA was 3.6 in the most recent report, and look where that falls on the scale above. We could guess that the "bend" in the graph is getting worse over time, and probably represents a nonlinear relationship between grades, standardized tests, and college grades. If you have a linear predictor (e.g. for determining merit aid), it would be good to back-test it to see where the residual error is.

First Year College GPA is related to the other variables through a correlation, taken below from the SAT report.

We're really interested in R-squared, the percentage of variance explained. In the very best case, if we take the biggest number on the chart and square it, is about 36%. That is, the other two thirds of first year performance is left unexplained by these predictors. Indeed, SAT-W is larger than the other two SAT components (even combined). Now why might that be? Is SAT-W introducing a qualitatively different predictor?

To pursue that thought, suppose Mr. Beckman is right, and the SAT-W is heavily influenced by how long the essay is. This leads to an interesting conjecture. Perhaps what is happening is that the SAT-W is actually picking up a non-cognitive trait. That is, perhaps in addition to assessing how well students write, it also assesses how long they stick to a task: their work ethic, so to speak. If so, I wonder if this is intentional. The College Board has a whole project dealing with non-cognitives , so it's certainly in the air (see this article in Inside Higher Ed).

My guess is that they figured out the weights in reverse, starting with a bunch of essays and trying out different measures to see which ones are the best predictors. And length was one that came up as significant. It's not an entirely crazy idea.

You can read the essays. I did not know this until I saw the College Board response to the ABC article:
Many people do not realize that colleges and universities also have the ability to download and review each applicant's SAT essay. In other words, colleges are not just receiving a composite writing section score, they can actually download and read the student's essay if they choose to do so.
So theoretically, you could do your own study. Count the words and correlate.

Finally, I have to say that the ABC News site is an example of an awful way to present information. It makes my head hurt just to look at the barrage of advertisements that are seemingly designed to prevent all but the most determined visitors from actually reading the article. I outlined the actual text of the article in the screenshot below.

The first thing you get is a video advertisement in the box. If you want to read the whole article, you have to click through four pages of this stuff. I didn't make it that far. Maybe it's a test of my stick-to-it-ness, and there's a reward at the end for those with the non-cogs to complete the odyssey. If so, I flunked.

Update: Here's a 2005 NYT article about SAT writing and length, thanks to Redditor punymouse1, who also contributed "Fooling the College Board."

Sunday, October 11, 2009

Michigan State Noncognitive Study

A "Report of the First-Year Follow-up of College Applicants at Twelve Universities" was forwarded to me by a colleague. According to him, the report is public now, but the appendices with survey instruments are not. I haven't seen the report itself online, and because of copyright don't feel like I can post it. But herein are some highlights.

The authors are Neal Schmitt, Abigail Billington, Juliya Golubovich, Jessica Keeney, Timothy Pleskac, Matthew Reeder, Ruchi Sinha, and Mark Zorzie, and the work was supported by the College Board, which is interesting.

The basics:
  • Participants: Earlham College, Furman University, Johnson & Wales University at Providence, Kenyon College, Lafayette College, Meredith College, Michigan State University, Ohio State University, Purdue University, University of North Carolina at Chapel Hill, University of Southern California, and University of Washington
  • 844 students provided enough data for analysis, initially through a College Board web survey of applicants on background, interests, and judgment, then through a follow-up for those who enrolled. The overall yield rate (enroll/app) was 26%. Of those who enrolled, 42% responded to the follow-up survey (there was a $20 incentive).
  • The kinds of data considered were biodata (background and life history, not blood pressure), a situational judgment test (SJT) representing behaviors, demographics, self-reported performance (BARS), citizenship behaviors (positive or negative), academic satisfaction, social satisfaction, grades and standardized test scores, an inventory of "shock" events a student may have experienced, the big five personality traits, substance abuse, use of time, and self-reported withdrawal tendency.
The report is chock-full of tables of data with correlations and crosstabs, but I'll skip to the regression results. Here are some key findings, quoted from the executive summary (pg. 3):
  • First-year college GPA is predicted significantly by several biodata scales, most notably Knowledge, Ethics, and Perseverance, but HSGPA and SAT/ACT scores are much more predictive of college GPA than are biodata and SJT.

  • Self ratings of performance (BARS), Organizational Citizenship Behavior (OCB), and student self-reports of Deviance were especially well predicted by the biodata measures and SJT while HSGPA and SAT/ACT were relatively uncorrelated with these outcomes.
It's disappointing that first year GPA isn't better predicted by this gob of noncognitive variables. But completion is a better goal, and the study hasn't had time to mature to that point. For example, any actionable information about first-year retention would be worth its weight in undergrads. Stay tuned for the next report.

Update: Dr. Neal Schmitt gave me a link to a publications page, which will soon include the body of the report I cited.

Friday, October 02, 2009

Friday Potpourri

This morning my cognitive cogs are engaged in so many projects it seems appropriate to dish up a small buffet of assessment-related items:

1. SES gets you a seat: An article in InsideHigherEd this morning tells that as access to higher ed boomed in the US, so did demand, so that competition naturally favored those best able to negotiate the barriers to admission. Guess who that was? This quote is attributed to the researcher, Sigal Alon:
[T]esting became a more important factor in admissions at more institutions, and that wealthier families are much speedier to adapt to changes in admissions rules.
This sounds like the "economic value of error" stuff I was prattling on about a few weeks back. Although I can't find the source article online, a kind of appendix is here. She reiterates what everyone paying attention knows:
Test scores thus seem largely redundant with high school grades (Crouse and Trusheim 1988). Test scores do have some predictive value for college success, but is this improvement worth the negative impacts of testing?
She goes on to advocate a more aggressive program than simply using noncognitives to better identify talent in lower SES groups. She's probably right--if only because the high SES groups will learn how to negotiate the new system of merit just as well as they did the previous one.

2. Rubics are bad for your brain: Well, maybe I exaggerate, but two articles posted to the ASSESS listserv provide cautionary advice for those constructing rubrics. On another level, there is (I find) a tendency to put too much faith in the things to begin with. Maybe that's because a rubric is tangible evidence of a plan. It looks like science--with its dissecting of some learning outcome into dimensions of quality and component. I won't reiterate my own annoyances with this quasi-religious belief here. You can find better and more constructive comments in the papers themselves:
3. NSSE Dimensions: At a recent IR meeting locally, I saw a very nice analysis of one institution's NSSE data from several years. The question being considered was whether or not the scorecard-type ratings that NSSE reports out are valid. You can see what I'm talking about in the 2008 annual report here. I clipped out a bit from page 8 to show some of the "dimensions" in question:


Here you can see that Level of Academic Challenge, Active and Collaborative Learning, and Student-Faculty Interaction are indices consolidated, summarized, and reported out as being meaningful. There is a great power to words, which gets abused routinely by the testing and assessment profession (like misuse of the word "measurement"). To be fair, it happens anywhere there's room for doubt--opportunity to advertise something more than what's there without being caught. We humans seem to prefer certainty over doubt, even if it guarantees we're wrong. But I digress.

The putative dimensions listed (Enriching Educational Experiences and Supportive Campus Environment are the other two, not shown above) are drawn from the items on the NSSE. In order to validate such things, one would test the construct validity using factor analysis. You can find lots of reports at NSSE referencing validity here. I will have to go through them later, to see what I can find in this vein.

The presentation at the IR meeting used confirmatory factor analysis (CFA), which posits that the dimensions are defined according to given rules, and then tries to verify that using the results at hand. Ideally, the items related to Academic Challenge would all show up as elements of a single vector when the correlation matrix's singular value decomposition is calculated. The actual results were not convincing.

I poked around to see if there were other studies out there, and found this one. Here, researchers at James Madison University and Eastern Mennonite University used similar methods and found similar results. They summarize their findings:
[A] CFA of the benchmarks as specified by NSSE literature produced poor model fit. For the sample of students examined for this study, the five-factor benchmark model supported by NSSE was not upheld, thus a five-factor solution is uninterpretable for the sample. A comparison of benchmark scores from this sample to scores from a sample at another university should not be made, as the benchmark scores from this sample were not empirically supported. Until further studies can provide evidence for an interpretable set of benchmarks for this and similar samples, one should not make policy or programmatic decisions based on the benchmark scores.
The whole article is worth reading. It's well-written, and there are interesting details about particular loadings.

Monday, September 28, 2009

Predicting Success

Nature abhors a vacuum it's said. My daughter showed me the other day how her science class used this "principle" to determine the amount of oxygen in the air by using a candle and a test tube in water to measure before and after. Of course, it's isn't really abhorrence but air pressure that makes vacuum a chore to create down here where we live. But maybe economics really does abhor a passed-over opportunity. As in the old joke where one economist says "hey, there's a hundred dollar bill laying there on the ground," to which the other replies "can't be so--someone would have picked it up."

I've argued for a while that the low predictive validity of GPA + SAT creates market opportunities for those willing to experiment with other demonstrations of achievement. Here's a list of previous posts related to the topic:
In Malcolm Gladwell's Outliers, he has the following observation about the international math and science test called TIMSS . He notes that the test is accompanied by a 120-question survey, which many students don't complete. In his words:
Now, here's the interesting part. As it turns out, the average number of items answered on that questionnaire varies from country to country. It is possible, in fact, to rank all the participating countries according to haw many items their students answer on the questionnaire. Now, what do you think happens if you compare the questionnaire rankings with the math rankings on the TIMSS? They are exactly the same. (pg. 247)
He concludes a page later that "We should be able to predict which countries are best at math simply by looking at which national cultures place the highest emphasis on effort and hard work."

One could certainly ask for better analysis--why not correlate by student the number of survey items completed against the math score, rather than aggregating by country? But the sentiment is certainly the same expressed in noncognitive literature: there's more to success than the ability to do mental gymnastics.

In InsideHigherEd today there was an article "Next Stages in Testing Debate" talking about institutions that de-emphasize SAT in admissions decisions:
[A] common idea was that decreasing reliance on the SAT does not mean any loss of academic rigor and can in fact lead to the creation of classes that do better academically (and are more diverse).
This may require a rethink and additional training of admissions staff:
That report said that for too many admissions officers, the only training they receive on the use of testing may come from the technical training provided by testing companies, entities that have a vested interest in the continued use of testing.
Some of the other types of accomplishments sought by admissions officers are evidence of creativity, practical skills, wisdom about how to promote the common good (Tufts), and essays (George Mason). This idea is something that seems to be blooming. As Mr. Gladwell might say, it's blinking toward an outlying tipping point.

Evidence of this arrived in my in-box the other day: a forwarded email from the Law School Admisssion Council (LSAC) with the following news:
LSAC has funded research on noncognitive skills for some time. A study funded by LSAC—Identification, Development, and Validation of Predictors for Successful Lawyering by Marjorie Shultz and Sheldon Zedeck—identified 26 noncognitive factors that make for successful lawyering. The study included suggestions for assessments that might measure those factors prior to admission to law school.
I hadn't heard of Shultz and Zedeck, so I scurried off to the tubes that comprise the Internets to find out more. You can find the whole 100-page report and more here. The executive summary has an imposing, lawyerly warning on the front page:
NOT TO BE USED FOR COMMERCIAL PURPOSES NOT TO BE DISTRIBUTED, COPIED, OR QUOTED WITHOUT PERMISSION OF AUTHORS
I suppose this implies an argument that fair use somehow doesn't apply to this work. In any event, since I've obviously already violated the terms with the quote above, I may as well proceed.

Rather than focus on something easy like just predicting law school GPA, the researchers actually assessed job performance and got LSAT and law school performance data. Then they threw a bunch of tests at the problem, like Hogan Personality Inventory, Hogan Development Survey, Motives, Values, Preferences Inventory, and Self Monitoring Scale, a Situational Judgment Test, and Biographical Information Inventory. This on a sample of more than 1100 subjects. I think the word I'm looking for is "wow."

The results found successful indicators of effectiveness (per their definition) that seemed to be assessing independent characteristics, adding dimensionality to the standard predictors. In fact, the LSAT didn't seem to predict success at all. Go look at the executive summary for details--it's not very long.

All in all, this seems to be a solid study that shows that noncognitives are important--perhaps better--predictors of professional effectiveness than the traditional cognitive ones.

Other noncog stories in the news:
The College Board is even getting into the act. A 2004 publication talks about "individualized review" that includes factors beyond GPA and test. I am led to understand that they have a big project underway now on noncogs, but I can't find the website.

Okay, remember the vacuum? What if we managed to identify these noncognitive variables and start to use them? The LSAC folks outline what can happen next:
A major concern about developing an assessment for noncognitive factors is the possibility that the test would be so coachable that its results would be unreliable in the high-stakes environment of law school admissions.
I argued here (Zog's lemma), here, and here that any imperfect predictor invites error inflation for economic gain, but it doesn't take a genius to see that if checking "I'm a hard worker" gets me more financial aid, I'll be more inclined to overestimate my industriousness. This is a game theory problem with no solution that's likely to be mass marketed. Imagine if the assessment of prospective students had to be done laboriously by hand by highly trained admissions staff, instead of relying on a convenient test that cranks out a one-dimensional predictor. I'm not sure that's a bad thing.

Monday, August 31, 2009

Money, Genes, and College

The New York Times has a recent article on SAT related to family income here. It shows that incomes and scores have high positive correlation for 2009. I've reproduced the graph from the article below.
The article doesn't say that income causes higher scores (as I recently speculated), but the article is nevertheless criticised by an economics professor here. In his blog post, Prof. Mankiw suggests that a significant part of the slope is due to genes, with an argument along the following lines, which I've made more explicit here:
  1. IQ correlates positively with income
  2. The ability to perform well on an IQ test is influenced by heredity
  3. IQ correlates positively with SAT
Therefore, students of wealthy parents should have higher SAT scores merely because they are smarter. This is presented as a bias for explaining part of the slope of the curve evident above (which is greatly exaggerated by the scale used, please note). How much of the SAT bonus is due to IQ is not spelled out, other than the concluding note in Dr. Mankiw's article:
It would be interesting to see the above graph reproduced for adopted children only. I bet that the curve would be a lot flatter.
I interpret "a lot flatter" to mean that the IQ contribution accounts for a significant part of the slope. This is all reminiscent of the The Bell Curve and the controversy of genes vs. environment in the creation of intelligence.

Some analysis is in order. The inheritability of intelligence is a very political topic. The left would like to assume that all people really are created equal, and that the "blank slate" is there to be written upon. This legitimizes interventions that affect socio-economic status (SES). The right would like to believe that interventions are counter-productive because intelligence is fixed at birth. Both argue from the conclusions back to reasons for believing them, which is probably some kind of tragedy of the commons in the public realm: in order to stay in power, parties have to act sometimes in a way that is contrary to the common good. Since I don't have to be elected I can freely wish a pox on both their houses. Where is the science on the matter?

1. Income and IQ

There seems to be good evidence that IQ and income correlate positively. It's important to remember that IQ and intelligence are not the same thing. The first is a monological definition based on a particularly kind of cognitive test, and the second is a vocabulary item in common usage. In The Bell Curve and subsequent articles, the authors try to make the case that intelligence stratifies the employment landscape with the Very Dull at the bottom, hardly educated and hardly employable, and the Very Bright at the top. (Those words annoy me because dull is not the opposite of bright. Dim is.) In reading those arguments, it conjures up for me a kind of overarching social history that's implied--something on the order of what Marx sparked. In any case, it's a very strong conclusion they try to reach. You can spend a lot of time reading the debate about that book and related concepts.

It's noteworthy that physical attractiveness seems to also be linked to higher income.

2. Genes and Intelligence

The problem with an approach like The Bell Curve is that it's the wrong discipline. If you want to make conclusions about genetics, you need to actually look at some genes. This is especially true if you want to make big conclusions. The authors try very hard to control for environmental factors, but that doesn't substitute for identifying DNA that is causally linked to intelligence. That kind of work is being done by biologists, however. Here are some examples you can find on ScienceDaily.com:
Of course, there is science pointing to environmental influence over intelligence as well:
Largely ignored in the grand debate over environment vs. genes are new findings about epigenetics, which you can survey here and here.

I think a reasonable person has to conclude that genes, epigenetics, and environment all play a role. Because genes are discrete, it ought to be possible to identify consequent effects with more precision than the other categories. The third article linked in the list above pins a particular gene to an R^2 of 3% of IQ.

One of the puzzling things about IQ is that if it really describes a primarily genetic effect, how can we explain the dramatic rise in scores in the last decades? This is the so-called Flynn Effect, summarized in Wikepedia as the change over a 30-year period where:
  1. the mean IQ had increased by 9.7 points (the Flynn effect),
  2. the gains were concentrated in the lower half of the distribution and negligible in the top half, and
  3. the gains gradually decreased from low to high IQ.[reference]
What would a genetic-based explanation look like? I think you'd have to assume that the genes for Dullness became expressed less often, perhaps because they are going extinct. This would suggest a harsher survival environment for said genes. It seems to me, however, that's it's becoming easier to survive without developing intelligence, but that's just my take on it. Others have commented on this at great length. See a rather critical piece here.

I think it's also fair to conclude from the evidence we have that smarter parents on average will produce smarter offspring, all other things being equal, but that this is not fully deterministic. The question of how much intelligence is inheritable is still open.

3. IQ and SAT

The link between IQ and SAT is also controversial, at least from test-maker's point of view. See an overview here. There seems to be a strong correlation between the two, which you can read about in this research article.

SAT isn't very good at predicting first year college grades, but that's what it's designed for. It's even less good at predicting success beyond the first year. This isn't perhaps surprising: the content of the test resembles high school and college freshmen academic work. There are many other factors that influence success. I was interested to read here that:
Bates College, which dropped all pre-admission testing requirements in 1990, first conducted several studies to determine the most powerful variables for predicting success at the college. One study showed that students' self-evaluation of their "energy and initiative" added more to the ability to predict performance at Bates than did either Math or Verbal SAT scores.
Is IQ similarly limited in predicting college success? I couldn't find anything definitive about that in the time I had at my disposal, so I'll leave the question open.

Conclusion.

If we can reach any conclusion, it's a very weak one. Almost certainly income is indirectly linked to SAT scores through the associations advanced. However, how much that bonus is remains unclear. My blog post last time pointed out that it's not only that higher incomes get higher SATs, but that the year-to-year differential is also higher. If this is not some statistical artifact, it doesn't jibe with a purely genetic explanation, as it would require the gene pool to be evolving at a very high rate, implying extraordinary pressure from a fitness gradient.

I looked for more historical data on this increase as a function of wealth, but unfortunately it's only in the last two years that the SAT included the salary range from zero all the way to $200,000+. The old version only went to $100,000, and the most interesting part of the curve is above that. Having said that, I did not find any evidence for my theory about the economic value of error with the data that is available. We'll have to wait for next year's results.

There is, however, a direct causal explanation that is difficult to imagine away. It contrasts with the somewhat specious illustration Prof. Mankiw gives:
Suppose we were to graph average SAT scores by the number of bathrooms a student has in his or her family home. That curve would also likely slope upward.
Correlation and causation are different things. But consider another scenario. Suppose we were to graph SAT scores by the number of prep-tests a student attended, or the number of times they took the test (guaranteed on average to raise the maximum score), or the quality of the high school they attended. All of those have plausible causal connections to how well a student performs on a test of high-school cognitive material, no? And which students have access to the best high schools? Can afford to take the test multiple times? Can afford prep-tests that are advertised to raise scores 100 points?

The ironic thing is that the economic value of raising the SAT is real because of the usual policies in awarding financial aid. So the wealthiest are in the best position to reap that aid for both reasons of intelligence and the ability to buy error (over-prediction due to coaching or taking multiple tests). This trend has been documented for a long time here.

This line of thought gave me an interesting idea some time ago, which I hope to turn into a project. At my current institution we have embarked on a shift toward recruiting better students (students who have a better chance of success). But the conversation with the board, as well as public perception, includes the median SAT scores of the students we admit. My preference is to use better predictors (including noncognitive ones) to find the best students, but this is largely independent of their SATs, which creates a tension between actual goals and perceptions.

The idea is to create a grant-funded summer camp for the low-SAT that are nevertheless predicted to do well. In this camp they would receive SAT test-taking preparation and then retake the test immediately afterwards. This would raise median SATs without affecting our ability to get the students we want.

Update: here are some references:

Thursday, August 27, 2009

Amplification Amplification

There's something seductive about self-reference--a topic delightfully expounded in Douglas Hofstadter's book Metamagical Themas. Self-reference shook the foundations of math with the Russell Paradox and then Gödel's incompleteness weirdness. But that doesn't stop the fun--by no means! The latest, coolest application I've seen is the RepRap, which is a machine that can build other things, including itself. I suppose that if the conditions are right it will create a whole ecology of replicator replicators. That's what we are too, of course, just squishier and less prone to rust.

This gratuitous introduction is by way of explaining the title. I've spent some time thinking about Zog's Lemma about the amplification of error, and wondering if there's actually anything there. I think there may be, even in the cold light of the fluorescent sun. I will therefore try to amplify on my prior post.

The idea is this:
As economic value increases, for weak predictors the error of prediction can increase over time.
I opened up WinEdit to try a mathematical formulation, but I haven't gotten too far with that yet. The formality may get in the way anyway. We're considering the prediction of the value of some future valuation. The examples will make it clear, but think Prediction = Actual + Error. The last term could be negative. Consider a couple of scenarios.
1. Imagine pulling a dollar bill from your wallet or handbag. How much is it worth, in dollars? One, right? Well, if you got a fiver by mistake, it's worth five bucks. You're not likely to be convinced by a huckster that your fiver is only really worth one and fork it over for four quarters.
Here we have a situation where predictive error is very low and economic value is low too. The actual value of the dollar is realized when you buy something with it. Let's suppose that a fast food burger is a dollar. If you get a single, it was a dollar. If you get five, it had Lincoln on it.
2. Now imagine that you can't tell the difference between a one dollar bill or a five, or a twenty, or a hundred... You have to depend on "experts" for this. The problem is real because you don't want to walk to the burger stand with a hundred dollar bill; you're just not that hungry. This is not entirely hypothetical: US currency is all the same size, so if you're blind this is the way it is. Euros come in different sizes.
Now the predictor has dropped because you can't be sure whom to trust. There's a chance someone will lie to you when you show them the bill and try to take advantage of you.

Both of these scenarios are with relatively low stakes. The competition for your one dollar or five is not as great as if it were a million, right? What happens when we now consider situations where the economic value of the prediction is high?
3. Airplanes are pretty well understood--just watch this. Predicting what will happen under lots of conditions is achievable. Still, things go wrong once in a while. I think that we can agree that this is high stakes--no need to put a dollar figure on a successful landing.
In this situation, what has happened is that the error has gotten smaller over time. That is, by assiduously investigating every prediction failure, we learn how to make better predictions. This is just how science works, of course.

Now for the weird part. Try to imagine situations where the economic value of the predictor is high, but the accuracy isn't so great. What happens now?
4. You want to plant your crops as soon as possible for economic reasons. The shaman says that the gods have decreed that the last freeze will come late because of some indiscretion within the tribe.
This is deadly serious--if your crop gets frosted over, you may not last the next winter through. But the predictor isn't a very good one. What are the dynamics? In this case, the hapless farmer doesn't have many options. Without an anachronistic scientific approach, there is no way to reduce the error. But that's not the end of the story.

The Shaman has a vested interest in maintaining credibility, no? If it emerges that he's full of buffalo cakes, the tribe may put him out on his ear. Since P = A + E, and (despite his best efforts) he has no way to influence the actual outcome A, nor can he improve his predictions P, his only margin for improvement is E. But how can that be? True, he (or she) can't actually improve the error, but he can try to make it appear so with mystical mumbo-jumbo of sufficient impressiveness. So his best bet is to create the best fakery he can--double and redouble the pageantry and behavior that's so far out of norm that it MUST be profound. A successful faker can then get by without paying any attention to the real problem of prediction, and can just get better and better at pretending to.

The sticky point of this equation--what we need actual math for--is whether Herr Shaman could completely abandon accuracy. Maybe he has some residual knowledge passed down about the seasons, the moon, etc. Can the error rate actually naturally increase in such a situation? Only if the effort in maintaining the better predictor is not more than compensated by an equal effort spent toward new and improved humbug.

But what if we can affect the error ourselves?

5. Odysseus wants to sneak his men past the furious cyclops Polyphemus, who he has just blinded. He and his followers ride underneath sheep so as to fool the giant as he checks who exits by touch. (Homer I ain't.)
The predictor is the cyclops in this high stakes assessment. Odysseus (not being Circe) can't change the fact that his men are men, so the A is immutable. He can only affect the prediction by increasing the error, which he does in fine style. Notice the contrast between this and the previous example. In the previous one, disguising the size of E was done for economic value. Here, E is actually increased as much as possible for the same reason. Odysseus spends a lot of time increasing E, actually.

In common parlance, this is called cheating or manipulation, and it's common even in low-stakes situations as we all know.
But what about the amplification of error I advertised? Glad you asked.

Just imagine a recursive situation like number five above, where there are repeated rounds. Each time, Odysseus tries to sneak by the mutilated son of Poseidon, and each time the cyclops tries to detect him. There are lots of situations like this: they evolve. In biology it's sometimes called a Red Queen Race, after a bit from Alice in Wonderland about running as fast as you can just to stay in place. So we have a curious effect where the predictor P is precariously balanced against the error E, like a tug of war: not moving much perhaps, but with great forces involved.

But there's no reason to assume the playing field is fair. What if the conditions are better for creating E than for eliminating it? That might be the case for the SAT test. The chart below was clipped and pasted together from one in InsideHigherEd yesterday here. It lists year over year improvements in SAT score averages by income range.


It shows convincingly that more money gets you higher scores. Exactly why that's true is debatable, but remember this is an increase for one year. One hypothesis is that students of wealthier families have more access to test preparation and can pay to take the test more times, and probably have parents who enable all this as well as pay for it. Does all this effort increase the teen's actual ability to succeed in college (let's call it A)? Or is it responding to the economic value of the predictor (the SAT score) by increasing E instead?

If it's the latter case, then we have a pretty snapshot of an imperfect estimator demonstrating prediction error on the increase because of economic forces. This would imply that the psychometricians who create and score the SAT haven't found a way to counter the increased error that's hypothesized. Zog's moment in the sun?

Tuesday, August 25, 2009

Zog's Lemma: Assessment and the Amplification of Error

The concept of outcomes assessment is like a Swiss Army spatula: it's used in all kinds of ways. As opposed to the proverbial Russian Army hardware, pictured below (ubiquitous on the Internets):
To some, outcomes assessment is a touchy-feely endeavor of encouraging the practitioners of higher education to do the right thing, close the right loop, bring a glowing smile to the visiting team. All that. I'm comfortable with that.

To others, it's a more serious matter, more scientific in approach, rather like making tick-marks on the door's threshold on a child's birthday to signify evidence of growth. We know that is serious business: with shoes or without, and where exactly is the top of the head? Does hair count or not? It grows too, after all.

The scientific approach requires belief in theory: that assessments are valid to the purposes we employ them for. In reality we usually have no real way to know that based on a solid physical theory. The belief has to be defended by what statistics can be summoned to make a case. (In truth, beliefs don't need to be defended at all; they only need to be believed.) What comprises this validity? I'd like to address that question sideways. The discussion of what constitutes validity is a groove cut deeply in the literature of psychometrics, and I'd rather ask a more important question: of what use is it to believe in the validity of an assessment?

It's easy to make hay from the fact that the definition of validity itself isn't settled, but this is unfair, I think. It's a difficult philosophical nut to crack. For my purposes, predictive validity is the most important aspect of assessment. If an assessment doesn't tell us anything about what's going to happen in the future, I don't see much use in it. The Wiki article on predictive validity contains an interesting observation that is the crux of the value of an assessment:
[T]he utility (that is the benefit obtained by making decisions using the test) provided by a test with a correlation of .35 can be quite substantial.
Utility is a concept from economics that allows for a weighting of outcomes to balance outcomes that would otherwise be numerically indistinguishable. A classic example is the fact that $1000 means more to a poor person than it does to a rich person, even though it won't buy any more for the former. The relative worth to the penniless is more. The point is that even imperfect predictive assessments are useful.

An example of this is the college admissions process. Let's suppose that with the data gathered on the application, the institution can estimate the probability of "success" of student. Success could be retention to second year, or GPA > 2.5, or graduation, or whatever you deem important. The statistics can be generated with a logistic regression, which will yield a model with a certain amount of predictive power. No model is perfect, and even in retrospect (feeding the original data back into the model) it will not correctly classify applicants as "predict success" or "predict failure" with 100% accuracy. There will be some proportion of false positives and false negatives. This can be visualized on a Receiver Operating Characteristic (ROC) curve. You can see a bunch of them on google images. The usefulness of the ROC curve is that it lets you visually explore the decision of where to set a threshold for decision-making. If you set it too high, you get too many false negatives (reject too many qualified candidates). Too low and you admit too many false positives.

For an institution, such a tool is obviously useful to believe in: one can work backwards from the desired number size of the entering class to see where the threshold should be set. You could even estimate the number of false positives. Any power to discriminate between successful and unsuccessful students is better than none. Of course there are other factors, such as ability to pay, that make this more complicated. To keep things simple, I won't consider these distractions further.

One can imagine a utopia springing from this arrangement, where the assessments continually get better and the applicants learn to distinguish themselves by giving signals that the assessments can recognize. The first doesn't seem to be happening, and the second has an unfortunate twist. An analogy from biology serves us well here.

The story of the peacock, according to evolutionary biologists, is that a showy mating display is worth the biological cost of making all those pretty feathers because of the payoff in reproduction. The similarity is that when mates choose each other they have limited information to go on--an assessment we could characterize with a ROC curve if we had all the facts. This produces a distortion that favors any apparent advantage in a potential mate, or in the case of the peacock, creates over time a completely artificial means of assessment. The peacock is saying "look, I'm so healthy I can drag around all this useless plumage and still escape predators."

So, if the value of partial assessments is clear from the institutional vantage, it's a different picture altogether from the applicant's point of view. It's interesting that InsideHigherEd has an article this morning on this very topic. According to the article, in response to increasing competitiveness at "top institutions":
[H]igh school students could respond to the pressure by taking more rigorous courses and studying more -- or they could focus their attentions on gaming the system and trying to impress.
A study by John Bound, Brad Hershbein, and Bridget Terry Long shows that while this perhaps motivates students to take a more rigorous curriculum, it also prompts them to spend more time in test preparation, or in games like trying to engineer more time to take the test. Peacock plumage? It's hard to say without seeing actual success rates. It could be that having the willingness to spend all that extra effort is itself a noncognitive predictor of success (that is, not related to the scores themselves, but to the personality traits of the applicant). As Rich Karlgaard put it in Forbes magazine, a degree from an elite institution is valuable because:
The degree simply puts an official stamp on the fact that the student was intelligent, hardworking and competitive enough to get into Harvard or Yale in the first place.
What is clear is that those with the means to game the system are better off than those who do not. So we see things like test-prep for kindergartners at $450/hr in the upper crust of society, but probably not so much in housing projects. This likely produces more false positives among the select group: it's as simple as money buying better access. [Update: see this InsideHigherEd article for some dramatic numbers to that effect. Year over year, SAT scores increased 8-9 points for $200,000+ families, and 0/1 points for the poorest group.]

The economic demand for false positives has become an industry under No Child Left Behind, and that doesn't seem to be changing. The New York Times has published letters from teachers on this topic. Some quotes:
  • [T]he use of test data for purposes of evaluating and compensating teachers will work against the education of the most vulnerable children. It is a mistake to conceptualize education as a “Race to the Top” (as federal grants to schools are titled) — for children or schools. (Julie Diamond)
  • Linking teacher evaluations to faulty standardized tests ignores the socioeconomic impact on a nation that is both rich and poor. Can a teacher confronting the poverty of some children in Bedford-Stuyvesant be made to compete with a teacher instructing affluent children in Scarsdale?(Maurice R. Berube)
  • My job went from teaching children to teaching test preparation in very little time. Many of our nation’s teachers have left their profession because the focus on testing leaves little room for passion, creativity or intellect. (Darcy Hicks)
  • Standardized tests are, by their nature, predictable. Most administrators and teachers, fearing failure and loss of position and/or bonuses, de-emphasize or delete those parts of the curriculum least likely to be tested. The students sense this and neglect serious studying because they know that they will be prepped for the big exams. (Martin Rudolph)
The economics of false positives is clear: test prep is an industry, teachers and administrator and schools are rated by how many they generate. Of course the object is not to create false positives, that's just the result of so much emphasis on an imperfect assessment.

Less attention is paid to false negatives. Research points to noncognitive traits like grit, planning for the future, and self-assessment as being important to actual success, but these are not directly accounted for in standardized assessments. To be sure, college admissions officers look at extra-curricular activities to try to add value to SAT, GPA, and curriculum, but I think it's safe to say that the cognitive assessments are primary. How badly do we underestimate actual performance?

That question comes up when we evaluate or create a predictive model for applicants. The resulting predicted (first year) GPA can be used for admissions decisions and financial aid awards. In the Noel-Levitz leveraging schema applicants are sorted into "low-ability" to "high-ability" bins for individual attention. The question of false negatives is the same as the question "how accurate is the description 'low-ability', based on the predictors?" Not very good, as it turns out.

Every time I ran the statistics I got the same answer: at the very bottom end of our admit pool--those Presidential and provisional admits who were supposed to have the hardest time--about half of them performed well (I usually use GPA > 2.5 for that distinction). Since we reject students below that line, we should assume that about half of those students just below the cutoff would have done as well too. If this is typical, there are a LOT of false negatives.

Remember, this line of thought applies to all outcomes assessments to one degree or another. Let's take a concrete example. Suppose we want to assess vocabulary knowledge of German language students with an exam. Memorizing a single word (including declensions for nouns and conjugations for verbs) is low-complexity. The total complexity is that for one word times, let's say 10,000 total items, less any compressibility of this data. All told, this is a lot of complexity (measured in bits) for a human. So it's a suitable subject for assessing, and the low complexity per item means that each of them can be assessed with confidence.

Supposing that we do not have the resources to test our learners on all 10,000 items of vocabulary, we'll have to sample randomly and hope that the ratios are representative (or else spend a lot of time checking correlations and such). Maybe our assessment is only on 100 items, chosen at random from the 10,000. If Tatiana actually only knows K of the items, then there is a chance of K/10,000 that she will know an individual word, and theoretically score K/10,000 on average (on any sized test). Given this arrangement, what is the trade-off between false positives and false negatives?

We will assume we can use the normal distribution to estimate these Bernoulli trials. This only requires that K not be too large or too small--in those cases the chance of an error diminishes anyway. We already know the mean is p=K/10,000, and the standard deviation is given by SQRT(p(1-p)). Two standard errors is about 3% for large and small values of p, ranging to about 5% for those in the middle. These would decrease for a number of items larger than 100, and increase for a smaller test.

The conclusion is that if we set our cutoff for passing in a usual spot, say 70%, we should expect at least half of the scores in the range 65-75% to be either false positives or false negatives. For example, if Tatiana actually knows 70% of the vocabulary items, she has a 50% chance of passing the test because the distribution of her scores over all possible tests is (almost) symmetrical and centered on 70%.

The actual number of false positives and negatives depends on where the skill ranges of the test-takers lie in relation to the cutoff value. The more there are close to the cutoff, the more errors there will be. I actually witnessed a multiple-choice placement test being graded one time, and inquired about the cutoff to find that the grade to pass was actually less than the average result expected by chance! This certainly reduces false negatives, but I'm not sure about the overall result.

The vocabulary example is a best case, where the items themselves are of low complexity and (we hope) not subject to a lot of other kinds of error. When the predictive validity is actually quite low (as in the example where we explain a small percentage of the variance in the outcome), the proportion of errors in both directions is far worse. What does this mean, then when we "believe" in a test like the SAT and adopt it as a de facto industry standard? Even together with high school GPA, these can typically explain only about 33% of the variance of first year college grades.

First, it is of benefit to the institution to be able to imperfectly sort students by desirability. It is costly if the institution has to bid for the apparent best applicants with institutional aid, and there is incentive to look at noncognitives and better predictors, but I can't see a lot of progress in that direction. So the false positives get bid up with the rest. Meanwhile, the false negatives--those applicants who would succeed but don't show it on the predictor--get passed over for admission or merit aid.

A tentative conclusion is that a weak predictor gets amplified in a competitive environment. This is probably true in lots of domains, like the evolutionary biology example of the peacock. Probably some economist has his/her name attached to it: Zog's Lemma or something. I'll have to ask around. (I just googled it--apparently there is no Zog's Lemma.)

For the assessment types, there are two lessons. First, there are going to be errors in classification. Second, those errors get amplified as the assessment gains importance. This is an argument against high-stakes assessments, I suppose. I've always gotten good results by keeping assessments free from political pressures like instructor or program review. "Accountability" creates problems as it tries to solve them. Acknowledging that would be a fine thing.

PS, if you're interested in other kinds of amplifiers in science, read this.

UPDATE: see Amplification Amplification

Saturday, March 07, 2009

SAT Validity

The College Board has a nice page on Data and Reports, with a report on SAT Validity showing correlations between SAT and high school grades (HSGPA) and first year college grades (FYGPA). The table below from the report shows these statistics.

There is also a report on demographic breakdowns called Differential Validity and Prediction of SAT, which shows differences between race and gender groups. On the whole, predictive validity doesn't vary much, especially considering that what we're interested in usually is R-squared.

Since I've been interested in using noncognitive indicators to predict achievement, I looked for some statistics to guide me. All I've found so far is in the paper "Predicting the Academic Achievement of Female Students Using the SAT and Noncognitive Variables" by Julie R. Ancis and William E. Sedlacek. In the study used in the paper, two of the noncognitive dimensions seemed to be the best predictors (my interpretation): realistic self-appraisal and community service. There wasn't information about the residual when GPA and SAT are also taken into account, but putting these together, one can create an approximate "best case" scenario based on the statistics. This is shown in the chart below.



The contribution of the noncognitives was to me disappointingly low, and it is likely even smaller because of the correlation with SAT and GPA, which is unknown here. That is, the green slice might actually overlap with the red or blue ones.

It's interesting to note that that "real-life" correlations like the one in the Ancis and Sedlacek study are lower between SAT and FYGPA than in the College Board's research reports. This has been my experience too, and I can't explain the difference. I came across a rather impassioned argument for dropping the SAT as a predictor in a 1992 proposal from Jonathan Baron at University of Pennsylvania. His correlations are even lower than mine for SAT and FYGPA, and he argues that there's not enough value added by the SAT (after taking into account other predictive variables the university uses) to justify its continued use. He makes an interesting point about the emphasis on SAT in the admission process:
College admissions criteria have major effects on high-school
education. College-bound high-school students do what they think
will help them get into a good college. Students now spend a
considerable amount of time preparing for the SAT. (When my son
was taught how to take multiple-choice exams in Kindergarten,
when he was in a group of children who could already read, I
complained that this was an inappropriate activity, and I was
told that it's never too early to start preparing for the SAT!)
If they were told that the SAT was not important but their grades
and their achievement test scores WERE important, they might
spend more time trying to learn something and less time trying to
learn how to appear to be intelligent on a test. This might be
reason enough to drop the SAT, even if it were somewhat useful
for prediction.
The trend toward standardized testing as a measure of minds is not limited to SAT. Now we have the whole No Child Left Behind apparatus and an increasing appetite at the Department of Education to infect higher education with this philosophy using the likes of the CLA. I hope that instruments using noncognitive assessment can get more attention and be developed into something useful.

Monday, February 09, 2009

Beyond the Big Test

This weekend I got my copy of William E. Sedlacek's text with that title, which is also subtitled Noncognitive Assessment in Higher Education. Anyone who's read my blog in the last year will know I've scratched my head over the evident value that admissions processes typically put on ACT/SAT, and to a lesser extend high school grades (see this article, for example). Our work with surveys like the CIRP shows that there are behavioral and attitudinal traits that make good predictors of, for example attrition. These are generally ignored in the admissions process. Or it might be better to say they're not looked for.

Well, silly me. Apparently this subject has a long history in the literature and is a quite well developed concept. Some major schools like North Carolina State University have used these methods successfully. So my research program to find variables for use in predicting academic success can accellerate considerably--we just have to customize the work others have done.

I have not finished the book, but can tell already that it's a wonderful resource. The background and history of noncognitive assessment is given, as well as solid research findings, actual survey instruments, and examples of how to coach staff in looking for these traits in interviews, on existing application materials, essays, etc.

The focus of the book is toward evening the playing field for what the author calls non-traditional students. In his usage, this means anyone who isn't a white male. My purpose is more targeted, but the material is no less useful for it.

The specific noncognitive traits identified, with some help from factor analysis in the process of ascertaining construct validity of surveys, is as follows (pg. 7):
  1. Positive self-concept
  2. Realistic self-appraisal
  3. Successfully handling the system
  4. Preference for long-term goals
  5. Availability of strong support person
  6. Leadership experience
  7. Community involvement
  8. Knowledge acquired in a field
These are described in detail, of course, along with methods of assessing them.

I'll be recommending that our university proceed full speed ahead with this project, to catch what we can of the current cycle. This will undoubtably make Mr. Sedlacek happy, as it will entail buying many more copies of his book for distribution.

Thursday, January 29, 2009

Admitting the Right Students

All of us who teach like to have students in the classroom who want to learn. In a recent conversation this topic came up. The math teacher I was speaking with said he'd take a hard-working curious student over an indifferent student any day--regardless of academic preparation of the latter. I would too. The real reward of teaching, after all, is to see growth. One of my proudest (shared) accomplishments is that one of my MAT 100 (really a remedial course) students fell in love with the topic and went on to graduate with a BA in mathematics. She now works in a technical profession and is very successful.

All this by way of introduction to my topic: how to judge applicants for admissions purposes? Academic preparation isn't enough if we care about things like motivation and perseverance. I've thought about and written about different angles of this topic for a long time, and now find myself in a position to do something about it.

The conventional wisdom, if there is such a thing, is that grades and SAT are reasonable predictors of achievement. These are called cognitive measures for reasons best known to the psychometricians. As I recall high school, there was a lot more to grades than cognition, but never mind. SAT is a favorite target for those who don't like this narrow thinking, and I'd count myself with the critics. I took the ACT in 1980 in Illinois, and remember liking standardized tests then. I think I got a 28 on the ACT, whatever that means. I also took the ASVAB military battery (in the sense of tests, not artillery), and actually found the results last summer when I was going through stuff in the garage. I remembered taking the test, driving to a National Guard armory in East Saint Louis with my friend Mark, following the lines marked on the floor to the various stations. One was an attitudes and behavior survey, where they discovered that I'd never smoked pot. They made fun of me for that, or else it was some kind of awe. This was 1978, and the stuff was everywhere. One station was the ASVAB, complete with number 2 pencils.

In the Army's eyes, I was being rated for potential by this test, and they took it very seriously. Frankly, I loved taking standardized tests because I always did reasonably well on them, and hence had no stress about the results. What's interesting about the ASVAB is the one domain where I did not do well at all. You can see on the image that there's a 55% bar in the middle. When I pulled this out of the box in the garage last summer I had to squint to read the faded type of the explanation. If you look on the far left, there's a CL designator: that's for "clerical work." The battery decided I'm no good at it. If I recall correctly, this part of the test included such things as counting how many Cs were in a line of Os. Something like this:
OOOOOOOCOOOOOOOOOOCOOOOOOOOCOOOOOOOOOOOOOOOOOOOOOOOOO

That's the sort of thing that drives me crazy. I will tip my hat to the test designers: I am in fact not well suited for mindless clerical work. Did the rest of the battery rate me accurately? I guess that depends on what the expected outcomes were. I wasn't at that time very good leadership potential, was a bit lazy academically from never having had to work hard in school, and didn't have a lot of ambition to go change the world. Yes, I was only 16 or so, but looking at the test results would still have vastly overestimated what kind of officer I would have made at that point. In short, there were a lot of imporant things that the test didn't measure. This anecdote underlines my main thesis about standardized tests: they are good for simple tasks but not complex ones. And there are many complexities to what makes a succesful college graduate. The specifications are likely different from institution to institution as well.

All this by way of introduction to an article I found while researching the link between SAT and race. The article from InsideHigherEd by Scott Jaschik dates from September 2008, and opens with:
[C]ritics of standardized testing — and especially of the SAT — have said that these examinations fail to capture important qualities [...]
What's interesting is that the College Board--creator of the SAT--agrees, and has an active research program for developing new approaches. This is good news and bad.

First it's good, because the old SAT "cognitive test" will lose some of its halo and the market place for talent can become more efficient. This is good for all of us. It's also good because it will put standardized testing--the modern day equivalent of phrenology, in my opinion--under more of a microscope. If policymakers have more sophisticated ways to think about achievement, it's beneficial to everyone.

It's also bad in a very selfish sense. This is because the institutions that are already acting on this market inefficiency will see their lunch being shared around the table. The article mentions a couple of these. Tufts University and Oregon State University are using non-traditional approaches to admissions according to the article. Of course, I'm not really serious about this--I'm very happy that there's competition for the meme that SAT is a real achievement score, and I'm confident that a real research program can keep us on the cutting edge of finding the most suitable students for our institution. Ultimately it's good to have competition, and for reasons outlined below, I think each institution will have to find its own solution anyway.

So what are the new methods looking for besides "cognitive" processes? The Group for Research and Assessment of Student Potential (GRASP) has twelve so-called dimensions they consider:
  1. Knowledge, learning, mastery of general principles
  2. Continuous learning, intellectual interest and curiosity
  3. Artistic cultural appreciation and curiosity
  4. Multicultural tolerance and appreciation
  5. Leadership
  6. Interpersonal skills
  7. Social responsibility, citizenship and involvement
  8. Physical and psychological health
  9. Career orientation
  10. Adaptability and life skills
  11. Perseverance
  12. Ethics and integrity
You can read descriptions of these by following the link to their site. In addition to GRASP, SAT is working on modifications to their instrument, which may result in a whole new test. I'm a bit dubious about this project, however.

There are already ways to game the SAT. How much more will this be true when the 'right' answers are clearer? That is, it's much easier to appear to have attractive behaviors and attitudes than it is to actually possess them. I can imagine a new preparation industry springing up to coach test-takers who can afford it. Ultimately I think the industrialization of this metric is doomed for this reason. That leaves individual admissions policies to find ways to gather information in ways that are less likely to be faked. The article makes the same point.

The results of SAT's trials with their new test items are interesting. By de-emphasizing academics in favor of the College Board's experimental "biodata" and "situational judgment", traditionally under-served minorities showed significant increases in enrollment. Particularly notable is an almost 6% increase in black enrollments at the highly selective institutions that participated in the experiment.

There are problems, like gaming the system, but the payoff is worth it for individual institutions. The article quotes Pamela T. Horne, who has been working with GRASP results. She gives voice to my feelings on the topic:
“This is mission-driven,” she said, noting that colleges don’t define their missions as “enroll students with high SAT scores,” but they do prize leadership, artistic vision and various other qualities that might now be measured.
Unfortunately, if you asked many college presidents and board members, they probably would say that high SAT scores are an institutional priority. Why else would there be all the concern about which scores are reported (first, last, all, average?). Ultimately, more enlightened institutions can take advantage of the limitations of the SAT and other standardized predictors by individualizing their own processes. It's an exciting challenge.

Update: Reading the comments on the article, I found a reference to a textbook on the subject of non-cognitive assessment. Here's the link on Amazon.com. Beyond the Big Test: Noncognitive Assessment in Higher Education (Jossey Bass Higher and Adult Education Series)
by William E. Sedlacek (Feb 26, 2004)

Thursday, January 08, 2009

The Talent Bubble

I've argued before that the last decade has seen tuition increases and discount rate increases driven in part by a red-queen's race for the most talented students. The best applications are often seen to be those with high SAT, high high school GPA, and extras like co-curricular activities. Competition for these is fierce, and institutions with the highest endowments, or otherwise can offer the best aid packages, are in a commanding position. The Internet facilitates multiple applications, and the price (from a college's point of view) gets bid up as in an auction.

I was part of a conversation today with an enrollment professional who put the proportion of second-generation African-American applicants at 15%. For an HBCU, this means the pool of "good" applications (in the standard recruiting definition) is tiny. It becomes expensive to create attractive packages for these students. There will be pressure to sacrifice need-based aid in order to buy talent.

This is a lousy business model in the short term. With a decade-long perspective, it's attractive to have a growing pool of successful alumni, but you can bankrupt yourself in the process. It's a zero-sum game--there are only so many really good applications. But is that really true?

In my study of an admission matrix at one institution, the student enrolled at the bottom end (provisionally) succeeded about half the time. That is, there's a 50% chance that an applicant that looks lousy on paper is going to defy expectations on the upside. This isn't really surprising when you consider that predictions of first-year GPA based on grades and SAT aren't very good. This is especially true for low social-capital applications, such as first-generation students.

I've argued that we overprice high SATs at the cost of under-pricing some of the low SATs. If we could tell which low SAT students would succeed, this would be a gold mine for any institution. I actually wrote to ETS years ago and suggested that they develop and alternative instrument, but never heard back.

Here's how it would work. In addition to high school GPA (and SAT if you absolutely have to have it), find other indicators of success. Things like high school attendence records ought to be useful, but there are surely surveys that can be developed that would help. The CIRP is very helpful, for example, in post-facto analysis of attrition. A few attitude and behaviour question slipped into the application form might be enough to get started. The danger is that applicants figure out what combination of responses will help them, and 'game' the responses. There may be some way around that.

If you crack open that puzzle, you find yourself with the 85% of the African-American students (or 65% for caucasian) who are not being bid for aggressively. Of these, perhaps only 20% are of interest to you, but if you can zoom in on that 20% you've done yourself a real favor: found good students who don't cost as much as the high SAT crowd.

This is a project I'll be engaged in soon. I'll start with the app form and add some questions of a the type identified from an analysis of CIRP responses compared to college GPAs. Once we've identified a few indicators of success, we'll focus a few questions on those topics.

Thursday, November 27, 2008

SAT Redux

I'm teaching an Introduction to Statistics course this term, and we just did linear regression. I used the opportunity to dig out a study of freshman grades, and showed the the students how to make a predictive model using SAT and high school GPA. The data was from several hundred students, with ACT mapped to SAT scores according to standard formulas. High school grades alone explained 33% of the variance in first year college grades in this sample. Adding SAT bumped it up to 37%. This isn't much added information, given all the trouble and expense of the SAT. Maybe it's just our type of students, but for us, it's not worth the effort. Better would be an attitude and behavior survey like the CIRP.