Showing posts with label writing. Show all posts
Showing posts with label writing. Show all posts

Tuesday, October 12, 2010

Assessing Writing

Over the last week I've had the pleasure of revisiting the assessment plans we put in place at my prior institution, as it prepares to submit its SACS fifth-year report. I pitched in by doing some number crunching. The topic of the Quality Enhancement Plan (a SACS requirement for a program to improve teaching and learning) is writing effectiveness. This is a popular topic for QEPs, and I tried to make a list of such institutions a while back. A common problem is how to assess success.

In this case, the program spanned three initiatives with a range of assessment activities, including the NSSE, internal surveys, and qualitative assessments. For assessing writing, there are multiple types of assessments, but I'm just going to focus on the "big picture" assessment here: the Faculty Assessment of Core Skills (FACS) piece. I've written about the general method on this blog many times, and you can find an overview in the manuscript Assessing the Elephant, although the most recent results aren't in there yet.

The FACS surveys faculty opinions about individual students' writing abilities, provided that they have opportunity to observe such (not necessarily teach it or even count it for a grade, however). The scale for reporting is tied to the idealized college career (pre-college work, fresh/soph level work, jr/sr level work, work at the level we expect of our grads), and represented here on a 0-3 point scale. Getting the data is trivially easy and basically free. We started in fall 2003, and by now there are over 25,000 individual observations recorded on over 3,000 students (about a fourth of these on writing).


The graph above shows three cohorts, controlled for survivorship, each over four years. The error bars are two standard errors. One trend is that the first two years have plateaus, after which growth looks linear. In order to look at the quality of the data, I also graphed the average minimum and maximum ratings, combining the three cohorts.

This shows a consistent half-point average difference across eight semesters of attendance. That's not bad, and reliability statistics show that raters match exactly about half the time, far more than could be the case randomly. At my current institution, I've been getting even better numbers for some reason.

Although these graphs are nice, they don't actually show the effect of the QEP. That is, how do we know this growth wasn't happening anyway? This is the problem that will bedevil most QEP assessment efforts. In this case, one of the programs was to increase use and quality of the college's writing center.


This graph isn't mine; I took it with permission from the draft report.  It shows the dramatic growth of the writing center use. (Student body size is around 1100, for comparison). The use of the writing center also gives us a kind of control group for studying increase in writing skill. It's not perfect, because conventional wisdom is that the students who use the writing center tend to be those who are told they need to, meaning their skills are perceived to be lower than their peers in general. We can compare the users versus non-users using FACS:


This shows that indeed writing center users started with about equal or slightly less assessed skill, but overtook and exceeded their peers over four years. It gets even more interesting if we disaggregate by entering (high school) GPA.

This is the majority of students, and it shows that, in fact, for this "B" and better students, use of the writing center corresponds to their being seen as better writers within a year, and that this persists. On the other hand, for those less-prepared students (per HSGPA predictor), the story is different.


Here, according to FACS scores, the conventional wisdom is true: these students really do start off with lower perceived skill level, and it takes a year to reach near-parity with their peers. But by the fourth year, they have surpassed them. Note the numbers on the scale: even with the jump at the end, these students are rated far below their HSGPA>3 peers, writing center or not.

The slopes of the lines show something we've noticed before: a so-called Matthew Effect, whereby the most able students learn the fastest. Compare the blue lines (non-writing center users) in the two graphs above. The higher HSGPA students increased by .81, whereas the lower HSGPA group increased by only .32. Use of the writing center for this latter group more than doubled this increase, to .84.

I'm generally skeptical of assigning causes and effects without a lot more information, but these results are very suggestive, and certainly do nothing to contradict a conclusion that the writing center use is pushing the better students to higher performance, while enabling the less-prepared students to steadily and dramatically increase their skill.

Tuesday, May 18, 2010

The Philosophical Importance of Beans

Pythagoras was against them: you simply couldn't be a Pythagorean and eat beans.  Nor should you pick up anything that has fallen.  Bertrand Russell lists fifteen prohibitions in A History of Western Philosophy.  The one that baffles me the most is the one about not eating from a whole loaf.  Are you supposed to wait around until someone nibbles from it first, and then have at it?  This catch-22 reminds me of teaching in Shanghai. I was teaching English conversation to young college students who'd been weened on the language.  They had questions I couldn't answer to their satisfaction.  What's the difference between "I went to the store," "I have gone to the store," and "I did go to the store"?  One day they had the left and right panels of the blackboard filled up with vocabulary, leaving me only the middle panel to write on.  They got terribly upset when they thought I'd erase it, so I left it the way it was.  But when I left I made a big box in the middle of the board that looked like this:

The students found this hilarious.  Normally, a class officer would make sure the board was cleaned, but when I came to class the next day the left and right panels were clear, but my bad luck box was still there.  I used it to talk about game theory and Pascal's wager, which they found interesting.  A weekend passed, and the following Monday the whole board was clean. The students had found a janitor who didn't know any English to erase it.  But I digress.

More recently, the subject of beans was a topic for a discussion on the assessment list server about the importance of agreeing on ratings (sometimes called centering or leveling) before assigning rubric-based scores. I was trying to find a topic where assessment was clearly subjective:
Suppose we ask 10 people to taste our green beans and ask them to rate as "good" or "not good."  Assuming we don't tip the scales somehow, we could legitimately report out to the public that of the 10 people who sampled the beans, 8 said they were good. 

But if we have the raters center first, the 2 holdouts are going to have to be convinced that the beans really are good, or else we ignore them.  Either way, we stamp GOOD on the can and give it a diploma, which may be misleading to 20% of the consumers.
A response from Hugh Stoddard:
Your green beans would be far more appealing to consumers if the raters discussed what constitutes a "good bean" and whether their judgement should be based on color, shape, size, sweetness, crunchiness, or nutritional value.
And a response to that from Joan Hawthorne:
Now we're getting into food where I feel a great deal of expertise after oh-so-many years of very dedicated eating.  I know in general that green beans are nutritious and I'm glad about that.  But hypothetical attributes of green bean goodness matter not one whit when it comes to my pleasure (or lack of pleasure) in eating a particular green bean dish.  Ditto for chocolate, wine, coffee, whatever.  If you're doing a shared rating of goodness for shared discussion, criteria may be helpful.  But I don't care if the top chocolate -- according to rater scoring based on criteria definition -- is this one or that one.  I care about which one appeals most to ME.
So who is right?  I originally assumed that green bean tasting would be so subjective that the idea of using rubrics was clearly nonsense.  But it took a while for me to really understand why.  Here's my answer in a nutshell: more doesn't equal better, and life isn't really a linear combination.

More is worse. A little salt is good.  More is bad.  A little crispness is good, more is too much.  Suppose Stanislav and Tatiana determine how salty they prefer their green beans.  The graph might look like this:

Here, the horizontal axis is increasing saltiness going to the right, and the vertical scale is pleasure in eating the beans, as subjectively reported (this is all made up data, in case you missed that point). 

Clearly, it's no good telling Stanislav that he's a wimp for not being able to tolerate more salt, or claiming that the "correct" value for optimum saltiness is somewhere between the two peaks.  Notice that there are two conditions here that we imagine away when we work with rubrics and ratings:
  1. There may be legitimate differences in opinions and tastes that can't be "centered" away
  2. more of an expressed trait isn't necessarily better
Wait--how can more of something be worse on a learning rubric scale?   Let's take something that ought to be straightforward: writing correctness.  Is it possible for writing to be too correct?  Here are some cases where it might be:
  • poetry, where one may bend the rules a little for style or form
  • advertising, where space is at a premium
  • dialogue or other less formal speech, where correctness can sound stilted.
  Suppose we "correct" Lewis Carrol's "The Jabberwocky," to take an extreme example. 
’Twas brillig, and the slithy toves
Did gyre and gimble in the wabe;
All mimsy were the borogoves,
And the mome raths outgrabe.
About the only correctness present here is that the words take familiar forms (nouns, verbs, adjectives) despite not making sense.  Is this poem fully correct with respect to use of language?  Clearly not.  Is that bad?  No.  It would lose its whole meaning if you corrected it to something like:
It was brilliant and the swaying boughs,
Did gyrate and gleam in the wind,
etc.
That made me ill.  We could get out of this by defining special kinds of correctness for each type of writing, but that still leaves the awkward question of how to assign a rating to something that is too much of something, AND it invalidates comparing one type of correctness to the other.

Tastes differ. Not everyone likes to read the same authors.  We have different tolerances for style and tone, and even correctness.  If I see a single "it's" where the author means "its," I usually stop reading; I take it as an indication of sloppiness.  In academic writing, style is tricky, and tastes will vary.  Too much informality, and you will not be taken seriously.  Too much formality, and few people will slog through what you've written.  But it's not necessary that we all agree on these things.  Indeed, it's probably impossible. 

I think Hugh Stoddard makes a good point when he suggests rating specific aspects of the green beans.  You could apply that to saltiness, crispness, and so on.  We still won't get agreement, but it would focus the conversation.  There is, however, a snag with this project too:  combinations aren't necessarily linear. 

We see this effect in analyzing student evaluations from classes, where students fill in bubbles on a Likert scale to answer such questions as "Were the goals of the course clearly stated?"  Students who rate the instructor low tend to also rate everything else low--the textbook, for example.  Similarly, someone who has a low tolerance for salt may be turned off by the mushiness of the beans simply due to too much salt.

In writing this is true also.  I worked with a colleague who would fail any paper with a single spelling error, regardless of any other merits the paper might have.  For him, correctness and style were tightly linked--errors ruin the style.  Or imagine a well-reasoned essay about the Battle of Britain, where the student misspells the names of all the important leaders in Britain, Germany, and the US.  Is that correctness or content or both?

These effects tend to complicate the clean-looking lines of rubrics, and the way people tend to think of them.  They are not precision instruments, but guides.  Their most valuable product is not numbers but ideas from focused conversations.  Rubrics are not measurement tools in any reasonable definition of measurement.

The point is not that rubrics are bad or that centering is bad. Rather, the point of this discussion is to consider the complexities of the situation when designing assessments and interpreting results.  If we "agree to agree" we lose information, and sometimes that's probably okay.  But other times it might be useful to try to understand what the subtleties are.  Imagine if we forced a diverse audience to watch the newest summer movie and then kept them in the room until they all rated it the same way.  If we understand the complexities of what it is we're trying to do, we will report more honestly (more modestly probably), and more usefully.  If you use a writing rubric across different disciplines, blend all the ratings into one average and then advertise that "we have measured the writing ability of our students to be X," it's misleading and not very useful.  Unfortunately, enough of that happens that administrators, accreditors, and the department of education seem to think that such sentences are meaningful. That's a bad thing. 

Aren't beans interesting?

Wednesday, January 13, 2010

Writing Projects

In the Southern Association (SACS) region, a Quality Enhancement Plan is now part of the decennial accreditation reaffirmation process. This is a project to improve student learning. At Coker we focused on writing, and I've stayed interested in the idea of how to better teach and assess writing. After bumping into several others at the annual SACS meeting with similar challenges in this area, I decided to try to make a list of writing QEPs. This is necessarily incomplete. If you have others I can add to the list, please email me.

The hyperlinks are to QEP documents where I could easily find them. I will update this list as I get more information.

Auburn University-Montgomery (WAC site)
Caldwell Community College & Technical Institute
Catawba Valley Community College
Central Carolina Community College
Clear Creek Baptist Bible College
Coker College
Columbus State University
Judson College
King College
Liberty University (pdf)
Lubbock Christian University (pdf)
South College
Texas A&M International University
The University of Mississippi
University of North Carolina Pembroke (pdf)
University of Southern Mississippi (pdf)
Virginia Military Institute (qep) (core curriculum)

One source: List of 2004 class QEPs from SACS (pdf)

My blog posts on writing assessment


Tuesday, March 17, 2009

Generalized Silliness

I often use the Indian story made famous by Godfrey Saxe about blind men from Indostan who encounter a elephant and try to describe it. The intent is to show how assessing anything complex can depend heavily on one's point of view. The moral of Saxe's poem is:
So oft in theologic wars,
The disputants, I ween,
Rail on in utter ignorance
Of what each other mean,
And prate about an Elephant
Not one of them has seen!
This was my first impression of the debate about student learning outcomes of a complex nature, like 'critical thinking.' A March 16, 2009 article in The Chronicle nicely illustrates a "elephant problem," that of assessing writing. The article's introduction is thus:
Writing instructors need to do more to prevent colleges from adopting tests that only narrowly measure students' writing abilities, several experts said last week at the annual meeting of the Conference on College Composition and Communication.
I've used such simple measures myself. An administrator was desperate for a quick solution to assessing writing quality, and had the English faculty come up with a simple exercise to test all of our students. Each was given a paragraph with errors to correct. The errors ranged from their/there sort of things to incorrect parallel structure. Several faculty members scored the results, including myself. One notable phenomenon was that students responded in ways that showed the were trying to answer the 'correct' way and not the natural way they spoke and wrote. A good example in the "real world" is The Clash song "Should I stay or Should I go?", which includes the line:
The underlining is courtesy of MS Word's grammar checker. There's actually a website praising the band for their correct use of grammar. I've scratched my head trying to figure out who's right in this case, but what's clear is no one would actually talk that way unless they were trying to be 'correct.' In Assessing the Elephant I write about what I call monological definitions--those imposed by an outside system and not subject to debate. Grammar rules are like that. So people act differently when they think they will be judged on the 'right' answer rather than the one they'd normally give. This is bound to bias the results on any simple test like this. This is just one tiny example of how an artificial test like this leaves much to be desired when assessing writing. Even simply measuring a student's ability to write correctly--the simplest possible assessment--is fraught with problems.

As an aside, I have to relate my favorite response on the exercise I described. The paragraph the students were asked to correct was on the topic of life at the south pole. One of the sentences was something like "The fish living under the water of the arctic are adapted to the near-freezing conditions." This student circled "under the water" and penciled in "where else would the fish live?" Indeed.

According to the quoted article, finding more sublte forms of assessing writing is a problem needing a solution:

Charles Bazerman, the organization's chair, said in an address to members that the move toward formal writing assessment is "inevitable," and that writing instructors need to furnish the evidence and the expertise to help improve how students are judged.

"It is up to us to provide more subtle forms of assessment," Mr. Bazerman said.

If "subtle" doesn't include authentic, I think this will be a very difficult proposition.

The article goes on to describe how writing assessment was implemented at one university, in an ad hoc manner (emphasis added):
Initial designs for the new assessment, put together by administrators at the last minute, were met with "confusion, ignorance, even outrage" among the writing faculty
This illustrates a central conceit about assessment in general, viz, that it's easy and can be done in straightforward and obvious ways. Heck, even an administrator without training or knowledge of the field can do it (my extrapolation). Contrast this from the faculty's stated opinion, taken from the article:
[G]rowth in writing is complex, slow, often nonlinear, intrinsically contextual, and not necessarily immediately visible.
Context is key to authenticity. Writing a lab report is different from writing a poem, which is different from writing a mathematics article. Experts in those areas can judge writing samples and rate them under the right conditions, but just because Stanislav can write a decent math paper does not mean that he can write a good review of a lit crit article. This is so obvious that it hardly bears mentioning, except that this kind of mistake gets made over and over.

I've used the following example before, but it serves well to illustrate the point. A supervisor is generally capable of judging whether employee Stanislav does a 'good job' at executing his responsibilities. Can we therefore make the leap to assume that there is a general test we can construct for someone doing a 'good job'? This test would have to work equally well for, say, a nuclear engineer, an airline pilot, a stock trader, and a science teacher. Obviously not. I had a math professor who called the subject of mathematical Topology "generalized silliness." That's a good description of what a general test of 'doing a good job' would be. Ditto for a general test of 'good writing' or 'critical thinking.'

The power of induction (generalizing from examples to create categories and rules) is powerful and easy to misuse. Just because there is 'writing' and there are 'tests' and there is 'exellence' does not mean that there are general tests of writing excellence. To assume that such a thing exists and proceed on that basis is, well, silliness of a general kind.