Showing posts with label SACS. Show all posts
Showing posts with label SACS. Show all posts

Saturday, January 28, 2012

Assessing a QEP

On Wednesday, Guilford College hosted a NCICU meeting about SACSCOC accreditation. I had volunteered to do a very short introduction to my experience with the Quality Enhancement Plan (QEP) at Coker College, since I had seen the thing from inception to impact report. I got permission from Coker to release the report publicly, so here it is:


The whole fifth year report passed with no recommendations, and the letter said nice things about the QEP, so it's reasonable to assume that it's an acceptable exemplar to use in guiding your own report.

Assessment of the QEP program is an important part of the impact report, and this is a good place to record how that worked. The QEP at Coker was about improving writing effectiveness in students, and we tried several ways of assessing success. Only one of these really worked, so I will describe them in enough detail so you don't repeat my mistakes. Unless you just feel compelled to.

Portfolio Review.
I hand-built a web-based document repository (see "The Dropbox Idea" for details) to capture student writing. After enough samples were accumulated, I spent a whole day randomly sampling students in four categories: first year/fourth year vs day/evening. There were 30 of each, for 120 students. Then I sampled three writing samples from each to create a student portfolio. There was some back and forth because some students didn't have three samples at that point. I used a box cutter to redact student names, just like I imagine the CIA does. Each portfolio got an ID number that would allow me to look up who it was. The Composition coordinator created a rubric for rating the samples, and one Saturday we brought in faculty, administrators, adjuncts, and a high school English teacher to rate the portfolios. We spent a good part of the day applying the rubric to the papers, and many of the papers were rated three times. All were rated at least twice by different raters.

The results were disappointing. There were some faint indications of trends, but mostly it was noise, and not useful for steering a writing program. In retrospect, there were two conceptual problems. First, the papers we were looking at were not standardized. It's hard to compare a business plan to a short story. Second, the rubrics were not used in the assignments, but conjured later when we wanted to assess. It's essential for rubrics to be effective that they be as integrated as possible into the construction of the assignment.

So this was a lot of work for a dud of a report, most of which is probably my fault.

Pre-post Test
One of the administrators decided we should apply a writing placement test, which we already had data for, as a measure of writing gain by giving it again as a post-test after students took the ENG 101 class. The assignment was to find and correct errors in sample sentences. The English instructors told us it wouldn't work and it didn't. More noise.

Discipline-Specific Rubrics
We did, in fact, learn something from the rubric fiasco. We allowed programs to create their own rubrics, which could be applied to assignments in the repository. So an instructor could look at a work, pull up the custom rubric, and rate it right there and then. Since the prof knew the assignment, this seemed like a way to get more meaningful results. I think this would have worked, but by the time we got all the footwork done, the QEP program was a couple of years under way. I left the college before it was possible to do a large-scale analysis of the results that were in the database. In summary: good idea, executed too late.

Direct Observation by Faculty
Back in 2001, when I got the job of being SACSCOC liason, I got a copy of the brand new Principles and started reading. The more I read, the more I was terrified. And nothing frightened me more than CS 3.5.1, the standard on general education. I didn't know at the time that the standard said one thing, but everyone interpreted in a completely different way (it was written as a minimum standard requirement, but everyone looked for continuous improvement). So I was one of those people you see at the annual meeting who look like they are on potent narcotics, drifting around with a dazed look at the enormousness of the challenge. (Note: I think they should hand out mood rings at the annual meeting so you can see how stressed someone is before you talk to them.)

In an act of desperation, I led an effort to create what we now call the Faculty Assessment of Core Skills (FACS), which is nothing more than subjective faculty ratings of liberal arts skills demonstrated by students in their classes. The skills included writing effectiveness. At the end of the semester, each instructor was supposed to give a subjective rating to each student taught for observed skills on the list. You can read all about this in the Assessing the Elephant manuscript, or on this blog, or in one of the three books I wrote chapters for on the subject.

Because we had started the FACS before the QEP, we had baseline data, plus data for every semester during the project's life. Thousands and thousands of data points about student writing abilities. When we started the FACS I didn't have much hope for it--it was a "Hail Mary" pass at CS 3.5.1. But as it turns out, it was exactly what we needed. We were able to show that FACS scores improved faster for students who had used the writing lab than those students who didn't. Moreover, this effect was sensitive to the overall ability of the student, as judged by high school grades.  See "Assessing Writing" for the details.

I have given many talks about the FACS over the years, and get interesting reactions. One pair of psychologists seemed amazed that anything so blatantly subjective could be useful for anything at all, but they were very nice about it. When I post FACS results on the ASSESS-L list serve, you can hear the crickets chirping afterwards. I guess it doesn't seem dignified because it doesn't have a reductionist pedigree.

So I was shocked at the NCICU meeting, when SACSCOC Vice President Steve Sheeley said things like (my notes, probably not his exact words) "Professors' opinions as professionals are more important than standardized tests," and "Professors know what students are good at and what they are not good at."

The reason for my reaction is that when one hears official statements about assessment, it's almost always emphasized that it has to be suitably scientific. "Proven valid and reliable" is a standard formula, and certainly "measurable" figures (see "Measurement Smesurement" for my opinion on that). However it is stated, there isn't much room for something as touchy-feely as subjective opinions of course instructors. I do give good arguments for both validity and reliability in Assessing the Elephant, but FACS is never going to look like a psychometrician's version of assessment. So it was a shock and a very pleasant surprise to hear a note of common sense in the assessment symphony. I think when Steve made that remark, he assumed that this special knowledge professors acquire after working with students was simply inaccessible as assessment data. But it's not, and by now Coker has many thousands of data points over more than a decade to prove it. And it turned out to be the key to showing the QEP actually worked.

I have implemented the FACS at JCSU, and created a cool dashboard for it. I showed this off at the meeting, and you can download a sample of it here if you want. The real one is interactive so you can disaggregate the data down to the level you want to look at, even generating individual student reports for advisors. Setting up and running the FACS is trivial. It costs no money, takes no time, and you get rich data back that can be used for all kinds of things. Everyone should do this as a first, most basic, method of assessment.

Tuesday, May 03, 2011

SACS Changes to the Principles Proposed

A few weeks ago I posted about recommended changes to our SACS-COC accreditation standards. Today the review committee's recommendations for change were announced. You can find the document here. There is a comment period on this document open until May 18.

Friday, April 29, 2011

Small College Initiative

I attended the SACS Commission on Colleges Small College Initiative this week. You can find the slides on the SACS web site here. Below are some of my notes and observations.

Mike Johnson talked about CR 2.5 (Institutional Effectiveness) and pointed out the difference between assessment and evaluation. In my interpretation of his remarks, the former is gathering data and the latter is using it to draw conclusions for action. In my experience, this gap is where many IE cycles break down. Signs of this are "Actions for Improvement" that:
  • Are missing altogether
  • Are too general or vague to be put into practice
  • Report that everything is fine, and no improvements are necessary
  • Suggest only improvements to the assessments
I have a lot more to say about this (there's a surprise), and am preparing a talk for the Assessment Institute and a paper for NILOA on related topics.

Another important point is that IE cannot be outsourced or even in-sourced to a director. The whole point is that is is a collaborative exercise in striving to achieve goals. I think results are proportional to participation. In a similar vein, Mike noted that computer software can help organize reporting, but doesn't magically solve the problem of generating quality IE loops. Garbage in = garbage out.

A wonderful suggestion was to use the creation of "board books"as a way to encapsulate IE reports in a natural way that's already being done. Mike's larger point here is that we already have many real IE processes--all institutions that manage to survive use data one way or another--and there's no need to create an artificial one for reporting. I saw this during a review, where the institution had wonderful processes in place, but didn't include that documentation in the compliance certification, and instead reduced all that rich information into a four-column grid that "looks like it's supposed to." Of course, one problem here (in my opinion) is that there doesn't seem to be a standardized way to look at IE processes. If we were serious about it, we'd do inter-rater reliability studies and create tight rubrics with lots of examples in a library, showing what's acceptable and what's not. I think this would go a long way toward reducing the number of out-of-compliance findings. Way back when--over a decade ago--I heard a SACS VP complaining that even back then, IE had been around a long time and college should know what to do by now. That's true as far as it goes, but it should be acknowledged that: 1) it's very hard to satisfy committees, and 2) it's not entirely clear what is acceptable and what's not. Part of the problem is that while the theory of IE loops is easy to understand, practice is far more difficult. Sort of like Socialism.

There was a clarification that it is acceptable for institutions to sample programs for reporting 3.3.1.1 in lieu of reporting outcomes for every single one. There is supposed to be a policy statement about this on the web site, but I couldn't find it after several minutes of searching the list for 3.3.1, effectiveness, outcomes, sampling, etc. If someone finds it, please let me know. The main thing is that it should be representative and not look like it was cheery-picked (e.g. reporting only programs that have discipline-based accreditation). 

It was noted that CS 2.10 implicitly has learning outcomes reporting requirements, making it a pseudo-IE standard. I included this in my recommendations for 'fixes' to the Principles in my letter to SACS, posted here for comment. Not many institutions seem to be flunking it, though, unlike 3.3.1.1 (see below).

The fifth year report was highlighted in a break-out session. You can find additional slides on this topic on the website here.  Out of 39 institutions, 28 were cited on 3.3.1.1, and alarmingly, the number of citations on the QEP Impact Report is 33%. Although this says that the review process is no  piece of cake (which is good--it should be meaningful), it points to a problem. In fact, the rationale for the Small College Initiative is to help address this problem, which is particularly acute for small schools. As a side note, over lunch I talked to an IR director who speculated that there is a bias against citing large schools, particularly ones with high rankings. It would be really interesting, in conjunction with the inter-rater reliability study I fantasized about above, to have blind reviews of 3.3.1.1. Given the growing emphasis on student learning outcomes (including the new credit-hour rules), a whole separate system for learning outcomes may need to be developed. One of the challenges on the horizon, in my view, is the contradiction of grades. On the one hand they are the basic unit of merit for courses, with a vast bureaucracy behind them. By contrast, grades are not seen as 'real' assessments. This needs to be fixed. I don't know if Western Governors University's model is the answer, but what we have now makes no sense, and it is impossible to explain to the public.

Reasons given for flunking a QEP included:
  • Bad planning, which leads to a bad report. One kind of bad plan is one that's too broad. 
  • Failure to execute it, e.g. if a new administration comes in and lacks enthusiasm for the old project
  • Not talking about goals and outcomes in the report. Hard to believe.
  • Not describing the implementation (just narrating the creation, perhaps)
  • Not collecting or using data
  • Bad writing. Ironic, since so many QEPs are about writing.
 Tips for writing QEP impact reports:
  • Follow the directions given in the SACS policy
  • Address all the elements
  • Keep narrative to 10 pages. (You can apparently link out to other documents, which I hadn't heard before. I thought everything had to be in 10 pages.) [Edit: see the update below]
  • Use data, but include analysis--don't just put in graphs with no explanation.
Networking over lunch, I gleaned a couple of nifty ideas. At one institution, faculty contracts include a 'gotcha' clause, which stipulates that if assessment reports are not done by date X, then the prof has to stick around until date Y to finish them. This provides an incentive to get them done. Also, the reports are broken down into phases across the academic year, so that not everything is done at once. Smart.

Update: Mike Johnson posted a note to the list server saying that the links in the 10 page (max) Impact Report can only be internal to the report itself, which does not allow 'extra room'. In his words:
Links within a disk or flash drive are okay as long as the documents that are part of the link are included in the ten page maximum length. So please do not use hyperlinks to documents as a means to lengthen the report.

Tuesday, March 15, 2011

Improving the Principles of Accreditation

If you are in the SACS/CoC region, you know all about the Principles of Accreditation, the document that outlines accreditation standards, and which every institution must use to report compliance every ten years and (with fewer sections) every five years in between.  Until March 31, the Commission is accepting comments that will inform a review of said document. This is an excellent opportunity to make your voice heard in what is, after all, a peer-review process.

I have some observations I will post here for comment before sending them off to the Commission, in order to see if others agree or can suggest better approaches. I will restrict my comments to the institutional effectiveness (IE) sections. So here goes.

Note: CR = Core Requirement, CS = Comprehensive Standard, and FR = Federal Requirement

CR 2.10  The institution provides student support programs, services, and activities consistent with its mission that promote student learning and enhance the development of its students. (Student Support Services)
Although this isn't in the IE sections formally, it has a requirement that student support services "promote student learning and enhance the development of its students." This is a clear IE requirement, and is exceptional in that no other functional units, including academic programs, are required to pass this level of detailed IE review as a core requirement. Taken literally, an institution can be sanctioned severely for not assessing learning for student support services, which seems out of line with the more strategic level requirements that comprise the CR sections. I would suggest moving the IE language to 3.3.1.1, quoted below in the current version:
3.3.1 The institution identifies expected outcomes, assesses the extent to which it achieves these outcomes, and provides evidence of improvement based on analysis of the results in each of the following areas: (Institutional Effectiveness)
3.3.1.1 educational programs, to include student learning outcomes
3.3.1.2 administrative support services
3.3.1.3 educational support services
3.3.1.4 research within its educational mission, if appropriate
3.3.1.5 community/public service within its educational mission, if appropriate
Note that "student support services" doesn't appear. There are administrative and educational support services, but not student support services. The nomenclature needs to be cleaned up so we know exactly what we're talking about. One approach is to assume that any service that has learning outcomes is an educational support service, and therefore the modification could be as follows:
  1. End the statement of CR 2.10 after "mission," omitting the IE component.
  2. Change 3.3.1.3 to read "educational support programs, to include student learning outcomes."

CS 3.5.1  The institution identifies college-level general education competencies and the extent to which graduates have attained them. (College-level competencies)
This is the general education assessment requirement. Note that it doesn't appear in 3.3.1 unless general education is defined as a program by the institution. On the other hand, 3.5.1 is NOT an IE requirement--there is no statement about use of results to improve, just that you assess the extent to which students meet competencies.This is a knotty puzzle, so let me take it one part at a time.

First, the phrase "the extent to which" is ambiguous. Does it mean relative to an absolute standard or a relative one? This is by no means splitting hairs. If it means the former, then the institution MUST define what an acceptable competency is for each outcome, and presumably report out percentages that meet the standard. If it's a relative "extent to which" then simply reporting raw scores of a standardized test against national norms would work.
Example (absolute): Graduates will score 85% or more on the Comprehensive Brain Test. In 2010, 51% of graduates met this standard.

Example (relative): In 2010, graduates averaged 3.1 on the Comprehensive Brain Test, versus a national average of 2.9.
The standard is silent about the complexities of sampling graduates too. Are ratings from all graduates to be included? I would assume not, since this standard generally doesn't apply to IE processes because of the impracticably of it.

My sense of this is that we should allow institutions to define success in either absolute or relative terms, as best suits them, and include this standard with the other 3.3.1 sections so that it formally becomes part of IE. This will resolve the ambiguity about whether or not general education is program, and require that general education assessments actually be used for improvements. It also would broaden the scope to include students generally, not just graduates, who may have been out of general education courses for two years by then.

The modification could simply be to:
  1. Add: CS 3.3.1.6 general education, to include learning outcomes
  2. Delete: CS 3.5.1

Finally, let's look at 3.3.1.1 itself. The first issue is subtle. It concerns the meaning of the language "provides evidence of improvement based on analysis of the results". This can mean two different things, and I've seen it interpreted both ways, causing confusion.

The first interpretation is that it means that you have to demonstrate that improvement happened. This is a very high standard. It means things like benchmarking before changes are occurred, and then assessing the same way later on to see what impact occurred. When I hear speakers talk about QEP assessment, this is generally assumed, but it leaks over into the other IE areas too. Anytime you take a difference between two measurements it amplifies the relative error--this is a basic fact from numerical analysis. So you have to have very good assessments and they have to be objective (for reliability) and numerous (small standard error) and scalar (so you can subtract and still have meaning). Also, pre-post tests are the only method that can be used with any resemblance to scientific method. That is, you can't survey two different populations to compare unless you think you can explain all the variance between the populations (read Academically Adrift to see how problematic that is even for educational researchers). My objections to this are:
  1. No small program could ever meet this standard; the N will never be big enough.
  2. Many subjective assessments are very valuable, but are useless in this interpretation
  3. We don't yet have the technology to create scientific scalar indices of things like "complex reasoning" or other fuzzy goals, despite the testing companies' sales literature.[1]
  4. Random sampling is often impossible, which introduces biases that are probably not well understood
  5. Pre-post testing is quite limited, and severely restricts the kinds of useful assessments we might employ
The other interpretation is simply that we use analysis of assessment data to take actions that would reasonably be expected to improve things, but we don't have to prove it did. This is the standard I've seen most widely applied, except perhaps for QEP impact. Arguably the higher standard should apply to QEP, but  that's beyond the scope of my comments here.

In practical terms, the second interpretation is the most useful. It has to be borne in mind that assessment programs are mostly implemented by teaching faculty, who are probably not educational researchers by training, and in my experience tend to become frozen with a sort of helplessness if they think they are expected to track learning outcomes like a stock ticker tracks equity prices. I blogged about this a while back. Too much emphasis on proofs of improvement is paralyzing and counterproductive. On the other hand, free-ranging discussions that include the meaning of results, subjective impressions from course instructors, and other information that speaks to the learning outcome under consideration is a gold mine of opportunities to make changes for the better.

The most powerful argument against the strict (first) interpretation, however, is that there is simply no way to guarantee improvement on some index unless (1) one cheats somehow, manipulating the index, or (2) the index reflects some goal that is so obviously easy to improve that it's trivial. Either way the meaningfulness of the program vanishes, and we are left with many programs that are out of compliance (not showing improvement) or in compliance in name only (showing fake improvement or showing trivial improvement). 

My recommendation is to clarify the meaning of the language to read (bold emphasizes the change):
CS 3.3.1 The institution identifies expected outcomes, assesses the extent to which it achieves these outcomes, and takes actions based on analysis of the results in each of the following areas: (Institutional Effectiveness)
It should be obvious that the actions are intended to effect positive change, since this is rather the whole point of IE.

There is another issue with 3.3.1 that deserves attention. This concerns goals or outcomes that are not about student learning. It seems to me from the interpretations I've seen of 3.3.1 by reviewers and practitioners that the methods that apply to learning outcomes leak over to administrative areas, and that the expectations for compliance are tilted out of plumb. For example, there seems to be an expectation that every unit do surveys and have goals related to those, and that moreover this is necessary and sufficient.

Part of the problem may be the language "the extent to which," which almost insists on a scalar quantity.  But in fact, the most important goals for a given unit's effectiveness may have little to do with surveys and not be naturally a scalar.

One example I saw takes issue with a compliance report that presented "action steps" as goals. An action step might be "Approval of architectural drawings of the new library by 5/1/11." This sort of thing shows up all over Gantt charts:

(image courtesy of Wikipedia)

 This method of tracking complicated goals to completion is very common and very effective. The bars may or may not represent progress toward completion--they are simply timelines. So, for example, an OK stamp from the city engineer on your electrical plans is not a "percent to completion" item--it's either done or not. In other words it's Boolean.

Reviewers can be allergic to Boolean outcomes like:
  • Complete the library building on time and on budget.
  • Implement the new MS-Social Work program by Fall 2012
  • Gain approval via the substantive change process for a full online program in Dance by 2013 (good luck!)
  • Maintain a balanced budget every fiscal year.
For some reason, there seems to be a bias against this kind of goal, and I can't figure out why. These are obviously important operational items, key to the effectiveness of respective units. But a unit that includes these and doesn't have a satisfaction survey may be cited, whereas the reverse may be true too.

It may be that I'm making a mountain out of a termite hill here, but I think the language of "extent to which" could be changed to make it clearer that Boolean objectives are sometimes the natural way to express effectiveness goals.

For example, 3.3.1 could read (change bolded):
CS 3.3.1 The institution identifies expected outcomes, assesses success in a manner appropriate to each outcome, and takes actions based on analysis of the results in each of the following areas: (Institutional Effectiveness)
I know the non-parallel structure will offend the grammarians, but someone more expert can consider that conundrum.

One final nit-pick is that outcomes may not really be expected, but simply striven for--aspirational outcomes, in other words. If they are expected, then by nature they are less ambitious than they might otherwise be. I also put the odd dangling " in each of the following areas" at the front where it belongs, so my final version is this:
CS 3.3.1 In each of the following areas, the institution identifies aspirational outcomes, assesses success in a manner appropriate to each outcome, and takes actions based on analysis of the results: (Institutional Effectiveness)

[1] I've written much about assessing complex outcomes before, and this is not the page to rehash that issue. The short version is that learning happens in brains, and unless we understand how brains change when we learn, we are not able to speak about causes and effects as they relate to the physical world. See this article about London taxicab drivers to see a study that links learning to physiological changes to the brain. I don't mean to imply that assessing complex outcomes is useless, but just that we should be modest about our conclusions. It's called complex for a reason.

Tuesday, December 07, 2010

SACS gets a Back Channel

Last year's tweets from the SACS/COC December meeting were almost non-existent. That has been remedied this year thanks to contributions from several indefatigable tweeters. You can see them at http://twitter.com/search?q=%23sacs.


I have notes to write up and post here, but will have to wait until the weekend. The big news for me was that Coker's Fifth Year report sailed through with no recommendations. Congrats to Kaye and Pat and Daniel and everyone I don't know about who made that happen.

Here's a gem I found in the resource room, while picking over the 3.3.1 sections (paraphrased):
Objective: 60% of the students taking the test will be in the 100th percentile
I know what they meant, but I got a good laugh out of it. It puts Lake Wobegon to shame!

Friday, December 03, 2010

SACS 2010

I'm off to Louisville tomorrow morning for the December SACS meeting. Last year's backchannel was about zilch, and it looks like #SACS means something in another context, judging from the Twitter search for #SACS. I'll tweet to #SACS anyway, from my phone or iPad. Send me an email if you want to have a coffee and compare notes.

Wednesday, November 10, 2010

SACS: QEP Change

Last week at the NCICU meeting hosted at Guilford College, we were treated to a very nice lunch. (You can see where my priorities are.) We also heard presentations on finances, institutional effectiveness, QEP, and substantive change, as these pertain to our regional accreditation (SACS).Only those in the SACS region are likely to find the following significant.

I was particularly interested in the topic of creating and assessing a quality enhancement plan (QEP). Most of what I heard aligned with the notes I took from annual meeting, but there was one new twist related to us by SACS VP Mike Johnson, viz., a change in the wording in the Principles of Accreditation that reflects a broadening of what a QEP can be. Note that what follows is my unofficial interpretation, and I welcome corrections.

Here's the old version of the description of process, from the original Principles, pages 9-10.

The language to pay attention to is the part about "issues directly related to improving student learning." From the same document, the actual section 2.12 from the requirements:
Compare to the 2010 version of each. The first thing to notice is that the new version has more text.



The new language included in both references includes an option that wasn't there before. In addition to a focus on learning outcomes, it's possible to focus on the environment supporting student learning. My interpretation of the discussion at the meeting is that this also allows for some flexibility with regard to assessment, given that an environment is generally affective, not specific like, say a writing tutorial.

The specific example mentioned was that of a first-year experience, which may have many components working together to enhance student success. Because it's diffuse, the assessment of particular learning outcomes, with a "before and after" benchmark approach may not be reasonable. This seems like a good thing, because it will allow for more creative options with regard to both the construction and assessment of the QEP.

Sources: See the SACS website at www.sacscoc.org for full documentation. Also, there is an email list for SAC-related questions, which you can subscribe to here.

Update: In addition to the changes noted above, there is a new comprehensive standard under institutional effectiveness that applies to the QEP. My understanding is that this allows reviewers to find an institution out of compliance with a comprehensive standard rather than a core requirement. The effect of the former can be remediated through  the usual report process, but a core requirement failure would be much more severe. So this gives institutions the same chance to fix things that other comprehensive standard failures do, which wasn't the case before. Here's the text:

It's interesting what this focuses on. Capability to execute, broad involvement, and assessment are on the list, but nothing about an appropriate focus on learning outcomes. Apparently if you get that wrong, it's still a deadly sin.

Tuesday, October 12, 2010

Assessing Writing

Over the last week I've had the pleasure of revisiting the assessment plans we put in place at my prior institution, as it prepares to submit its SACS fifth-year report. I pitched in by doing some number crunching. The topic of the Quality Enhancement Plan (a SACS requirement for a program to improve teaching and learning) is writing effectiveness. This is a popular topic for QEPs, and I tried to make a list of such institutions a while back. A common problem is how to assess success.

In this case, the program spanned three initiatives with a range of assessment activities, including the NSSE, internal surveys, and qualitative assessments. For assessing writing, there are multiple types of assessments, but I'm just going to focus on the "big picture" assessment here: the Faculty Assessment of Core Skills (FACS) piece. I've written about the general method on this blog many times, and you can find an overview in the manuscript Assessing the Elephant, although the most recent results aren't in there yet.

The FACS surveys faculty opinions about individual students' writing abilities, provided that they have opportunity to observe such (not necessarily teach it or even count it for a grade, however). The scale for reporting is tied to the idealized college career (pre-college work, fresh/soph level work, jr/sr level work, work at the level we expect of our grads), and represented here on a 0-3 point scale. Getting the data is trivially easy and basically free. We started in fall 2003, and by now there are over 25,000 individual observations recorded on over 3,000 students (about a fourth of these on writing).


The graph above shows three cohorts, controlled for survivorship, each over four years. The error bars are two standard errors. One trend is that the first two years have plateaus, after which growth looks linear. In order to look at the quality of the data, I also graphed the average minimum and maximum ratings, combining the three cohorts.

This shows a consistent half-point average difference across eight semesters of attendance. That's not bad, and reliability statistics show that raters match exactly about half the time, far more than could be the case randomly. At my current institution, I've been getting even better numbers for some reason.

Although these graphs are nice, they don't actually show the effect of the QEP. That is, how do we know this growth wasn't happening anyway? This is the problem that will bedevil most QEP assessment efforts. In this case, one of the programs was to increase use and quality of the college's writing center.


This graph isn't mine; I took it with permission from the draft report.  It shows the dramatic growth of the writing center use. (Student body size is around 1100, for comparison). The use of the writing center also gives us a kind of control group for studying increase in writing skill. It's not perfect, because conventional wisdom is that the students who use the writing center tend to be those who are told they need to, meaning their skills are perceived to be lower than their peers in general. We can compare the users versus non-users using FACS:


This shows that indeed writing center users started with about equal or slightly less assessed skill, but overtook and exceeded their peers over four years. It gets even more interesting if we disaggregate by entering (high school) GPA.

This is the majority of students, and it shows that, in fact, for this "B" and better students, use of the writing center corresponds to their being seen as better writers within a year, and that this persists. On the other hand, for those less-prepared students (per HSGPA predictor), the story is different.


Here, according to FACS scores, the conventional wisdom is true: these students really do start off with lower perceived skill level, and it takes a year to reach near-parity with their peers. But by the fourth year, they have surpassed them. Note the numbers on the scale: even with the jump at the end, these students are rated far below their HSGPA>3 peers, writing center or not.

The slopes of the lines show something we've noticed before: a so-called Matthew Effect, whereby the most able students learn the fastest. Compare the blue lines (non-writing center users) in the two graphs above. The higher HSGPA students increased by .81, whereas the lower HSGPA group increased by only .32. Use of the writing center for this latter group more than doubled this increase, to .84.

I'm generally skeptical of assigning causes and effects without a lot more information, but these results are very suggestive, and certainly do nothing to contradict a conclusion that the writing center use is pushing the better students to higher performance, while enabling the less-prepared students to steadily and dramatically increase their skill.

Wednesday, January 13, 2010

Writing Projects

In the Southern Association (SACS) region, a Quality Enhancement Plan is now part of the decennial accreditation reaffirmation process. This is a project to improve student learning. At Coker we focused on writing, and I've stayed interested in the idea of how to better teach and assess writing. After bumping into several others at the annual SACS meeting with similar challenges in this area, I decided to try to make a list of writing QEPs. This is necessarily incomplete. If you have others I can add to the list, please email me.

The hyperlinks are to QEP documents where I could easily find them. I will update this list as I get more information.

Auburn University-Montgomery (WAC site)
Caldwell Community College & Technical Institute
Catawba Valley Community College
Central Carolina Community College
Clear Creek Baptist Bible College
Coker College
Columbus State University
Judson College
King College
Liberty University (pdf)
Lubbock Christian University (pdf)
South College
Texas A&M International University
The University of Mississippi
University of North Carolina Pembroke (pdf)
University of Southern Mississippi (pdf)
Virginia Military Institute (qep) (core curriculum)

One source: List of 2004 class QEPs from SACS (pdf)

My blog posts on writing assessment


Friday, January 08, 2010

Assessment Committees

In planning for a review of SACS 3.3.1 compliance university-wide this spring, we had to consider how to structure the process and committee structure. After discussion, it occurred to me that there are really two different things going on, and that they might be profitably separated. One is the ubiquitous Assessment Committee, which has been useful over the last year as a kind of R&D group--thinking up ways to more effectively assess learning outcomes. One subcommittee worked on technology (eportfolios, for example), and another on results and meaning (I called it the epistemology committee). Both of these are composed of mostly faculty.

But the process of review is something else. For one thing, assessment is only one component, and arguably not even (gasp) the most important. In order for the assessments to be useful, several things have to go right. Things like thinking about what assessments would be meaningful in advance of other planning, the organizational follow-through, and use of results while bearing the big-picture in mind. To me, this sounds like a job for department chairs. So, I will see if I can get a small group of chairs to for an Academic Effectiveness Committee to compliment one on the administrative side. The Assessment Committee can still do the R&D, but the actual review of program reports will be done by the new creature.

Speaking of which, there is an interesting discussion on the SACS-L listserv (see this post) about what the standard of success should be for 3.3.1. For the non-SACS folks, this is accreditation speak for the requirement to close the loop in effectiveness planning, expressed for learning outcomes thus:
3.3.1 The institution identifies expected outcomes, assesses the extent to
which it achieves these outcomes, and provides evidence of
improvement based on analysis of the results in each of the following
areas:

3.3.1.1 educational programs, to include student learning outcomes
etc.
At the heart of the discussion is the meaning of the language in 3.3.1, which hasn't changed much since 2006, when I wrote to the authors to ask that they resolve the ambiguity (posted here).

It occurred to me that there is an odd thing about 3.3.1, if interpreted most strictly. It essentially asks every institution and every program to conduct independent scholarly research on the learning of students. To do this right is obviously impractical, so it starts to resemble China's Great Leap Forward:
Mao encouraged the establishment of small backyard steel furnaces in every commune and in each urban neighborhood. Huge efforts on the part of peasants and other workers were made to produce steel out of scrap metal. To fuel the furnaces the local environment was denuded of trees and wood taken from the doors and furniture of peasants' houses. Pots, pans, and other metal artifacts were requisitioned to supply the "scrap" for the furnaces so that the wildly optimistic production targets could be met. Many of the male agricultural workers were diverted from the harvest to help the iron production as were the workers at many factories, schools and even hospitals. [wikipedia]
It would make more sense to give programs a choice. Either they could sign up to do real research on outcomes, OR simply adopt proven techniques that resulted from actual research at a well-funded institution. Why do we need to run thousands of ill-designed experiments on the same subject in parallel instead of a few good ones, and just use the results of those? Of course, this has a dark side, since the big psychometric industry would love to lock up this business. Still, if learning outcomes is our goal, why aren't we focused more on using techniques that are known to work instead of continually trying to discover them?

Tuesday, January 05, 2010

Happy New Year! Oh, and SACS

Dear blogospherians, cyberdenizens, and hypertextuals,

Here's hoping your 2010 is a happy and productive one! I've had a lovely and relaxing two week break, in which to read and write and do the inevitable projects around the house. At some point I came across an article about why we humans behave inconsistently. For example, we may have fabulous will power one day to stick to a diet, and then throw it out the window the next. The explanation in this psychology piece was that we have what you might call different personalities that inhabit our cranial chambers at different times. This sounds creepy and like something out of a horror film, but it was meant in a mild sense. Whatever the case, I decided to embrace that idea and just read sci-fi novels and not think about work for two weeks, and boy did it feel good. I discovered Richard K. Morgan, Charles Stross, and Jack McDevitt. Now I have a stack of unread books by these gentlemen that will take me months to read. At the bitter end I did some writing myself on a novel that progresses asymptotically (i.e. only to be finished if there is infinite time available).

SACS
If you are in the south, the Southern Association needs no introduction. There is now a list serv to share ideas about accreditation issues thanks to Patrick S. Williams, PhD, Associate VP for Institutional Effectiveness, University of Houston-Downtown. This is very welcome! You can sign up here (instructions copied from Pat's email):
To subscribe:
- Send an email to LISTSERV@LISTSERV.UHD.EDU (Caps not required; I'm
only using them for clarity.)
- You can leave the subject line blank.
- In the body of the email write SUBSCRIBE SACS-L YOURFIRSTNAME
YOURLASTNAME (Be sure to substitute YOUR first & last names.)
- It's best (but not essential) to delete your signature or anything
else that follows the SUBSCRIBE message.
Note that the listserv isn't actually moderated by SACS officials--they tend to hint and nudge rather than come out and say anything that's not official policy (and for good reason). But the crowdsourcing of experts from member institutions is a wonderful way to solve problems. Ultimately, it would be really nice to have something like StackOverflow for accreditation issues. It's a very smart interface for group problem solving. The overhead is considerably higher than that for a listserv, however, and it requires a large body of consistent users to be of much good.

In other news, there's an interesting comment exchange on my post "Assessment and Automation". Scroll to the bottom to see it.

Thursday, December 10, 2009

SACS 2009

The annual Commission on Colleges meeting was in Atlanta this year, and I drove down from Charlotte with a colleague under blue skies. Drove back in the rain. In between, the days and nights were packed. I had optimistically taken along one book on higher ed, one text on complexity theory, two sci-fi novels, and a briefing book of articles on college enrollment (is there a word for fear of being stuck somewhere without something to read?) I got through a total of about six pages, I think.

The sessions were mostly good, and the early (as in 7:30am) round tables were even better. It's an open secret that the round table discussions can be the best part of the conference. I tried to find sessions on the SACS five year report, which is a new requirement for reporting mid-cycle. Here are some of the more interesting bits from my notes:
  • You can change the QEP mid-stream. Radical change, like revising the goals isn't recommended, but it is accepted that institutions change, and plans don't always work as intended. Liberty University gave a presentation on this.
  • After the five year report, the QEP is no longer relevant to accreditation. Some projects may be over after five years. Others are property of the institution, to do as it pleases. Some will become institutionalized, others quietly dropped. Basically, after the report, it's time to start thinking about the next one.
  • For the pilot institutions--the first to come under the Principles of Accreditation--no one flunked the impact report on the QEP except for institutions who simply didn't execute it at all. On the other hand, several sections of the limited compliance certification that goes with the report were problematic, including documenting policy for handling student complaints and properly addressing distance learning programs.
  • Some institutions have prepared drafts of the impact report and are willing to share. For example, University of West Florida has a wealth of public documents about their QEP here.
  • On the SACS website, under Institutional Resources, there are the official documents about the report. The direct page is here, which includes report instructions and timeline.
  • SACS voted to change the rules to make it easier to pass the QEP (section 12 of the Principles). This is a technical change that allows recommendations to be made about the QEP proposal during the decennial reaffirmation process without triggering a full-fledged punitive sanction. This doesn't have anything to do with the five-year report, but signals that SACS is being reasonable about the requirements.
  • The five year report is encouraged to be electronic (CD or website, but don't do a website), but should be self-contained. We are not supposed to submit in both paper and electronically, unlike my experience with the compliance certification, where I learned at the last minute that they wanted both (in addition to renumbering all the sections). Any electronic report should be user friendly. I'd like to underline that. The typical higher ed admin is NOT tech-savvy. Use low-tech solutions. I still like paper, based on my experiences with review committees, and don't see any advantage to risking electronic submissions. We were advised that whatever we present should not depend on links to the main university site--the report has to be self-contained, even if electronic.
There's much more from the conference, which I'll get to later. Let me sign off with some proof of my assertion that SACS attendees aren't much into technology. I used one of the commons computers at the conference to browse to Twitter to see what the backchannel chatter was. Here's all that's there:
One lousy post. Compare that to #educause...