Showing posts with label FACS. Show all posts
Showing posts with label FACS. Show all posts

Saturday, January 28, 2012

Assessing a QEP

On Wednesday, Guilford College hosted a NCICU meeting about SACSCOC accreditation. I had volunteered to do a very short introduction to my experience with the Quality Enhancement Plan (QEP) at Coker College, since I had seen the thing from inception to impact report. I got permission from Coker to release the report publicly, so here it is:


The whole fifth year report passed with no recommendations, and the letter said nice things about the QEP, so it's reasonable to assume that it's an acceptable exemplar to use in guiding your own report.

Assessment of the QEP program is an important part of the impact report, and this is a good place to record how that worked. The QEP at Coker was about improving writing effectiveness in students, and we tried several ways of assessing success. Only one of these really worked, so I will describe them in enough detail so you don't repeat my mistakes. Unless you just feel compelled to.

Portfolio Review.
I hand-built a web-based document repository (see "The Dropbox Idea" for details) to capture student writing. After enough samples were accumulated, I spent a whole day randomly sampling students in four categories: first year/fourth year vs day/evening. There were 30 of each, for 120 students. Then I sampled three writing samples from each to create a student portfolio. There was some back and forth because some students didn't have three samples at that point. I used a box cutter to redact student names, just like I imagine the CIA does. Each portfolio got an ID number that would allow me to look up who it was. The Composition coordinator created a rubric for rating the samples, and one Saturday we brought in faculty, administrators, adjuncts, and a high school English teacher to rate the portfolios. We spent a good part of the day applying the rubric to the papers, and many of the papers were rated three times. All were rated at least twice by different raters.

The results were disappointing. There were some faint indications of trends, but mostly it was noise, and not useful for steering a writing program. In retrospect, there were two conceptual problems. First, the papers we were looking at were not standardized. It's hard to compare a business plan to a short story. Second, the rubrics were not used in the assignments, but conjured later when we wanted to assess. It's essential for rubrics to be effective that they be as integrated as possible into the construction of the assignment.

So this was a lot of work for a dud of a report, most of which is probably my fault.

Pre-post Test
One of the administrators decided we should apply a writing placement test, which we already had data for, as a measure of writing gain by giving it again as a post-test after students took the ENG 101 class. The assignment was to find and correct errors in sample sentences. The English instructors told us it wouldn't work and it didn't. More noise.

Discipline-Specific Rubrics
We did, in fact, learn something from the rubric fiasco. We allowed programs to create their own rubrics, which could be applied to assignments in the repository. So an instructor could look at a work, pull up the custom rubric, and rate it right there and then. Since the prof knew the assignment, this seemed like a way to get more meaningful results. I think this would have worked, but by the time we got all the footwork done, the QEP program was a couple of years under way. I left the college before it was possible to do a large-scale analysis of the results that were in the database. In summary: good idea, executed too late.

Direct Observation by Faculty
Back in 2001, when I got the job of being SACSCOC liason, I got a copy of the brand new Principles and started reading. The more I read, the more I was terrified. And nothing frightened me more than CS 3.5.1, the standard on general education. I didn't know at the time that the standard said one thing, but everyone interpreted in a completely different way (it was written as a minimum standard requirement, but everyone looked for continuous improvement). So I was one of those people you see at the annual meeting who look like they are on potent narcotics, drifting around with a dazed look at the enormousness of the challenge. (Note: I think they should hand out mood rings at the annual meeting so you can see how stressed someone is before you talk to them.)

In an act of desperation, I led an effort to create what we now call the Faculty Assessment of Core Skills (FACS), which is nothing more than subjective faculty ratings of liberal arts skills demonstrated by students in their classes. The skills included writing effectiveness. At the end of the semester, each instructor was supposed to give a subjective rating to each student taught for observed skills on the list. You can read all about this in the Assessing the Elephant manuscript, or on this blog, or in one of the three books I wrote chapters for on the subject.

Because we had started the FACS before the QEP, we had baseline data, plus data for every semester during the project's life. Thousands and thousands of data points about student writing abilities. When we started the FACS I didn't have much hope for it--it was a "Hail Mary" pass at CS 3.5.1. But as it turns out, it was exactly what we needed. We were able to show that FACS scores improved faster for students who had used the writing lab than those students who didn't. Moreover, this effect was sensitive to the overall ability of the student, as judged by high school grades.  See "Assessing Writing" for the details.

I have given many talks about the FACS over the years, and get interesting reactions. One pair of psychologists seemed amazed that anything so blatantly subjective could be useful for anything at all, but they were very nice about it. When I post FACS results on the ASSESS-L list serve, you can hear the crickets chirping afterwards. I guess it doesn't seem dignified because it doesn't have a reductionist pedigree.

So I was shocked at the NCICU meeting, when SACSCOC Vice President Steve Sheeley said things like (my notes, probably not his exact words) "Professors' opinions as professionals are more important than standardized tests," and "Professors know what students are good at and what they are not good at."

The reason for my reaction is that when one hears official statements about assessment, it's almost always emphasized that it has to be suitably scientific. "Proven valid and reliable" is a standard formula, and certainly "measurable" figures (see "Measurement Smesurement" for my opinion on that). However it is stated, there isn't much room for something as touchy-feely as subjective opinions of course instructors. I do give good arguments for both validity and reliability in Assessing the Elephant, but FACS is never going to look like a psychometrician's version of assessment. So it was a shock and a very pleasant surprise to hear a note of common sense in the assessment symphony. I think when Steve made that remark, he assumed that this special knowledge professors acquire after working with students was simply inaccessible as assessment data. But it's not, and by now Coker has many thousands of data points over more than a decade to prove it. And it turned out to be the key to showing the QEP actually worked.

I have implemented the FACS at JCSU, and created a cool dashboard for it. I showed this off at the meeting, and you can download a sample of it here if you want. The real one is interactive so you can disaggregate the data down to the level you want to look at, even generating individual student reports for advisors. Setting up and running the FACS is trivial. It costs no money, takes no time, and you get rich data back that can be used for all kinds of things. Everyone should do this as a first, most basic, method of assessment.

Thursday, April 28, 2011

Motivation and Intelligence

Angela Duckworth et al have a new article in the Proceedings of the National Academy of Sciences (PNAS) entitled "Role of test motivation in intelligence testing." The abstract reads, in part:
Intelligence tests are widely assumed to measure maximal intellectual performance, and predictive associations between intelligence quotient (IQ) scores and later-life outcomes are typically interpreted as unbiased estimates of the effect of intellectual ability on academic, professional, and social life outcomes. The current investigation critically examines these assumptions and finds evidence against both. [...] After adjusting for the influence of test motivation, however, the predictive validity of intelligence for life outcomes was significantly diminished, particularly for nonacademic outcomes.
The press on the paper includes Science Daily's "Motivation Plays a Critical Role in Determining IQ Test Scores," and Discover's blog post "IQ scores reflect motivation as well as 'intelligence'."

We include 'Effort' in our faculty-assessed end-of-semester FACS survey, and have found a link between grades and this effort rating. Of course, it could just be that professors who thing students work hard also tend to give them higher grades, so over the summer we will look at multi-year correlations to eliminate that confounding factor.

The graphs below (courtesy of Google charts) shows GPA in red, and hours completed in blue above the distribution bars for rated student effort across all classes. The heights of the bars give the percent of the distribution that received that rating. The left one is our first survey, Spring 2009. The one on the right is from Fall 2010. The sample size has increased as we've gotten better participation.

      

The drop in credits earned is due to more first year students being included in the sample. The year-by-year story is similar, except that the overall averages have an interesting shape as ecological samples from first year to fourth:


The first year students in the graph are the first class to fall under the new (much higher) admissions standards. The number is the average effort rating on a scale of zero (minimum effort) to three (great effort). This is for N=1403, Fall 2010. Note that there is a survivorship bias, so that we'd expect the averages to grow as the time-in-school increases. I don't yet have true longitudinal data.

Inter-rater reliability was measured by finding the frequency of exact matches for two instructors rating the same student. There were 385 instances of this, with a match rate of 50.7%. It's not hard to find the rate of pure-chance matches (dot-product the distribution with itself), but I haven't done that. In the past, the chance of matching randomly has been around 35%. See this source for more on that.

Thursday, January 20, 2011

Individual FACS reports

I have about two dozen web pages marked to write articles about, but haven't found the time. I'm trying to wrap up part three of my novel (see that blog), and still working on writing up my research. A couple of new things for me:

  • Our non-cognitive research is going well. Preliminary results based on one semester's grades indicate that a non-cog survey has the potential to add information to the usual enrollment inputs. We'll have much better data in the fall, with a whole year of grades and year-to-year retention data.
  • We built and launched an early alert system for tracking student who have academic difficulty early in the semester. It's in the testing phase right now.
  • We're developing a new web site, and one of the most important, and (to all appearances) ignored sections is the description of academic programs. We're spending a large effort there to develop superb pages that will sell programs to students. Stay tuned. Our Google Analytics shows that these are the most frequented pages by outside visitors. Not surprising, since academics is the product a university sells.
  • I've done a lot of development on reporting FACS scores. There is now a self-serve site to generate reports by term, program, or class. This week I added the ability to drill down to the individual student level, for use by advisors. Here's an example:

The red lines and bars show where this second year student is performing relative to other students, based on faculty assessments from the prior semester. At at glance you can see that this student is performing well below par, and according to professors is also not putting forth effort at the same level as peers. The sample sizes are necessarily small (although they will increase as the assessment becomes institutionalized). Note, however, that all three raters agreed that creative thinking was demonstrated at the pre-college "developmental" level.

Off to the Southern Education Foundation today to attend an assessment meeting in San Antonio.

Saturday, October 23, 2010

Do Students Add Value?

You hear "value-added" a lot these days; just google it. One article from RAND Corp tries to sum it up ("EvaluatingValue-Added Models for Teacher Accountability"), but in reading the paper I was drawn to one of the references from 2002, with the lengthy title "What Large-Scale, Survey Research Tells Us About Teacher Effects On Student Achievement: Insights from the Prospects Study of Elementary Schools" by Brian Rowan, Richard Correnti, and Robert J. Miller at the Consortium for Policy Research in Education at University of Pennsylvania's Graduate School of Education.

I don't propose to do a review of either of these papers here, but one quote struck me from the latter. It should first be noted that the authors strike a cautious note about the nature of such research in the introduction on page 2:
[O]ur position is that future efforts by survey researchers should: (a) clarify the basis for claims about “effect sizes”; (b) develop better measures of teachers’ knowledge, skill, and classroom activities; and (c) take care in making causal inferences from nonexperimental data. 
The point from the paper that struck me was this quote from page six:

Two important findings have emerged from these analyses. One is that only a small percentage of variance in rates of achievement growth lies among students. In cross-classified random effects models that include all of the control variables listed in endnote 4, for example, about 27- 28% of the reliable variance in reading growth lies among students (depending on the cohort), with about 13-19% of the reliable variance in mathematics growth lying among students. An important implication of these findings is that the “true score” differences among students in academic growth are quite small [...]
Let's think about that for a moment. The variation in "learning" (as numerically squashed into an average of a standardized test result, I think) is mostly not due to the variation among students. This struck me as absurd at first, but then I realized it's just a statement about how variable students are: to what degree the phenotypes sitting in our classrooms differ in their respective abilities to learn integral calculus (in my case).

As a reality-check, I pulled up my FACS database and classified students by their first semester college GPA and looked at their average trajectory in writing, as assessed by the faculty. Here it is.




The top line is 3.0+ students, then comes 2.0-2.99 students and so on over eight semesters, showing average writing scores. The improvement semester by semester does seem pretty constant, regardless of how "talented" the students are. Note that this particular graph isn't controlled for survivorship.

The numbers hide most of the real information, however. Is the improvement of the lowest group really comparable to that of the uppermost? There are different skills involved in teaching fast learners versus slow (which is considerably related to how hard the students work, a non-cognitive). If one substitutes standardized test scores for actual learning, this problem can only get worse. 

No, I still don't like averages. Here's the non-parametric chart for the 3.0+ group over eight semesters.


The blue portion of the bar is the proportion of these students who receive the highest rating, and so forth. The red ones are "remedial" ratings.

Update: the effect size of differences between student in my FACS scores is small, but it seems to be real. Especially if you look at particular treatments like that of the writing lab in an earlier article. As a first approximation, perhaps student abilities don't matter as to how much they learn, but I strongly suspect that conclusion is vulnerable on a number of fronts. See my more recent article on Edupunk and the Matthew Effect.

Tuesday, October 12, 2010

Assessing Writing

Over the last week I've had the pleasure of revisiting the assessment plans we put in place at my prior institution, as it prepares to submit its SACS fifth-year report. I pitched in by doing some number crunching. The topic of the Quality Enhancement Plan (a SACS requirement for a program to improve teaching and learning) is writing effectiveness. This is a popular topic for QEPs, and I tried to make a list of such institutions a while back. A common problem is how to assess success.

In this case, the program spanned three initiatives with a range of assessment activities, including the NSSE, internal surveys, and qualitative assessments. For assessing writing, there are multiple types of assessments, but I'm just going to focus on the "big picture" assessment here: the Faculty Assessment of Core Skills (FACS) piece. I've written about the general method on this blog many times, and you can find an overview in the manuscript Assessing the Elephant, although the most recent results aren't in there yet.

The FACS surveys faculty opinions about individual students' writing abilities, provided that they have opportunity to observe such (not necessarily teach it or even count it for a grade, however). The scale for reporting is tied to the idealized college career (pre-college work, fresh/soph level work, jr/sr level work, work at the level we expect of our grads), and represented here on a 0-3 point scale. Getting the data is trivially easy and basically free. We started in fall 2003, and by now there are over 25,000 individual observations recorded on over 3,000 students (about a fourth of these on writing).


The graph above shows three cohorts, controlled for survivorship, each over four years. The error bars are two standard errors. One trend is that the first two years have plateaus, after which growth looks linear. In order to look at the quality of the data, I also graphed the average minimum and maximum ratings, combining the three cohorts.

This shows a consistent half-point average difference across eight semesters of attendance. That's not bad, and reliability statistics show that raters match exactly about half the time, far more than could be the case randomly. At my current institution, I've been getting even better numbers for some reason.

Although these graphs are nice, they don't actually show the effect of the QEP. That is, how do we know this growth wasn't happening anyway? This is the problem that will bedevil most QEP assessment efforts. In this case, one of the programs was to increase use and quality of the college's writing center.


This graph isn't mine; I took it with permission from the draft report.  It shows the dramatic growth of the writing center use. (Student body size is around 1100, for comparison). The use of the writing center also gives us a kind of control group for studying increase in writing skill. It's not perfect, because conventional wisdom is that the students who use the writing center tend to be those who are told they need to, meaning their skills are perceived to be lower than their peers in general. We can compare the users versus non-users using FACS:


This shows that indeed writing center users started with about equal or slightly less assessed skill, but overtook and exceeded their peers over four years. It gets even more interesting if we disaggregate by entering (high school) GPA.

This is the majority of students, and it shows that, in fact, for this "B" and better students, use of the writing center corresponds to their being seen as better writers within a year, and that this persists. On the other hand, for those less-prepared students (per HSGPA predictor), the story is different.


Here, according to FACS scores, the conventional wisdom is true: these students really do start off with lower perceived skill level, and it takes a year to reach near-parity with their peers. But by the fourth year, they have surpassed them. Note the numbers on the scale: even with the jump at the end, these students are rated far below their HSGPA>3 peers, writing center or not.

The slopes of the lines show something we've noticed before: a so-called Matthew Effect, whereby the most able students learn the fastest. Compare the blue lines (non-writing center users) in the two graphs above. The higher HSGPA students increased by .81, whereas the lower HSGPA group increased by only .32. Use of the writing center for this latter group more than doubled this increase, to .84.

I'm generally skeptical of assigning causes and effects without a lot more information, but these results are very suggestive, and certainly do nothing to contradict a conclusion that the writing center use is pushing the better students to higher performance, while enabling the less-prepared students to steadily and dramatically increase their skill.