Showing posts with label attrition. Show all posts
Showing posts with label attrition. Show all posts

Sunday, October 27, 2019

Thinking in Predictors

There are some ideas that help us see the world more clearly. For example, the idea that a flipped coin has no memory of past events helps us realize that a recent string of heads does not increase the likelihood of tails (if the coin-flipping is done fairly). Thinking in predictors is one of those ideas that can make complex problems a little easier to think about.

This post is an introduction to how to visualize predictive power in the simplest case, where the outcome we want to know about only has two values. Examples include predicting first year student retention, student graduation, and applicants enrolling after being admitted. Having only two outcomes means there are only two possible predictions in a particular case--we predict that the outcome is one or the other of the two possibilities.

Let's suppose that you want to predict which of the new first-year students will return for their second fall. Suppose for a moment that I have a crystal ball and know the answer for each of your students. I randomly choose one student who will retain and one who will admit, and put their ID numbers under two cups on the table. Your job is to guess which one is the retained student.

With no additional information to go on, it's easy to see that your chances are 50% of choosing the correct cup.
Guess which is the cup with the retained student identifier
What if I tell you that the student associated with the right cup has a higher high school grade average than the one on the left? If your college is like mine, then predicting the right cup will have the retained student in it raises your chances of being correct.
Now you have some information to use as a predictor. 
We can quantify how much this new information matters in increasing the probability of correct prediction. For a recent retention analysis, I calculated this probability for a few dozen possible success indicators. Here's part of that list, sorted by the best predictors.

List of predictors of two year retention
The first column is the name of the potential predictor variable. The next two columns are the predictive power over men and women respectively, and the last column is how complete the data set is--the fraction of students for which we have that information. In this case, we're predicting two-year retention, and the first variable on the list (Attend) is the retention status itself.

Knowing the Attend variable would be having the answer before we have to choose the cup. Since the outcome variable can predict itself perfectly, the probability is 1 in the two measure (AUC) columns.

The second row (GPA) is the student's college grade average. If we know that information for each of the two students with IDs under the cups, and if we guess that the student with the higher GPA is the one who sticks around for two years, we'd be correct 68% of the time for males and 69% of the time for women students in the historical data. The data completeness rate is 99% because a few students leave before earning any grades. So as a measure of predictive power over two-year retention, grades have an AUC of about .68.

One of the more interesting predictors on the list is TFS_Transfer. This is a response to a question on the HERI freshman survey that asks about likelihood to transfer. This turns out to be a useful predictor for us, although we can only link IDs to 70% of the students.

Next time I'll explain why the columns with the probabilities are called AUCs and how you go about calculating them.






Saturday, November 22, 2008

The Value of Retrospection

When we first started getting serious about finding the causes of retention a few years ago, we decided to implement an "assessment day" survey to gather information about behaviors and attitudes of our students. This has taken place during the last four Octobers in a specially-added class date (one extra was added to the academic calendar). Departments are encouraged to use this time to assess their programs, but IR claims the most popular time slot to administer this shot-gun survey. We get almost half the student body at 9:30 T or Th, and it's a suitable mix of classes, so that works pretty well. The instrument itself is a 100-question scantron survey of questions that were solicited from various academic and administrative units. The form has been reviewed by the IRB, and asks students for their student ID number, which most give voluntarily.

The original idea was to have some retrospective data to look at for retention purposes. It works like this. In fall 2008 we had, of course, a group of students who had attended in fall 2007 but didn't graduate and didn't return: our attrition pool. Because the survey forms are tagged by student IDs, we can look back and see what indicators there might have been. This has proven to be very useful. We have since started using the CIRP for the same purpose--and it's really been a great source of information. It's essential to get as many student IDs as possible, however. Otherwise it's much less useful because you don't know who left and who stayed without the ID.

Here's one method of mining the data. Create your database of student IDs--I'll use just 1st year students from fall 2007 here--and use the trick I outlined here to get a 0/1 computed variable called 'retain' to denote attrit/retain. Add that as a column to the CIRP data or your custom survey by connecting student IDs. You can add other information too, like athlete/non-athlete, or zip code or whatever. Load this data set into SPSS and do an ANOVA, as shown below:
I've had issues loading directly from Excel, and usually end up saving a table as a .csv file--it seems to import better that way. You can only add 100 variables at a time, and the CIRP is longer than that, so it has to be done in chunks. Each will look something like this:
When the results roll in, look for small numbers in the significance column. I usually use .o2 as a benchmark. Anything less than that is potentially interesting. Of course, this depends on other factors, like sample size and such.

Now that's interesting--the ACCPT1st and CHOICE variables are very significant, meaning that they have power in distinguishing between those who returned and those students who didn't. Since I already had this data set in Access, I did a simple query and used the pivot table view to look at the CHOICE variable. For reference, the text of the survey item is:
Is this college your:
1=Less than third choice?
2=Third choice?

3=Second choice?

4=First choice?
Here are the results.

Students for whom the institution was their first choice were the first to leave. Not only that, but these are the majority. This turned out to be a critical piece of information. By performing another ANOVA with CHOICE as the key variable, and then using the 'Compare Means -> Means' SPSS report, we can identify particular traits of these 'First Choicers' as we have come to call them. We corraborate this with other information taken from the Assessment Day surveys, and a picture of these students emerges. I also geo-tagged their zip codes to see where they came from. More on First-Choicer characteristics will come in another post.

This was the beginning of the Plan 9 attrition effort, which is deep in the planning phase now. The bottom line is that we discovered that many of our students don't understand the product they're buying, and we don't understand them very well either. It's not the kind of thing one can slap a bandaid fix on, but will require a complete re-think of many institutional practices.

Friday, November 21, 2008

Targetting Aid

At the 2008 Assessment Institute I spent my time in sessions on first year seminars and retention strategies. I learned some interesting stuff. One was that at one institution where assignments were tracked, it wasn't the quality of student work that predicted retention, but the amount of it. I decided to test that at my institution by looking at our home-grown portfolio system statistics. I compared the average number of paper submissions by students who were retained to those who left for a single semester. There was no significant difference in our case.

Most strategies I saw targeted student engagement in one way or another. Activities like learning communities or work study increased the likelihood of student success. I asked questions about how retention committees worked with financial aid offices to fine tune awards. This was based on my own work here showing that grades and money are the two big predictors of attrition. No one I talked to had done such a thing, citing institutional barriers to efforts. Well, we've done it here, and had some limited success. Here's what we did.

The graph below shows the student population divided into total aid categories in increments of $3000. This is plotted against retention (dark line) for that group and (on the right scale) GPA for that group (magenta).

It's obvious that both grades and retention increase with financial aid, which says some interesting things about they way we recruit students, grant institutional aid, and provide academic support services. Ultimately this has led to a comprehensive retention plan I called Plan 9. But what we did immediately was focus on the group of students that have a decent GPA, but are historically showing low retention. That's the group circled on the graph. We targeted these students with small extra aid awards, and saw retention for that group sore to over 80%. I did a follow-up survey of the students receiving this aid to ask if it made a difference. My response rate was low, and the results lead me to believe there were other factors at work too. Maybe we just got lucky. But the results were so good we're trying it again this year.

Plan 9 includes lots of engagement stuff, and is really a comprehensive look at retention from marketing all the way through a graduate's career. A big part of it will focus on re-engineering aid policies. It always amazes me how budget discussions in the spring focus so much attention on tuition policies, when for private colleges at least, aid policies are much more important.