Saturday, May 24, 2025

Revisiting A Nation at Risk

In 1981, Reagan's Secretary of Education commissioned a report that became the 1983 A Nation at Risk: the Imperative for Educational Reform. It opens with a jump scare: Our nation is at risk. Political, but not polemical, the report bypasses academic framing and goes straight for national alarm to ensure that the reader comes away with the right conclusion. Nation's diagnosis was a "rising tide of mediocrity," and the resulting fervor led to reform efforts, both in K-12 and higher education. But the report avoids blaming schools directly. 

[Schools] are routinely called on to provide solutions to personal, social, and political problems that the home and other institutions either will not or cannot resolve. We must understand that these demands on our schools and colleges often exact an educational cost as well as a financial one.

We might call this problem "departmentalism," and it probably started in 2000BC Mesopotamia. Here's how it goes. (1) identify a problem and give it a name, like "emergency management," (2) create an organization with staff, budget, and paperwork assigned to handle the problem, and--here's the problem--(3) blame the department for any failures. I suppose this is natural: once you have a sign hanging on your door that reads "educator," any failure of education must originate there, right? Departmentalism ignores external causes of problems and continually adds new issues to the portfolio of the department as they arise.

But Nation dodges that trap by warning against departmentalism: it's not just the schools. A short time later the report dives into the national angst of the era: being out-competed by Asia and Europe. The old world had been rebuilt and America no longer had a lock on production. The "malaise" of the Carter years lingered on. Yet the appeal for action was deeply patriotic, channeling Thomas Jefferson.

A high level of shared education is essential to a free, democratic society and to the fostering of a common culture, especially in a country that prides itself on pluralism and individual freedom.
For our country to function, citizens must be able to reach some common understandings on complex issues, often on short notice and on the basis of conflicting or incomplete evidence.

The rhetoric would warm the heart of any president of a liberal arts college. The means is what we might call "guided bootstrapping."

All, regardless of race or class or economic status, are entitled to a fair chance and to the tools for developing their individual powers of mind and spirit to the utmost. This promise means that all children by virtue of their own efforts, competently guided, can hope to attain the mature and informed judgment needed to secure gainful employment and to manage their own lives, thereby serving not only their own interests but also the progress of society itself.

There are two strands of classical liberalism here (see Fukuyama's Liberalism and its Discontents). One is that a fair society gives everyone a chance. The other is more optimistic, that everyone can achieve the same level of competence. The political winds had shifted so that the first version is adopted throughout Nation. That doesn't seem to have been the case a few years earlier.

Prelude to Nation

The 1966 Coleman Report was a research project on education during the LBJ administration. It had reached the politically-incorrect conclusion that school conditions alone could not account for differences in student achievement; a gap at first grade expanded over time under the same conditions. 

[O]ne implication stands out above all: That schools bring little influence to bear on a child’s achievement that is independent of his background and general social context; and that this very lack of an independent effect means that the inequalities imposed on children by their home, neighborhood, and peer environment are carried along to become the inequalities with which they confront adult life at the end of school.
This Matthew Effect seems to be common in education, e.g. in college student writing. A recent example concerns the economic benefits of college degree: "going to college has become regressive, offering more to kids the richer they are," as reported here. This early stratification of futures seems vaguely unamerican and is counter to the utopian end of the classical liberalism spectrum, where social engineering levels the playing field.

One retrospective notes a secondary  effect of the Coleman report.

Before Coleman, a good school was defined by its “inputs”—per-pupil expenditure, school size, comprehensiveness of the curriculum, volumes per student in the library, science lab facilities, use of tracking, and similar indicators of the resources allocated for the students’ education. After Coleman, the measures of a good school shifted to its “outputs” or “outcomes”—the amount its students know, the gains in learning they experience each year, the years of further education graduates pursue, and their long-term employment and earnings opportunities.
It's worth reading that whole article if you have time; it provides nuances to the original report as well as an update on data relevant to the central claims. My reading of the history is that LBJ's agenda was school integration, and he was looking for support that amounts to departmentalism: integrating schools will largely close the racial achievement gap. 

I'd reframe the basic question that we still face in education as how do we use data analysis to understand the Matthew Effect and use government to maximize the potential of every student? If we only focus on schools, not--for example, poverty--there's only so much than can be done to close gaps. In fact, the later we start in a young person's development, the more our effective efforts might widen achievement gaps.  To see this imagine an oversimpified model where each student i has achievement X_i, and new policies produce a maximum achievement five years later of Y_i = L*X_i. The rich still get richer, and at a faster rater because of the boost. Everyone ends up better than they would have before the change, but the Matthew Effect is a result. I'm not saying this is inevitable, or that proportional growth is always the right model, but it's plausible.

Academic Expectations

The next section of Nation gives statistics of decline in learning. This is known to be untrustworthy, for example failing to account for selection effects. A wider swath of high-schoolers took the SAT in 1980 than in 1963, so the test-taking population was less elite. It's not surprising that the average scores declined. Cathy O'Neil has a chapter in Weapons of Math Destruction on this topic. A little later the report's authors illustrate this flaw in their logic without irony.

[T]he average graduate of our schools and colleges today is not as well-educated as the average graduate of 25 or 35 years ago, when a much smaller proportion of our population completed high school and college.
After the statistics there's a quaint-sounding appeal to the importance of the humanities, warning against an "overemphasis on technical and occupational skills," appealing to the liberal arts. Then the authors describe what they heard from interviews, including this. 

What lies behind this emerging national sense of frustration can be described as both a dimming of personal expectations and the fear of losing a shared vision for America.
The lowering of expectations is easy to understand--it's straight out of Ibn Khaldun, where the fall of a civilization stems from the laxity that civilization provides. The "fear of losing a shared vision for America," on the other hand, gave me a flash of nostalgia. That idea seemed more wholesome in the 1980s, or maybe that's just the glow of my college years. However, unlike the current politicization of education--as as shared vision problem--the emphasis in Nation was on rigor and the economic consequences of dumbing-down. 

[T]he knowledge base continues its rapid expansion, the number of traditional jobs shrinks, and new jobs demand greater sophistication and preparation.

That sentiment would have be recognizable for decades to come, through Obama's college for all push, which arguably lasted until 2024. This argument has flipped; see the Lightcast report on the future of labor, for example, and the flawed but fascinating The Case Against Education that argues college degrees are nothing more than market signals.  

Nation was arguing for quantity and quality both, but there's evidence that academic excellence has not been maintained in the intervening years. There's a case to be made that college is getting easier, flooding the labor market with a surfeit of bachelor's degrees, many of whom end up underemployed. As the signaling power of the degree erodes, dissatisfaction increases. If so, this goes back to Coleman's Matthew Effect (he didn't call it that) and the idea that we can have either high standards or a lot of graduates with lower average abilities. If we want a lot of graduates, then we have to reduce the bar so they can graduate (~40% still don't graduate, so the bar could be lower yet). The limitation is that many students in first grade are already set on a path that's unlikely to lead to completion of a rigorous college degree. 

Matching Abilities to Outcomes

A fair and efficient education system lies somewhere between the extreme "every student tracked from birth into fixed opportunities" to "no filtering or tracking at all." The latter describes what we have in the US, a caveat emptor philosophy that fails those on the lowest rungs of the economic ladder. Those least prepared for college are also less savvy at picking a college. By contrast, Germany's education system nudges students into career tracks. Given where American higher ed ended up today, it seems likely that incentives have been misaligned with outcomes. In particular, I suspect that colleges and universities got addicted to enlarged student populations and lowered filters to accommodate. See the brand new book on this topic Capitalizing on College. From the blurb:

[The book] reveals how three of the strategies these schools adopted--growing a traditional endowment, pioneering a periphery market, or even creating a network of multiple markets--were initially successful but ultimately fell short in raising enough revenue to support operating a residential campus. Only a fourth accelerated strategy of going to scale raised the necessary funds--but at the cost of undercutting their mission by leading them to view students as dollars.

While Nation has a political message, it's not after simple solutions, seeking to "avoid the unproductive tendency of some to search for scapegoats among the victims, such as the beleaguered teachers." Here again we see the wise resistance to what I called Departmentalism. The scapegoating would unfortunately return with a vengeance under G. W. Bush and No Child Left Behind. The prescription for educational problems involved students as much as institutions.

At the level of the individual learner, [excellence] means performing on the boundary of individual ability in ways that test and push back personal limits, in school and in the workplace.
This is nuanced in that it tacitly admits the problem Coleman identified and leans into the idealism of the country's boot-strapping myth and Reagan's cowboy image. It's calling for a new generation of strivers and dreamers.

Excellence characterizes a school or college that sets high expectations and goals for all learners, then tries in every way possible to help students reach them. Excellence characterizes a society that has adopted these policies, for it will then be prepared through the education and skill of its people to respond to the challenges of a rapidly changing world.

The prescription to schools is "challenge and support," which is a reasonable middle path between the student tracking poles. It means we are flexible with college entry requirements, for example. To err on the side of over-admitting applicants, and then make up for any deficiencies that leak through with extra support. At least, that's a practical way to read it from a school's perspective. Indeed, if we look at the rhetoric produced by colleges and those who lend them credibility, this is apparently the model we have now. The problem with the last forty years of this challenge-and-support idea is that the support part was "solved" with Departmentalism (i.e. formally, but not actually), and the challenge was not maintained; it went the other way to accommodate unprepared students. The federal engineering to make colleges competitive by awarding aid to students directly means that students are often seen as dollars, and admissions requirements are set by financial goals. 

Risk and Reward

Nation is pointing to an ethical trade-off between university finances and the risks borne by students. This is more of a problem now than it was in the 1980s because of the cost of college. To implement the challenge-and-support philosophy effectively, many colleges would have to make fundamental changes. We'd have to admit that while remediation may be part of the solution, equally important is matching learning to motivation and ability. Colleges shouldn't admit students without a clear idea of what success for that student looks like. For many colleges currently, this would be financial suicide. 

In Nation we find a commitment to value diversity and maintain equity and excellence at the same time. It translates to a goal that implicitly acknowledges the Matthew Effect, but seeks to optimize outcomes within that constraint.

Our goal must be to develop the talents of all to their fullest. Attaining that goal requires that we expect and assist all students to work to the limits of their capabilities. 

In other words, it doesn't claim that all students have the same capabilities, but that we should help each reach his or her potential. To accomplish this without placing the whole burden on teachers would require social engineering. LBJ had his Great Society. Reagan's proposed version was the Learning Society. 

At the heart of such a  society is the commitment to a set of values and to a system of education that affords all members the opportunity to stretch their minds to full capacity, from early childhood through adulthood, learning more as the world itself changes. Such a society has as a basic foundation the idea that education is important not only because of what it contributes to one's career goals but also because of the value it adds to the general quality of one's life.
The authors could hardly have imagined the power of the internet, online classes, and AI for life-long learning. The means to implement the vision described in that passage have become abundantly available. The cultural change--that Learning Society--is not obviously in evidence nowadays.  It wasn't in evidence in 1983 either. The report is short on diagnosis, but does suggest that too many students do the minimum necessary, and "In some colleges maintaining enrollments is of greater day-to-day concern than maintaining rigorous academic standards," which is prescient in pointing out the ethical dilemma I mentioned above.

Call to Action

Nation's call to action is to everyone.

Thus, we issue this call to all who care about America and its future: to parents and students; to teachers, administrators, and school board members; to colleges and industry; to union members and military leaders; to governors and State legislators; to the President; to members of Congress and other public officials; to members of learned and scientific societies; to the print and electronic media; to concerned citizens everywhere. America is at risk.
Despite this generality, much of the detail about student learning naturally centers on schools, because that's where the data comes from. There are a couple of pages of findings about a curriculum that's too loose, declining rigor, class time, and study time. The findings for classroom teaching are specific and actionable.

The Commission found that not enough of the academically able students are being attracted to teaching; that teacher preparation programs need substantial improvement; that the professional working life of teachers is on the whole unacceptable; and that a serious shortage of teachers exists in key fields.
The report mentions low salaries as a contributing factor. And the diversity of the student body is associated with tracking.

We must emphasize that the variety of student aspirations, abilities, and preparation requires that appropriate content be available to satisfy diverse needs. Attention must be directed to both the nature of the content available and to the needs of particular learners. The most gifted students, for example, may need a curriculum enriched and accelerated beyond even the needs of other students of high ability. Similarly, educationally disadvantaged students may require special curriculum materials, smaller classes, or individual tutoring to help them master the material presented. 
This is followed by a general statement about high expectations appropriate to each group. In light of the Coleman Report and more general instances of the Matthew Effect, this isn't a strategy to minimize learning gaps between groups, but to increase the learning of each group to its maximum potential. Policy recommendations support this in the pages following.

We recommend that schools, colleges, and universities adopt more rigorous and measurable standards, and higher expectations, for academic performance and student conduct, and that 4-year colleges and universities raise their requirements for admission. This will help students do their best educationally with challenging materials in an environment that supports learning and authentic accomplishment.

As I observed above, the success of this program would tend to enlarge achievement gaps because of the Matthew Effect. To close gaps, changes would have to start before students are born.

The report's general statement about academic excellence is followed by specific recommendations. The first two are striking in the context of higher ed's current state.

1. Grades should be indicators of  academic achievement so they can be relied on as evidence of a student's readiness for further study. 
2. Four-year colleges and universities should raise their admissions requirements and advise all potential applicants of the standards for admission in terms of specific courses required, performance in these areas, and levels of achievement on standardized achievement tests in each of the five Basics [English, math, science, social studies, and computer science] and, where applicable, foreign languages.
The first of these should be read in light of the recent research of Jeff Denning and colleagues, who found that recent grade inflation is linked to lower academic standards, so that colleges were graduating more students, but who had learned less on average. Incredibly, for decades accrediting agencies have actively discouraged the use of grade data to inform learning improvement efforts. This was a huge lost opportunity.

Other recommendations support academic rigor through standardized tests, improved textbooks (including addressing learner diversity), and the development of pedagogy and teaching materials using good research.  At the societal level, the recommendations support higher pay and status for teachers. Much of this is tasked to state and local governments, with some federal support.

[W]e believe the Federal Government's role includes several functions of national consequence that States and localities alone are unlikely to be able to meet: protecting constitutional and civil rights for students and school personnel; collecting data, statistics, and information about education generally; supporting curriculum improvement and research on teaching, learning, and the management of schools; supporting teacher training in areas of critical shortage or key national needs, and providing student financial assistance and research and graduate training. 
As noted earlier, the authors avoided the Departmentalism trap, and acknowledge that although government can make important changes, that's not the whole picture. The report then addresses students and parents directly. 

You have the right to demand for your children the best our schools and colleges can provide. Your vigilance and your refusal to be satisfied with less than the best are the imperative first step. But your right to a proper education for your children carries a double responsibility. As surely as you are your child's first and most influential teacher, your child's ideas about education and its significance begin with you. You must be a living example of what you expect your children to honor and to emulate.

What's obviously missing in the prescription is the role of the government in leveling the field. The Zeitgeist was personal responsibility and small government and reaction against The Great Society. The war on poverty had become a Sitzkrieg. Ultimately, Departmentalism and oversimplification crept back in. Technology took culture into a world out of science fiction. Ibn Khaldun could hardly have imagined the means of instant gratification that education must compete with. The Matthew Effect in educational and economic potential has not been tamed. The Jeffersonian sine qua non of an educated citizenry has new relevance.

 

Concluding Thoughts

I thought that 2023's 40-year anniversary of Nation would have led to retrospection, but the only example I found was from the Hoover Institution, an edited volume that gives a brief history of the intervening decades and provides various perspectives. The rest of what I found are op-eds or relatively shallow pieces. 
 
My interest in Nation was sparked by the follow-up report on higher ed called Involvement in Learning, which made recommendations for administrations, faculty, and accreditors. I wrote about the impact on the assessment movement in the Update. My initial impressions of Nation were also colored by Cathy O'Neil's book, which described some of the shoddy statistics cited as evidence in the report. Despite this, it seems like the authors got some important things right. They avoided scapegoating, acknowledging that educational quality is the consequence of the society's priorities as a whole, and they acknowledged the diversity of students in productive ways. Some of the recommendations seem prescient now, like the need for higher teacher pay and the emphasis on excellence. They navigated the contradictions of classical liberalism with practical strategies rather than scapegoating or utopian ideals. 

But, as the saying goes, culture eats strategy for breakfast. Nation wasn't the first call for educational excellence. A 1958 report The Pursuit of Excellence sounded many of the same themes. A commentary in The New York Times from 1983 quoted this passage.
Teachers tend to be handled as interchangeable units in an educational assembly line. The best teacher and the poorest in a school may teach the same grade and subject, use the same textbook, handle the same number of students, get paid the same salaries and rise in salary at the same speed to the same ceiling.
It was this thread--treating education as an assembly line--that has had the most staying power. Its attraction is magnified when combined with the emphasis of measuring outcomes (sparked by Coleman). Rather than customizing education to the potential(s) of the student, a minimum competency mindset seems to have prevailed. It probably isn't coincidental that this is the cheapest way to deliver education. Consequently, the problems identified in Nation (and preceding research) still exist, and gaps still exist. 
 
This review has three important lessons for any college's IR office. First, track and report academic rigor using grades. Second, don't just focus on average outcomes. Identify and understand the Matthew Effects as they exist at your institution. Finally, take a hard look at the risk/reward for incoming students. How far is financial pressure leading the institution to load up risk on students?

Friday, August 16, 2024

A Canticle for Bloom*

Introduction

Stephen Jay Gould promoted the idea of non-overlaping magisteria, or ways of knowing the world that can be separated into mutually exclusive domains, where each "holds the appropriate tools for meaningful discourse and resolution." The tension Gould was trying to resolve was between religion and science: 

Science tries to document the factual character of the natural world, and to develop theories that coordinate and explain these facts. Religion, on the other hand, operates in the equally important, but utterly different, realm of human purposes, meanings, and values—subjects that the factual domain of science might illuminate, but can never resolve. -- Stephen Jay Gould from Rock of Ages

I'll call these "ways of knowing" or WOKs, which seems more down to earth than "magisteria." Each WOK contains cross-checks on knowledge that are particular to the domain.  Scientific questions are judged by scientific standards. Personal choices are based on experience and usually don't have "correct" answers but degrees of validation.

Should you give a friend a loan? That's a lived experience question. Are your car's spark plugs failing to ignite? That's a science question. Should you take a recommended drug despite severe side effects? That's somewhere in between. 

There's a third WOK we need to talk about: ideology. Steven Mintz provided a crisp definition recently in his blog on InsideHigherEd

Ideologies simplify, clean up and package reality into something easily consumable, palatable and appealing to a mass audience. In doing so, ideologues discard the messy, complex and often unpleasant aspects of reality, presenting only what fits neatly within their framework. Ideology, thus, distorts reality by filtering out anything that does not conform to its narrative.

Ideologies are particularly powerful when associated with a utopia. For example, the way that Marxism evolved into an intellectual justification for Stalin's USSR. Lysenkoism enforced ideology on biological research.  

Here's a diagram of our three WOKs with what we might find in the overlaps.

Galileo was interested in the physical reality of the cosmos (among other things), at the center of the diagram, creating new methods for WOK 2. But the correct way of speaking about the universe needed to adhere to church doctrine (ideology, WOK 3). This is the "correctness" overlap, only some of which corresponded to reality (geocentricism did not). See Steven Shapin's The Scientific Revolution for a nuanced narrative; it's not as simple as the usual telling of Galileo vs the Church. Jennifer Michael Hecht would rightfully insist on putting "ritual" in the overlap between ideology and lived experience, and argue that the ritual can fulfill a social and personal need. Ideology isn't bad; we need it. It can just overlap in odd ways with other WOCs.

SLO Assessment 

Yesterday I participated in a conversation with a small group of experienced assessment directors as part of an ongoing project to fix accreditation standards. The discussion echoed a theme I heard last year when I interviewed a dozen peer reviewers from various accreditors, that there's value in the formal kind of assessment that gathers data and does statistics on it, but it's more common to see success by just getting faculty together to talk about learning goals. We also talked about accreditation standards. I suggest that there are three important WOKs in assessment:
  1. Professional judgment and collaboration
  2. Educational measurement and inferential statistics
  3. Adjudication of accreditation policy
If we find that first-generation students have abysmal pass rates in math (WOK 2), this will affect conversations about pedagogy, support services, course prerequisites, and so forth that would happen in a department meeting (WOK 1). Conversely, if the math faculty all agree that Calculus 1 isn't adequately preparing students for Calculus 2 based on classroom experiences (WOK 1), it might prompt a more formal analysis of grades and test scores (WOK 2).

Perhaps in a perfect world, assessment offices would operate within these two WOKs, with a large faculty support role to facilitate conversations and share knowledge (WOK 1), with a separate function to gather and warehouse data, do research, and connect with the wider research community to bring ideas back (WOK 2). It's not clear that universities would fund such an outfit, however. Assessment offices are expensive and only exist because of accreditation requirements, to get the reports done (WOK 3).

 

The Third Circle

Any bureaucracy draws from a kind of ideology, at least implicitly. Paperwork and procedure serve to "simplify, clean up and package reality into something easily consumable," in Mintz's formulation. Behind the paperwork is a purpose: the DMV's goal is safer roads. The EPA's is  a clean environment. The validation of WOK 3 uses a formalized classification of the world (e.g. driver's test) to assign cases to policy distinctions (driver's license granted or denied).
 
Accreditation requirements provide the motivation to run assessment operations, but they also impose a particular ideology. I described its origins and effects in "Assessment standards are broken." In short, the ideology can be abbreviated as "define-measure-improve" and is a version of "scientific management," descendants of Taylor and Drucker and others, of which Six Sigma is a variation. There is an intended overlap with WOK 2, almost subsuming the scientific WOK within the ideology, with the goal "we're going to require you to use science to improve the state of education."
 
Robert Birnbaum catalogs variations of this idea, calling the phenomenon 
 
a paradox of complexity and simplicity. Its central ideas may appear brilliantly original. Yet at the same time they are so commonsensical as to make us wonder why we had not thought of them ourselves, and so obviously reasonable as to defy disagreement. 
-- Management Fads in Higher Education, pg 5.
 
The result is that the internal validation of knowledge in WOK 3, which is done by trained peer review teams, assumes the preeminence of WOK 2 (scientific knowledge). 
 
Any new idea has to compete with existing ones, and faculty tradition was the natural enemy of the define-measure-improve protocol, despite the obvious overlap with what teachers do on the job, and what they want to accomplish. In the Sturm und Drang of 1983's A Nation at Risk, educators took a lot of heat (a theme in US politics). In higher education, teaching work had to be mapped from existing practices, seen as inefficient, to ones that aligned with the scientific management principles.
  • Defining learning

    • Old: choose textbooks, write syllabus, approve curriculum, create tests or other assessments

    • New: write statements of student learning objectives, often in a hierarchy of course, program, institution

  • Measurement

    • Old: grade tests, writing samples, performances, etc. Build a shared sense of acceptability via faculty consensus and constant exposure to students, assign summative course grades

    • New: use only a few approved methods, including specific assignments, papers associated with rubrics (no grades!). Learning is seen as distinct from more general student success. Emphasis on outputs instead of inputs.

  • Improvement

    • Old: Professional growth in teaching practice, department or institution level consensus on change (WOK 1), using data summaries like grade or test averages or pass rates to identify needs for improvement (WOK 2).

    • New: Averages or frequencies of approved data sources to find deficiencies, then imagine a way to remedy them
The requirements to write reports using the new methods lobotomized WOK 1 for faculty. To comply with the reporting requirements they had to start over using approved replacement methods in order to be allowed to "know" how their students were doing. Naturally they resented this. Hated it, even, and have produced a genre of articles complaining about assessment. We still deal with the effects.
 
It's worth repeating that what assessment directors say works best is WOK 1--what the scientific approach was intended to replace--and because that's what actually works, the accreditation reviews have relaxed over time to allow more room for WOK 1. Your situation depends on your accreditor, but it's still an awkward fit because of the need to perform the other rituals (defining and gathering data in the approved way) in order to validate WOK 1, which doesn't really need that extra work to function.
 
As I described in "Assessing for Student Success" the science project falls apart immediately, because what students learn in a college curriculum is a lot of detailed topics with interconnections. I estimated several hundred topics (SLOs if you like) in a math curriculum. These can't be described, let alone measured, within the parallel framework the accreditors created. 
 
It's worth noting that in the industrial setting, where these ideas were formed, it is possible to define and measure everything important, and have real-time data from instruments on an assembly line. That doesn't translate well to education, where there's little standardization.

The accreditation requirements strengthen the main thing that works (WOK 1) by creating more opportunities for faculty to talk about student learning, especially when facilitated by a good assessment director. But the requirements also diminish the effectiveness of faculty work by heaping on artificial requirements in the name of science. This poor attempt to mandate WOK 2 fails the validity checks within WOK 2: sample sizes are too small, too noisy, and the causal models used are too simple. 

This collision between science and policy gets resolved ad baculum: you will be beaten with the stick of non-compliance until you at least pretend to believe that rubrics are always valid measures. In short, the SLO accreditation requirements co-opt the authority of science, but replace scientific standards with ideologically-correct ones. This has created a self-sustaining culture of compliance, abetted by consultants, vendors, and peer review training that maintains this closed garden of bureaucracy as science.

 

A Litmus Test

A few years ago, I concluded that the best way to illustrate how accreditation standards drove us into an epistemological ditch was to spotlight their allergy to course grades. In the age of big data, the idea that we'd arbitrarily throw out millions of data points covering the whole history of students at our institution is absurd. It only makes sense within the accreditation bubble, and so it puts a spotlight on the difference between the scientific claims of accreditors versus the reality. To heighten the contradictions, so to speak. You can find my summary of research on grades and learning here. Starting on page 23 you'll find a list of standard objections to using course grades as data about student learning. 
 
I won't rehash that material here. It suffices to quote an article that cites Bloom (the taxonomy guy) from 1976, eight years before define-measure-improve kicked off:
 
Perhaps the most productive use of GPA is as a covariate. GPA has the potential to explain nearly half the variance in education research models (Bloom, 1976), thus shedding light on the variance explained by other variables of interest, such as changes in course or curriculum design.  
 
-- Bacon, D. R., & Bean, B. (2006). GPA in Research Studies: An Invaluable but Neglected Opportunity. Journal of Marketing Education, 28(1), 35–42.
 
 
When I have the opportunity, I ask accreditors what they think of course grades as data. This is a litmus test for self-reflection. So far they've all failed it. The wise ones don't want to come out too strongly against grades, because it implies they don't think transcripts are meaningful. They generally hedge, calling grades "indirect measures." But there's no test for directness in the vocabulary of WOK 2. You won't find a way to create p-values on directness in the manual on educational measurement, because the idea isn't statistical. We already have a robust vocabulary on reliability and validity, and there's no need to confuse the issue with "directness." 

When pushed on that point, one accreditor representative cited an anecdote about how a student didn't feel like the grade reflected learning. This would be ironic if the scientific claims of the define-measure-improve protocol were serious about the science. After dismissing faculty consensus as opinion, reciting "the plural of anecdote isn't data," and distributing buttons at conferences that read "show me the data," a high priest of the order can simply use an anecdote to dismiss the whole research record on course grades. The remarkable thing is that this doesn't seem to cause any cognitive dissonance. 

The point of this illustration is that the overlapping circles in the three WOKs don't presently stand the light of public exposure. The accreditors will look ridiculous. We need to fix it before that happens, to create a more sensible overlap of the WOKs.

There's More

This article is long enough, and I'm going to stop here. I have not discussed the effective use of WOK 2 (educational measurement and inferential statistics) as a successful assessment tool. That may be addressed in a future post. Suffice to say that that requirements of a good research project (e.g. large data sample, tests of reliability) aren't feasible in 99% of department-level accreditation reports. From a practical point of view, it's a lot of extra work to do a real research project to get a tiny amount of credit, if any at all (none at all if you research retention instead of learning). The irony is that WOK 3 assumes that it contains WOK 2, when in fact the intersection is nearly empty.
 

Tuesday, June 11, 2024

Why the Student/Faculty Ratio is a Bad Metric

The student/faculty ratio, which represents on average how many students there are for each faculty member, is a common metric of educational quality. The ratio shows up in the Common Data Set (CDS) and college guides, presumably so prospective students can compare colleges. 

The standard way to count students and faculty for the CDS calculation is to equate the total number of students into a smaller number of artificial standardized units. That's because some students may take a single class, while others take ten or more an academic year, and faculty teaching loads similarly vary. This is usually done by converting each population into Full Time Equivalent (FTE) units. The idea is that a part time student isn't the same as a full time student, but we might count, say, three part time students as equal to one full time student. This is the CDS approach, which we can write as FTE = FT + PT/3. 

The part time conversion formula is ad hoc, chosen for convenience instead of meaningfulness. It considers all part timers (students or faculty) as averaging a third of a full load, when that will vary across institutions and across time. Additionally, "full time" is defined by policy, and this also varies by institution. A full time faculty load at a research institution is probably less than the load at a teaching college. The full-time definition for students is usually a range of numerical values, e.g. a full load is somewhere between 12 and 20 credits. There's a big difference between 12 and 14 or 16 or 18 credits, when averaged over the whole student population, because it directly impacts how many course sections need to be taught. So the student FTE works okay as a measure of revenue (paying tuition for full load), but not as a demand measure (how many classes do we need to teach). 

Similarly, faculty members who get release time from teaching or conversely teach overloads may be "full time" for contractual purposes, but not reflect their classroom presence. There are complicated adjustments in the CDS definition, which refers to the AAUP definition of full time faculty. For example, a faculty member on leave for research (e.g. sabbatical) still counts even though they are not in the classroom, but if they have a replacement hired, that replacement should not be counted. A certain amount of judgment is required to decide if a hire is a replacement or not, resulting in "house rules" for counting faculty. 

These effects combine to erode the meaningfulness of an FTE-based student/faculty ratio. But we might take a step back and ask what the ratio is intending to do anyway.

Deriving the Raw Ratio

If we focus on the student classroom experience, we might think of a student in a class as the basic unit for counting. How many of these are there? If there are \(N_s\) students, and on average each student has a class load of \(L_s\) each academic year, then there are a total of \(N_s L_s\) units of "student-classes." That's how many of these experiential units were consumed.
 
From the faculty perspective--the production side--if there are \(N_f\) faculty members, with an average teaching load of \(L_f\) classes, and average class size of \(A\), then the total student-classes is those three numbers multiplied together. Since student-classes taught (produced) must equal student-classes taken (consumed), these can be set against each other as
 
$$ N_s L_s = N_f L_f A $$ 

Using these raw counts (not FTEs) of students and faculty, the ratio is then 

 $$ \frac{N_s}{ N_f} = A \frac{L_f }{ L_s} $$ 

The average loads for faculty and students are largely determined by policy. For example, if six classes per year is the contractual load for a full-time faculty member, and ten courses per year is necessary to complete a bachelor's degree in four years, then the load fraction is 6/10. Because of part-timers and overloads, the measured load averages won't be exact, but they should be close to that and relatively stable over time for an institution--as long as policies don't change. 
 
The raw student-faculty ratio is then measuring average class size times an index of institutional policy, the load ratio \(L_f / L_s\), including the  prevalence of part-timers and overloads. This load ratio won't be the same between institutions, so including it as a factor is not appropriate if we want a comparable index. If we drop the load ratio from the right side, we have average class size, which is a comparable index of student experience. It's crude--a distribution would be better--but as a single metric it's not terrible. 
 

The FTE Ratio

 
With that understanding, we can now see what the FTE ratio is all about. Suppose we create a "true" FTE calculation for faculty by dividing the total number of classes taught by the policy's specification for a full load. So if 3000 courses are taught in an academic year, and the faculty handbook says the load for a full-time faculty member is six courses, then the FTE faculty is 3000/6 =  500. We can do a similar calculation with students, e.g. using twelve credits as a full time student load.

In the terms defined earlier, the number of classes taught is \(N_f L_f\), which we need to divide by the faculty policy-defined load \(P_f\) to get \( \text{FTE}_f = N_f L_f / P_f \). Similarly for students, \( \text{FTE}_s= N_s L_s/ P_s \). In each case, the \( L/P\) fraction is expressing the measured average load as compared to the policy load. If there are a lot of adjuncts teaching, the average load per faculty member might be 5.1, whereas the full-time load is 6. Putting this together we have
 
$$ \frac{\text{FTE}_s}{\text{FTE}_f} = \frac{N_s L_s}{N_f L_f} \cdot \frac{P_f}{P_s} $$ 

The derivation in the previous section shows that 
 
 $$ A = \frac{N_s L_s}{N_f L_f} $$ 
 
so the FTE ratio is: 

$$ \frac{\text{FTE}_s}{\text{FTE}_f} = A \frac{P_f}{P_s} $$ 

The difference between the raw ratio and the FTE ratio is that the former uses empirical average loads for students and faculty, whereas the FTE version uses policy-defined loads. In both cases, these vary by institution.

In the actual CDS calculations, the house rules and approximations, like FT = PT/3, will add error, so you probably won't get exactly the formula above.

Discussion

The student/faculty ratio conflates two types of educational quality. Average class size crudely evaluates the quality of in-class instruction, as a measure of accessibility to faculty while teaching. The load ratio (empirical or policy-based) assesses how much time faculty have to devote to each class, as well as how much students must spread themselves around to cover the required load. These are related to classroom experience, but are different dimensions. I suggest we adopt a rule for measures of "one dimension per dimension," in which case we could describe classroom quality as (1) average class size, (2) average teaching load, and (3) average student load. Attempting to combine all that information into one metric just creates a mess. 

I've never seen the above derivation before, and until I did it I didn't know what was going into the ratio calculation. I suspect that most producers and consumers of the metric don't really understand it, and probably misuse it. For many purpose, the average class size is a convenient summary measure of student experience that also represents operational efficiency: financially it represents both revenue and cost in the same scale. It can be aggregated or disaggregated to whatever level of analysis you care do to do. And it's easily understood and communicated.

Bottom line: use average class size instead of student/faculty ratio.

Thursday, October 26, 2023

Assessment Institute 2023: Grades and Learning

I'm scheduled to give a talk on grade statistics on Monday 10/26, reviewing the work in the lead article of JAIE's edition on grades. The editors were supportive in not just accepting my dive into the statistics of course grades and their relationship to other measures of learning, but they recruited others in the assessment community to respond. I got to have the final word in a synthesis. 

That effort turned into a panel discussion at the Institute, which will follow my research presentation. The forum is hosted by AAC&U and JAIE, comprising many of the authors from the special edition who commented. One reason for this post is to provide a stable hyperlink to the presentation slides, which are heavily annotated. The file is saved on researchgate.net, and to cite the slides directly use:

If you want to cite the slides in APA format, you can use:

Eubanks, D. (2023, October 26). Grades and Learning. ResearchGate. https://www.researchgate.net/publication/374977797_Grades_and_Learning

The panel discussion will also have a few slides, but you'll have to get those from the conference website. 

Ancient Wisdom

As I've tracked down references in preparation for the conference, I've gone further back in time and would like to call out two books by Benjamin Bloom, whose name you'll recognize from taxonomy fame. 

The first of these was published more than sixty years ago.

Bloom, B. S., & Peters, F. R. (1961). The use of academic prediction scales for counseling and selecting college entrants. Crowell-Collier.
 
It starts off with a bang. 
 
The main thesis of this report is that there are three sources of variation in academic grades. One is the errors in human judgment of teachers about the quality of a student's academic achievement. Testers have over-emphasized this source of variation and have tended to view grades with great suspicion. Our work demonstrates that this source of variation is not as great as has been generally thought and that grade averages may have a reliability as high as +.85, which is not very different from the reliability figures for some of the best aptitude and achievement tests.
 
The treatment then goes on to describe the typical correlations between high school and college grades of around .6, which is nearly exactly what I see at my institution (it's slowly declining over time). The methods include grade transformations to improve predictiveness. That's still a useful topic, whether predicting college success or building regression models with learning data. 

The second book is:

Bloom, B. S. (1976). Human characteristics and school learning. McGraw-Hill.
 
 From page 1:

This is a book about a theory of school learning which attempts to explain individual differences in school learning as well as determine the ways in which such differences may be altered in the interest of the student, the school, and ultimately, the society.
 
The ideas here are more sophisticated than what we find in the advice books on assessment practice in that the focus is "individual differences" not just group averages. This turns out to be the key in understanding the data we generated and is reviewed in the Assessment Institute presentation. I wish I had read this book before having to reinvent the same ideas.That's partly my own limitation, because I don't have a background in education research. And I increasingly realize that assessment practice--at least all those program reports we write for accreditors--doesn't resemble research at all. Why is that?

Assessment Philosophies

 
There's an upcoming edition of Assessment Update (December, 2023) that will feature several articles critiquing the report-writing focus that concerns so much of assessment practice. In it I'll describe my understanding of the history of that practice and why it fails to deliver on its promise. A big part of that is because of the disconnection between report-writing and the theory and practice of educational research. Most recently we can see that in the non-impact that the Denning et. al work has made on assessment practice or accreditation reporting. That work should have caused a lot of soul-searching, but it's crickets. The conference program book doesn't mention Denning, and although there are several mentions of graduation rates, nothing directly links grades, graduation, and learning (other than my own presentation). When I've had occasion to talk to accreditation VPs, they seem to have not heard of it either. Nor was there a discussion on ASSESS-L. 

I recently came across Richard Feynman's quote that "I would rather have questions that can't be answered than answers that can't be questioned." Here are the two ways of looking at assessment that correspond to those cases.

1. Product Control by Management

I think the reason for the gap between research on learning and assessment practice is due to the founding metaphor of the define-measure-improve idea that drives assessment reports. As I'll mention in the Update article, that formula for continual improvement is a top-down management idea (e.g. related to Six-Sigma) that prioritizes control over understanding the world as it is. The impetus came from a federal recommendation in 1984 and became accreditation standards that causes universities to set up assessment offices that tell faculty how to do the work. There's a transparent flow from authority to action. Although educational research appears in the formula, it's not the priority and in practice it's impossible to do actual research on every program and every learning outcome at a university.

So assessment has become a top-down phenomenon that emphasizes product control (learning outcomes) with emphasis on control. This led to proscriptions on grades and other "answers that can't be questioned."

2. Educational Research

Managing hundreds of SLO reports for accreditors is soul-destroying, and the work that many of us prefer to do is (1) engage with faculty members on a more human level and lean on their professional judgment in combination with ideas from published research, or (2) do original research on student learning and other outcomes.

The research strand has a rich history and a lot of progress to show for it. The early days assessment movement, before being co-opted by accreditation, led to Scholarship of Teaching and Learning and a number of pedagogical advances that are still in use. 

Research often starts with preconceived ideas, theories, and models of reality that frame what we expect to find. But crucially, all of those ideas have to eventually bend to whatever we discover through empirical investigation. There are many questions that can't be answered at all, per Feynman. It's messy and difficult and time-consuming to build theories from data, test them, and slowly accumulate knowledge. It's bottom-up discovery, not top-down control.

What's new in the last few years is the rapid development of data science: the ease with which we can quickly analyze large data sets with sophisticated methods. 

The Gap

There's a culture gap between top-down quality control and bottom-up research. I think that explains why assessment practice seems so divorced from applicable literature and methods. I also don't think there's much of a future in SLO report-writing, since it's obvious by now that it's not controlling quality (see the slides for references besides Denning). 

This is a long way to go to explain why all those cautions about using grades have evolved. Grades are mostly out of the control of authorities, so it was necessary to create a new kind of grades that were under direct supervision (SLOs and assessments). It's telling that one of the objections to grades is that they might be okay if the work in the course is aligned to learning outcomes. In other words, if we put an approval process between grades and their use in assessment reports, it's okay--it's now subject to management control.

Older Posts on Grades

You might also be interested in these scintillating thoughts on grades:

Update 1

The day after I posted this, IHE ran an article on a study of placement: when should a student be place in a developmental class instead of the college introductory class. You can find the working paper here. From the IHE review, here's the main result:

Meanwhile, students who otherwise would have been in developmental courses but were “bumped up” to college-level courses by the multiple-measures assessment had notably better outcomes than their peers, while those who otherwise would have been in college-level courses but were “bumped down” by the assessment fared worse.

Students put in college-level courses because of multiple measures were about nine percentage points more likely than similar students placed using the standard process to complete college-level math or English courses by the ninth term. Students bumped up to college-level English specifically were two percentage points more likely than their peers to earn a credential or transfer to a four-year university in that time period. In contrast, students bumped down to developmental courses were five to six percentage points less likely than their peers to complete college-level math or English.

Note that the outcome here is just course completion. No SLOs were mentioned. Yet the research seems to lead to plausible and significant benefits for student learning. I point this out because, at least in my region, this effort would not check the boxes for an accreditation reports for the SLO standard. It fails the first checkbox: there are no learning outcomes!

If you want to read my thoughts on why defining learning outcomes is over-rated, see Learning Assessment: Choosing Goals.


Update 2

I had to race back from the conference to finish up work on an accreditation review committee, and I'm finally catching my breath from a busy week. Here are my impressions from the two sessions at the 2023 Assessment Institute.
 
 
This was my research talk, which was well-attended, and had great participation. I went through most of the slides to highlight empirical links between grades, learning, course rigor, student ability, and the development paths in learning data. One question I got was about using the word "rigor" or "difficulty" to describe the regression coefficients from a class section, which is really just a positive or negative displacement from the average expected grade. In the literature there's a more neutral term 'lift' that is used, which is better. I have avoided using 'lift' only because it's one more thing to explain, but I think that was an error. 
 
The reason that lift is better is because it doesn't come with meanings we associate with rigor or difficulty. The question was about the potential difference between a course that's academically challenging, which might be good for learning engagement, versus one that's poorly taught. I don't think we can tell the difference directly with the grade analysis I'm using (random effects models with intercepts for students and course sections). That's a great opportunity for future research--what additional data do we need, or is there some cleverer way of using grade data, like a value-added model by instructor. The Insler, et al piece referenced in my slides might be a starting point.
 
A metaphor that nicely contrasts the research I summarized versus the state of SLO compliance reports is horizontal versus vertical. The SLO reports are vertical in the sense that the chop up the curriculum into SLOs that are independently analyzed; there's no unifying model that explains learning across learning outcomes, and consequently few opportunities to find solutions that apply generally instead of per-SLO. So we get a lot of findings like "critical thinking scores are low, so we'll add more of those assignments," and not general improvements to teaching and learning that might affect many SLOs, students, and instructors at once. The latter is the horizontal approach that I emerged from the research on grades. 

Specifically, it's clear that GPA correlates with learning outcomes across the curriculum, so it's more efficient to think about how students with different ability levels or levels of engagement navigate the curriculum than it is to look at one SLO at a time and hope to catch that nuance. Additionally, improving one instructor's teaching methods can affect a lot of SLOs and students. 

One of the questions got to that point--it was about how a rubric might be used to disambiguate aspects of an SLO, perhaps in contrast to the more general nature of what grades tell us. In the vertical approach, the assessment office would conceivably have to oversee the administration of dozens or hundreds of such rubrics--kind of what we are asked to do now. In the horizontal approach, we put the emphasis on supporting faculty to take advantage of new teaching methods in all their classes. This entails supporting faculty development and engaging faculty as partners rather than as supervisors, which is contrary to the SLO report culture, but it's bound to be more effective. 

Some of the questions were important, but too high level to really address from the results of the empirical studies, because of the layers of judgment and politics that are involved. For example, how can the empirical results be used to inform a well-structured curriculum. I think this is quite possible to do, but it partly depends on the culture in an academic program, and their collective goals. In principle, we could aim to make all courses in a curriculum have the same lift (i.e. difficulty --see above), but that may not be what a department want or needs. Perhaps we need some easier courses so that students can choose a schedule that's manageable. We're not having those conversations yet, but they are important. Should some majors be easier than others to permit pathways for lower-GPA students? If so, how do we ensure that these aren't implicitly discriminatory? It gets complicated fast, but if we don't have data-informed conversations, it's hard to get beyond preconceptions and rhetoric.

This forum was the culmination of several years of work on my part, and perhaps more importantly due to the vision and execution of the editors of J. Assessment and IE. Mark was there to MC the panel, which largely comprised authors who responded to my lead article on grades with their own essays that are found in the same issue. Kate couldn't make it to the panel, but she deserves thanks for making space for us on the AAC&U track at the conference. We didn't get a chance to ask her what she meant by "third rail" in the title. Certainly, if you have tried to use course grades as learning assessments in your accreditation reports, you run the risk of being shocked, so maybe that's it.

Peter Ewell was there! He's been very involved with the assessment movement since the beginning and has documented its developments in many articles. I got a chance to sit down with him the day before the panel and discuss that history. My particular interest was a 1984 NIE report that recommended to accreditors to adopt the define-measure-improve formula that's still the basis for most SLO standards, nearly forty years later. My short analysis of the NIE report and its consequences will be found in December, 2023 edition of Assessment Update. Peter led that report's creation, so it was great to hear that I hadn't missed the mark in my exegesis. We also chatted about the grades issue, how the prejudice got started (accident of history), and the ongoing obtuseness (my words) of the accrediting agencies. 

One of my three slides in my introductory remarks for the panel showed the overt SACSCOC "no grades" admonition found in official materials. I could have added another example, taken from  a Q&A with one of the large accreditors, (lightly edited)

Q: [A program report] listed all of the ways they were gathering data. There were 14 different ways in this table. All over the place, but course grades weren't in the list anywhere, and I know from working in this for a long time that there has been historically among the accreditors a kind of allergy to course grades. Do you think that was just an omission, or is there active discouragement to use course grades?

A: We spend a considerable amount of time training institutions around assessment, and looking at both direct and indirect measures, and grades would be an indirect measure. And so we encourage them to look at more direct ways of assessing student learning.

It's not hard to debunk the "indirect" idea--there's no intellectual content there, e.g. no statistical test for "directness." What it really means is "I don't like it." and "we say so."

I should note, as I did in the sessions, that this criticism of accreditor group-think is not intended to undermine the legitimacy of accreditation. I'm very much in favor of peer accreditation, and have served on eleven review teams myself. It works, but it could be better, and the easiest thing to improve is the SLO standard. Ironically, it means just taking a hard look at the evidence.

I won't try to summarize the panelists. There's a lot of overlap with the essays they wrote in the journal, so go read those. I'll just pull out one thread of conversation that emerged, which is the general relationship between regulators and reality. 

A large part of a regulator's portfolio (including accreditors, although they aren't technically regulators) is to exert power by creating reality. I think of this as "nominal" reality, which is an oxymoron, but an appropriate one. For example, a finding that an instructor is unqualified to teach a class effectively means that the instructor won't be able to teach that class anymore, and might lose his or her job. An institution that falls too short may see its accreditation in peril, which can lead to a confidence drop and loss of enrollment. These are real outcomes that stem from paperwork reviews.

The problem is that regulators can't regulate actual reality, like passing a law to "prove" a math theorem, or reducing the gravitational constant to reduce the fraction of overweight people in the population. This no-fly zone includes statistical realities, where the trouble starts for SLO standards. Here's an illustration. A common SLO is "quantitative reasoning," which sounds good, so it gets put on the wish list for graduates. Who could be against more numerate students? The phrase has some claim to existence as a recognizable concept: using numbers to solve problems or answer questions. But that's far too general to be reliably measured, which is reality's way of indicating a problem. Any assessment has to be narrow in scope, and the generalization of a collection of narrow items to some larger skill set has to be validated. In this case, the statement is too broad to ever be valid, since many types of quantitative reasoning depend on domain knowledge that student's can be assumed to have. 

In short, just because we write down an SLO statement doesn't mean it corresponds to anything real. Empiricism usually works the other way around, by naming patterns we reliably find in nature. As a reductio ad absurdum consider the SLO "students will do a good job." If they do a good job on everything they do, we don't really need any other SLOs, so we can focus on just this one--what an advance that would be for assessment practice! As a test, we could give students the task of sorting a random set of numbers, like 3,7,15,-4,7, 25,0,1/2, and 4, and score them on that. Does that generalize to doing a good job on, say, conjugating German verbs? Probably not. If we called the SLO "sorting a small list of numbers" we would be okay, but "doing a good job" is a phantasm.

So you can see the problem. When regulators require wish lists that probably don't correspond to anything real, we get some kind of checkbox compliance, and worse: a culture devoted to the study of generalized nonsense.

Two Paths to Relevance

I am always impressed at these conferences how nice the assessment community is. There is enormous potential energy there to do good in the world, and my thesis throughout these talks was that their work is impeded by outdated ideas forced on us by accreditation requirements. 

I don't think the current resource drain caused by the compliance culture is sustainable, and it behooves us to look for alternatives in order for assessment offices to stay relevant in a world of tighter budgets and chatGPT-generated assessment reports. As I describe more fully in the upcoming Assessment Update, and as I mentioned at the conference, I think there are two paths forward. One is to do more data science on student learning and success, extending the type of work that IR offices do. I don't think we need a lot of these, assuming that we can generalize results. For example, if there's a demand for sophisticated grade statistics like I presented, software vendors will line up to sell them to you, and it's probably cheaper than hiring someone to build the software from scratch. 

The second strand is where we probably need to most effort, and that's to extend the successful work that assessment offices do now to engage with faculty members on improvement projects. If we abandon the SLO reporting drill, the reason for faculty resentment vanishes, and we can spend time reading literature on pedagogy, become familiar with discipline standards (like ACTFL), and instead of relying on dodgy data, engaged in a collaborative and trusting relationship with faculty to rely on whatever combination of organic data (like final exam scores, course grades, rubrics where they already exist) and their professional judgment.

You might object that faculty are working with us now because of the SLO report requirements, and if that goes away they may stop taking our calls. This is true, but it has administrative and regulatory implications. If accreditation standards take a horizontal approach and ask for evidence of teaching quality, it's a natural fit to a more faculty development style of assessment office. And it will likely produce better results, since an institution can currently place no value on undergraduate education and still pass the SLO requirement with the right checkboxes.

But I don't claim to have all the answers. It's encouraging that the conversation has begun, and I think my final words resonated with the group, viz. that librarians and accountants set their own professional standards. After more than three decades of practice, isn't it time that the assessment practitioners stopped deferring to accreditors to tell them what to do and set their own standards?

Saturday, July 29, 2023

Average Four-Year Degree Program Size

"How much data do you have?" is an inevitable question for program-level data analysis. For example, assessment reports that attempt understand student learning within an academic major program typically depend on final exams, papers, performance adjudication, or other information drawn from the seniors before they graduate: a reasonable point in time to assess the qualities of the students before they depart. Most accreditors require this kind of activity with a standard addressing the improvement of student learning, for example SACSCOC's 8.2a or HLC's 4b. 

The amount of data available for such projects depends on the number of graduating seniors. As an overall assessment of these amounts I pulled the counts reported to IPEDS for 2017 through 2019 (pre-pandemic) for all four-year (bachelor's degrees) programs. These rows of data each come with a disciplinary CIP code, which is a decimal-system index that describes a hierarchy of subject areas. For example 27.01 is Mathematics, and 27.05 is Statistics. Psychology majors start with 42. 

We have to decide what level of CIP code to count as a "program." The density plot in Figure 1 illustrates all three levels: CIP-2 is the most general, e.g. code 27 includes all of math and statistics and their specializations. 

There are a lot of zeros in the IPEDS data, implying that institutions are reporting that they have a program, but it has no graduates for that year.  In my experience, peer reviewers are reasonable about that, and will relax the expectation that all programs produce data-driven reports, but your results may vary. For purposes here, I'll assume the Reasonable Reviewer Hypothesis, and omit the zeros when calculating statistics like the medians in Figure 1.


 
Figure 1. IPEDS average number of graduates for four-year programs, 2017-19, counting first and second majors, grouped by CIP code resolution, with medians marked (ignoring size zero programs).
 
CIP-6 is the most specific code, and is the level usually associated with a major. The Department of Homeland Security has a list of CIP-6 codes that are considered STEM majors. For example, 42.0101 (General Psychology) is not STEM, but 42.2701 (Cognitive Psychology and Psycholinguistics) is STEM. The CIP-6 median size is nine graduates, and it's reasonable to expect that institutions identify major programs at this level. But to be conservative, we might imagine that some institutions can get away with assessment reports for logical groups of programs instead of each one individually. Taking that approach, and combining all three CIP levels effectively assumes that there's a range of institutional practices, and enlarges the sample sizes for assessment reports. Table 1 was calculated under that assumption.
 
Table 1. Distribution of average program sizes with selected minimums. 

Size Percent
less than 5 30%
less than 10 46%
less than 20 63%
less than 30 72%
less than 50 82%
less than 400 99%  

Half of programs (under the enlarged definition) have fewer than 12 graduates a year. Because learning assessment data is typically prone to error, a practical rule of thumb for a minimum sample size is N = 400, which begins to permit reliability and validity analysis. Only 1% of programs have enough graduates a year for that. 

A typical hedge against small sample sizes is to only look at the data every three years or so, in which case around half the programs would have at least 30 in their sample, but only if they got data from every graduate, which often isn't the case. Any change coming from the analysis has a built-in lag of at least four years from the time the first of those students graduated. That's not very responsive, and would only be worth the trouble if the change has a solid evidentiary basis, and is significant enough to have a lasting and meaningful impact on teaching and learning. But 30 samples isn't going to be enough for a significant project either.

One solution for the assessment report data problem is to encourage institutions to research student learning more broadly--starting with all undergraduates, say--so that there's a useful amount of data available. The present situation faced by many institutions--reporting by academic program--guarantees that there won't be enough data available to do a serious analysis, even when there's time and expertise available to do so. 

The small sample sizes lead to imaginative reports. Here's a sketch of an assessment report I read some years ago. I've made minor modifications to hide the identity.

A four-year history program had graduated five students in the reporting period, and the two faculty members had designed a multiple-choice test as an assessment of the seniors' knowledge of the subject. Only three of the students took the exam. The exam scores indicated that there was a weakness in the history of Eastern civilizations, and the proposed remedy was to hire a third faculty member with a specialty in that area. 

This is the kind of thing that gets mass-produced, increasingly assisted by machines, in the name of assuring and improving the quality of higher education. It's not credible, and a big part of the problem is the expected scope of research, as the numbers above demonstrate. 

Why Size Matters

The amount of data we need for analysis depends on a number of factors. If we are to take the analytical aspirations of assessment standards seriously, we need to be able to detect significant changes between group averages. This might be two groups at different times, two sections of the same course, or the difference between actual scores and some aspirational benchmark. If we can't reduce the error of estimation to a reasonable amount, such discrimination is out of reach, and we may make decisions based on noise (randomness). Bacon & Stewart (2017) analyzed this situation in the context of business programs. The figure below is taken from their article. I recommend reading the whole piece.

Figure 2. Taken from Bacon & Stewart, showing minimum sample sizes needed to detect a change for various situations (alpha = .10). 

Although the authors are talking about business programs, the main idea--called a power analysis--is applicable to  assessment reporting generally. The factors included in the plot are the effect size of some change, the data quality (measured by reliability), the number of students assessed each year, and the number of years we wait to accumulate data before analyzing it. 

Suppose we've changed the curriculum and want the assessment data to tell us if it made a difference. If the data quality isn't up to that task, the data quality also isn't good enough to tell us that there's a problem that needs to be fixed to begin with--it's the same method. Most effect sizes from program changes are small. The National Center for Education Evaluation has a guide for this. In their database of interventions, the average effect size is .16 (Cohen's D, or the number of standard deviations the average measure changes), which is "small" in the chart. 

The reliability of some assessment data is high, like a good rubric with trained raters, but it's expensive to make, so it's a trade-off with sample size. Most assessment data will have a reliability of .5 or less, so the most common scenario is the top line on the graph. In that case, if we graduate 200 students per year, and all of them are assessed, then it's estimated to take four years to accumulate enough data to accurately detect a typical effect size (since alpha = .1, there's still a 10% chance we think there's a difference when there isn't). 

With a median program size of 12, you can see that this project is hopeless: there's no way to gather enough data under typical conditions. Because accreditation requirements force the work to proceed, programs have to make decisions based on randomness, or at least pretend to. Or risk a demerit from the peer reviewer for a lack of "continuous improvement."  

Consequences 

The pretense of measuring learning in statistically impossible cases is a malady that afflicts most academic programs in the US because of the way accreditation standards are interpreted by peer reviewers. This varies, of course, and you may be lucky enough that this doesn't apply. But for most programs, the options are few. One is to cynically play along, gather some "data" and "find a problem" and "solve the problem." Since peer reviewers don't care about data quantity or quality (else the whole thing falls apart), it's just a matter of writing stuff down. Nowadays, ChatGPT can help with that. 

Another approach is to take the work seriously and just work around the bad data by relying on subjective judgment instead. After all, the accumulated knowledge of the teaching faculty is way more actionable than the official "measures" that ostensibly must be used. The fact that it's really a consensus-based approach instead of a science project must be concealed in the report, because the standards are adjudicated on the qualities of the system, not the results. And the main requirement of this "culture of assessment" is that it relies on data, no matter how useless it is. In that sense, it's faith-based.

You may occasionally be in the position of having enough data and enough time and expertise to do research on it. Unfortunately, there's no guarantee that this will lead to improvements (a significant fraction of the NCEE samples have a negative effect), but you may eventually develop a general model of learning that can help students in all programs. Note that good research can reduce the number of samples needed by attributing score variance to factors other than an intervention, e.g. student GPA prior to the intervention. This requires a regression modeling approach that I rarely see in assessment reports, which is a lost opportunity.
 

References

Bacon, D. R., & Stewart, K. A. (2017). Why assessment will never work at many business schools: A call for better utilization of pedagogical research. Journal of Management Education, 41(2), 181-200.