Showing posts sorted by date for query complexity. Sort by relevance Show all posts
Showing posts sorted by date for query complexity. Sort by relevance Show all posts

Friday, August 16, 2024

A Canticle for Bloom*

Introduction

Stephen Jay Gould promoted the idea of non-overlaping magisteria, or ways of knowing the world that can be separated into mutually exclusive domains, where each "holds the appropriate tools for meaningful discourse and resolution." The tension Gould was trying to resolve was between religion and science: 

Science tries to document the factual character of the natural world, and to develop theories that coordinate and explain these facts. Religion, on the other hand, operates in the equally important, but utterly different, realm of human purposes, meanings, and values—subjects that the factual domain of science might illuminate, but can never resolve. -- Stephen Jay Gould from Rock of Ages

I'll call these "ways of knowing" or WOKs, which seems more down to earth than "magisteria." Each WOK contains cross-checks on knowledge that are particular to the domain.  Scientific questions are judged by scientific standards. Personal choices are based on experience and usually don't have "correct" answers but degrees of validation.

Should you give a friend a loan? That's a lived experience question. Are your car's spark plugs failing to ignite? That's a science question. Should you take a recommended drug despite severe side effects? That's somewhere in between. 

There's a third WOK we need to talk about: ideology. Steven Mintz provided a crisp definition recently in his blog on InsideHigherEd

Ideologies simplify, clean up and package reality into something easily consumable, palatable and appealing to a mass audience. In doing so, ideologues discard the messy, complex and often unpleasant aspects of reality, presenting only what fits neatly within their framework. Ideology, thus, distorts reality by filtering out anything that does not conform to its narrative.

Ideologies are particularly powerful when associated with a utopia. For example, the way that Marxism evolved into an intellectual justification for Stalin's USSR. Lysenkoism enforced ideology on biological research.  

Here's a diagram of our three WOKs with what we might find in the overlaps.

Galileo was interested in the physical reality of the cosmos (among other things), at the center of the diagram, creating new methods for WOK 2. But the correct way of speaking about the universe needed to adhere to church doctrine (ideology, WOK 3). This is the "correctness" overlap, only some of which corresponded to reality (geocentricism did not). See Steven Shapin's The Scientific Revolution for a nuanced narrative; it's not as simple as the usual telling of Galileo vs the Church. Jennifer Michael Hecht would rightfully insist on putting "ritual" in the overlap between ideology and lived experience, and argue that the ritual can fulfill a social and personal need. Ideology isn't bad; we need it. It can just overlap in odd ways with other WOCs.

SLO Assessment 

Yesterday I participated in a conversation with a small group of experienced assessment directors as part of an ongoing project to fix accreditation standards. The discussion echoed a theme I heard last year when I interviewed a dozen peer reviewers from various accreditors, that there's value in the formal kind of assessment that gathers data and does statistics on it, but it's more common to see success by just getting faculty together to talk about learning goals. We also talked about accreditation standards. I suggest that there are three important WOKs in assessment:
  1. Professional judgment and collaboration
  2. Educational measurement and inferential statistics
  3. Adjudication of accreditation policy
If we find that first-generation students have abysmal pass rates in math (WOK 2), this will affect conversations about pedagogy, support services, course prerequisites, and so forth that would happen in a department meeting (WOK 1). Conversely, if the math faculty all agree that Calculus 1 isn't adequately preparing students for Calculus 2 based on classroom experiences (WOK 1), it might prompt a more formal analysis of grades and test scores (WOK 2).

Perhaps in a perfect world, assessment offices would operate within these two WOKs, with a large faculty support role to facilitate conversations and share knowledge (WOK 1), with a separate function to gather and warehouse data, do research, and connect with the wider research community to bring ideas back (WOK 2). It's not clear that universities would fund such an outfit, however. Assessment offices are expensive and only exist because of accreditation requirements, to get the reports done (WOK 3).

 

The Third Circle

Any bureaucracy draws from a kind of ideology, at least implicitly. Paperwork and procedure serve to "simplify, clean up and package reality into something easily consumable," in Mintz's formulation. Behind the paperwork is a purpose: the DMV's goal is safer roads. The EPA's is  a clean environment. The validation of WOK 3 uses a formalized classification of the world (e.g. driver's test) to assign cases to policy distinctions (driver's license granted or denied).
 
Accreditation requirements provide the motivation to run assessment operations, but they also impose a particular ideology. I described its origins and effects in "Assessment standards are broken." In short, the ideology can be abbreviated as "define-measure-improve" and is a version of "scientific management," descendants of Taylor and Drucker and others, of which Six Sigma is a variation. There is an intended overlap with WOK 2, almost subsuming the scientific WOK within the ideology, with the goal "we're going to require you to use science to improve the state of education."
 
Robert Birnbaum catalogs variations of this idea, calling the phenomenon 
 
a paradox of complexity and simplicity. Its central ideas may appear brilliantly original. Yet at the same time they are so commonsensical as to make us wonder why we had not thought of them ourselves, and so obviously reasonable as to defy disagreement. 
-- Management Fads in Higher Education, pg 5.
 
The result is that the internal validation of knowledge in WOK 3, which is done by trained peer review teams, assumes the preeminence of WOK 2 (scientific knowledge). 
 
Any new idea has to compete with existing ones, and faculty tradition was the natural enemy of the define-measure-improve protocol, despite the obvious overlap with what teachers do on the job, and what they want to accomplish. In the Sturm und Drang of 1983's A Nation at Risk, educators took a lot of heat (a theme in US politics). In higher education, teaching work had to be mapped from existing practices, seen as inefficient, to ones that aligned with the scientific management principles.
  • Defining learning

    • Old: choose textbooks, write syllabus, approve curriculum, create tests or other assessments

    • New: write statements of student learning objectives, often in a hierarchy of course, program, institution

  • Measurement

    • Old: grade tests, writing samples, performances, etc. Build a shared sense of acceptability via faculty consensus and constant exposure to students, assign summative course grades

    • New: use only a few approved methods, including specific assignments, papers associated with rubrics (no grades!). Learning is seen as distinct from more general student success. Emphasis on outputs instead of inputs.

  • Improvement

    • Old: Professional growth in teaching practice, department or institution level consensus on change (WOK 1), using data summaries like grade or test averages or pass rates to identify needs for improvement (WOK 2).

    • New: Averages or frequencies of approved data sources to find deficiencies, then imagine a way to remedy them
The requirements to write reports using the new methods lobotomized WOK 1 for faculty. To comply with the reporting requirements they had to start over using approved replacement methods in order to be allowed to "know" how their students were doing. Naturally they resented this. Hated it, even, and have produced a genre of articles complaining about assessment. We still deal with the effects.
 
It's worth repeating that what assessment directors say works best is WOK 1--what the scientific approach was intended to replace--and because that's what actually works, the accreditation reviews have relaxed over time to allow more room for WOK 1. Your situation depends on your accreditor, but it's still an awkward fit because of the need to perform the other rituals (defining and gathering data in the approved way) in order to validate WOK 1, which doesn't really need that extra work to function.
 
As I described in "Assessing for Student Success" the science project falls apart immediately, because what students learn in a college curriculum is a lot of detailed topics with interconnections. I estimated several hundred topics (SLOs if you like) in a math curriculum. These can't be described, let alone measured, within the parallel framework the accreditors created. 
 
It's worth noting that in the industrial setting, where these ideas were formed, it is possible to define and measure everything important, and have real-time data from instruments on an assembly line. That doesn't translate well to education, where there's little standardization.

The accreditation requirements strengthen the main thing that works (WOK 1) by creating more opportunities for faculty to talk about student learning, especially when facilitated by a good assessment director. But the requirements also diminish the effectiveness of faculty work by heaping on artificial requirements in the name of science. This poor attempt to mandate WOK 2 fails the validity checks within WOK 2: sample sizes are too small, too noisy, and the causal models used are too simple. 

This collision between science and policy gets resolved ad baculum: you will be beaten with the stick of non-compliance until you at least pretend to believe that rubrics are always valid measures. In short, the SLO accreditation requirements co-opt the authority of science, but replace scientific standards with ideologically-correct ones. This has created a self-sustaining culture of compliance, abetted by consultants, vendors, and peer review training that maintains this closed garden of bureaucracy as science.

 

A Litmus Test

A few years ago, I concluded that the best way to illustrate how accreditation standards drove us into an epistemological ditch was to spotlight their allergy to course grades. In the age of big data, the idea that we'd arbitrarily throw out millions of data points covering the whole history of students at our institution is absurd. It only makes sense within the accreditation bubble, and so it puts a spotlight on the difference between the scientific claims of accreditors versus the reality. To heighten the contradictions, so to speak. You can find my summary of research on grades and learning here. Starting on page 23 you'll find a list of standard objections to using course grades as data about student learning. 
 
I won't rehash that material here. It suffices to quote an article that cites Bloom (the taxonomy guy) from 1976, eight years before define-measure-improve kicked off:
 
Perhaps the most productive use of GPA is as a covariate. GPA has the potential to explain nearly half the variance in education research models (Bloom, 1976), thus shedding light on the variance explained by other variables of interest, such as changes in course or curriculum design.  
 
-- Bacon, D. R., & Bean, B. (2006). GPA in Research Studies: An Invaluable but Neglected Opportunity. Journal of Marketing Education, 28(1), 35–42.
 
 
When I have the opportunity, I ask accreditors what they think of course grades as data. This is a litmus test for self-reflection. So far they've all failed it. The wise ones don't want to come out too strongly against grades, because it implies they don't think transcripts are meaningful. They generally hedge, calling grades "indirect measures." But there's no test for directness in the vocabulary of WOK 2. You won't find a way to create p-values on directness in the manual on educational measurement, because the idea isn't statistical. We already have a robust vocabulary on reliability and validity, and there's no need to confuse the issue with "directness." 

When pushed on that point, one accreditor representative cited an anecdote about how a student didn't feel like the grade reflected learning. This would be ironic if the scientific claims of the define-measure-improve protocol were serious about the science. After dismissing faculty consensus as opinion, reciting "the plural of anecdote isn't data," and distributing buttons at conferences that read "show me the data," a high priest of the order can simply use an anecdote to dismiss the whole research record on course grades. The remarkable thing is that this doesn't seem to cause any cognitive dissonance. 

The point of this illustration is that the overlapping circles in the three WOKs don't presently stand the light of public exposure. The accreditors will look ridiculous. We need to fix it before that happens, to create a more sensible overlap of the WOKs.

There's More

This article is long enough, and I'm going to stop here. I have not discussed the effective use of WOK 2 (educational measurement and inferential statistics) as a successful assessment tool. That may be addressed in a future post. Suffice to say that that requirements of a good research project (e.g. large data sample, tests of reliability) aren't feasible in 99% of department-level accreditation reports. From a practical point of view, it's a lot of extra work to do a real research project to get a tiny amount of credit, if any at all (none at all if you research retention instead of learning). The irony is that WOK 3 assumes that it contains WOK 2, when in fact the intersection is nearly empty.
 

Friday, November 19, 2021

Learning Assessment: Choosing Goals

Introduction

I recently received a copy of a new book on assessment, one that was highlighted at this year's big assessment conference:

Fulcher, K. H., & Prendergast, C. (2021). Improving Student Learning at Scale: A How-to Guide for Higher Education. Stylus Publishing, LLC.
 
This is not a review of that book; I just want to highlight an important idea I came across, from pages 60-63 of the paperback edition. The authors contrast two methods, described as deductive or inductive,for selecting learning goals that form the basis for program reporting and (ideally) improvement.Here's how they name the two methods, along with my suggested associations (objective, subjective) in italics:
  • Deductive (Objective): "A learning area is targeted for improvement because assessment data indicate that students are not performing as expected." (p. 60).

  • Inductive (Subjective): "[P]otential areas for improvement are identified through the experiences of the faculty (or the students)." (p. 60).

Although it's not in the book, I suggested the associations to objective/subjective because objectivity is often seen as superior to subjectivity in measurement and decision-making. See, for example Peter Ewell's NILOA Occasional paper "Assessment, Accountability, and Improvement: Revisiting the Tension."

In the early days of the assessment movement, campus assessment practices were consciously separated from what went on in the classroom. This separation helped increase the credibility of the generated evidence because, as “objective” data-gathering approaches, these assessments were free from contamination by the subject they were examining.(p. 19)
The desire to be free from human bias is related to assessment's roots as a management method--the same family tree as Six-Sigma, which helped Jack Welsh's General Electric build refrigerators more efficiently. It's related to the positivist movement's emphasis on definitions and strict classifications, and further back to the scientific revolution. However, the history of human beings "objectively measuring" one another includes dark and tragic episodes, as recounted in part in the 2020 president's address to the National Council on Measurement in Education (NCME). From the abstract:

Reasons for distrust of educational measurement include hypocritical practices that conflict with our professional standards, a biased and selected presentation of the history of testing, and inattention to social problems associated with educational measurement.
While the methods of science may be mooted as objective, the uses of those methods are never unbiased as long as humans are involved. 
 
A subtler critique of objective/deductive decision-making comes from artificial intelligence research. Deduction requires rules to follow, whereas induction depends on educated guesses. For a fascinating demonstration of these two approaches, see Anna Rudolf's analysis of two chess-playing programs. The contest pitted an objective/deductive algorithm (Stockfish) that used rules and points to assign values to positions, against an inductive/subjective program (Google's AlphaZero) that wasn't even initially taught the rules of the game! It played by sense of smell. And won. The video highlights AlphaZero's "inductive leaps" that have it making plays that look ill-advised according to the usual rule-based understanding of play, but turned out to be a winning strategy.

So putative objectivity can be challenged in at least two ways, viz. that human involvement means it's not really objective, and that objective methods aren't necessarily better than subjective ones in solving problems beyond a certain complexity level. 

In their book, Fulcher and Prendergast describe another compelling reason to consider an inductive/subjective approach: practicality.

Learning Goals

Accreditation requirements for assessment reporting insist that academic programs declare their learning goals up front. Measurement and improvement stem from these definitions, which is why some assessment advocates philosophize about verbs. My thesis here is that programs could do more productive work if the accreditors didn't force them to pre-declare the goals as a prerequisite to an acceptable report. I am not claiming that it is a bad idea to think about learning objectives and describe them in advance; rather that if we only take that approach and ignore the rich subjective experience of faculty outside the those lines, we unnecessarily diminish our understanding and ability to act. Accreditation rules are too restrictive.

For a faculty member, however, this pre-declaration of goals in approved language is where cognitive dissonance begins. The curriculum's aims are already amply described in course descriptions, syllabi, and textbooks. A table of contents in an introductory text will list dozens of topics, many worth at least one lecture to cover. In some disciplines, like math, these topics are associated with problem sets (i.e. detailed assessments). The pedagogy and assessments in these textbooks has evolved with experience; we don't teach calculus anything like what's found in Newton's Principia, and for good reason. Teachers develop subjective judgments about what material is likely to be difficult for students, and learn tricks to help them through it. This is an "AlphaZero" method: inductive, depending on human neural nets to assemble and weigh data rather than relying on preset rules like "if average student scores are below 70%, I will take action."

In a 2021 AALHE Intersection paper, I counted learning goals in a few textbooks and estimated that there are at least 500 for a typical program. A first calculus course contains learning goals like "computing limits with L'Hôpital's rule" and "differentiating polynomials." By contrast, a typical assessment report for accreditation asks a program to identify around five learning goals. That's a factor of 100 difference, and we should forgive a faculty member for incredulity at this point: can five categories significantly cover those 500 goals? Even in six-sigma applications, imagine how many measurements there are to specify a refrigerator. It's a lot more than five. And each of those measures, like bolt size and thread width, are specific to a particular state of assembly, where it can be tested--just like usual classroom assessments via assignments and grading. You don't want to build a whole "cohort" of refrigerators and then find out that the screws holding it together are all the wrong size.

This cognitive dissonance between formal systems and common sense is familiar. Anytime you've filled out a form and found that the checkboxes and fill-ins don't make sense in your case, that's a collision between the deductive/objective and inductive/subjective worlds. We live primarily in the latter, reasoning without decision rules. We follow the rules of the road--which lane to drive in, obeying signals and signs, but the rules don't tell us where to go or why. The most significant decisions in life are subjective/inductive.

Examples

Some of the following examples are my interpretations of selected assessment work drawn from the literature, conferences, and practice.

Interview Skills

In a 2018 RPA article, Keston (co-author of the book cited above) and colleagues documented a program improvement in Computer Information Systems.

Lending, D., Fulcher, K. H., Ezell, J. D., May, J. L., & Dillon, T. W. (2018). Example of a program-level learning improvement report. Research & Practice in Assessment, 13, 34-50 [link]

The same story is told less formally in a collection of assessment stories here. It's short and worth reading, but here's the apposite part--a faculty member's realization based on recent direct observation of student work:

We realized that nowhere in our curriculum did we actively teach the skill of interviewing. Nor did we assess our students’ ability to perform an effective [requirements elicitation] interview, apart from two embedded exam questions in their final semester.

So although although there existed an official learning outcome intended to cover this important interview skills, it was assessed in a way that hadn't elicited the important information that students weren't doing it well. This illustrates the flaw in assuming that a few general goals can adequately cover a curriculum in detail. But faculty are immersed in these details, and good teachers will notice and act on such information. So in this case, the deductive machinery was all in place, including an assessment mechanism, but it failed where the inductive method worked. 

An important point illustrated here, and mentioned in the book, is that having faculty enthusiasm for a project is conducive to making changes.

Discussion-Based Assessment

At the 2021 online assessment meeting put on by University of Florida (Tim Brophy and colleagues), I saw Will Miller's presentation. From the abstract (my underlining):

Jacksonville University decided to design and implement a discussion-based approach to assessment for the 2020-2021 academic year. This decision—which was made to minimize the bureaucratic feeling of sending out templates and providing deadlines—has led to enhanced faculty and staff understanding of assessment without the stress of prolonged back and forth reviews and evaluations. [...]

[T]he discussion-based approach has led to expressions of cathartic relief. Faculty and staff in this model are provided with immediate feedback, an opportunity to reflect on their accomplishments in this unparalleled time in higher education, and to hear affirmations of value and contribution. Ultimately, we have been able to collect information that is both deeper and wider than in previous assessment cycles while gaining countless insights into the efforts of the campus community to ensure success in the face of adversity. 

An example I remember from the talk is that the Dance faculty faced special challenges teaching classes virtually, as one might imagine they would. They noticed that students were getting injured at alarming rates when practicing on their own. So they addressed the problem. This didn't happen because they had pre-declared a learning outcome about dance injury, and then created benchmark measures and so on, but because faculty were paying attention.

Notice the indications in the abstract (which I underlined) that the inductive/subjective approach provides more timely and useful information and is more natural than the formal goal-setting. 

Teaching Statistics

In Spring 2021 I taught an introductory statistics class to working adults as an online-only section. It was my first completely online (synchronous) class, and although I'd taught the material before, I had to rethink everything for the new format. 

A key concept in statistics is that estimates like averages and proportions are associated with error, and estimating that error is an important learning goal. Typically error bounds like confidence intervals are obtained using a look-up table of values, often found in a textbook's appendix. You can see one here, and a portion is reproduced below.
 
It's hard enough to teach students to use these tables and the rules that go with them, even when I can walk around the room and point to the relevant sections. It seemed like a poor use of our online time to try to replicate that. Instead, I built an app that students could use to find the critical values by moving sliders around to represent the problem parameters.

The results of this switch were immediately gratifying. Students grasped the concepts more quickly, got the answer correct more often (in my subjective judgment), and some of them even developed an intuition for what the critical value would be without looking it up. I was so happy with the results that I shared the app with others at the university. 
 
At the end of a course, instructors are invited to submit assessments of students.  There was more than one preset learning goal that might apply in this case. I've abbreviated the descriptions to focus on the topic of confidence intervals, p-values, and t-scores.
  •  Analytical reasoning. [...] explain the assumptions of a model, express a model mathematically, illustrate a model graphically, [...], and appropriately select from among different models. 
  • Quantitative methods. [...] articulate a testable hypothesis, use appropriate quantitative methods to empirically test a hypothesis, interpret the statistical significance of test statistics and estimates [...] 
  • Formal reasoning. (General education) [...]  mastery of rigorous techniques of formal reasoning. [...] the mathematical interpretation of ideas and phenomena; [...] the symbolic representation of quantification,
Even without formal assessments, like comparing before-and-after averages, it was clear that students had learned more about this topic, but those gains could easily be counted in all of the learning goals bulleted above. There is no chance that someone analyzing the summative data from those assessments could work backwards to figure out what was going on in detail. Conversely, if those summative scores were lacking, there would be no way to track it down to the issue of finding critical values.

The inability of general learning assessments to tell us much useful about detailed work is why we see so many assessment reports that "add more critical thinking exercises to the syllabus" or something equally anodyne, just to make the peer reviewer check the box. 

Student Success

There are many examples of deductive/objective assessments that do work. The formal approach works exactly when statistical and measurement theory say it will work, with sufficient samples of reliable data, analytical expertise, and where we can reasonably hypothesize cause-and-effect. 

Good research in student learning falls in this category, but it is difficult to do, and not likely to occur in program-level accreditation reports. Examples using grades, retention, graduation rates, license exam pass rates, and so on, are more common. One example is Georgia State University's success over time with graduation rates. 

Unfortunately, the most voluminous and reliable data available to understand learning is course grades, which are generally banned by accreditors as a primary data source. So the most fruitful lines of inquiry for deductive/objective methods are denied.

Discussion

In the stats example, students learned better when I switched from a rules-based "look up the value on a table" method to one where they had to interact with the probability distribution, monitor changes in shapes and values, and narrow in on the parameters needed--an inductive/subjective method. It was subjective because the app has limited precision, so students had to decide when a value was "good enough" for the application. It was inductive because finding the critical value was done by trial and error, moving sliders around.

Teaching students rules-as-knowledge has limited benefits even in math. It's better to build intuition and understanding of interacting components. This is just as true for an academic program trying to understand student learning. Intuition, experience, professional judgment, and the great volume of observation that informs these cannot be dismissed as "merely subjective" without losing most of the actionable information available.

This is not to say that pre-declaring that we want our students to be able to write at a college level is a useless activity; it's just that such formulas fail to account for 99% of what happens in education, or else cover it at such dilution that it can't be used for assessment of the curriculum in a meaningful way. This conflation with declaring goals and assessing them was bemoaned by Peter Ewell in a 2016 article: "One place where the SLO movement did go off the rails, though, was allowing SLOs to be so closely identified with assessment." (Here SLO = student learning outcome.)

Defenders of learning goal statements as the foundation for assessment seem to dismiss most of what happens in college classrooms as mere "objectives" or "stuff on the syllabus," without recognizing that such material comprises the bulk of a college education. The only justification is that accreditors require it (proof by we-say-so), continuing the doom-loops of peer review, consultant, and software contracts in a largely vain attempt at relevance. Ewell, in the same article, connects the SLO-assessment problem as central to these accreditation issues (third paragraph from the end). 

An example of a practical learning goals statement for a physics program can be found at UC-Berkely's site," where program goals are divided into knowledge, skills, and attitudes. This is a nice overview that is suitable to inform students and help faculty reach consensus about curriculum and student outcomes after graduation. Such "ground rules" for the program can also help the faculty identify what should change. This might be related to student learning as perceived by faculty, but it could also be in response to changes in the physics discipline or the job market.

Such general goals as "broad knowledge of classical mechanics" have limited use in informing data-gathering activities for the reasons already mentioned: most learning goals are very specific and best dealt with in the context of courses. An exception would be the goal of cultivating particular behaviors over time, like "successfully pursue career objectives." That requires a new kind of data collection, like 1. when are students aware of their career options?, 2. when do they start taking actions, like seeking a mentor, selecting graduate programs, asking for letters? How are these behaviors associated with the social capital that students arrive with (e.g. are first-generation students less likely to engage with careers early?). It's a significant amount of work to take such a project seriously. 
 
One type of longitudinal study is quite easy: when courses have prerequisites we can compare the grades in the first class to the grades in the second. If the second mechanics course shows low grades despite the same students having high grades in the first course, there might be a problem. As noted, such analysis is off limits for assessment reports because of accreditors' predispositions against grades as data.

Conclusions

The analysis here, stemming from Fulcher & Prendergast's observations, is compelling because it explains facts that are otherwise puzzling. Why is it that after all these years, the assessment reporting process doesn't seem to have produced much, and continues to be locked in a perpetual start-over mode? Part of the explanation is surely that pre-specifying learning goals and then attempting to measure performance prior to any action is not practical in most cases.

If assessment offices focused on helping teachers improve their attention to specific learning goals and their in-class assessments, those teachers will take away new habits that benefit students in many learning goals. This is more efficient than the one-at-a-time rules-based approach, and--as we've seen--more likely to result in meaningful change.
 
 

Thursday, June 19, 2014

A Cynical Argument for the Liberal Arts, Part Fifteen

Previously: Part Zero ... Part Fourteen

After being derailed last time by a quotation, I'll take up the question of constructive deconstruction: how can the corrosive truth-destroying effect of Cynical "debasing the coin of the realm" lead to improvements within a system? Note that a system (loosely defined: society, a company, a government) is the required  'realm' in which to have a 'coin.' My modern interpretation is official or conventional signalling within a group. Waving hello and smiling are social coins of the realm. Within organizations, a common type of signal is a measure of goal attainment, like quarterly sales numbers. It's this latter, more formal type of signal, that is our focus today.

A formalization of organizational decision-making might look like this (from a slide at my AIR Forum talk):


Along the top, we observe what’s going on in the world and encode it from reality-stuff into language. That’s the R → L arrow. This creates a description like “year over year, enrollment is up 5%,” which captures something we care about via a huge reduction in data. R → L is data compression with intent.

As we get smarter, we build up models of how the world behaves, and we reason inside our language domain about what future observations might look like with or without our intervention. We learn how to recruit students better under certain conditions. When we lay plans and then act on them, it is a translation from language back into the stuff of reality—we DO something. By acting, we contribute to the causes that influence the world. We don’t directly control the unfolding of time (red arrow), and all pertinent influences on enrollment are taken into account by the magical computer we call reality. Maybe our plans come to fruition, maybe not.

Underlying this diagram are the motivations for taking actions. Our intelligence, whether as an individuals or as an organization, is at the service of what we want. It's an interesting question as to whether one can even define motivation without a modicum of intelligence--just enough to transform the observable universe (including internal states) into degrees of satisfaction with the world. As animals, we depend on nerve signals to know when we are hungry or in pain. These are purely informational in form, since the signals can be interrupted. That is, there is no metaphysical "hunger" that is a fundamental property of the universe--it's simply an informational state that we observe and encode in a particular way. In a sense, it's arbitrary.

For organizations, analogs to hunger include many varieties of signals that are presumed to affect the health of the system. Financial and operational figures, for example. These are often bureaucratized, resulting in "best salesman of the quarter" awards and so forth. An essential part of the bureaucracy is the GOAL, which is an agreed-upon level of motivation achievement, measured by well-defined signals. An example from higher education is "Our goal for first year student enrollment is an increase of 6% over three years."

A new paper from the Harvard Business School "Goals gone wild," by Lisa D. Ordóñez, Maurice E. Schweitzer, Adam D. Galinsky, and Max H. Bazerman, provocatively challenges the notion that formal goals are always good for an organization. This itself is a cynical act (challenging established research by writing about it), but the paper itself is a great source of examples of how Cynical employees can react to a goal bureaucracy in two ways:

  • Publicly or privately subverting goals through actions that lead to real improvements in the organization, or 
  • Privately debasing the motivational signals that define goals for personal gain.
The first case is desirable when the goals of an organization, when taken too seriously, are harmful to it. Too much emphasis on short-term gains at the expense of long-term gains is an example. This positive reaction to bad goals is not the topic of the paper, but proceeds from our discussion here. The second case is ubiquitous and unremarkable: any form of cheating-like behavior that inflates one's nominal rank, reaping goals-rewards while not really helping the organization (or actually harming it).


Earlier, I referenced a AAC&U survey of employers that claimed they wanted more "critical thinkers." Employees who find clever ways to inflate their numbers is probably not what they have in mind. Before the Internet came around, I used to read a lot of programming magazines, including C Journal and Dr. Dobbs. One of those had a (possibly apocryphal) story about a team manager that decided to set high bug-fixing goals for the programmers. The more bugs they found and fixed, the higher they were ranked. Of course, being programmers, they were perfectly placed to create bugs too, and the new goals created an incentive for them to do just that: create a mistake, "find" it, fix it, get credit, repeat. This is an example of organizational "wire heading." 

From the Harvard paper's executive summary, we can see problems that an ethical Cynic can fix;
The use of goal setting can degrade employee performance, shift focus away from important but non-specified goals, harm interpersonal relationships, corrode organizational culture, and motivate risky and unethical behaviors.
Armed with the liberal-artsy knowledge of signals and their interception, use, and abuse, AND prepared with an ethical foundation that despises cheating, an employee has some immunity to the maladies described above. Of course, this depends on his or her position within the organization, but generally, Cynical sophistication allows the employee to see goals as what they really are--conventions that poorly approximate reality. This is the "big-picture" perspective CEO-types are always going on about. Moreover, an ethical Cynic who is in a position of power is less likely to misuse formal goal-setting as a naive management tool. 

The paper itself is engagingly written, and is worth reading in its entirety. With the introduction above, I think you'll see the potential positive (and negative) applications of Cynical acts. Informally, the advice is "don't take these signals and goals too seriously," which is a point the original Cynics repeatedly and dramatically made. More formally, we can think of a narrow focus on a small set of goals as a reduction in computational complexity, in which our intelligent decision making has to make do with a drastic simplification of the world. Sometimes that doesn't work very well, and the Cynics are the ones throwing the plucked chickens to prove it.

Next: Part Sixteen

Thursday, June 05, 2014

A Cynical Argument for the Liberal Arts, Part Thirteen

Previously: Part Zero ... Part Twelve

Here we continue the discussion of Cynicism in higher education, where I continue to use the charge "debase the coin of the realm" as the tool for analysis. In a previous installment I said that any statement of fact was a violent act. Allow me to pick up that idea here, since much of college consists of listening to statements of fact.

We constantly observe the world through our senses of internal and external affairs, and some of these we encode into language, as in "the cat just knocked over the vase." This entails data compression, since we don't have time to describe everything about the cat and the vase or the irrelevant features of the situation. But it's such a convenience that we may forget that it's just language, just a crude approximation of what was actually observed in that moment. If you'll allow me a neologism, I'd like to refer to compressed data as 'da'. So we pass da back and forth, decompress it internally, ignore the inevitable errors, and this becomes a high value "coin of the realm." If you tell me "the bridge is washed out ahead, but if you turn at the gas station you can get to town that way," it's useful da. However, passing falsehoods along debases this coinage, and lying is something we tell kids not to do.

A statement like "seasons are due to the fact that the Earth's axis of rotation is tilted relative to its orbital plane" require work to decompress. Much work in college is just to build up an ontology that lets a student associate new ideas in useful ways, sometimes by translating them into pictures. A movie that illustrates the principle in the sentence will be more effective than that sentence because moving pictures can directly simulate motion of the Earth and sun, and a viewer can create his or her own da from it. I assume that much of the time students don't know what we professors are talking about. It's just so much da-da, and that represents a failure to teach or learn or both.

Statements of fact that are inaccessible for technical reasons may not be threatening simply because it's a foreign language, and easily dismissed. On the other hand, much of what students learn in the humanities will challenge the way they think about the world. Religious fundamentalism running up against modern biology is an obvious example. But the critical theory that accompanies much of the humanities is more insidious, and can undermine the whole project of reasoning and knowing with a challenge from relativism. A student who internalizes this may not know what to believe by the end of it. There is something good in this idea that there is not always a single correct answer to a question, and that the process of reasoning is more valuable than any one product of it. It seems that the opposite is generally trained into students in the typical K-12 curriculum, where advancement may depend on a standardized test.

To explicate the process of deconstructing truth, allow me to re-use the example of Diogenes tossing the plucked chicken at the feet of Socrates, after the latter declared man to be a featherless biped. How does this trick work?

The Cynical method of attack in this case comprises an action that demonstrates a contradiction between the real world and the way we are told to perceive it. A modern example is found in "The Serial Killer Has Second Thoughts: The Confessions of Thomas Quick," which I will take at face value for the purposes of this argument (feel free to debase that value). A mentally confused man confesses to brutal crimes, and becomes a celebrated serial killer. Then it turns out that he couldn't have committed the crimes he confessed to. Many people have been convicted of crimes that they didn't commit. But it seems that he willingly participated in this fiction, which makes him a fowl-flinging Cynic. The coin of the realm debased here is the operation of the justice system and media (in Sweden in this case) and its accurate representation of reality. Like most public Cynical acts, it comes at a high price.

In order to construct such an epistemological attack on a system, one needs to find categories that are incompatible and by acting produce examples that occupy both categories. For example, the conception of worker-as-machine, which is an employer's point of view, conflicts with worker-as-human, which is the employees' point of view. A worker strike is an incompatible categorization: workers who are not working, and therefore a Cynical chicken toss. Once you see the trick, other examples become obvious: art that is not art (Duchamps), slave-as-beast versus slave-as-human, animal-as-meat versus animal-as-living-creature, Earth-as-resource versus Earth-as-home.

Some of the major discoveries of the world are paradoxical category-defiers. The discoverers of the calculus used infinitesimals, which are quantities 'infinitely small but not zero.'  The only thing infinitely small can mean is zero, so this is contradictory, and yet it is exactly the idea needed to unlock differential calculus. Or the Copenhagen interpretation, in which reality acts like probability waves collapsing into real particles. Or Einstein's category-breaking conceptions of space and time (how can time operate at different rates, when time determines what a rate is?). Or the Dirac Delta, which is essentially a box that has zero width, is infinitely tall, and contains one unit of area! The Liar Paradox "this statement is false" in the hands of Russel and Turing and Gödel caused a rethinking of the fundamentals of mathematics and set logical limits on what we can know. These we should probably consign to 'cynical' rather than 'Cynical' status, since they are arguments rather than physical acts, but this may be splitting hairs.

More common than public abuse of our feathered friends is the secret use of contradicting categories for personal gain. When these become public, they may serve as inadvertent Cynical examples. For example, Slate's "Here’s the Awful 146-Word “Essay” That Earned an A- for a UNC Jock". Quote:
The University of North Carolina–Chapel Hill has already been embroiled in a scandal for allowing its athletes to enroll in fake courses for easy credit. 
A university's motivation for recruiting a top athlete is different from its motivation for recruiting a top student, yet NCAA rules require that athletes also be successful students (there is no requirement the other way around). When forced to coincide for purposes of classification, the overlap "student-athletes" seems to often include examples that fit the latter but not the former category. Actually producing these examples is a Cynical act, even though it's not intended to be public (quite the opposite). There is a straightforward way to debase the coin of the realm (the meaning of college grades). Unfortunately for these Cynics, sometimes people find out make it public.

It seems that private Cynicism for gain, as opposed to public Cynicism, needs a name. I will refer to it as crypto-Cynicism, meaning hidden. The name seems appropriate, since this is a lot of what spies do, like assuming the identity of someone else, forging documents, getting people to trust them who shouldn't, or hiding a real message within an apparent one.

The examples show that Cynicism (as I have interpreted it) remains a powerful force for change, and that it doesn't come with ethical instructions. Colleges that teach the students the deepest forms of subversion, like deconstruction, relativism, critical theory, and so on (these overlap) hand over to the young minds solvents that can turn the world--and their own minds--to goo. Since becoming secular, colleges shy away from construction part, which is unfortunate. We could be more intentional about the existentialist project that should follow the deconstruction--finding a personal meaning to construct out of the goo. But aside from that, there is a more troubling question.

Perhaps what employers want is really two things: (1) a class of graduates who are smart enough to pick up new tasks of varying complexity and fill the technical and social demands of being a machine-part in an organization, and (2) a smaller cadre of crypto-Cynics who are willing to break rules in order to get ahead. If so, then the unsettling conclusion is that a two-class system is exactly what's needed. The first is trained to follow rules and keep their heads down, mastering jobs that real machines will eventually take. The second provides motivation and vision uncluttered by the norms of society, its laws, science, or even a common understanding of the world, to serve as leadership. The spectrum of this second group would range from normal humans to psychopaths and mystics, who succeed or fail depending on their environment. The contrast between these two types (1) and (2) can perhaps be seen in stories like this one from the Huffington Post: "For-Profit College Enrolls, 'Exploits' Student Who Reads at Third-grade Level" in these two quotes:
A librarian at a southern California campus of Everest College abruptly resigned last week, deeply upset that the for-profit school had admitted into its criminal justice program a 37-year-old man who appears to read at a third-grade level.  
versus
Everest is owned by for-profit giant Corinthian Colleges, which is facing a lawsuit for fraud by the attorney general of California and is under investigation by 17 other state attorneys general and four federal agencies. [A senior administrator] at Corinthian, told me today that the campus believed it was appropriate to take a chance on admitting the student.
There's a natural fix to this particular problem: require that the leadership of the company must come from its own graduates.

[Part Fourteen]