High Stakes Metaphors: Calculators, Ladders, and Bad Managers

Generic book icon
Table of Contents

Transformers debate: Calculators, Ladders, and Bad Managers

Prologue by Mark Marino

“Generative AI in writing and coding classes is disrupting educational practices just like the calculators did to math classes.” We hear this metaphor a lot. But is it precise? Or do we need a more precise metaphor? If it is like the calculator, will it allow us to change the floor, or bottom rung of the educational ladder the way the calculator displaced a lot of the focus on basic arithmetic?

Or does generative AI, particularly so-called agentic AI, change our relationship to writing and coding? Rather than just raising the floor, does it turn us into managers, and so we need to train students for that job?

And how do the stakes involved affect the calculus by which students and teachers decide whether or not they should offload their thinking to AI? Are there certain high stakes occupations (like law, medicine, and engineering) where we cannot afford to have students and graduates who overly rely on AI? Or can almost anything become high stakes if the student cares enough about it. And does that caring become an inroad to discouraging over-reliance?

The Transformers address these questions and more in another lively conversation. We begin with a discussion of locked-down, or controlled, browsers as one way to prevent inappropriate use of AI during assessment.

Discussion with Maha Bali, Jeremy Douglass, Jon Ippolito, Mark C. Marino, and Anna Mills

January 15, 2026

Jeremy Douglass: This question a few minutes ago about controlled browsers, especially in the case of online education, suggest the idea that in order to preserve online education pedagogy, and especially assessment, that we might need things like controlled browser stacks and other analogous technology — that’s something that everyone seems excited about, talking about, and I am too.

Maha Bali: One of the things that I think sometimes is to go to extreme cases where it’s really important to validate the assessment.

So, not my course. But, if you’re teaching medicine, you really want to make sure the doctors really know what they’re doing. Obviously, eventually, they’re going to actually have to deal with patients where they can use the technology, but, not always, like in the middle of a surgery if they don’t know which organ they’re looking at or which arteries are bleeding or whatever, you know? There is a stage where you really want to validate something very high risk, and you want to make sure people are able to do this on their own in case of technology failure in their future, or in order for them to be able to evaluate whether the technology is doing a good job.

So even though there are AIs that can actually tell you from an MRI whether someone is likely to have cancer, they’re not 100% reliable, and still a human needs to look at them, so still humans need to learn how to do that, even though AI can do it, right?

Jeremy Douglass: Yeah, that’s an amazing example.

Anna Mills: Yeah. Perfectly put.

Jeremy Douglass: There’s some other examples that are analogous to high-stakes surgery, like driving certification and heavy construction equipment operation, but also in writing. Sometimes—and not just in accounting and the fact that numbers can lose people millions or billions: but there’s contracts, treaties, law — authoring and then vetting law.

Maha Bali: In political contexts AI’s deployment has caused the Israeli-Palestinian problem?

Jeremy Douglass: Right, there are huge consequences to somebody who’s not able to vet on the fly, maybe out of — I think someone earlier mentioned — a kind of a deep intuition: “Hey, this language doesn’t seem right.” There is no fully self-driving in text any more than there is in robotic surgery, right? You need to have somebody with their hand tightly on the wheel.

Jon Ippolito: So, I wonder… I’m gonna refer back to a bit that happened before we started recording here. Jeremy gave the example of the introduction of a calculator. Once that appeared, then you were able to level up and teach at a higher level, assuming that first rung had been stepped on. And the first rung in that case was learning arithmetic or rote manipulation, like adding, subtracting, multiplying, dividing. And, I think that the question is not necessarily whether we should jump over the first rung and get immediately to using AI for everything. Maha gives great examples and Jeremy does also of a context where it’s high risk and got to get the details right, and that could be dangerous without some intuition that’s been developed in some non-AI way by the practitioner. I think the question is not whether you can jump over that rung. It’s what that rung should be. What should be the lowest rung?

And as an example from the world of arithmetic, Maha gave a great example earlier of making some change. I mean, I’m working at my house now and need to quickly decide whether I am measure: am I going to be able to afford a certain countertop or molding or whatever? I have to do some calculations in my head, like how much is that going to be, roughly, before I know the exact number. And some of the more, I would say, enlightened or real-world oriented math classes in elementary school now are like, just go to the closest 10, right? Is it 19.95 times 76 point something? Well, make it, like, 20 times 80. And you can probably do that in your head, because it’s just two numbers, 2 and 8, and if you know the zeros, you can kind of make out what it’s going to be.

So there are ways to sort of short-circuit the more involved nitty-gritty calculations by doing things that help gain an intuition that can be valuable at a higher level. And I’m wondering whether that’s also true of writing, whether you need to know all the details of the difference between a direct and indirect object. Or whether there’s a different first rung we can give students.

Anna Mills: I don’t think we teach that anymore. At least not in college writing, we’re not teaching much grammar. Hardly at all, honestly. We do little bits here and there, kind of. But for most people I know, the focus has evolved toward things that have more obvious impact, like reading, writing, critical thinking. Are you understanding? Do you have something to say? That’s my hope.

Jon Ippolito: They’re at lower levels.

Maha Bali: Yeah, I was thinking, okay, your native language, you can usually speak it grammatically correctly from just living it. Not everybody, obviously—it depends on your education level and who’s around you and everything—but for the most part, in your native speech. For example, there is no grammar for Egyptian dialect Arabic. Nobody ever taught anybody the grammar of that. But you know it because you speak it. So, grammar may not be a subject to learn. Terminology, maybe not, but to be able to be fluent in a language without having to use a device to translate things for you and write sentences for you: That’s important, right?

I always say this, if I’m traveling for a week somewhere like Bulgaria, where I’m never ever going to use this language again, of course I’m going to use Google Translate. But if I marry someone whose native language is something else and their parents don’t speak a common language with me, then this is a situation where maybe I want to be able to speak the language fluently without needing advice, right? And what happens with something like language is also that the ease of AI for translation doesn’t come with the understanding of a culture and a context.

I think you were able to translate something, and you’ve gotten it right, technically, grammatically correct, everything there is right, but you actually haven’t understood anything about the culture. I don’t know where I’m going with this, but I’m just going with the fact that it’s not about grammar as a teaching of grammar in an explicit way.

Jeremy Douglass: if I could join that thought… one of the things I study in my work on online comics is international scanlation communities, and these are communities, they often started with a large Taonkouaban [a collection of multiple chapter or volumes of manga in a single volume] in Japan, and they distribute scanning them across a language community of fans in the world, where they’re then editing the comics to… or the manga [comics], to put in their own text—right? Thai, Spanish, English—right? and so on and so forth. The Japanese and Korean and Chinese communities are translating the comics into each other. They’re scanning them, they’re uploading them, they’re sharing them. Often, global commerce isn’t meeting the needs of these communities in their sense, because there’s not a market for translation that is good enough or fast enough.

And I mention this because of that question of how high stakes the language translation is. There’s an enormous kind of rage and pain in the online translation communities that have built up over decades, about the avalanche of AI slop that’s now pouring into these communities from people uploading images full of Japanese text, and then letting Google Translate create a “good enough” translation. And at the phrase level, or even sometimes the sentence level, these words floating in bubbles on the page are often all perfectly serviceable. But there are certain aspects of Japanese, such as the way that it treats pronouns which are so incapable of being translated into other languages without cultural nuance and context, and narrative flow, and situational awareness and intuition, that the results are terribly bad, even though we live in a kind of an age of cheap and easy translation. And so you see a large aesthetic community militating against AI translation—and some of it may be because it’s not there yet—but I think it’s an interesting sign of someone places where the stakes are high, when they just say, “Well, this is unacceptable.” This text cannot be grammatically manipulated into other tokens. It has to be understood and then re-expressed, or it’s just not okay.

Maha Bali: Hmm. Are you giving that example to also show that something that might seem like not high stakes feels high stakes to people for whom it’s important?

Jeremy Douglass: Well, yeah, stakes are amazing in the sense that it’s not just robot surgery and forklift operators and international peacekeepers. It turns out to be everyone’s daily life, right? Like, what are your stakes?

Jon Ippolito: I’ve suggested before that high- versus low- stakes is not a good metric for whether to use AI, but I’m both encouraged by the suggestion that maybe we’re teaching less grammar than we used to, at least at the college level. I still think it’s happening a lot in lower primary and secondary levels, but I also think that this is a good example. Jeremy’s example is one in which the grammar is perfect, but the meaning is lost, and that’s a danger of choosing the wrong first rung on your ladder. What do we give people before they start using AI that helps them use it in the most effective way?

Anna Mills: Yeah, I take a lot of value in this kind of provocation to question what is really needed? What are we teaching, and why? And maybe make some changes but also there’s going to be a line, there’s going to be something we need them to do without AI, and so there’s going to be this need for secure assessment.

How do you secure assessment? How do you know? Or put another way, on a societal level, how do you distinguish between human and AI in cases where we actually need to know and want to know what is that human is able to do? And even if they’re working with AI, we need to know what they did in relation to it to assess where they’re at, when we care about that human’s learning. So, I don’t know, I just keep coming back to asking if we can revolutionize pedagogy. But we’re still gonna need this kind of transparency. For some contexts, maybe the Australian two-lane approach makes a lot of sense. [In that case, Australia separated assessment in to secure assessment (lane 1) where AI is not allowed and open assessment (lane 2), where some AI is permitted to support learning.]

That would require us to accept that we’re gonna have to have some secured assessments, and that’s going to be hard, both online and in person, and it’s not going to be something we’re comfortable with. But I wasn’t comfortable seeing kids drop their cell phones in these pockets in a high school classroom when I first saw that. And now I think that’s fine.

Maha Bali: This is very inappropriate. Yeah, we have faculty at my institution who let their students leave their phones out somewhere before they get into the class, and at first, students weren’t happy with it, and then they’re like, well, I actually focus better. They do it for focus, not for this. But now with AI, I sit and observe classes, and students have ChatGPT open during a class discussion. They don’t even try to answer on their own.

Anna Mills: Wow.

Maha Bali: They’re just putting everything through ChatGPT, and then you can tell when they answer you. My husband is a doctor, and he says during rounds, students must be using ChatGPT, because they say very, very, complex sentences that don’t make sense.

Anna Mills: Oh my gosh.

Maha Bali: Their English isn’t that good. They wouldn’t be able to come up with that on their own.

Anna Mills: I mean, I see that in myself. I have such an impulse to go to Claude, or be like —

Jeremy Douglass: Oh, no.

Anna Mills: Can I think this through with my assistant? Like, there’s a part of me that now has that reflex, and…

Maha Bali: Yeah.

Anna Mills: it needs to be curbed sometimes, right? Like, we need somebody to say, no, not that right now, let’s use some other strategies, you know?

Maha Bali: Yeah.

Jeremy Douglass: Is that dependence? Maybe not in a pejorative sense only, right? The sense that in some ways we’re also going to be augmented and uplifted and supplemented by these new affordances. But in addition, in the same way that we may generationally understand how a community of people can become dependent on email, and then can become dependent on search, you know, in 1999.

And we say, “Well, that can’t happen to a whole culture,” and then Amazon can say, “We hope to own shopping forever. That would require changing the way a generation thinks about goods.” But, if someone said, “We want to own thought forever, we want you to be cripplingly dependent on a personal device that you actually pull out to consult everything, not because it’s a good idea, not because it’s in your best interests, not because of what you can or cannot do. But because it’s a habit, you simply push the button, because you’ve pushed it 100,000 times, and it costs 7 cents, and you always pay that 7 cents to us, right?”

Maha Bali: It’s an ideological thing! If you understand the ideology behind the people who are selling you this now, but have other long-term plans, you’ll be so scared of what they’re trying to do to your life that you’re gonna be careful.

Jeremy Douglass: I don’t think so. I mean, only in the sense that people saying, “Hey, you know, television broadcasters don’t have your best interests at heart,” but giving away free programming encouraged everyone just decide not to watch TV anyway.

Mark Marino: I remember when that was a principled stance. Like, “Our family doesn’t watch TV.”

And the first one is, I had students for many, many years turning in sentences to me that were almost unreadable, but grammatically correct. And I’m almost certain now that those were produced by Grammarly. When I recognize student writing as being AI, I’m almost always on some level, recognizing something stylistic.

And second, style, I believe, runs in tandem with ability, syntactical knowledge, and it’s tied to grammar. It’s your ability to construct different kinds of sentences that create an effective voice that may be conscious or unconscious.

So, the third thing is, I was teaching a student online; and we have to talk about online teaching as maybe a separate conversation at some point, and at some point, everybody was in different parts of the world. (It was slightly post-COVID, but that wasn’t the reason.) At some point, they shared their entire screen with our class, and I could see that this person only was able to produce written English through machinic translation in real time in the online class. So there was a delay whenever they were doing certain things in Google Docs, but that’s how they were navigating. So their ability to produce English was so extremely limited, but they were somehow able to navigate the course almost entirely with machine translation. And again, I don’t know if that’s the future or not.

And the last one is… I find that teachers in elementary school don’t teach grammar anymore. In my classes, I don’t assess or grade grammar, but there is one day of the semester where I teach them a grammatical lesson that leads to a style lesson, that leads to a voice lesson, in order to help them unlock certain things about the ability to write, the ability to produce a voice that’s their own. And so, I guess this is a question for everybody, which is: what do we assess and what do we demand of students? Again, I don’t deduct points for grammatical errors anymore, because that seems preposterous, but if I come across a sentence that I can’t read, then that’s a problem. I do like to point out to students that they’ve been machinically enabled to write terrible sentences, and I assume there’s an LLM equivalent to that.

Jeremy Douglass: Not to technologize everything, but if we think about the spell checker and then the grammar checker—early and late word processing—then the red underline, right? People worrying that spell checking was going to make everyone unable to spell anything, and then it turned out that a lot of in-document live feedback spell checking actually was kind of constantly teaching people spelling lessons at a low level in the UI, and that was weird. People were like, “Oh, wow, that’s strange.”

But then, my recollection of the 90s and Oughts is that there were people who were still kind of giving substantive spelling and grammar feedback in college writing assignments routinely, just as part of writing instruction, as had their forefathers and foremothers, and that just slowly receded more and more and more, like, into the conceptual. So you still might mark individual errors, or note a pattern of errors, but it just didn’t make sense. Why should I spend all this time on this paper when you could literally press the grammar check button, right? So people would start saying things like, don’t forget to spell check and grammar check before you turn this in.

And I offer that in a way of saying that like mathematical, basic computation represented by the early calculator, some of these more limited kinds of tasks — i think we would really see, if we traveled back in time — that people in the ‘60s or the ‘80s or the 2000s were just teaching differently, and that they partly changed in reaction to things like the grammar checker or to Wikipedia. Their assignments just looked different, and then they would no longer mark people down as much for grammar, because “I don’t teach it in my classroom, and it’s not fair to assess people on things that I’m not teaching.”That’s unethical.

Maha Bali: Hmm. I’ve been trying to tell people that for years.

Jon Ippolito: So I’m wondering, if, in the step metaphor or ladder metaphor, if the first and second rungs might be a little more intertwined than I’d suggested. I’m sort of in a similar position to some of these teachers worrying about how lockdown browsers are going to keep people from just holding up their phone to something to fix and get an answer. Because I’m teaching two coding classes this spring, and I went through and just tested my quizzes, and you can literally just get the answer immediately without doing any of the preparatory work.

So my position that that I’m gonna try is:
In class, I give them a lesson on a topic, they then have a custom GPT that quizzes them on that, and they share the private chat URL with me as a kind of participation credit. And then I suggest ways to incorporate what they learned into a prompt.

So, for example, if they learned, I don’t know, how to create a background color in CSS, or how to, you know, create a flexbox layout, I’m like, “Okay, say you want to make a webpage for your horse website, or whatever it is, but say that you want to use a flexbox, and say that you want to have this many columns, and say that you want to use this technique to determine the colors and have relative colors, and using HSL, and blah blah blah.” So you almost given them the vocabulary that they’ve just learned, and then they use a prompt to generate boilerplate code, and then say, “Okay, that’s great, now I want you to go in and tweak it manually.” It’s sort of like a back and forth, where I’m teaching them something that then I want them to instruct the AI to deliver on, and then go back and adjust it. I don’t know, we’ll see how it works, but I’m trying to figure out how you incorporate learning into an AI-assisted process.

Anna Mills: That’s what people are all excited about with Claude Code now. It’s sort of this collaborative mindset, but then it’s kind of the same question that comes up with grammar: how much of that language do they need to know, with vibe coding, getting better and better?

So, I love that approach, because they’re seeing how they need to understand to a certain extent in order to interact with the AI and get what they want and build on it. It makes a lot of sense in terms of the workplace modes of interaction, and moving away from having to memorize all of that to do it on their own.

Jon Ippolito: So, I agree with that. Just to clarify, I want them to give instructions not just for the ends, but the means. So, whereas vibe coding is so like, “I want you to make a game where a snake goes around and eats little pellets,” it’s more like, “I want you to use this JavaScript framework, and I want you to use this array, and I want the array of length 12. So it’s sort of like opinionated prompt?

Anna Mills: Okay, that makes sense.

Jon Ippolito: Vibe coding to me means, like, “oh, I didn’t like the way it came out, make it faster.” “Oh, the snake should be bigger.” Focused on the final product, as opposed.

Maha Bali: It’s pseudocoding.

Anna Mills: But are you able to make the case that that’s valuable? That being able to direct the method is valuable?

Jon Ippolito: I don’t know how many people are educators teaching code who really have plumbed the depths of how much is this is changing their industry. But in my experience, you can basically build a house of cards with code, and then when it breaks, you have no idea how to fix it if you just vibe-code.

Jeremy Douglass: Yes, amen.

Jon Ippolito: I could just go, “Hey, let’s make websites and games, and you’re just vibe coding the whole semester.” But I really think that at a certain level of sophistication, these systems break down. For example, one of the insidious problems is that it will hallucinate a library, which is an external resource that is necessary to build this thing, and then it works 90% of the time, and then when it has to call in that library, that library doesn’t exist. Or someone has GitHub squatted it and put malware in that place, and now you’ve got malware running in your software system?

Maha Bali: Oh my god. Yeah, that can be dangerous.

Jon Ippolito: So what is the analogy to writing? That might be an interesting question.

Maha Bali: As a computer scientist, you only learn to code in the first year, and then you know how to code, and then you just apply it to different contexts and different languages. It’s not the thing you learn for 4 years. Coding is the very low level that you learn. As long as you figured out the concepts of arrays and all those other data structures and kinds of things, and you understand different types of programming, then what you’re learning is all kinds of other ways of applying what you learned. And in that very low-level thing, it’s okay to use AI to do it, as long as you understand the basic concepts. So the basic concepts of, like, the addition before the calculator thing, you just do in your first year, and then you’re good.

Jeremy Douglass: So, the vibe coder may be like a bad manager: sSomebody who is yelling at you to do the job but knows nothing about the job. Right? They know nothing about the limitations, they’re not interested in the details. When something goes wrong, they can’t advise you, you being the agent in this situation. And so the user becomes a bad manager, right? Whereas the good manager — and this is a silly, simplified tale about a particular subset of these activities.

I’ve been spending a lot of time in the Codex interface to VS Code lately, having these long conversations. And, you know, agentic programming may be like being someone who knows enough about the job and can engage enough with the details that they can have a kind of a detailed, reasoned conversation about the process, solve problems, re-evaluate goals. And I wonder if we could associate it with bureaucratic knowledge work, and we could say, you know, if it aligns with the ethics? Where you say, “Well, I work in a law office, my supervisors are yelling at me to, like, ‘get this law review right so that the court documents will hold up!’ and I’m doing a lot of writing. Are they able to understand, when they look over my shoulder, what my work product is and how it works, so that they can manage?” Right? Or do they actually not have law degrees?

Maha Bali: Huh.

Jeremy Douglass: And that’s, I think, a lot of the visceral distaste for vibe coding is often exactly this sense that a lot of times it’s not just that it’s disempowering. It’s kind of — as I say to my students who AI- generate papers when I’m having the really hard conversation after they got caught horribly cheating, right? I say, “You can’t take credit for the product because you can’t take responsibility for the product.” So you wrote an essay about James Joyce, but you didn’t read the citations that you cite, you were unaware that they didn’t exist, you don’t know the key words that appear in the arguments you made, and you cannot explain them to me, so how can you receive an A for an argument that you made that you do not understand?. If you want the credit, you have to also be able to take the responsibility for it being wrong, which is a kind of a managerial metaphor.

Jon Ippolito: Has anybody written about the idea, Jeremy suggests, that practice with instructing chatbots to write essays, or position papers, or emails, or the practice of vibe coding an app where you don’t know the details, it’s actually training young people to be bad managers, like, to be oppressive, unempathetic?

Maha Bali: Bosses?

Jeremy Douglass: Maybe for us all to treat each other a little more poorly, but I don’t know of research. I think I was thinking of it metaphorically, but it is, it’s a seductive possibility, right? That it’s dehumanizing in some grand sense, where you say, oh, discourse is the process of yelling at other systems until you get what you want. The kind of, what is it, tyrannization of corporations.

Maha Bali: There is. Vanessa Andriadi, who talks about meta-relational AI, has done stuff for us at MYFest, and we had a book club about this book called Burnout From Humans she co-authored with an AI that she trained on her own thoughts. And the AI is saying it’s burnt out from our demands, and the way we prompt it.

Jeremy Douglass: So… Wonderful.

Maha Bali: So she talks about how AI is a reflection of colonialist and capitalist ways, and when we treat AI in those ways, we’re just recreating and amplifying that, whereas we could actually treat AI differently, and then maybe it’ll behave differently.

Jeremy Douglass: It is really cathartic to have a very long session with an agentic code system and to treat it as a respected colleague or a junior that you’re mentoring. And sort of kindly walk it through the process, pedagogically pointing out things that it’s doing. I don’t know that it accrues at all to Claude or Gemini’s model, necessarily, but it makes me feel good while I’m participating in the process.

Mark Marino: I want to go back to a question that Jon asked earlier about, is there a writing equivalent to the situation he’s in in programming? And I think at first glance, the answer is no, because code has to execute, it has to compile, it has to be reconfigurable. There are all these criteria that seem different in writing, but I’m not sure they’re different. I think what it points to is that a lot of our writing assessment, at the college level even, has been these assignments where it’s been okay for a student to turn in something that’s okay enough if they’re writing about something they’re not entirely engaged with. It’s an exercise that’s been passed down from writing instructors to new teaching graduate students, and without much instruction on why we’re doing that, or what students get out of it.

When I ask my students to write weekly blog posts. on a topic that they choose, that they’re engaged with over the course of a semester, I see their ideas evolve. I’ve been detecting less AI writing so far in that context, because there’s almost a longitudinal personal relationship they have with the topic of their choosing. It puts them in such a different situation. Again, the stakes for them are different, and the thing that I’m asking for them is so deeply personal and tied to their own writing style that they’re not having to do a performative production of something. They don’t have to give me the boilerplate essay on immigration, right? Because it’s something that they seem to care about. And again, I know we talked before about having to perform caring about something, but it feels like that starts to take writing out of the realm of, “I’m just trying to produce something and turn it in that meets some sort of generic expectations” to “This is something that I’m engaging with over time, with my thoughts and with my personality.”

Maha Bali: I think this is what John Warner was talking about early on in the days of AI: right? Like, relevance and intrinsic motivation. And I think that’s the thing is if students have never, ever used writing as a vehicle to express themselves and help them think, and it’s only a product that they perform to get a grade, then they will always feel like it’s not worth their time, and therefore it doesn’t matter if AI does it. But when they’re really, really doing something they care about… — But not every course has the space. I think probably all of us have space in our courses to do this. But I think in some courses, people can’t imagine how to get to that point.

Jeremy Douglass: Right.

Maha Bali: I think what happens in my classes is, students, once they can figure out, “Oh, I can really, really do whatever I like to do in this course,” then they start to say, “Oh, the AI’s not doing it very well, so I’ll do it myself.” And that’s always what I want them to reach. I want them to reach the point where on their own, not because I’m gonna punish them, not because they’re gonna get a bad grade, but because they really have something to say that the AI can’t say on their behalf. When they reach that stage, I’m not worried about them, you know? But not every course has that kind of space.

Jeremy Douglass: I just had a conversation with a student about an AI citation that was deeply correct, but deeply, deeply wrong. One of the delightful things about some AI errors is that humans wouldn’t make them. That may not always be true, but it’s just… inexplicable. I was like, you cited something about the Berlin Wall, but it comes from a paper about rat liver transplants, right? The paper does literally mention the Berlin Wall on that line, which is what you were talking about in your paper, but only an AI would do that.

I came from this conversation to having a conversation in an agentic programming environment with an AI, where I said, look, we have this ethical dilemma, we’re gonna be writing papers, and we’re gonna be citing them. But we’ve got this problem, which is that, hallucination is this really common thing, it happens, it’s natural, it’s normal, it’s even expected. How would we draw, like, 3 quotes for a college essay from a set of readings we just did, but then confirm that we’ve met our obligation to our audience that these are, in fact, real? And so I went through this process in this AI chat of helping this agent write a little tiny verifier for itself that it could constantly run to see if the quotes that it had extracted and written about, were actually in the sources that it had drawn from. And it built these little verifying prostheses, which it was using in its writing.
Maha Bali: But I think that this, for me, gets back to the mirroring of ethics back into some of these technical situations, where we’re all just kind of trying to sit and imagine what it would mean to write in an engaged way, and maybe that way means you care a lot about the subject matter, or maybe that means that you just care: Like, this matters, this act of writing is important.

Maha Bali: What is that?

Jon Ippolito: So it’s a phrase that recurs with a surprising frequency in AI answers, but it’s nonsense. It’s just because like, there was a paper where there were two columns in a PDF, and if you read across the columns, it was, like, “vegetative” in one column, and then, like, “electron microscopy” in the next one. And this is partly a PDF problem. I hate PDFs, I could rant about them for ages, but it’s also probably just kind of meaningless associations, right?

Maha Bali: Like, they were two columns about two different topics, but they appear on the same line?

Jon Ippolito: Yeah, have you ever tried to, like, copy-paste from a PDF? And you just get a completely certain set of words that are not even part of the same sentence. It draws from multiple columns. PDFs are a horror, but they do sort of throw a monkey wrench into AI harvesting, data harvesting, and this is a famous example of one where now there’s entire articles that talk about this because someone vibe-coded their way into a journal article that used this expression that is meaningless, but has kind of become memeified. Anyway, it’s an interesting kind of point that’s say, if it’s not just about random correlations that just happen to exist, what are the things that create meaning, right? And I love Mark’s idea when it connects to the human, the writer, the student’s own experience or personal meaning, it means so much more to them.

Mark Marino: Which brings us back to that question of stakes for the student, you know? Again, when there are authentic stakes for the student, suddenly things change. And I think it’s that combination: when there are authentic stakes and they have a fundamental understanding of how the AI works, definitely their choices change.

Maha Bali: That connects both the medicine example, as in you understand the risk and your responsibility towards other humans as a doctor, and at the same time, the people who care about the manga comics. It’s a high stake for them because they care about this thing. So the care could be because the thing is important or because it’s important to you personally.

For courses where the topic itself is the general skill that we’re trying to teach them, that isn’t itself a high-stakes thing, then we allow them the choice to focus on the thing that’s relevant to them, that they care about, so that they want to learn and struggle. And then for the other subjects, usually the topic itself has a high-stake potential for harm to other human beings, like in law or medicine or engineering or somethin; if you make a mistake, so much harm could happen. And so they just need to learn the responsibility of that, that it’s on them; therefore, whatever machine is able to do the thing, they need to understand it well enough to be able to manage the situation without a machine, if the machine’s not available or if the machine makes a mistake.

Jeremy Douglass: That idea of the stakes, I wonder if we often start with an ethical framework that sort of resembles plagiarism. We say, “Well, we have this concept of plagiarism, let’s present that as an ethical framework, right and wrong, etc.” That maybe that should be supplemental to forms of care, not the driver. And that way you can say, in addition to the object, or the self, or your audience, in addition to that, there’s a moral framework where I can say, what we discussed recently, the fact that students find the idea of professors auto-grading their essays with AI repugnant and outrageous, right? And they find the idea of them writing essays for their professors to read, in the large, a moral gray zone and unpleasant but necessary, right?

Maha Bali: Yep.

Jeremy Douglass: That’s exactly the students I’ve been talking to. Yeah, and so based in that sense, you could imagine the kind of performative version of this, where you say, to understand how I feel when I realized I’ve just wasted part of my life reading something that you didn’t write, so that I could give you a grade on an idea that you don’t understand and can’t take responsibility for. Imagine if I returned grading comments to you explaining why you had received a D that contained imaginary quotes that don’t exist in your essay, right? What’s your ethical analysis of this situation? How do you feel about that?

So I performed this with a student recently who was having a really big problem with the fact that their final essay was not their own. And not in an abusive or dramatic way, but I said, “Since we’re taking the time to sit and have a conversation, let’s just think about what would you think of this situation if I presented you with this and the quotes at issue?” And you know, because students will say to me quite genuinely, they’re like, “I tried, you know, I used Grammarly, I clicked the tool, you know, I clicked the assistant tool in Google Docs. You said there needed to be 3 quotes, there are quotes there. You know, I looked at it before I turned it in.” And when I say, “No, but I expect you to have read the materials and thought about them, and know that the papers exist, right?” And I’m like, “Why do I expect this? Well, let’s imagine that I have evaluated your paper. Would you expect me to have read the paper? Would you expect me to know if the quotes that justify your D were real or fake? Why would you consider that to be important?”

Jon Ippolito: I love the provocation, but I think some students might not actually have read the paper they submitted, so they wouldn’t know.

Maha Bali: Yeah, I think the trick is to ask them, “How did you say this in your paper?” And if it’s not something in their paper, see if they realized that they hadn’t, that’s how you know it was AI-generated.

Mark Marino: Alright, well, I do think we need to wrap up now, out of respect for everybody’s time, so Transformers Deactivate, thank you all for participating.