I wouldn't say Pangram is broken, but I would say that it's brittle
It seems way too easy to provoke Pangram into contradicting its own results. So let me get to the nut of this thing before I do my usual meandering.
Recently, someone accused me of using AI to write this old post from about a year ago, specifically highlighting the section about Ta-Nehisi Coates. I replied by saying that the post contained no AI writing; nothing I publish has been written or edited by an LLM. My accuser proceeded to say that the AI detection tool Pangram flagged that portion as 100% AI written with high confidence, a result I replicated. (Pangram report.) As the man who wrote that piece, I knew that I wasnât guilty, but I also knew that wasnât going to fly with other people.
I was sufficiently annoyed that I paid for a Pangram license and found that, when I entered the entire essay, Pangram cleared it as 100% human written with high confidence. (Pangram report.) That is to say, the essay of about 5,000 words was declared 100% human written with high confidence even though it contains the section of about 300 words which was declared 100% AI written with high confidence.
Subsequently, Iâve found that I can produce the same sort of inconsistency in the opposite direction - I can break the section that was called 100% AI into pieces that are then called 100% human written. (Pangram report of an example.) So pieces that Pangram call 100% human written are part of a larger piece that Pangram calls 100% AI written which in turn is part of an essay that Pangram calls 100% human written, a Russian doll of contradictory results.
If nothing else, thereâs a fundamental inconsistency here: it doesnât make sense for a system to say that it has high confidence that an essay is 100% human written but also that a piece of that essay is 100% AI-written, again with high confidence, and also to say that the piece itself has subsections that are 100% human writtenâŠ. These claims are mutually exclusive and must necessarily erode our confidence in the instrument.
Experimenting, Iâve also found that I can pretty reliably get Pangram to declare large sections of text to be AI-written or human-written depending on the larger textual context I place those sections in - that is, I can take AI-generated text and embed it in human-generated text and have Pangram declare it 100% human, and I can take human-generated text and embed it in AI-generated text and have Pangram declare it 100% AI. Iâve also found that itâs pretty easy to induce false positives, that is, to write texts myself that Pangram identifies as AI-written.
All of this makes me hesitant about the Pangram tool, and in particular the way many people use it, as a one-shot gotcha machine. In a purely self-interested sense, the most obvious thing for me to do would be to point to this 100% human written outcome for the whole essay, say âSee???,â and get indignant about having been accused. But I think thereâs a lot to talk about here, and I think that people who write for a living have to have these conversations. These tools are being used right now, and with consequences, and we have to work this stuff out.
The percentage meter seems clearly broken
One frustration of discussing Pangram results is that many people seem to think that the percentage thatâs expressed is the confidence level - that is, that a 100% AI result is saying âwe are 100% sure that this text is AI generated.â But thatâs not whatâs being measured there. The percentage is supposed to indicate what portion of the text is suspected to be LLM written. The degree of confidence is flagged there underneath âAI Generated,â although annoyingly the flag only appears when confidence is high, with no âConfidence Lowâ flag that ever appears, in my experience.
This is all fine and good - the percentage AI generated and confidence numbers are different things that should be easy to interpret. The first problem is that many, many people are clearly interpreting â100% AIâ to mean â100% confidence.â The second problem is that the percentage meter just does not appear to work!
Anecdotally, a really suspiciously high portion of all Pangram outputs are 100%, whether 100% AI or 100% human. This is suspicious not only because a lot of LLM writing that gets published is likely hybrid, integrated with a writerâs own words, but also because the basic processes through which detectors like Pangram function seem likely to produce fewer polar outcomes. And a detector that spits out a lot of 100%s seems easier to break.
So letâs break the percentage meter
This is easy to do. Hereâs a paragraph Iâve copy and pasted from a piece I wrote in 2017, which I hope you will accept as given was not written with the help of an LLM. I have appended three sentences to the end of the paragraph that were written by ChatGPT, using the original paragraph as a prompt, highlighted in red. The human-written, pre-LLM section is 239 words, while ChatGPTâs portion is 71 words long.
The gambler's fallacy is when you expect a certain periodicity in outcomes when you have no reason to expect it. That is, you look at events that happened in the recent past, and say "that is an unusually high/low number of times for that event to happen, so therefore what will follow is an unusually low/high number of times for it to happen." The classic case is roulette: you're walking along the casino floor, and you see the electronic sign showing that a roulette table has hit black 10 times in a row. You know the odds of this are very small, so you rush over to place a bet on red. But of course that's not justified: the table doesn't "know" it has come up black 10 times in a row. You've still got the same (bad) odds of hitting red, 47.4%. You're still playing with the same house edge. A coin that's just come up heads 50 times in a row has the same odds of being heads again as being tails again. The expectation that non-periodic random events are governed by some sort of god of reciprocal probabilities is the source of tons of bad human reasoning - and journalism is absolutely stuffed with it. You see it any time people point out that a particular event hasn't happened in a long time, so therefore we've got an increased chance of it happening in the future. But unless the process has some actual mechanism of correction or reversion, the mere passage of time changes nothing. Droughts do not make rain âdue,â losing streaks do not create future wins, and a long period without a crisis does not by itself make a crisis more likely tomorrow. The past can tell you something about the underlying probability of an event, but it cannot compel randomness to balance its books.
So, ideally, Pangram would flag this as 23% AI generated; that is, indeed, what the exact percentages are by wordcount. Hereâs what actually happens (report):
So, yeah, not great. That passage is three-quarters human written, dominantly human written, and yet Pangram says with high confidence that itâs 100% AI.
I know there are a lot of people who would endorse a âone dropâ rule with LLM assistance, but Pangram itself is saying that its technology operates a certain way, and it doesnât appear to operate that way; certainly you can break it quite consistently. Pangramâs own examples suggest that it can return mixed results with a specific percentage. Look, they even give you human vs robot %: And yet in dozens of trials Iâve never gotten a mixed report like this. Theyâve all been either 100% human or 100% AI, even though Iâve been Frankensteining human and LLM writing together constantly for this exercise.
For whatever reason, Pangram seems to want to find 100% results, even though users are clearly trusting it to detect what portion of a text is LLM generated. This all seems quite bad to me! This technology is being used in a way that has severe professional consequences for some people. It should work the way itâs supposed to work; it should offer a âpercentage AI writtenâ thatâs actually a reflection of how much of the text was LLM generated, and if it canât do that, they should drop the percentiles and just give a quantitative, numeral confidence figure.
Hereâs where we can do a âyou try it at homeâ exercise
If you have a Pangram account, take human written text, append LLM-generated text, and see if you get a 100% AI result. With enough naked ChatGPT text youâll get there. And if youâd like, you can try and go the other direction - if you have a long enough human-written piece, you can stick LLM-generated text somewhere in it and see what you get back. I think youâll find that itâs easy to pass such text through as 100% human. This additionally means that you can easily get specific passages into texts flagged as 100% human or 100% AI, again fairly consistently.
I just donât think thatâs ideal, and it disturbs me how cavalier some people are about it.
Conflicting results and what to do
I have no idea what you do with conflicting results like these. I suppose many would say that the much longer overall essay has many more datapoints and the whole-essay analysis should thus be considered more authoritative than the 300-word passage analysis. But I can assure you that this wonât stop people who donât like me from declaring some sort of victory!
This is what I worry about with this whole controversy - weâre very likely to end up with everybody confirming their priors. People who donât like me will say that the short passage analysis proves that that portion was LLM written, while people who do like me (I assure you, a few such people exist) will say that the whole-essay analysis has more data and is obviously the right level of perspective. Nobody will be convinced of anything.
The nightmarish possibility is that people will start cutting up essays that pass Pangramâs test into small pieces, looking to find parts that get flagged as AI, so that nobody ever feels confident about anything. If 50 words is not sufficient to have confidence in a result, you shouldnât be able to put 50 words in the detector.
Several people on Notes said some version of âWell of course nothing is 100%, but the more data that Pangram has, the more accurate it will be, and that 300 word chunk that got flagged just wasnât long enough.â But the minimum amount of text Pangram will accept is 50 words, so it should work on 50 words. Iâm obviously not equipped to discuss the technical specifics here, but since I heard some version of this many times, I want to say clearly: if a particular selection of text provides insufficient data to get accurate results, you shouldnât be able to use the detector on that selection of text. At the very least, if the confidence in a conclusion is low, that low level of confidence should be prominently displayed in the results.
Even a very small error rate will produce false outcomes regularly
Given enough repetitions. Iâve written more than 5000 blog posts, newsletter entries, and freelance essays in my career; I canât be sure because my old fredrikdeboer.com blog was lost years ago, but it may be closer to 6000. But letâs narrow down: since ChatGPT was released to the public in 2022, Iâve published about 1200 posts here, and those are the ones that are relevant.
Now, one frustration I have with this discussion is how often people quote error rates that come straight from Pangram, which doesnât seem particularly rigorous to me. But letâs be generous and use the companyâs own numbers. Pangram reports a false positive rate of 0.19%, and claims something closer to one in ten thousand on academic essays. (That is, the kind turned in by college students, presumably Pangramâs most common usage.) At those rates, 1,200 posts of mine yields an expected two or three spurious âAI writtenâ verdicts - small, but not zero.
But! Even that considerably understates the exposure to potential error, because ânewsletter postâ or âessayâ or âdocumentâ isnât the only unit these tools actually operate on. Pangram returns sentence-level âhighlightsâ and a separate AI-assist judgment, so every paragraph becomes its own opportunity to be wrong. Thereâs thirty paragraphs in the post that got flagged. Letâs say my average is half that. At that rate, my post-ChatGPT archive represents something like 18,000 chances to get flagged, and at one in ten thousand you would expect a few passages in there to come back marked as machine-written no matter what I did or didnât do.
This is why a âone drop ruleâ doesnât work. An accusation should come from a large-scale investigation of a writerâs corpus. But thatâs hard and time intensive and, most importantly, doesnât fit with the desire to own oneâs enemies online, which drives a lot of this stuff.
And, of course, âFreddie deBoerâ is not the right unit of analysis, but rather the whole community of people who write. That community produces a lot of text, and so even at 1 in 10,000, there will be an immense number of false positives being generated, if youâre testing everybody. And as I suggested above, some peopleâs natural writing style is more likely to be flagged than that of others, even with no LLM assistance. Whatever Pangram keys on is going to get keyed on over and over again; thereâs inevitably a bias against whoever happens to write near the boundary.
My own writing doesnât appear to have that problem, but other peopleâs will. Canât you imagine the outrage if itâs eventually found that, say, Chinese international students are more likely to produce false positives with their writing, thanks to underlying second language consistencies? Those consistencies exist; back when I taught Chinese undergrads at Purdue, they would very often write ânow a dayâ instead of ânowadays,â including when they were writing live and unscripted in the computer lab.
It all worries me. Despite Pangramâs reputation for avoiding false positives, itâs very easy to induce them. A lot of Pangram enthusiasts boast that the system is very unlikely to produce a false positive-
Comments
No comments yet. Start the discussion.