It's just one damn thing after another

RSS
The enduring je ne sais quoi of LLM writing#

You’re trying to manipulate me by doing easy things that you know will work, and the fact that it slightly does work just makes it even more galling!

- Scott Alexander

 

Most people know LLM text when they see it, so why are detection programs so ineffective? And why have LLM coding abilities run so far ahead of their (perceived) linguistic abilities?

In 2019 Sarah Constantin pointed out that if only conscious, focused logical thought can detect a bot, maybe some people will become more aware of when they’re thinking actively vs not. The examples in the article (from the just-released GPT-2) are worth looking at, because the magic eye technique for detection works better after you see how the trick works, and what the LLM was doing back then was... a little bit more obvious:

We believe this project is the first step in the direction of developing large NLP systems without task-specific training data. That is, we are developing a machine language system in the generative style with no explicit rules for producing text.

Skimming through it looks fine. Sarah points out that only by concentrating on the logic in the sentence do you realise that a rule for producing text is something like "if the next word starts with a vowel then use an an" - totally different to task specific training data (a collection of books on rocks). The next year GPT-3 was able to generate full paragraphs without obvious logical inconsistencies inside them, and the glitches in the matrix moved to higher levels over time until now it's easy for even experts in an area to be quickly lulled into a Gell-Mann Amnesia state. But the tell is always there, the model writing cheques its logic can't cash and its ephemeral ontological categories collapsing on their first contact with a counterexample.

Despite the difficulty in spotting their lack of a consistent world model, LLMs are failing the Turing Test hard in a way that humans can detect better than AI text detectors. In an ACL 2025 study1 five people who used LLMs heavily in their own writing judged which half of nonfiction articles were produced with frontier models, and the majority got 299/300 right. This was better than all of the automated detectors tested except Pangram, which matched them at 99.3% but wrongly flagged 3 of the human written articles against the experts' 0. Articles rewritten by the models based on expert feedback fooled the detectors more, humans got all 30, GPTZero 14, Fast-DetectGPT 7, Binoculars 2, and RADAR none.

The expert explanations suggested that structure, excessive formality, thin originality, an oddly uniform clarity, and recurring AI vocabulary were the tells. A year later the ones I currently notice most are contrasting negation, stilted metaphors, and meaning signifiers: Here’s the thing, that sentence was empty calories, not a balanced meal. When writing code comments LLMs love to relitigate history (no theory of mind to look at context as a bystander), develop their own jargon heavy dialect for discussing changes, and pontificate on unimportant aspects (low insight to word ratio is a symptom of not understanding meaning). When writing short stories I would characterise the tells as emotional over-signalling and overall aesthetic cowardice. And generally just acting like a complete simp in conversation.

The two best-known slop word lists disagree with each other, each one flagging whichever kind of writing it happened to be built from: one says academics are the worst offenders, the other says Jane Austen. The obvious word tells are squished over time, for example when was the last time you saw a model say delve even though that was very common a few years ago. Statistical detection degrades with each model version but the expert reader's judgement does not2. What separates machine from human now is mostly sentence architecture rather than vocabulary, and humans are better at picking that up.

Human writers drop two-clause replacive (it's not a bug, it's a feature) sixfold when they switch from arguing to explaining. Across 392,000 words by fourteen pre-2021 essayists it runs at 73.6 per million but in 854,000 words of research papers it falls to 12.6. A 2026 frontier model doesn't make that switch, across 1.36 million words of its own technical prose the same construction runs at 60 per million, delivering the human argumentative rate in an expository register. Feed the same opening sentences to a small 7B base model before and after instruction tuning, and the argumentative phrase "rather than" rises 8.7×, the copular replacive 3.8×, the em-dash 20×. That suggests that rather than being lots of ad copy fed into pretraining (the majority of the training corpus for modern models is web text which has two-clause replacive at around 30/M) this argumentative form is being selected for in post-training.

The contrasting negation in “The fault, dear Brutus, is not in our stars, But in ourselves, that we are underlings” is elegant and draws the mind down from fate and the heavens to Cassius. But Shakespeare was the OG, if you try writing like that today the literati will call you a Bromidiom. LLMs don't care - they have an overwhelmingly reliance on cliche. Everything is a shadow, an echo, a whisper, a void, a heartbeat, a pulse, a river, a flower—you see it spinning its Rolodex of 20-30 generic images and selecting one at random, one of the techniques that nostalgebraist calls flashy eyeball kicks. RLHF training works well for coding, but here it rewards literary sounding descriptions above the model's ability to deliver them tastefully, as before the model still writing cheques its abilities can't cash.

A lot of greatness in creativity happens at the ends of the distribution, but a model optimizing for the most likely reward drifts toward the average. With more specific guidance from a high taste user the LLM has more options other than obvious rhetorical tricks. But if you have ever pushed a diffusion image model with more and more rules you'll see it improves for a while and then things get... weird. My theory is that you can't control the default ick at the same time as generating a clever story - constraints on the ick force prose subtlety which selects for folk wisdom style ideas that flow well. An idea generation planning phase helps, but the constraint of having to then write the ideas subtly without diluting them still ends up both softening the structural choices and leaking LLM tells into the output. I have much better results with a first "ideas" pass with lots of context and few restrictions to generate a smart but ham-fisted story, and then a second "style" pass which takes whatever story the first created and rewrites it based on a number of style rules. The stories that I make multiple course corrections to are still better.

Let's go back to 2019 again to say I'm from the future and this all goes horribly wrong see Neal Stephenson describing Autonomous Proxies for Execration - programs that would overwhelm the internet with thousands of conflicting automated lies about a target and fill the ecosystem with so much noise that the average human user is forced to doubt the veracity of any information, performing a sort of cultural inoculation against AI text. We are still trying to fight that, but AI detection depends on probing things which are expensive to imitate unless you understand meaning, which is something that humans still do best. But also - our brain is sculpted by reinforcement learning and runs a dopamine prediction-error-based algorithm, so y'know - we do often fail or wildly succeed with the same lazy shortcuts as an LLM. Their tics will continue to be squashed, but equally humans are going to co-evolve into communicating at a higher level denser LLM style because once you have network effects it works, and this alone will have as much influence on future language as our OG Shakespeare did.

  1. 1- Russell, Karpinska & Iyyer (2025), People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text, arXiv:2501.15654. (arXiv)
  2. 2- Yu, Yu, Liu, Chen, Zhang, Yu & Shao (2025), EvoBench, Findings of ACL 2025, 14605–14620. (ACL Anthology)
The ARC-AGI of history#

Writing something set in the future? Superintelligent computer says no, and don’t forget your scheduled chem spray in 2030. How do pros post post-superintelligence prose?

The Expanse doubles down on AI as normal tech. Fine - assuming the Anthropic Principle works for me, and I happily suspended disbelief until about the third mention of expert systems. Further out in the future I really admire the Herbert judo move of making the removal the setting (Butlerian Jihad and the replacement of thinking machines with human mentats), or Vinge dividing the galaxy into zones of thought and banishing superintelligence to the edge, but back in the near future we are mostly left with shuffling superintelligence aside with legal restrictions (the Turing Police in Neuromancer) or some sort of balance of power containment (The Quantum Thief).

Putting the future behind us, the current forecasting consensus1 is “stock up on canned goods and shotguns just in case, or use the current small window before AGI to coordinate a Plan A shaped solution to get us to a world of AIs pointed at each other in a Mexican standoff”. So thats what - Quantum Thief wins? The End. But actually how well does that sit intuitively, and does intuition even matter when facing a really unprecedented situation?

The AI futures model (AIFM) might be grossly and irresponsibly simplified to a basic growth rate boosted by a progress multiplier and divided by a bottleneck factor. Growth increases as more compute comes online, but soon the progress multiplier dominates as the AI improves itself, and pretty much every objection you could ever think of and some more besides have been bundled into the bottleneck factor. Despite this, the bottleneck does not scale proportionately with the growth, allowing an AI takeoff.

However, my intuition when I was writing Synthetic Control Measures and friends back in 2023 was that in fact we were likely going to get various forms of proportional “progress friction”. That still seems directionally correct today even after all the changes in the last 3 years, although with the amount of ink spilled since on AI I am sure this gut feel formulation has already been explicitly dismissed somewhere. No harm in another pass though, so let's sketch out what the mechanism behind this intuition might look like:

  1. 1.Verification becomes increasingly expensive. Efficiency gains and baking in increasingly detailed reality maps gives us better and better verification, but this gets outraced by the complexity of building the next even more detailed layer of reality verification. Verifier production becomes verifier limited.
  2. 2.Ignoring 1 costs more in slow-feedback domains. You can build features on a pile of buggy (i.e. out of touch with reality) software (vide Apple post 2013) at some extra cost. The cost of a pile of buggy theories giving you the wrong tensile strength for an actively powered building is much higher, and in such domains the more the capability must expand to meet the expanding needs of the capability.
  3. 3.Cost controls the rate of advance. Advances compound in areas with easy wins and easy verifications, and these diffuse quickly through the economy which pulls in more funding. Compounding advances in fast-feedback domains also increase progress in slow-feedback domains at a lesser rate, but because of 2 these are the first domains where ROI turns negative and causes a bust. Exposure contaminates the more productive fields into a mini AI winter, until compounding gains in the productive fields pulls money back in. As long as 1 and 2 hold this results in more boom and bust cycles.

Does this look like an Outside Context Problem for AIFM, or can we work it in as a prior and tweak a parameter somewhere? Where even are these new falsifiable predictions that we can all laugh at later?

I suspect that AIFM does account for this in a different way with its periodic capability slowdown multipliers, but I don’t have a good enough feel for the model to see if that works better than my assumption that the progress multiplier and the bottleneck are inversely related by friction from the verification process. What I do know is that in the absence of confounding factors a coupling between variables should be quite testable. Sadly in a macro forecast confounding factors abound, but lets not let that get in the way of some good ol predicticating for the next 5 years:

  1. 2027Trust in AI outcomes increases as frontier models become genuinely funny and persuasive rather than the unlikeable slopmeisters of today. Median cost per token of the hundred most popular models falls under US$0.80/M input tokens, and we see per-outcome rather than per-token billing become more widespread.
  2. 2028Absolute measures of AI ability are becoming increasingly decoupled from deployment outcomes, and relative measures between AIs similar to an Elo rating system develop, along with independently human “vibe based” measurement systems. US labs voluntarily sign up to a scheme to monitor compute use with a clause about future ratcheting, and a subjective capability cap is agreed between the US and China
  3. 2029The proportion of compute devoted to verification is now a separate line item on the majority of frontier labs financial reporting. US job losses in software development from AI reduce headcount to under 2026 levels. Open weight models are still only around 6 months behind the frontier, and OpenAI reduces spend on model training and new DCs, causing a debt crisis which ends Oracle in its current form. The voluntary scheme to ratchet compute use has still not had a measurable impact and is considered mostly ai-washing
  4. 2030The software R&D multiplier is still arguably in the 2-3x range but it is getting harder to estimate, and one of the major AI labs has either gone bust, been acquired, or completely stopped training new models. Job losses in software accelerate and begin in other fields, but it is hard to distinguish from a wider economic downturn partially caused by the first big pullback in AI spend. An anti AI backlash in the US causes the first legally enforceable limits on GPU use, and China independently implements something similar
  5. 2031A new AI developed AI paradigm becomes widely available and is more useful in robotics that LLMs. New AI labs emerge to commercialise this in different sectors of the economy. Parts of the US economy are growing 5-10% largely based on new AI developed technologies and productivity growth, while other parts are still contracting. The legal limits on GPU use are partially circumvented by the new paradigm, and there is little legislative appetite to draft anything to replace them.

This has now run a long way ahead of anything suggested by the model and is close to pure sci-fi, and we haven't even got to the mid 30s where superlinear recursive self improvement should be undeniable if I am wrong. Otherwise, my world is one where progress at the frontier stalls a few times a decade, where the only regulation which happens would be too little too late for under AIFM scenarios, where we don’t end up with clever cryptographic checks and balances, but despite that we still get away with our containment efforts and have a functioning economy involving humans for a long time to come. I’d love it if I was wrong and we got our act together earlier, because I could also be wrong about AI not being a near term existential risk.

Another consideration - is it a good idea to write about non perfectly behaved AIs at all, given it all ends up in their training data? The AI Hyperstition crowd says no, but I wonder whether writing more not particularly different good futures ends up strengthening the Waluigi Effect of their complement more than it strengthens the good. Perhaps a solution is to write radically different good futures which don't have evil complements in the literature, and give the AI more good options in a range of different situations. I literally summoned a deus ex AI to resolve a dumb-humans-cause-AI-takeoff crisis both because that seemed narratively more satisfying, but also in the hope that even when future humans screw up, the AI has an inkling from its training data that perhaps we are still worth saving.

1 - Of course once you move out from the easier eval questions towards the macro impacts that we care about here there is no consensus, but for simplicity I choose to narrow the field to the https://www.aifutures.org/ vs https://www.normaltech.ai/p/ai-as-normal-technology forecasts. They mostly agree on the early shape but become polar opposites later, so seemed like good proxies for the wider range of options, and I weight the consensus towards the former