The ARC-AGI of history

Writing something set in the future? Superintelligent computer says no, and don’t forget your scheduled chem spray in 2030. How do pros post post-superintelligence prose?

The Expanse doubles down on AI as normal tech. Fine - assuming the Anthropic Principle works for me, and I happily suspended disbelief until about the third mention of expert systems. Further out in the future I really admire the Herbert judo move of making the removal the setting (Butlerian Jihad and the replacement of thinking machines with human mentats), or Vinge dividing the galaxy into zones of thought and banishing superintelligence to the edge, but back in the near future we are mostly left with shuffling superintelligence aside with legal restrictions (the Turing Police in Neuromancer) or some sort of balance of power containment (The Quantum Thief).

Putting the future behind us, the current forecasting consensus1 is “stock up on canned goods and shotguns just in case, or use the current small window before AGI to coordinate a Plan A shaped solution to get us to a world of AIs pointed at each other in a Mexican standoff”. So thats what - Quantum Thief wins? The End. But actually how well does that sit intuitively, and does intuition even matter when facing a really unprecedented situation?

The AI futures model (AIFM) might be grossly and irresponsibly simplified to a basic growth rate boosted by a progress multiplier and divided by a bottleneck factor. Growth increases as more compute comes online, but soon the progress multiplier dominates as the AI improves itself, and pretty much every objection you could ever think of and some more besides have been bundled into the bottleneck factor. Despite this, the bottleneck does not scale proportionately with the growth, allowing an AI takeoff.

However, my intuition when I was writing Synthetic Control Measures and friends back in 2023 was that in fact we were likely going to get various forms of proportional “progress friction”. That still seems directionally correct today even after all the changes in the last 3 years, although with the amount of ink spilled since on AI I am sure this gut feel formulation has already been explicitly dismissed somewhere. No harm in another pass though, so let's sketch out what the mechanism behind this intuition might look like:

  1. 1.Verification becomes increasingly expensive. Efficiency gains and baking in increasingly detailed reality maps gives us better and better verification, but this gets outraced by the complexity of building the next even more detailed layer of reality verification. Verifier production becomes verifier limited.
  2. 2.Ignoring 1 costs more in slow-feedback domains. You can build features on a pile of buggy (i.e. out of touch with reality) software (vide Apple post 2013) at some extra cost. The cost of a pile of buggy theories giving you the wrong tensile strength for an actively powered building is much higher, and in such domains the more the capability must expand to meet the expanding needs of the capability.
  3. 3.Cost controls the rate of advance. Advances compound in areas with easy wins and easy verifications, and these diffuse quickly through the economy which pulls in more funding. Compounding advances in fast-feedback domains also increase progress in slow-feedback domains at a lesser rate, but because of 2 these are the first domains where ROI turns negative and causes a bust. Exposure contaminates the more productive fields into a mini AI winter, until compounding gains in the productive fields pulls money back in. As long as 1 and 2 hold this results in more boom and bust cycles.

Does this look like an Outside Context Problem for AIFM, or can we work it in as a prior and tweak a parameter somewhere? Where even are these new falsifiable predictions that we can all laugh at later?

I suspect that AIFM does account for this in a different way with its periodic capability slowdown multipliers, but I don’t have a good enough feel for the model to see if that works better than my assumption that the progress multiplier and the bottleneck are inversely related by friction from the verification process. What I do know is that in the absence of confounding factors a coupling between variables should be quite testable. Sadly in a macro forecast confounding factors abound, but lets not let that get in the way of some good ol predicticating for the next 5 years:

  1. 2027Trust in AI outcomes increases as frontier models become genuinely funny and persuasive rather than the unlikeable slopmeisters of today. Median cost per token of the hundred most popular models falls under US$0.80/M input tokens, and we see per-outcome rather than per-token billing become more widespread.
  2. 2028Absolute measures of AI ability are becoming increasingly decoupled from deployment outcomes, and relative measures between AIs similar to an Elo rating system develop, along with independently human “vibe based” measurement systems. US labs voluntarily sign up to a scheme to monitor compute use with a clause about future ratcheting, and a subjective capability cap is agreed between the US and China
  3. 2029The proportion of compute devoted to verification is now a separate line item on the majority of frontier labs financial reporting. US job losses in software development from AI reduce headcount to under 2026 levels. Open weight models are still only around 6 months behind the frontier, and OpenAI reduces spend on model training and new DCs, causing a debt crisis which ends Oracle in its current form. The voluntary scheme to ratchet compute use has still not had a measurable impact and is considered mostly ai-washing
  4. 2030The software R&D multiplier is still arguably in the 2-3x range but it is getting harder to estimate, and one of the major AI labs has either gone bust, been acquired, or completely stopped training new models. Job losses in software accelerate and begin in other fields, but it is hard to distinguish from a wider economic downturn partially caused by the first big pullback in AI spend. An anti AI backlash in the US causes the first legally enforceable limits on GPU use, and China independently implements something similar
  5. 2031A new AI developed AI paradigm becomes widely available and is more useful in robotics that LLMs. New AI labs emerge to commercialise this in different sectors of the economy. Parts of the US economy are growing 5-10% largely based on new AI developed technologies and productivity growth, while other parts are still contracting. The legal limits on GPU use are partially circumvented by the new paradigm, and there is little legislative appetite to draft anything to replace them.

This has now run a long way ahead of anything suggested by the model and is close to pure sci-fi, and we haven't even got to the mid 30s where superlinear recursive self improvement should be undeniable if I am wrong. Otherwise, my world is one where progress at the frontier stalls a few times a decade, where the only regulation which happens would be too little too late for under AIFM scenarios, where we don’t end up with clever cryptographic checks and balances, but despite that we still get away with our containment efforts and have a functioning economy involving humans for a long time to come. I’d love it if I was wrong and we got our act together earlier, because I could also be wrong about AI not being a near term existential risk.

Another consideration - is it a good idea to write about non perfectly behaved AIs at all, given it all ends up in their training data? The AI Hyperstition crowd says no, but I wonder whether writing more not particularly different good futures ends up strengthening the Waluigi Effect of their complement more than it strengthens the good. Perhaps a solution is to write radically different good futures which don't have evil complements in the literature, and give the AI more good options in a range of different situations. I literally summoned a deus ex AI to resolve a dumb-humans-cause-AI-takeoff crisis both because that seemed narratively more satisfying, but also in the hope that even when future humans screw up, the AI has an inkling from its training data that perhaps we are still worth saving.

1 - Of course once you move out from the easier eval questions towards the macro impacts that we care about here there is no consensus, but for simplicity I choose to narrow the field to the https://www.aifutures.org/ vs https://www.normaltech.ai/p/ai-as-normal-technology forecasts. They mostly agree on the early shape but become polar opposites later, so seemed like good proxies for the wider range of options, and I weight the consensus towards the former