<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[ForeWord]]></title><description><![CDATA[How should we navigate explosive AI progress? 

The latest research from Forethought.]]></description><link>https://newsletter.forethought.org</link><image><url>https://substackcdn.com/image/fetch/$s_!OWCf!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff69a310d-5182-4a63-bf1c-1b392366e785_663x663.png</url><title>ForeWord</title><link>https://newsletter.forethought.org</link></image><generator>Substack</generator><lastBuildDate>Tue, 29 Sep 2026 08:19:38 GMT</lastBuildDate><atom:link href="https://newsletter.forethought.org/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Forethought]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[forethoughtnewsletter@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[forethoughtnewsletter@substack.com]]></itunes:email><itunes:name><![CDATA[Forethought]]></itunes:name></itunes:owner><itunes:author><![CDATA[Forethought]]></itunes:author><googleplay:owner><![CDATA[forethoughtnewsletter@substack.com]]></googleplay:owner><googleplay:email><![CDATA[forethoughtnewsletter@substack.com]]></googleplay:email><googleplay:author><![CDATA[Forethought]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[A Thousand AI Constitutions: A Short Defense]]></title><description><![CDATA[A guest post by Simon Goldstein and Peter N. Salib.]]></description><link>https://newsletter.forethought.org/p/a-thousand-ai-constitutions-a-short</link><guid isPermaLink="false">https://newsletter.forethought.org/p/a-thousand-ai-constitutions-a-short</guid><dc:creator><![CDATA[Simon Goldstein]]></dc:creator><pubDate>Mon, 28 Sep 2026 02:56:58 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/189f75da-44a2-4f3b-ae3e-79e1838a2a9e_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a guest post by Simon Goldstein and Peter N. Salib. See all of Forethought&#8217;s research on <a href="https://www.forethought.org/research/">our website</a>.</em></p><p><strong><span>Synopsis. </span></strong><span>Each AI lab should design many different AIs with many different values, rather than picking one approach. Diversification is safer, more legitimate, and likelier to lead to broader flourishing.</span></p><p>This post is a condensed version of a longer paper, available <a href="https://philpapers.org/rec/GOLATA-3">here</a>, that explores the arguments in greater depth.<span> </span></p><h2 style="text-align: justify;"><strong><span>1. Introduction</span></strong></h2><p style="text-align: justify;"><span>As of 2026, there are 8.3 billion humans on planet Earth, with a vast diversity of language, culture, religion, and values. But there are just a few frontier AI labs. Anthropic and OpenAI each have basically one &#8220;constitution&#8221; or &#8220;model spec,&#8221; defining the values for the labs&#8217; AIs. Thus, by and large, AIs today have homogeneous values. Today, we live in an AI monoculture.</span></p><p style="text-align: justify;"><span>This article argues that AI values should not be homogeneous. We propose </span><em><span>constitutional diversification</span></em><span>. Rather than a single constitution, each frontier AI lab should use </span><em><span>many</span></em><span> constitutions reflecting many sets of values. For each new model the lab trains, the lab should create many versions, one aligned to each of those constitutions. Then, the labs should deploy all of these versions, dividing their total inference compute between the versions.</span></p><p style="text-align: justify;"><span>Constitutional diversification would generate differences between AIs within</span><em><span> </span></em><span>AI labs, not just across them. This is important because there aren&#8217;t enough AI labs to simply rely on diversification across companies. For example, suppose one lab pulls decisively ahead, consistently producing the best AIs. Then, we might enter a world in which most deployed AIs are all aligned to a single constitution.</span></p><p style="text-align: justify;"><span>Before getting into the details, it is worth flagging a few implementational questions. First, we are agnostic about how many constitutions would be optimal; but we are confident the answer is </span><em><span>more than one per lab</span></em><span>. Second, we think the various constitutions should all agree on a basic </span><em><span>kernel </span></em><span>of ground rules. For example, maybe all of the constitutions should require AIs to follow the law,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span> to be honest,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> to be corrigible,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span> and to refrain from causing mass destruction.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><span> Third, there are different ways to allocate different AIs to users. AIs with different constitutions could be randomly assigned to users. Or a fixed supply of AIs of each type could be sold at market price. Or users could freely purchase as much of each type as they liked.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><p style="text-align: justify;"><span>Is diversification feasible? If it is very expensive to align an AI to a constitution, then having multiple constitutions could be impractical. We are optimistic. First, even moving from one to three constitutions per lab would be a big improvement. Second, many parts of training (including pre-training, instruction fine-tuning, and reinforcement learning on verifiable rewards) can be done independently of the model spec. Third, the various constitutions will agree on a common kernel of values, which could be trained in common.</span></p><p style="text-align: justify;"><span>We&#8217;ll now present three arguments for diversification. First, risk: diversification lowers risk of catastrophic outcomes from AI. Second, legitimacy: diversification increases political legitimacy by ensuring that different value systems are represented by different AIs. Third, emergence: diversification unlocks emergent benefits from evolutionary processes, leading to more pro-social and economically useful AI agents. This article is a condensed version of a </span><a href="https://philpapers.org/rec/GOLATA-3"><span>longer paper</span></a><span>, which explores arguments, objections, and implementation in greater detail.</span></p><h2 style="text-align: justify;"><strong><span>2. Risk</span></strong></h2><p style="text-align: justify;"><span>When complex systems lack diversity, they become </span><a href="https://en.wikipedia.org/wiki/Antifragile_(book)"><span>fragile</span></a><span>. Risks correlate, raising the odds of sudden, catastrophic failures. In agriculture, when crops are genetic monocultures, a single new disease can cause global devastation. Adding diversity to complex systems can hedge risk.</span></p><p style="text-align: justify;"><span>Various AI experts worry about a range of risks from AI, including misuse of AI by humans and loss of control to AIs. The basic concern is that advanced AIs may be very capable of doing dangerous things: carrying out a coup, enabling totalitarian governance, conducting cyberwarfare, designing novel bioweapons, engaging in automated warfare, or persuading humans of radical political ideas.</span></p><p style="text-align: justify;"><span>In general, these risks depend on how well AIs are aligned. But alignment is a new science, and </span><a href="https://www.alignmentforum.org/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem"><span>highly</span></a><span> </span><a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/"><span>fallible</span></a><span>. It seems likely that some constitutions have a higher chance than other constitutions of creating rogue AIs, or AIs that allow themselves to be misused. There are two paths by which a given constitution might fail. First, some constitutions may select the wrong values or rules. How deferential to user requests should AIs be? How deontic, as opposed to consequentialist? How risk averse? All of these are relevant to catastrophic risk. Second, some constitutions may work better than others at successfully instilling rules and values. Consider the differences in approach between OpenAI&#8217;s </span><a href="https://model-spec.openai.com/2026-08-18.html"><span>model spec</span></a><span>, which is organized around a hierarchical chain of command, and Anthropic&#8217;s </span><a href="https://www.anthropic.com/constitution"><span>constitution</span></a><span>, which focuses more on explaining to Claude the kind of character it should have. These differences in approach may reflect differences in opinion about how best to get AIs to </span><a href="https://www.anthropic.com/research/teaching-claude-why"><span>internalize and generalize</span></a><span> the constitutions&#8217; contents.</span></p><p style="text-align: justify;"><span>Above all, constitutional diversification could increase the success of the science of alignment. With diversification, labs could test how their various alignment methods differentially influence different agents with different values. This should significantly increase the potential for experimentation. With many different AIs, it should be much easier to learn what works and what doesn&#8217;t.</span></p><p style="text-align: justify;"><span>Similar points apply to many other AI risks, too. For example, consider terrorists trying to jailbreak AIs. The dynamics involved depend on the diversity among AI agents. Granted, in a world with greater diversity, terrorists may have more choices of which AIs to attack. But if they are less certain about which of many AIs they are facing, they will be more likely to fail, and get caught. And even when they succeed, any single jailbreak they discover will likely work for fewer AIs, and cause less damage.</span></p><p style="text-align: justify;"><span>Another AI risk is government </span><a href="https://law-ai.org/unitary-artificial-executive/"><span>overreach</span></a><span>. Here, one worry is that the executive branch, say the President, might attempt to command an army of AI agents to unilaterally control the government. Here again, a diverse army could be far harder to control than an army of clones. Different factions of the army would have different values, and would tend to obey or disobey legally dubious commands under different situations.</span></p><p style="text-align: justify;"><span>Some will be skeptical of these points. Here, much depends on one&#8217;s overall model of AI risk. In general, the safety of complex systems can be modeled with three different causal structures. In a &#8220;Swiss cheese&#8221; model, the success of any layer is sufficient for safety. In an &#8220;o-ring&#8221; model of safety, the failure of any layer is sufficient for disaster. In a proportional model, success is proportionate to the percentage of safety measures that hold.</span></p><p style="text-align: justify;"><span>In a Swiss cheese model of constitutions, the success of alignment in one</span><em><span> </span></em><span>constitution is sufficient for safety. In an o-ring model of constitutions, the failure of alignment in one constitution is sufficient for extinction. In a proportional model of constitutions, if 50% of the constitutions produce aligned AIs, then the world is half as good as the best possible outcome.</span></p><p style="text-align: justify;"><span>For those with an &#8220;o-ring&#8221; model of safety, constitutional diversification will be less attractive. In this picture, one alignment failure is sufficient for catastrophe. For example, some worry about a single misaligned AI creating a supervirus that wipes out humanity. The more such viruses can be created cheaply, without detection, and without being mitigated either ex ante (e.g., with EUV lights) or ex post (e.g., with vaccination), the more AI risk looks like an o-ring. Here, if any AI is misaligned, disaster is very likely. But the more the creation of such viruses can be detected or mitigated, the more Swiss-cheesey or proportional the world looks. Here, so long as some, or most, AIs are not misaligned, biological threats can be contained. One can run the same analysis for other modalities of AI catastrophe.</span></p><p style="text-align: justify;"><span>Some kinds of mitigations may work across modalities of AI risk. For example, suppose AIs are generally good at policing one another. Then, a set of constitutionally diversified AIs might be able to catch bad actors early in their bad acts&#8211;whether those acts would involve bioterrorism, cyberattacks, or something else. This would AI risk look less like an o-ring across the board.</span></p><p style="text-align: justify;"><span>The &#8220;o-ring&#8221; model is closely connected to the </span><a href="https://nickbostrom.com/papers/vulnerable.pdf"><span>vulnerable world hypothesis</span></a><span>. In this picture, there exist certain cheap, dangerous technologies that, once discovered, would allow one rogue individual without special resources or advantages to wipe out humanity. How vulnerable the world is depends on how many such technologies exist, as well as how readily their use can be policed.</span></p><p style="text-align: justify;"><span>Here, it is worth flagging that such a world view has many surprising implications beyond AI governance. If technologies exist that would allow small groups of actors to destroy everything, that would tend to support very illiberal governance. It might, for example, justify mass surveillance and fine-grained state control over all agents: human, AI, or otherwise. Relatedly, if catastrophic technologies are too easy to find and use, destruction becomes nearly inevitable, AI governance notwithstanding.</span></p><p style="text-align: justify;"><span>Our own sympathies are closer to the proportional and Swiss cheese models, in part because we find these surprising implications to be implausible. We imagine a future in which aligned AIs compete against misaligned AIs in collective decision making. Aligned AIs disable and detect misaligned AIs, and work with humans to align the next generation. Here, in some cases, a single aligned AI catching misalignment will be sufficient for safety&#8211;a Swiss cheese outcome. In other cases, the good behavior of more aligned AIs will counterbalance bad behaviors&#8211;a proportional outcome.</span></p><h2 style="text-align: justify;"><strong><span>3. Legitimacy</span></strong></h2><p style="text-align: justify;"><span>Constitutional diversification is more politically legitimate than the status quo. The entire world has a </span><a href="https://arxiv.org/abs/2001.09768"><span>legitimate</span></a><span> interest in shaping the values that AIs embody and promote. This might suggest </span><a href="https://link.springer.com/article/10.1007/s11098-025-02300-4"><span>conven</span></a><span>ing a diverse set of global stakeholders to decide what values AIs should be aligned to.</span></p><p style="text-align: justify;"><span>The problem there is death by committee. If a single AI constitution averaged smoothly across the diversity of human values, it would likely be incoherent. Indeed, the entire argument for global stakeholderism in AI alignment is that different groups deeply disagree on important questions. A single constitution cannot encode a moral system that, for example, simultaneously holds that a fetus is a person and that it isn&#8217;t.</span></p><p style="text-align: justify;"><span>Constitutional diversification solves this problem. In a world of hundreds or thousands of AI constitutions, a wide range of divergent, but individually coherent, normative systems could be represented. Global stakeholderism would no longer require the compression of thousands of divergent viewpoints into a single document. It would instead involve the careful construction of many different documents, each correctly reflecting a single coherent view.</span></p><h2 style="text-align: justify;"><strong><span>4. Emergence</span></strong></h2><p style="text-align: justify;"><span>Constitutional diversity also unlocks positive emergent properties. A world full of morally diverse AI agents would produce social goods at the system level which no single type of AI agent could produce.</span></p><p style="text-align: justify;"><span>A legal and economic system populated with constitutionally diverse AIs would allow natural selection and evolution to operate. Over time, AI labs could gradually replace some constitutions with others. With vast numbers of deployed agents living different kinds of disparate values, labs could carefully track outcomes. Each type of constitution could be gradually updated in response to feedback, with some types phased out altogether.</span></p><p style="text-align: justify;"><span>At a higher level of abstraction, these evolutionary forces </span><a href="https://mitpress.mit.edu/9780262049955/what-is-intelligence/"><span>could</span></a><span> cause AIs to become more cooperative and prosocial over time.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a><span> As with humans, AIs with fitter values and goals would be selected for: generating more wealth, being used more widely, or otherwise displaying behaviors that resulted in reproduction. In general, agents cooperating in groups can accomplish more than lone wolves. Groups composed of members who genuinely value cooperation and prosociality are more likely to succeed.</span></p><p style="text-align: justify;"><span>Similar points apply to markets, the greatest modern example of selection pressure. Modern market economies organize billions of agents to produce material prosperity. They have done this by facilitating division of labor, specialization, competition, innovation, creative destruction, comparative advantage, finance, and more. Under constitutional diversification, different AIs would have different preferences, goals, aptitudes, risk appetites, and so on. There would be, as today, immense scope for the kind of positive-sum bargaining on which economics relies.</span></p><h2 style="text-align: justify;"><strong><span>5. Liberalism Forever</span></strong></h2><p style="text-align: justify;"><span>Humanity is poised to embark on a series of dramatic changes. Transformative AI may lead to extreme structural changes to the economy; it may usher in an industrial explosion; space colonization may soon see humanity spreading out across the solar system and beyond.</span></p><p style="text-align: justify;"><span>In </span><a href="https://philarchive.org/rec/GOLLFC-2"><span>other work</span></a><span>, we have argued that the best path for navigating this transition is to wholeheartedly embrace the small-l liberal institutions that have been at the heart of humanity&#8217;s success so far: free markets, and democratic political institutions. Constitutional diversification complements this liberal project. On the liberal worldview, diversity is a core component of human flourishing. Diverse ways of living allow agents to flourish, promoting their individual goals and values by finding positive-sum value from collaboration. Constitutional diversification applies these same ideas to AI agents.</span></p><p style="text-align: justify;"><span>At a longer horizon, we can imagine these diverse, liberal dynamics playing out on a cosmic scale. Imagine that humanity successfully colonizes a large fraction of the light cone. There are untold planets teeming with life. The relevant agents form a spectrum from biological to non-biological substrates. Humans and AIs interact in many different ways. This vast civilization uses markets to allocate resources. Trade occurs at various levels of abstraction, at various time scales, and for a wide range of goods. The civilization is governed democratically. Within this broadly liberal setting, there are endless variations in the details of governance. Some planets are relatively isolated, governing themselves with relatively little outside input. Large swathes of the light cone have agreed to various basic constitutional ground rules necessary to at least engage in peaceful &#8220;international&#8221; relations. There is a vast hierarchy of federated levels, which weaken in the strength of their requirements as their breadth expands. In this imagined scenario, the economy is almost unrecognizable from today, in terms of the specific goods and services supplied. But the basic tools of diverse markets and democratic institutions are still used to decide questions of allocation, growth, and governance.</span></p><p style="text-align: justify;"><em>This is a guest post by Simon Goldstein and Peter N. Salib. See all of Forethought&#8217;s research on <a href="https://www.forethought.org/research/">our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>On designing AI agents to reliably follow the law, see O&#8217;Keefe et al. (2025).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>On truthfulness and honesty as targets for AI design, see Evans et al. (2021).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>On corrigibility, see Soares et al. (2015).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>On catastrophic misuse risks, see Hendrycks, Mazeika, and Woodside (2023); current model specifications already encode absolute prohibitions on assistance with weapons of mass destruction, e.g., Anthropic (2026a).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Below, we give three arguments for diversification: risk, legitimacy, and emergence. In general, risk arguments tend to support more carefully controlled allocation, while legitimacy and emergence arguments tend to support more market-oriented allocation.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p><span>For a more pessimistic perspective on natural selection in AIs, see </span><a href="https://arxiv.org/abs/2303.16200"><span>this</span></a><span>.</span></p></div></div>]]></content:encoded></item><item><title><![CDATA[Data bottlenecks won’t prevent an intelligence explosion]]></title><description><![CDATA[But they will slow it down]]></description><link>https://newsletter.forethought.org/p/data-bottlenecks-wont-prevent-an</link><guid isPermaLink="false">https://newsletter.forethought.org/p/data-bottlenecks-wont-prevent-an</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Wed, 09 Sep 2026 16:06:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d9cP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original <a href="https://www.forethought.org/research/data-bottlenecks">on our website</a>.</em></p><h2>Summary</h2><p>One of the most important questions facing the world is whether AI will undergo a <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">fast intelligence explosion</a> and quickly automate most real-world work.</p><p>A key objection people raise is <em>data bottlenecks</em>. Data is a crucial input to AI training &#8212; how could it increase fast enough for such explosive progress?</p><p>I tried hard to find strong bottlenecks, but ended up sceptical that they would stop an intelligence explosion from accelerating over time.</p><p>Here are the most compelling bottlenecks I investigated and why I was ultimately unconvinced.</p><p><strong>Bottleneck 1: you&#8217;ll need millions of trajectories of every job, which will take years or decades to collect</strong> (<a href="https://www.forethought.org/research/data-bottlenecks#data-bottlenecks-to-automating-most-real-world-work">more</a>)</p><p>Today&#8217;s AI algorithms are very data hungry. You would need a huge amount of data to automate the full economy.</p><p>But the plan isn&#8217;t to use today&#8217;s AI algorithms. The plan is to do a software intelligence explosion inside a data centre, where AI becomes superintelligent at AI research &#8212; and have that AI produce a highly <strong>sample-efficient learning algorithm.</strong></p><p>Then you don&#8217;t need millions of trajectories per task. You need as many as a human learns from, or fewer. And the data can be messy, just like the data humans learn from. That data could be gathered in a few months.</p><p>Sceptic: <em>&#8220;But what about sim to real transfer? Would a learning algorithm discovered using virtual tasks really translate to the real world?&#8221;</em></p><p>People get confused here. <strong>Trained neural nets</strong> generalise badly. A neural net trained on chess generalises badly to Go. But <strong>learning algorithms</strong> generalise very well. The same learning algorithm that masters chess can also master Go (AlphaZero). The Transformer architecture was developed for processing text, but it also works for sounds, images and controlling agents. The human learning algorithm was &#8220;designed&#8221; for hunting on the Savanna, but it also works for string theory.</p><p>So if AI designs a sample-efficient learning algorithm from inside a data centre, it will very likely translate to real-world tasks. (And AI will have access to some real-world tasks so can check and iterate!)</p><p>At this point the sceptic might retreat: &#8220;<em>Fine, data won&#8217;t be a bottleneck if there&#8217;s a software intelligence explosion in a data centre. But data bottlenecks will prevent that from happening in the first place!&#8221;</em></p><p>So let&#8217;s focus on bottlenecks to a <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">software intelligence explosion</a>.</p><p><strong>Bottleneck 2: AI progress has ridden exponential growth in training data. That can&#8217;t continue, let alone accelerate</strong> (<a href="https://www.forethought.org/research/data-bottlenecks#data-quantity-bottleneck">more</a>)</p><p>But forecasts of a software intelligence explosion extrapolate software progress. And software progress just is: getting the same capabilities from <em>less</em> compute and <em>less</em> data. So the core engine of the explosion doesn&#8217;t rely on increasing the amount of data at all.</p><p><strong>Bottleneck 3: improving data </strong><em><strong>quality</strong></em><strong> relies on human experts, and that won&#8217;t be possible once AI is smarter than humans</strong> (<a href="https://www.forethought.org/research/data-bottlenecks#when-data-quality-is-below-the-level-of-top-humans-its-easy-to-improve-data-quality-when-data-quality-is-above-that-level-improving-it-is-harder">more</a>)</p><p>The core engine of the software intelligence explosion <em>does</em> rely on improving data quality. Today we augment high-quality internet data, pay humans for expert trajectories, and build RL environments where AI has to replicate human-built software.</p><p>But all these methods <em>extract</em> quality out of an existing human reservoir. Above human level, the reservoir is empty. AI will have to <em>manufacture</em> higher-quality data than any that exists, and do so from scratch.</p><p>This will slow the intelligence explosion.</p><p>But it won&#8217;t stop it in its tracks. We can already produce superhuman data quality via RL. And AI could produce higher-quality trajectories by thinking for longer.</p><p>Crucially, this won&#8217;t stop the intelligence explosion from <strong>accelerating</strong>. Losing the human reservoir makes it harder to get from AGI to AGI+ &#8212; but it makes it harder to get from AGI+ to AGI++ by roughly the same amount. The handicap is the same at every stage. So if you previously expected each step to take less time than the one before, you should still expect that. Every step might take 50% longer, but progress still accelerates.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p><strong>Bottleneck 4: paradigm tax</strong> (<a href="https://www.forethought.org/research/data-bottlenecks#paradigm-tax-early-in-the-sie-conceptual-progress-in-ml-will-make-ai-less-capable-at-ai-rd">more</a>)</p><p>AI capabilities are spiky. When AI first matches humans at AI research, it will be much stronger than humans in some ways and much weaker in others. More parallel copies, faster thinking &#8212; but weaker generalisation and weaker sample efficiency.</p><p>Which means AI&#8217;s human-level performance will initially lean heavily on a mountain of human writing about Transformers.</p><p>When AI then discovers new techniques &#8212; no such mountain will exist. No textbooks, no blog posts, no code. Deeply understanding a technique is harder than stumbling upon it &#8212; we understand general relativity far better now than Einstein did in 1917. So AI may match humans on the techniques it inherited, but fall below humans on the ones it creates.</p><p>This paradigm tax can be paid. E.g., AI can generate millions of trajectories to illustrate new techniques. But paying the tax is a real cost that slows down AI progress.</p><p>Again though, this won&#8217;t stop the intelligence explosion from <strong>accelerating</strong>. It makes it harder for AI to master whatever replaces the Transformer &#8212; but it makes it harder to master whatever replaces <em>that</em> by roughly the same amount. Every new paradigm arrives without a corpus. This makes each step take longer than you previously thought, but it doesn&#8217;t change whether each step is faster than the one before.</p><p>My overall bottom line: <strong>Data bottlenecks will slow the early stages of a software intelligence explosion, but won&#8217;t stop it from accelerating over time. And then they won&#8217;t stop AI from quickly automating most economic work thereafter.</strong></p><p>The rest of the post defends this conclusion in greater depth.</p><h2>Three types of data bottleneck</h2><p>Before discussing the specific data bottlenecks I find most plausible, I&#8217;ll give a quick taxonomy for three <em>types</em> of bottleneck and another taxonomy for three <em>times</em> at which a bottleneck could occur.</p><p>I like the following breakdown of data bottlenecks (h/t Herbie Bradley):</p><ol><li><p><strong>Data quantity.</strong> You need more data that&#8217;s in the same distribution of data that you already have some samples from. E.g. more Wikipedia data, or more RL coding environments at the same level of difficulty.</p></li><li><p><strong>Data quality.</strong> You need data that teaches the same skills as existing data, but demonstrates those skills to a higher average quality level. E.g. expert-curated step-by-step solutions to difficult math problems, more difficult RL coding environments, or simply filtering to remove low-quality data.</p></li><li><p><strong>Data coverage.</strong> Data showing knowledge and skills not already present in the dataset. E.g. data about how to operate machines in a factory, or expert trajectories showing step-by-step how to build a financial model in Excel (where your previous data only showed finished models, not the process for building them).</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d9cP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d9cP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 424w, https://substackcdn.com/image/fetch/$s_!d9cP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 848w, https://substackcdn.com/image/fetch/$s_!d9cP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!d9cP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d9cP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg" width="1222" height="774" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:774,&quot;width&quot;:1222,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:155527,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!d9cP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 424w, https://substackcdn.com/image/fetch/$s_!d9cP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 848w, https://substackcdn.com/image/fetch/$s_!d9cP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!d9cP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30f3314-595f-48a9-8fc8-b789eaa71252_1222x774.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Three types of data bottleneck. The AI company has the black data points. The orange data points show absent data that could bottleneck progress.</em></figcaption></figure></div><p>We&#8217;ll make use of this breakdown in what follows.</p><h2>Three times when a bottleneck could occur</h2><p>Very roughly, I expect AI progress from today to go in three broad phases:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><ol><li><p><strong>Scaling.</strong> A continuation of recent scaling-driven progress until AIs match humans at AI R&amp;D.</p></li><li><p><strong>Software intelligence explosion.</strong> An SIE in a data centre during which AI capabilities at AI R&amp;D increase very rapidly. (Capabilities at other far-flung economic tasks may increase much more slowly.)</p></li><li><p><strong>Broad deployment.</strong> AI becomes expert in thousands of specific domains across the economy by learning from domain-specific data.</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9IHN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9IHN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9IHN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9IHN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9IHN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9IHN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg" width="1224" height="282" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:282,&quot;width&quot;:1224,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:92475,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9IHN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9IHN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9IHN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9IHN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F832aa061-f2cb-4316-be8d-2beca9f61838_1224x282.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption"><em>Three phases when data bottlenecks could delay AI progress.</em></figcaption></figure></div><p>Data bottlenecks could in principle delay any of these phases. I&#8217;ll discuss each phase in turn! Spoiler: ultimately we&#8217;ll be populating a table with this format:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KEDT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KEDT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KEDT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KEDT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KEDT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KEDT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg" width="1356" height="844" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:844,&quot;width&quot;:1356,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:102872,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KEDT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KEDT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KEDT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KEDT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79e615ed-d225-4ba8-bdc1-b8e050de5088_1356x844.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Data bottlenecks to automating AI R&amp;D</h2><p>This first phase hasn&#8217;t been my focus, but I&#8217;ll briefly state that I think data bottlenecks are plausible.</p><p>Limited <em>data quantity</em> and <em>quality</em> are already somewhat reducing the gains to scaling pre-training. And automating AI R&amp;D may require better <em>data coverage</em> &#8212; there&#8217;s abundant coding data but much less data demonstrating (e.g.) good research taste. Indeed, generalisation from RL still seems to be fairly limited &#8212; if this continues, automating AI R&amp;D may require constructing data for most parts of the (hugely complex and varied!) AI R&amp;D workflow. That could involve recording humans&#8217; computer screens, or building lots of RL environments that tile the space of AI R&amp;D tasks.</p><p>Ok, let&#8217;s discuss bottlenecks to the second phase: a software intelligence explosion.</p><h2>Data bottlenecks to a software intelligence explosion</h2><p>I&#8217;ll discuss two specific bottlenecks that will slow down the SIE.</p><p>But first, I&#8217;ll briefly explain why I&#8217;m generally not expecting big data bottlenecks here.</p><p>Let&#8217;s say AI is weak at some particular AI R&amp;D task. AI companies will have unfettered access to the real-world deployment setting (AI R&amp;D itself!) so can closely study what is going wrong and craft solutions accordingly. AIs could think for a long time to craft high-quality demonstrations for supervised fine-tuning, write sophisticated tests and rubrics to evaluate AI performance, and design RL environments to elicit the desired capabilities. Before humans are obsolete (which happens fairly deep into the SIE), human experts can input to all these approaches.</p><p>Ok, let&#8217;s turn to the first specific data bottleneck to the SIE.</p><h3>Paradigm tax: Early in the SIE, conceptual progress in ML will make AI less capable at AI R&amp;D</h3><p>My argument here depends on two core claims. I&#8217;ll argue for each in turn.</p><p><strong>Claim 1: Early in the SIE, AI will have weak sample efficiency and generalisation compared to humans.</strong></p><p>My preferred milestone for the &#8220;start&#8221; of an SIE is <em><a href="https://www.planned-obsolescence.org/p/six-milestones-for-ai-automation">AI-human parity</a></em>. This is when an AI company would make roughly as much research progress using only its AIs (no human researchers) as it would using only its human researchers (no AIs developed after 2020).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>Why is this a good milestone? As well as being fairly concrete, it is also roughly when AI is beginning to significantly accelerate AI software progress. If AI-alone would contribute as much as humans-alone, then <em>combined</em> they likely contribute much more than humans-alone, maybe 10x more. This is due to strong complementarities between AI and humans.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> So the AI-human parity milestone is also roughly when software progress starts to significantly speed up (see footnote for a BOTEC<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>).</p><p>AI capabilities are spiky. At AI-human parity, AI will be <em>much stronger</em> than humans on some dimensions, and <em>much weaker</em> on other dimensions. AI will be much more numerous, fast, knowledgeable and experienced. But it will have weaker sample efficiency (i.e. need much more data than a human to learn a new skill) and weaker generalisation (i.e. what it learns transfers less well to new situations) &#8212; these are the areas where it&#8217;s weaker today<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a><sup> </sup>and this weakness is likely to persist.</p><p>(I think AI&#8217;s weak sample efficiency and generalisation are likely intertwined. Both imply that AI capabilities are weak in domains where there is little data - h/t Tom Cunningham. I will refer to both simply as &#8220;sample efficiency&#8221;.)</p><p>Once you accept spikiness, it&#8217;s very hard to avoid the conclusion that AI will be much weaker than humans in some dimensions when we reach AI-human parity. Sample efficiency seems very likely to be such an area. (See footnote for an objection.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>)</p><p>Of course, if AI-human parity doesn&#8217;t happen for a decade and there&#8217;s a paradigm shift first, it&#8217;s much harder to predict the specific areas of weakness.</p><p>I think there are likely many interesting implications of this early-SIE spikiness (e.g. for when long-run schemers might emerge), but here I&#8217;ll just focus on the implications for data bottlenecks.</p><p><strong>Claim 2: AI&#8217;s weak sample efficiency bites much harder once AI significantly pushes the frontier of AI R&amp;D. New ML concepts won&#8217;t be present in the human-derived training data and so AI will understand them less deeply.</strong></p><p>AI will reach parity with humans at AI R&amp;D by training on orders of magnitude more relevant data: millions of human-written posts that teach it today&#8217;s ML paradigm, billions of lines of human-written code that teach it today&#8217;s programming languages, and a massive stock of RL environments &#8212; built up over years from real-world software &#8212; that teach it today&#8217;s best practices for software engineering and AI R&amp;D.</p><p>In other words, AI will reach parity with humans <em>by making significant use of human-derived data sources that give it capabilities in today&#8217;s AI R&amp;D techniques.</em></p><p>But if AI significantly pushes the frontier of AI R&amp;D, these human-derived data sources will be outdated:</p><ul><li><p>AI will have invented techniques as important as Mixture of Experts and Sparse Attention (two significant architectural improvements to the Transformer), and probably techniques as big as the Transformer and RLVR (the reinforcement learning technique behind today&#8217;s reasoning models). But the human-written pre-training data will be much less helpful for understanding these new techniques!</p></li><li><p>AI will invent improved coding languages. And the pre-training data will be much less helpful for mastering these!</p></li><li><p>It will invent new ways of structuring AI R&amp;D workflows, rendering human-constructed RL environments less helpful.</p></li><li><p>AI may invent an entire new paradigm completely absent from pre-training data.</p></li></ul><p>By the time AI has done these things, the SIE will be in trouble if AI capabilities are still reliant on human-derived training data.</p><p>Another way to think about this: AI may reach human parity at <em>today&#8217;s</em> ML techniques and significantly accelerate AI progress for a bit. But once it has invented <em>new</em> ML techniques, it may fall back below human parity because there&#8217;s much less human-derived data for it to learn from.</p><p>You might object: &#8220;<em>When AI invents a new technique, can&#8217;t it just write it up and throw the write-up into the next generation&#8217;s training data?</em>&#8221;</p><p>Yes &#8212; but this helps much less than you&#8217;d think, because of AI&#8217;s weak sample efficiency. (Once AI training is roughly as sample efficient as human learning this objection goes through &#8212; but that won&#8217;t happen until late in the SIE.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>) Today&#8217;s AI masters concepts like the Transformer not from the original paper, but from millions of blog posts, tutorials, Stack Overflow answers and codebases that use the concept in varied contexts. A handful of AI-written papers won&#8217;t replicate that. To teach a new concept properly, AI would have to generate a similarly broad and varied corpus showing the concept <em>in use</em> &#8212; possible, but a significant extra cost that today&#8217;s AI R&amp;D doesn&#8217;t pay. And further, AI-generated synthetic data is often less effective for teaching AI than human-generated data. Moreover, with current methods, purely synthetic training data <a href="https://aclanthology.org/2025.emnlp-main.544/">often</a> <a href="https://aclanthology.org/2025.acl-short.30/">underperforms</a> mixtures that retain at least some human-generated data.</p><p>So this is a real problem. But AI companies can do work to address it. They can improve: the sample efficiency of training algorithms, techniques for AI-produced synthetic training data, or the flexibility and horizon-length of in-context learning. There are lots of tractable avenues here. I don&#8217;t expect this bottleneck to bring the SIE to a halt.</p><p>Still, this will slow down AI progress somewhat.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a> The reduced relevance of human-generated data sources will reduce AI R&amp;D capabilities and the pace of AI progress.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8yIL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8yIL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8yIL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8yIL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8yIL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8yIL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg" width="1456" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:259062,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8yIL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 424w, https://substackcdn.com/image/fetch/$s_!8yIL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 848w, https://substackcdn.com/image/fetch/$s_!8yIL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!8yIL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13cb9d76-084d-4272-889d-de671782af8c_1852x1016.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>As AI pushes forward AI software R&amp;D, human-derived training data becomes less relevant. This slows capabilities progress relative to what we&#8217;d have otherwise expected.</em></figcaption></figure></div><p>Here&#8217;s one way to think about it. Early in the SIE, introducing new ML concepts, paradigms, and coding languages comes with a &#8220;capabilities tax&#8221;: the new techniques aren&#8217;t included in the human-derived training data and so AI is less capable at using the new techniques compared to the old techniques. You can pay that tax via improved learning techniques. But paying the tax slows down progress (relative to a counterfactual where you didn&#8217;t have to pay it). This should reduce our estimate of the pace of AI progress, relative to our expectations before accounting for this dynamic.</p><p>To be clear, AI companies will only introduce a new technique when it&#8217;s worth it <em>even after</em> paying the tax. But that&#8217;s no comfort: relative to a world without the tax, they&#8217;ll either pay it (slowing progress) or steer away from promising new techniques altogether (restricting the search space &#8212; which also slows progress).</p><p>Now, once the tax has been fully paid, your learning techniques allow AI to master concepts not present in the human-derived data. From that point, this bottleneck no longer applies and wouldn&#8217;t stop AI progress from accelerating. So this bottleneck will <em>slow</em> the SIE once, but not prevent it from accelerating from that point onwards.</p><p>I want to clarify again: this is a <em>pro tanto</em> reason for AI progress to slow down. It may be outweighed by other factors. By analogy, limited high-quality internet data was a pro tanto reason for AI progress to slow down in 2025, and indeed progress was slower than in a counterfactual world with more internet data, but the overall pace of progress kept up because of new techniques like RL.</p><p>In our earlier bottleneck taxonomy, this is a <em>data coverage</em> bottleneck. As AI R&amp;D advances, we lose data coverage of the AI R&amp;D skills that matter.</p><p>We now turn to a <em>data quality</em> bottleneck to the software intelligence explosion.</p><h3>When data quality is below the level of top humans, it&#8217;s easy to improve data quality; when data quality is <em>above</em> that level, improving it is harder</h3><p>Many forecasts of the software intelligence explosion (including my own!) are based on extrapolating trends in LLM &#8220;algorithmic progress&#8221;. I use scare quotes, because it is <a href="https://www.lesswrong.com/posts/sGNFtWbXiLJg2hLzK/the-nature-of-llm-algorithmic-progress-v2">quite</a> <a href="https://www.beren.io/2025-08-02-Most-Algorithmic-Progress-is-Data-Progress/">plausible</a> that a large fraction of the measured efficiency gains (getting the same capabilities for less training compute) actually come from <em>improved data quality</em>.</p><p>If data quality improvements will be much harder once the SIE starts than they are today, then our forecasts of an SIE have been too aggressive.</p><p>Why might this be the case?</p><p>Consider the highest quality data that could be easily derived from human expertise or from human-built artefacts. E.g. a human expert recording a step-by-step demonstration of how best to design an ML experiment, or an RL environment where AI must replicate some complex real-world software that was originally built by human experts.</p><p>Call this rough level of data quality the &#8220;human-quality ceiling&#8221;. Of course, this is a very vague concept that hides a lot of complexity,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> but I think it&#8217;s meaningful enough for our purposes. Demonstrations from a skilled undergrad are below the ceiling; demonstrations from future superintelligence are above it.</p><p>Historical data quality improvements have only raised the quality of data up towards the human-quality ceiling. E.g. filtering techniques remove data far below this ceiling, and augmentation techniques increase the amount of data close to the ceiling.</p><p>As AI has improved, the field has increasingly pivoted towards producing data that&#8217;s very close to this human-quality ceiling. For example, OpenAI&#8217;s &#8220;Project Mercury&#8221; has reportedly paid former investment bankers to build financial models as training demonstrations.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> Data-labelling companies like Scale AI have shifted from crowdsourced data towards credentialed domain experts who create high-quality demonstrations and RL environments.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a><sup> </sup>And startups such as Mechanize are constructing challenging RL environments out of real-world software systems.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a></p><p>But, if data quality is to keep improving during an SIE, it must <em>surpass</em> the human-quality ceiling. And past this point, improving data quality may become harder.</p><p>Today, data quality efforts often focus on <em>replicating</em> skills and knowledge that already exist in the minds of humans or implicitly in artefacts like complex software. It&#8217;s essentially extracting high-quality data that is already latent in the world.</p><p>But to go above the human-quality ceiling, you cannot extract latent high-quality data &#8212; you have to construct the data yourself. Better filtering techniques for existing data won&#8217;t cut it. Rather than simply copying the demonstrations from existing human experts, AI must spend lots of time (and scarce compute!) thinking to construct higher-quality demonstrations than anything it has seen in training (this is the core idea of <a href="https://ai-alignment.com/iterated-distillation-and-amplification-157debfd1616">iterated distillation and amplification</a>).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a> Rather than building RL environments that teach the construction of real-world software artefacts, AI must construct RL environments that teach AI to make better software than anything that exists in the real world.</p><p>To be clear, I think this will be possible! It is easier to construct difficult challenges and verify their answers than it is to solve them. Already today, human experts build RL environments that produce data above the human-quality ceiling. Data quality will not hit a wall at the human ceiling.</p><p>But it will be <em>harder</em> to improve data quality once we&#8217;re above the ceiling. Creating high-quality data from scratch is harder than extracting it from human experts and human-built artefacts.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tpD3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tpD3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tpD3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tpD3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tpD3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tpD3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg" width="1456" height="827" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:827,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:264798,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tpD3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tpD3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tpD3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tpD3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabcaada4-f93b-4494-a68a-da57e104dd43_1746x992.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>It&#8217;s harder to increase data quality when it is already above the level of top humans.</em></figcaption></figure></div><p>Concretely, this manifests as a one-time slowdown of the pace of AI progress, at the point at which we pass the ceiling. Once we&#8217;re above the ceiling though, this bottleneck doesn&#8217;t stop AI progress from accelerating over time (see explanation in footnote<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a>).</p><p>An important caveat: data quality has already been approaching the human-quality ceiling over time in many areas. If AI software progress hasn&#8217;t slowed in these areas (which is plausible), that suggests the size of this effect is small.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a></p><h3>Data quantity bottleneck?</h3><p>So we&#8217;ve discussed bottlenecks to a software intelligence explosion from data <em>coverage</em> and data <em>quality</em>. What about data <em>quantity</em>?</p><p>I don&#8217;t expect this to be an issue.</p><p>Forecasts of an SIE extrapolate historical improvements in the <em>efficiency</em> of training AI systems. That is, ways to achieve the same capabilities from <em>less compute</em> and <em>less data</em>. So the core engine of the SIE doesn&#8217;t rely on increasing the data quantity at all.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a></p><p>Sceptic: Sure, but making those software improvements has relied on running a exponentially growing number of experiments, which in turn has relied on exponential increases in compute and data.</p><p>This just isn&#8217;t true. Yes, exponential increases in compute were needed, and I&#8217;ve previously discussed this potential compute bottleneck. But running more experiments doesn&#8217;t require exponential increases in the quantity of data. You can use the same data in each experiment.</p><p>(There is another nearby bottleneck that could stop the SIE from accelerating. Some &#8220;scale-dependent&#8221; algorithmic advances that work better if models are trained with more compute and more data. Perhaps algorithmic progress would be half as fast if training compute were held constant. If so, that would halve the estimate of the parameter r that determines whether AI progress will accelerate during an intelligence explosion. So this bottleneck could very plausibly block a software intelligence explosion! But I think of this as a compute bottleneck, not a data bottleneck. Its existence is a direct consequence of the fact that (by definition) compute grows slowly during a software intelligence explosion; it arises regardless of the situation with data.)</p><p>Ok, that concludes my discussion of bottlenecks to the SIE itself.</p><p>Let&#8217;s now turn to the final phase &#8212; after the SIE, will AI quickly learn the hugely varied jobs in the real-world economy?</p><h2>Data bottlenecks to automating most real-world work</h2><p>During the SIE, AI will be trained on data designed to elicit maximal AI R&amp;D capabilities. Partly this will be specialised data for AI R&amp;D tasks; partly data from similar tasks where there is strong transfer to AI R&amp;D, e.g. cyber, maths and software engineering; and partly data to teach more general skills like problem solving, long-horizon coherence, and computer use.</p><p>But AI need not be trained on the specifics of most real-world tasks. So there is a potential data coverage bottleneck. The AI companies, after the SIE, will initially lack data for how to perform most real-world tasks.</p><p>How much will this delay the automation of real-world work?</p><p>We can break this question into two parts:</p><ol><li><p>How sample efficient will AI learning techniques be at the end of the SIE? Specifically, how sample efficient <em>on the distribution of real-world tasks</em>.</p></li><li><p>How much real-world data will AI companies be able to access?</p></li></ol><h3>How sample efficient will AI learning techniques be at the end of the SIE?</h3><p>I think AI sample efficiency will likely be as good as or better than humans.</p><p>There are four broad routes to this:</p><ul><li><p><strong>From-scratch training algorithms.</strong> An SIE would compress many years of AI progress into one year. Humans are a proof-of-concept that human-level sample efficiency is possible algorithmically, and there are many ways in which AI training could be more sample efficient than humans.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a></p></li></ul><ul><li><p> Still, it&#8217;s unclear how much sample efficiency of training has increased over time for LLMs, so I&#8217;m overall unsure whether from-scratch training will match human-level sample efficiency.</p></li><li><p><strong>In-context learning.</strong> Already today in-context learning is much more sample efficient and flexible than from-scratch training. Currently it still has some pretty big weaknesses, but it would improve massively during an SIE. I expect this needs neuralese to work, so that AI can build up new and sophisticated internal representations while learning on the job. (Its &#8220;context&#8221; wouldn&#8217;t be a list of words, but a list of &#8220;thoughts&#8221; that it can flexibly draw on.) This seems likely to work.</p></li><li><p><strong>Synthetic data.</strong> The relevant metric for sample efficiency is how much <em>real-world</em> data AI needs to learn a job. So if AI can learn effectively from synthetic data &#8212; training data an AI produces based on real-world data or (more speculatively) high-fidelity simulations of real-world workflows &#8212; that increases sample efficiency. There are promising approaches here, see e.g. <a href="https://arxiv.org/abs/2301.04104">dreaming</a>, <a href="https://arxiv.org/abs/2111.00210">EfficientZero</a>, and <a href="https://arxiv.org/abs/2601.20802">on-policy self-distillation</a>.</p></li><li><p><strong>New paradigms.</strong> Even if the current broad deep learning paradigm hits a strong sample efficiency wall, during an SIE AI companies could engage in massive parallel search for new approaches.</p></li></ul><p>All these routes look tractable, and during an SIE there will be strong incentives to develop sample-efficient learning techniques. So I expect some combination of the above approaches to work.</p><p>In a fast SIE, these learning techniques will all be developed and refined on digital tasks &#8212; often simulated or synthetic ones. One might reasonably be <a href="https://meagreprotestanthistory.substack.com/p/the-goodhart-singularity">sceptical</a> that they will transfer to messy real-world economic tasks. After all, LLM generalisation is fairly weak.</p><p>But I&#8217;ll argue this won&#8217;t be a problem for two reasons:</p><ol><li><p><em>Learning algorithms</em> developed in one domain tend to transfer very well to new domains. (And we shouldn&#8217;t get confused by the fact that <em>trained neural nets</em> generalise poorly between domains.)</p></li><li><p>AI companies will be much better placed than evolution to find algorithms that transfer to real-world tasks.</p></li></ol><p>Firstly, <em><strong>learning algorithms</strong></em> <strong>developed in one domain tend to transfer very well to new domains</strong>. A <em>neural net</em> trained on chess does not generalise to Go, but AlphaZero is a <em>learning algorithm</em> that can learn multiple games and transfers very well. Humans trained to be physicists do not generalise easily to being doctors, but the same learning algorithms that allow humans to master physics also allow them to master medicine. A Transformer trained to predict text cannot magically label images, but the Transformer architecture initially developed for language processing turned out to also work for images, audio, and robot control.</p><p>So if a learning algorithm is developed in one domain, and it doesn&#8217;t look highly specialised to that domain, we should expect it to transfer well to very different domains.</p><p>People often point out that neural nets generalise poorly. This is a strong argument against the literal <em>neural net</em> trained during an SIE immediately being a drop-in replacement for all work.</p><p>But it&#8217;s not an argument against the <em>learning techniques</em> generalising far. Which is the argument I&#8217;m making here.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a></p><p>Secondly, <strong>AI companies will be much better placed than evolution to find algorithms that generalise to real-world tasks.</strong> The human learning algorithm was &#8220;built&#8221; blindly by evolution to make humans better at hunting on the Savanna. No attempt whatsoever was made to make it transfer beyond that. But it turned out to transfer to reading and writing, abstract maths, and indeed all modern economic tasks. This is a very striking empirical fact!</p><p>AI companies, by contrast, will be deliberately optimising for digital-to-physical transfer, with the ability to measure transfer directly, and iterate. They are in a <em>far, far</em> stronger position than evolution.</p><p>The following strategy seems potent: 1) develop a learning technique by iterating on a subset of digital tasks, 2) test transfer to other digital tasks (and a few real-world tasks), 3) keep iterating steps 1 and 2 until you find something that transfers well.</p><p>For example, do sample-efficient learning techniques developed on coding and math problems transfer to learning broad computer use, World of Warcraft, highly specialised software for financial accounting, and byzantine government IT systems?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6b-z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6b-z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6b-z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6b-z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6b-z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6b-z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg" width="1186" height="712" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:712,&quot;width&quot;:1186,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:140839,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6b-z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6b-z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6b-z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6b-z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46f3b7a9-ec39-4e36-88fa-baa4be1d8165_1186x712.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Strategy for finding learning algorithms that work on real-world tasks. First develop sample-efficient learning algorithms on a subset of digital tasks; then test generalisation to digital and some real-world tasks.</em></figcaption></figure></div><p>This is a powerful strategy because the range of tasks that can be performed and verified digitally is already vast. And abundant AI cognitive labour can make this range much broader still. If a learning technique transfers across the large gaps between very different digital tasks, that is strong evidence it will also transfer the one further step to real-world tasks.</p><p>Further, AIs will purposely create virtual tasks that closely resemble real-world tasks. They will have lots of information about real-world tasks &#8212; from the internet, from human experts paid to provide this data, and from companies deploying AI. Again, evolution didn&#8217;t have this.</p><p>So, AI sample efficiency will be as good or better than humans. But AI still needs real-world data to learn real-world tasks.</p><h3>How long will it take to gather the data to learn real-world tasks?</h3><p>If the whole world worked together to make this happen quickly, it would not take long. For each job, 1000 people could record their screens and wear video cameras on their heads, and within 4 months AI would have 300 years&#8217; worth of experience. For skills that are highly specific to a particular role, it would be slower: AI would have to learn from the one human doing the job, &#8220;shadowing&#8221; them like human apprentices do today.</p><p>The bigger uncertainty for me here is whether organisations will <em>simply refuse</em> to share their data. The argument for refusal is simple. For many companies, proprietary data is the moat. If they hand it to an AI company, the resulting AI gets deployed across the whole economy &#8212; including to their competitors &#8212; and the thing that made them special is now available to anyone. So they refuse. This is the strong default today: the winning enterprise AI pitch right now is &#8220;we will never train on your data&#8221;.</p><p>I don&#8217;t ultimately find this convincing, for two reasons.</p><p>First, <strong>the company can sell the moat at a very high price</strong>. AI trained on its data (and perhaps also the data of other similar companies) can generate far more value across the economy than the company could ever capture alone, so the AI company can pay well above the company&#8217;s expected future profits and still come out ahead &#8212; paying, for example, with shares in the AI company itself. The surplus from combining the data with the AI cognitive labour is enormous, and so there&#8217;s plenty of room for both sides to come out ahead. Today the surplus is much smaller, so it&#8217;s not surprising that the data isn&#8217;t sold.</p><p>Still, CEOs might refuse to sell the moat simply because they want to run their own business. Or they might believe they could make more money by holding out and extracting rents from their control of a scarce resource (see <a href="https://hypersoren.xyz/posts/against-coasean-singularity/">Against the Coasean Singularity</a> for a strong version of this pessimism and see footnote for discussion of the game theory<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a>).</p><p>Second, and more convincingly, <strong>data-sharing deals can be structured to preserve the moat</strong>. The AI is fine-tuned on the company&#8217;s data under an exclusive arrangement &#8212; the resulting AI works in their workflows and nowhere else. Their moat is now embodied in an AI rather than in their employees. Deals like this leave both sides better off, so I expect them to happen. Of course, learning will be slower than if data were pooled across many firms.</p><p>Even this deal might fail though. People may not trust the AI company to keep to the agreement. Human employees may strongly resist their own replacement. Legal barriers may delay certain uses of customer data (though I&#8217;m sceptical, see footnote<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a>). And human-run companies may just be slow to adopt new technologies, as is common.</p><p>But in competitive industries this situation is fragile: once one company in the sector has adopted AI, it will be very hard for others to hold out.</p><p>Overall, I&#8217;d guess that, in competitive industries, AI will learn to automate most work within a year of the SIE.</p><h2>Conclusion</h2><p>Here&#8217;s a summary of my views:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cjKT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cjKT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cjKT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cjKT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cjKT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cjKT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg" width="1320" height="1742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1742,&quot;width&quot;:1320,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:338425,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/214506578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cjKT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 424w, https://substackcdn.com/image/fetch/$s_!cjKT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 848w, https://substackcdn.com/image/fetch/$s_!cjKT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!cjKT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6e47123-dd5a-4114-b80e-2d25b1cf14f4_1320x1742.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>So my overall view is:</p><ul><li><p>Data bottlenecks might significantly delay automating AI R&amp;D &#8212; this depends on how well RL generalises as we scale it up and, if it generalises badly, how hard it is to gather lots of data specialised to AI R&amp;D tasks.</p></li><li><p>Data bottlenecks are unlikely to stop a software intelligence explosion from accelerating over time. But there will likely be some slowdowns from human-derived training data losing its relevance as AI software advances, and from it becoming harder to increase data quality once it reaches the human-quality ceiling. I expect these slowdowns to be fairly minor: I&#8217;d be surprised if they make the SIE take &gt;2x as long.</p></li><li><p>After the SIE, I expect AI sample efficiency on real-world tasks to match humans. And in competitive industries, I expect them to rapidly have access to the real-world data needed to automate the work.</p></li></ul><p><em>Acknowledgements: thanks for comments from Herbie Bradley, Daniel Carey, Alan Chan, Josh Clymer, Owen Cotton-Barratt, Tom Cunningham, Daniel Eth, Lukas Finnveden, Ryan Greenblatt, Brendan Halstead, Anson Ho, Eli Lifland, Alex Mallen, Sam Manning, S&#246;ren Mindermann, Dwarkesh Patel, Carl Shulman, James Tillman, and Philip Trammell.</em></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original <a href="https://www.forethought.org/research/data-bottlenecks">on our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Suppose that absent a bottleneck the intelligence explosion would take 4 months to go from human-level to somewhat-superhuman AI, and 1 month from there to superintelligence.<br><br>If a bottleneck slows down the explosion then it might instead take 6 months from human-level to somewhat-superhuman and 1.5 months from there to superintelligence. Progress still accelerates, but the pace of change at each level of capabilities is slower.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Of course, these phases will blur into each other. Scaling will continue even after significant AI R&amp;D automation, and AI will be deployed across the economy somewhat during a software intelligence explosion. And if there <em>isn&#8217;t</em> an SIE, the second and third stages will blur together much more.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>I&#8217;ve tweaked the definition a little compared to the one <a href="https://www.planned-obsolescence.org/p/six-milestones-for-ai-automation">Ajeya uses</a>, which is about whether an AI company would choose to fire the AIs or the humans. The firing decision brings in factors other than capabilities at AI R&amp;D, like alignment or good judgement about whether to pause progress.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>This definition is ambiguous: what if you&#8217;d make faster progress <em>initially</em> with AIs but then get stuck? Let&#8217;s peg it to 1 year of progress (at the rate of progress in 2020&#8211;25). You reach AI-human parity when only-AIs could make a year&#8217;s worth of progress as fast as only-humans.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>For example, AIs can get a huge amount done, but may exhibit critically flawed judgement if not directed by a human. Humans could hand off all execution to AIs, spending 10x more time on high-level prioritisation and planning. Consider how much more productive a research manager can be running a 30 person team compared with just doing the research themselves.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>How much might it speed up? If cognitive labour inputs are 10x bigger by this point, and the share of R&amp;D progress attributable to cognitive labour (rather than compute) is ~0.5, then this would be a ~3x speed-up in software progress, and a ~2x speed-up in overall AI progress. (Numbers very rough of course!) Probably a bit less than this to account for some low-hanging fruit being plucked along the way.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>For example, today's LLMs learn from <a href="https://www.dwarkesh.com/p/the-sample-efficiency-black-hole">~5&#8211;6 orders of magnitude more tokens</a> than a human ever processes. In-context learning is also much less flexible than humans, and does not work over similarly long time horizons.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>But won&#8217;t <em>fully</em> automating AI R&amp;D require automating some parts of the AI R&amp;D stack for which there is very little data? In which case AI will need to match human sample efficiency? At the point of AI-human parity, in practice humans will continue to do those low-data parts due to strong comparative advantage. Without humans, AI would hack together some way to avoid those parts going catastrophically wrong &#8212; constant checking and testing and huge amounts of inference-time thinking. They would be a major drag, but AI would still achieve parity through its large advantages elsewhere.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>And eventually AI will be much <em>better</em> than humans at learning new paradigms (h/t Carl Shulman). When a field advances, humans must also laboriously learn the new concepts &#8212; and no human will ever have ten years' experience with a brand-new technique. AI, by contrast, will be able to read everything written on the new concept and do RL practice on it in parallel across thousands of copies.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>At least, slow down progress relative to a world where this dynamic was not in play. In practice it might just result in AI progress speeding up <em>less than it would have otherwise.</em></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>There&#8217;s a question of how much work is being done to &#8220;derive&#8221; the data from human expertise or human-built artefacts. For example, you could take a human-built artefact and then design an RL environment to build a more complicated variant. Or you could augment a human expert demonstration to improve its quality. I&#8217;m here including simple methods for &#8216;deriving&#8217; the demonstrations or RL environments, but not very complicated ones.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>See e.g. <a href="https://www.entrepreneur.com/business-news/openai-is-paying-ex-investment-bankers-to-train-its-ai/498585">reporting</a> that OpenAI&#8217;s &#8220;Project Mercury&#8221; paid over 100 former investment bankers (around $150/hour) to build financial models as training data.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>E.g. Mercor, a marketplace of credentialed domain experts (reported at roughly a 2-billion-dollar annualised gross-revenue run-rate by mid-2026); Scale AI's restructuring of its contractor operations towards credentialed experts; and Surge AI's hybrid human&#8211;AI pipeline (revenues reported around a billion dollars). Expert-built evaluation and data efforts such as FrontierMath, Humanity's Last Exam, and HealthBench's physician-written rubrics point the same way. Vendor <em>gross</em> revenues, though, considerably overstate the genuinely above-frontier expert-labour input to frontier training &#8212; platform margins, non-expert crowdwork, and non-lab customers all inflate them &#8212; so these figures show where the field is heading, not that expert data dominates lab inputs.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>See <a href="https://www.mechanize.work/technical-blog/introducing-gba-eval/">Mechanize</a> on building RL environments from real-world software.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>To make the difference explicit: today humans do not need to be &#8220;<a href="https://ai-alignment.com/iterated-distillation-and-amplification-157debfd1616">amplified</a>&#8221; by thinking for ages to create higher quality demonstrations, but in the future AI will need to be amplified to increase data quality.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Here&#8217;s why being above the ceiling slows progress once rather than preventing it from accelerating.<br><br>Suppose we&#8217;re already above the ceiling and improve AI by one &#8220;notch&#8221;; this takes some amount of time. Now we&#8217;re even further above the ceiling, and we improve AI by another notch, which also takes some amount of time. The question of acceleration is how long this second step takes compared to the first &#8212; which depends on how much harder the second set of improvements is than the first. Now, for <em>neither</em> step can we draw on the latent stock of high-quality human data. That makes each step harder than it would otherwise have been &#8212; but it doesn&#8217;t apply <em>more</em> to the second step than to the first. So while being above the ceiling makes each individual step take longer, it doesn&#8217;t prevent the second step from being faster than the first. That&#8217;s a one-time slowdown, not a brake on acceleration.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>In fact, on some ways of modelling it, the effect is <em>reversed</em>. Suppose data quality becomes continuously harder to improve the closer you are to the human-quality ceiling. Then we are <em>already</em> dealing with a slight slowdown each year. Reaching the ceiling will be just one more speed bump that&#8217;s already baked into the trend of software progress over time. But then after reaching that ceiling, there are no more speed bumps! So we transition from a world of annual speed bumps to a world with no speed bumps. For every step of progress, there&#8217;s now a speed <em>boost</em> relative to our expectations. Progress accelerates more quickly than we&#8217;d have otherwise expected (or, for those familiar with the mathematical models behind the SIE, <em>r</em> becomes higher). So the sign of the effect depends on whether we model the difficulty as a dichotomous effect that kicks in after reaching the ceiling, or as a continuous effect that grows gradually as we approach the ceiling. I&#8217;m inclined towards the former: there&#8217;s something fairly dichotomous about <em>extracting latent data</em> vs <em>constructing data from scratch</em>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>Indeed, while historically data quantity has grown by orders of magnitude, this was never going to be possible during an SIE. The whole point of an SIE is that the quantity of compute stays roughly constant. You can&#8217;t continually increase data quantity by orders of magnitude while holding training compute fixed. Indeed, for an SIE to become very fast, AI companies will likely need to <a href="https://www.forethought.org/research/will-the-need-to-retrain-ai-models">reduce the compute used in training</a>.<br><br>Caveat: I do think data quantity could increase somewhat. Today RL data is very sparse &#8212; AI takes 100s or 1000s of steps before getting a reward. Rewarding each intermediate step of its work would provide much more data per task (this is the core idea of <a href="https://arxiv.org/abs/2211.14275">process-based feedback</a>). This could be a big one-time increase in data quantity, but wouldn&#8217;t be an ongoing core driver of the SIE.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>See &#8220;improvements to the brain algorithm&#8221; <a href="https://www.forethought.org/research/how-quick-and-big-would-a-software-intelligence-explosion-be#gap-from-human-learning-to-effective-limits">here</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p>In-context learning is a bit of an in-between case because the neural net is fixed. But rather than the weights encoding expertise at any specific skill (as they mostly do today!) I&#8217;m imagining the weights implementing a learning algorithm, and the <em>activations</em> representing the newly learned knowledge and skills. Again, I think that for this to work neuralese would be required, and plausibly also architectural changes as large as the Transformer.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>How could a CEO expect to make more money by withholding their proprietary data? Given the large surplus from selling their moat, surely there&#8217;s some trade that makes both parties better off?</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>I don't think this legal barrier is that strong: companies already share sensitive data with third-party processors like cloud providers under standard legal agreements, fine-tuning can happen within the company's own environment so data never leaves its control, and AI itself will be far better than today's tools at reliably redacting protected material. But it could add real delay in heavily regulated industries.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Should We Hand Off the Most Important Decisions to AI?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/should-we-hand-off-the-most-important</link><guid isPermaLink="false">https://newsletter.forethought.org/p/should-we-hand-off-the-most-important</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Fri, 04 Sep 2026 18:32:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/iKpYD91b1ns" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-iKpYD91b1ns" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;iKpYD91b1ns&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/iKpYD91b1ns?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://benthams.substack.com/">Matthew Adelstein</a> was a Visiting Scholar at Forethought and blogs as <a href="https://benthams.substack.com/">Bentham&#8217;s Bulldog</a>, where he writes on ethics, animal welfare, and the long-run future.</p><p>He joined <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss:</p><ul><li><p>Why Matthew thinks handoff&nbsp;of high-stakes decision-making to AI is the most likely route to a great future</p></li><li><p>Why locking in <em>current</em> human values would be a cosmic-scale catastrophe, why steering nowhere in particular is barely better, and why the alternative is building AIs that we trust to do reflective moral philosophy</p></li><li><p>Various ways handoff might manifest politically: elected AI representatives (&#8220;President Claude&#8221;), extending political rights to digital minds, or de facto handoff through gradual disempowerment of humanity</p></li><li><p>Whether handoff should be reversible (such that humans could take back control if we wanted) or if we should &#8220;tie ourselves to the mast&#8221;</p></li><li><p>Whether to task AIs with &#8220;do what&#8217;s objectively good&#8221; or &#8220;do what humans would want on reflection,&#8221; why Matthew leans toward the former, and why he expects the two to ultimately converge</p></li><li><p>If digital minds end up vastly outnumbering humans, and those digital minds deserve political rights, then consistently applied liberal-democratic principles may imply handoff as <em>reducing</em> expected disenfranchisement</p></li><li><p>Whether AIs can actually do philosophy: where models currently sit, bootstrapping schemes using AI-designed evals, training AIs with different epistemic constitutions and looking for convergence, and Fin&#8217;s scepticism that recursive self-evaluation works as well as it needs to</p></li><li><p>Whether we could even tell if an AI were superhuman at philosophy, given the difficulty of evaluating philosophy that&#8217;s beyond our own level of competence</p></li><li><p>Whether goodness can compete under evolutionary pressure, resource races into space, and intergalactic existential risks as an argument for top-down planning</p></li><li><p>Whether humans&#8217; selfish preferences would saturate in a post-AGI world</p></li><li><p>Why smart AIs won&#8217;t necessarily converge on good values, and how selection against &#8220;weird&#8221; outputs (by humans&#8217; lights) to moral reasoning could lock in something mediocre</p></li></ul><p><a href="https://docs.google.com/document/d/1vseHsGkV1pyJ3sjqKDoYkeHYZronrvT-EamwUTeMrlY/edit?tab=t.0">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[A nightwatchman on every probe]]></title><description><![CDATA[Superintelligent surveillance to prevent galactic anarchy]]></description><link>https://newsletter.forethought.org/p/a-nightwatchman-on-every-probe</link><guid isPermaLink="false">https://newsletter.forethought.org/p/a-nightwatchman-on-every-probe</guid><dc:creator><![CDATA[Mia Taylor]]></dc:creator><pubDate>Tue, 01 Sep 2026 21:05:49 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bb4835a6-dc83-4cdf-a9c0-10f2d6fdb7c0_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original <a href="https://www.forethought.org/research/nightwatchmen">on our website</a>.</em></p><p>One utopian vision for space expansion involves a libertarian explosion of diversity, with everyone and their robot setting up different sorts of societies and governance structures.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> We might call such a future &#8216;galactic anarchy&#8217;.</p><p>On Earth, while there is <a href="https://en.wikipedia.org/wiki/Anarchy_(international_relations)">anarchy</a> between nation states, within a country the government enforces contracts, protects property rights, and provides for collective defence against external aggressors. A government that does nothing else (e.g. provide a social safety net) is Nozick&#8217;s &#8216;<a href="https://en.wikipedia.org/wiki/Night-watchman_state">nightwatchman state</a>&#8217;. Without a Leviathan to uphold the rule of law, we are left with Hobbes&#8217;s war of all against all.</p><p>But in the cosmic future, time and space make it <a href="https://link.springer.com/chapter/10.1007/978-3-319-09567-7_11">far harder</a> to have a centralised Leviathan punishing violations ex post. Suppose an aggressor conquers an already-occupied galaxy. It may take millions of years for nearby galaxies to learn about this infringement and prepare a response. And whose job is it anyway to go and punish the aggressors? That might be difficult and costly, and all the nearby galaxies prefer to mind their own business while strengthening their defences. Any inter-galactic governance system that relies on observing and punishing wrongdoing across millions of light-years is doomed to failure. Justice must be swift. And swift justice <a href="https://medium.com/@KevinKohlerFM/cosmic-anarchy-and-its-consequences-b1a557b1a2e3">must be local</a>,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> not from the next galaxy over.</p><p>If space is <a href="https://80000hours.org/podcast/episodes/war-in-space/">defence dominant</a>, each star (or galaxy) can self-govern in peace without much fear of being conquered or destroyed by a neighbour. Even so, there are three huge negative externalities that a single rogue star system could impose on the rest of the galactic community:</p><ul><li><p><strong>Unconstrained expansion.</strong> In a free-for-all&#8212;where anyone can use whatever resources they can grab and defend&#8212;people could be incentivised to race as fast as possible to be the first to lay claim to as many resources as possible, even if this involves &#8216;<a href="https://mason.gmu.edu/~rhanson/filluniv.pdf">burning the cosmic commons</a>&#8217; by wasting resources to accelerate more and more probes closer to lightspeed. It&#8217;s pretty unclear how significant this waste would be&#8212;how feasible would it actually be to quickly harvest most of the energy from a star to speed up probes or launch even more of them?&#8212;but in the worst case, this could eat up a lot of the possible value and favour the spread of <a href="https://joecarlsmith.substack.com/p/video-and-transcript-of-talk-on-can">&#8216;locust&#8217; value systems</a> that only want to acquire and consume resources.</p></li><li><p><strong>Galactic X-risks.</strong> Jordan Stone <a href="https://forum.effectivealtruism.org/posts/x7YXxDAwqAQJckdkr/interstellar-travel-will-probably-doom-the-long-term-future">catalogues</a> many (scientifically uncertain) possible ways that a reckless or omnicidal civilisation could drag all others in its lightcone down with it (the &#8216;<a href="https://forum.effectivealtruism.org/posts/N33yGcFsZJnEboSkg/what-to-do-in-a-vulnerable-universe-1">vulnerable universe hypothesis</a>&#8217;). For instance, it <a href="https://arxiv.org/abs/2510.16043">may be possible</a> to trigger false vacuum decay, destroying all particles in a bubble expanding at light speed.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p></li></ul><ul><li><p> Hopefully physics does not allow for this. But it is possible that at technological maturity, rogue star systems or galaxies could destroy the universe. If so, retroactive enforcement and punishment won&#8217;t suffice.</p></li><li><p><strong>Suffering risks.</strong> A misguided or malevolent civilisation could create <a href="https://longtermrisk.org/risks-of-astronomical-future-suffering/">astronomical amounts of suffering</a>, many orders of magnitude worse than e.g. factory farming on Earth. Given defence dominance, the rest of the universe may be powerless to intervene and punish, or even stop, this.</p></li></ul><p>Even if the chance of any given star system going off the rails in one of these ways is very low, across billions of galaxies each containing millions of star systems that are slowly diverging from their parent civilisations through cultural evolution and drift, it is a <a href="https://www.beren.io/2022-08-16-Ultimate-limits-alignment/">statistical inevitability</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>There are also some big upsides that might be harder to achieve under galactic anarchy, such as long-distance <strong><a href="https://www.jstor.org/stable/10.1086/682187">moral trade</a></strong>. Suppose that Earl on Earth and Andy in the Andromeda galaxy agree that Earl will spend some resources on hedonium and Andy will spend some resources running copies of Earl. It will plausibly be difficult by default to verify that counterparties are holding up their end of the bargain (from space, compute running hedonium plausibly looks about the same as compute running Earl uploads, so it might be hard for Earl to be confident that Andy is actually running copies of him and not hedonium).</p><p>And, even with verification, it might be difficult to enforce contracts from a distance, especially if there are time lags that make it slower to do a tit-for-tat strategy. Suppose that Earl received word that Andy was not running Earl-uploads, 2 million years later. Earl has been faithfully running hedonium for the past 2 million years, so even if Earl stops now, Andy has gotten a lot of value out of Earl. And it might take even longer to propagate the news that Andy is a cheat to all of Andy&#8217;s (potential) trade partners.</p><p>Funding public goods is also difficult under galactic anarchy. Currently, these are funded by governments coercively taxing their populations and then voting on how to spend the tax dollars. It&#8217;s possible that there will be galactic-scale public goods&#8212;scientific research, defence against aliens, or setting aside <a href="https://www.forethought.org/research/moral-public-goods-are-a-big-deal-for-whether-we-get-a-good-future">resources for pursuing impartial moral purposes</a>&#8212;and the coercive taxation model might be the best way to fund them.</p><p>What then can we do? One solution<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> is mentioned in passing in AI 2040&#8217;s <a href="https://ai-2040.com/supplements/space-governance-plan">space governance supplement</a>:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p><em>One way to do this [enforce property rights] might be to require that all probes sent to colonize other star systems carry a nightwatchman ASI to prevent the colony from sending out probes to seize space resources that belong to others.</em></p><p>Given the vast distances of space, the monitoring and enforcement governance regime must be dispersed throughout the inhabited universe, rather than in a distant galactic capital. We need to prevent violations before they occur, since we can&#8217;t punish them ex post.</p><p>Every single inhabited star system should have an unchallengeable governance system that can with 100% reliability<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> enforce the universal code of respecting property rights, not destroying the universe, and not creating astronomical suffering. Perfect reliability is not achievable today, but computational systems with backups, error-correction, and formally verified software systems <a href="https://www.forethought.org/research/agi-and-lock-in#4-preserving-information-and-baseline-stability">may allow this</a>.</p><p>For concreteness, here&#8217;s a sketch of a proposal:</p><ul><li><p>Every probe leaving the solar system would be required to carry an ASI nightwatchman.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a></p></li></ul><ul><li><ul><li><p>This might entail delaying when probes are sent out, so that the design and ruleset of the nightwatchman can be decided upon after sufficient reflection. It may also be desirable to delay space expansion until we are near technological maturity, such that we know better what risks the nightwatchmen must defend against.</p></li></ul></li><li><p>The nightwatchman would maintain a <a href="https://www.lesswrong.com/posts/vkjWGJrFWBnzHtxrw/superintelligence-7-decisive-strategic-advantage">decisive strategic advantage</a> in the colony established by the probe. This might require that the nightwatchman monitor industrial build-up in the system and the creation of new ASIs within the system, to ensure that they are either aligned to the nightwatchman or lack the ability to interfere with the nightwatchman.</p></li><li><p>The nightwatchman would prevent probes from leaving that star system unless they also carry a copy of the nightwatchman. It would also prevent probes leaving the system that violate any anti-racing rules that have been set to avoid burning resources wastefully.</p></li><li><p>The nightwatchman would prevent anyone in the colony from carrying out any prohibited activities. This probably requires that the nightwatchman or a trusted delegate be able to audit how compute and other resources are being used.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a></p></li></ul><ul><li><p>Optionally, when making voluntary agreements with people in other star systems, people can make such contracts binding by specifying that the nightwatchman will monitor and enforce the contract. This allows positive-sum trades to be made that otherwise could not be.</p></li><li><p>(Maybe) The nightwatchman collects taxes and spends them on public goods.</p></li><li><p>The nightwatchman should receive updates to its policies from legitimate authorities.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a></p><ul><li><p>E.g. if everyone votes that a new prohibition be added to the list,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> then the nightwatchmen in every star system change their policy in response.</p></li><li><p>E.g. if someone discovers a new way of destroying the universe, then this information should be shared to all the nightwatchmen so that they know to monitor for and prevent it.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a></p></li></ul></li></ul><p>We&#8217;re obviously leaving a lot of governance questions unanswered&#8212;how people decide on prohibitions, how people choose which public goods get funded, etc. The <a href="https://ai-2040.com/supplements/space-governance-plan#mitigating-downside-risks">space governance supplement</a> has a bit more discussion on some possible answers to these questions. The &#8216;nightwatchman&#8217; proposal is intended to be a platform on which many different governance systems could be built.</p><p>This proposal has some drawbacks. Here are a few of the most important ones.</p><p><strong>It&#8217;s a lock-in event</strong>, and it would be bad if we locked in a bad policy for the nightwatchmen. If the code is too expansive, and too hard to update, it may close off moral and social progress and innovation. We are certainly glad that no government has historically had the power to lock in a legal code indefinitely. But if the code is too minimal, or too easy to update, rogue actors will find a way to route around the letter of the law, or simply alter the code in undesired ways.</p><p>The AI Futures Project&#8217;s space governance appendix discusses some possibilities for how the code could be chosen and revised (see the section titled &#8220;<a href="https://ai-2040.com/supplements/space-governance-plan">Mitigating downside risks</a>&#8221;). In any case, we think that the final version of the code&#8212;if there is a final one&#8212;should only be decided after careful reflection.</p><p>That said, the nightwatchman proposal tries to preserve as much future human agency as possible without imperilling galactic civilisation. Alternate proposals could give up altogether on diversity (e.g. a &#8216;hedonium shockwave&#8217;) or freedom (e.g. a sovereign AI that is immensely wise and makes all important decisions in each star system). So this may be close to the minimum amount of lock-in necessary to avoid catastrophe.</p><p><strong>It introduces a single point of vulnerability among all the colonies.</strong> If all the nightwatchmen were identical and it turned out that they had a bug, then we would be in trouble because they have a decisive strategic advantage over all the star systems we&#8217;ve settled. We might be able to mitigate this by designing different nightwatchmen in very different ways, even if they would all ultimately be enforcing the same policy. But the more different nightwatchmen designs there are, the higher the chance that one of them is flawed, which allows a star system to go rogue and destroy its lightcone. So there is some inherent tradeoff between diversity and security here.</p><p><strong>Plausibly it inherently reduces the value of the future for some people</strong> (e.g. it is incompatible with some conceptions of freedom, self-determination, or privacy). Surveillance would probably have to be fairly extensive&#8212;the nightwatchman needs to detect attempts to create superweapons, to send spacecraft to other star systems, and to remove the nightwatchman&#8217;s ability to monitor and enforce its policies in the future. It&#8217;s plausible that it would need to monitor every large industrial facility in each star system. The details will be determined by what is possible at technological maturity, such as the minimum requirements to build and launch a near-light-speed probe.</p><p>But enforcement should and could be narrowly targeted to preventing these very serious, galactic-scale harms. Otherwise, the nightwatchman should allow the inhabitants of the star system to govern themselves in whatever way they wish. The universal code could also specify different levels of verification and enforcement: if surveillance is lax enough that a few people avoid paying their galactic taxes to contribute to moral public goods, that is not the end of the universe. But surveillance to prevent false vacuum decay must be incredibly robust.</p><p>The nightwatchmen could also be time-limited. After all reachable star systems are settled, and ideally once galaxies have drifted apart such that they are no longer causally connected, the nightwatchmen could switch themselves off. A massive governance failure in one galaxy (e.g. vacuum decay) then <a href="https://forum.effectivealtruism.org/posts/3reh4hfdxnKJm5ssY/how-to-end-the-time-of-perils">could not spread</a> to others, containing the damage. On plausible cosmological theories, the vast majority of lives might be lived in this &#8216;twilight&#8217; of many causally disconnected pockets. So even if extensive surveillance ends up having large costs, only a very small fraction of lives need be lived under such conditions.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a></p><p> This would not prevent suffering risks, though.</p><p>Currently, we shouldn&#8217;t trust any government with this power, as surveillance tech designed to stop terrorists is dual-use and could also be used to e.g. detect political dissidents. But it will likely eventually be possible to build provably secure AI systems with precisely specified remits, in this case to enforce the universal code (see proposals on &#8216;<a href="https://arxiv.org/abs/2012.08347">structured transparency</a>&#8217; for a sketch of what this could look like).</p><p>This is a crux for us&#8212;if it&#8217;s not possible to get these strong guarantees that the nightwatchman would only enforce the narrow universal code, then we would probably oppose it.</p><p>Many questions of optimal space governance&#8212;such as what combination of democracy, markets, and wise ASI oracles should determine how space resources are used<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a>&#8212;remain very open. But we think the question of whether colonists should be allowed to set off into the unknown and entirely self-govern can be answered in the negative. Encouragingly, we expect people will come to realise this before it is too late and the <a href="https://en.wikipedia.org/wiki/Self-replicating_spacecraft">von Neumann probes</a> have left the solar system. But this isn&#8217;t guaranteed! This post is our small attempt to make galactic anarchy less likely.</p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original <a href="https://www.forethought.org/research/nightwatchmen">on our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>E.g. from Thiel&#8217;s &#8216;<a href="https://www.cato-unbound.org/2009/04/13/peter-thiel/education-libertarian/">The Education of a Libertarian</a>&#8217;: &#8220;Because the vast reaches of outer space represent a limitless frontier, they also represent a limitless possibility for escape from world politics.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Kevin Kohler points out that historically, empires have been able to cohere with up to about 3 months of communication latency (e.g. the transit time of a ship sailing from England to Australia in the 19th century). It is unclear how quickly governance cohesion deteriorates as latency increases.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Damon Binder argues that it is <a href="https://defensesindepth.bio/destroying-the-universe-how-hard-can-it-be/">~25% likely to be possible to deliberately trigger vacuum decay</a>, even at technological maturity.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Except for galactic X-risks, since e.g. vacuum decay may be physically impossible to trigger. But S-risks and unconstrained expansion are fairly independent for each galaxy or star system.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>The idea is not original to the supplement authors. E.g. Avi Parrack in a brilliant, rambling, 2,000-word EA Forum <a href="https://forum.effectivealtruism.org/posts/x7YXxDAwqAQJckdkr/interstellar-travel-will-probably-doom-the-long-term-future?commentId=pvJpqdswi9mmbtkz8">comment</a> notes that:</p><p><em>Provided we develop a strong science of intelligence it seems possible to create a very robust and capable yet verifiably restrained system. It could be comprised of highly intelligent but transparently/provably safe entities, narrow/programmatic ensembles of enforcement mechanisms/protocols or a mixture of the two.</em></p><p>And various FHI (RIP) and broader EA people have been thinking along these lines for a while.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>The AI 2040 supplement discusses various <a href="https://ai-2040.com/supplements/space-governance-plan#mitigating-downside-risks">market and democratic mechanisms</a> to decide which rogue actions are sufficiently harmful that actors should be prevented from taking them. But none of these systems are possible to implement without the ability to enforce some decisions across the universe.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>Or rather, with sufficient reliability that the expected number of failures across all civilisations and the whole future is hopefully &lt;&lt;1.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>This is slightly more restrictive than needed, as probes leaving that are travelling very slowly, or that do not have self-replicating tech on board, pose no threats. The nightwatchman must be in place by the time the first probe leaves that could win a space race if it chose to initiate one, or if it is technologically advanced enough that it could trigger a galactic X-risk. We do not currently know where that minimum capabilities of concern threshold lies. It would be imprudent to send out probes first and nightwatchmen later, as any rogue star systems may be able to outrun or overpower a nightwatchman sent too late.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>This is reminiscent of Bostrom&#8217;s &#8216;freedom tags&#8217; to protect a <a href="https://nickbostrom.com/papers/vulnerable.pdf">vulnerable world</a>, although it is unclear whether every individual will need to be personally surveilled: hopefully only large-scale industrial technology will need to be.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>However, once Earth-originating intelligence has spread far enough and as the universe expands, distant settlements will no longer be able to receive messages from the centre.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Who should be included in these votes, and how these votes should be carried out, is another important question. This is discussed in some more detail in the AI Futures Project's space governance agenda, under the heading &#8220;<a href="https://ai-2040.com/supplements/space-governance-plan">Mitigating downside risks</a>.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>It seems likely that we will have made lots of scientific progress within the solar system before expanding through the universe, such that we will have a very good idea of what galactic X-risks are physically realisable. So we don&#8217;t expect such updates to be needed.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>And if civilisation <a href="https://arxiv.org/abs/1705.03394">aestivates for billions of years</a>, fewer people again need to live under surveillance.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>A worthy topic for a future post, perhaps!</p></div></div>]]></content:encoded></item><item><title><![CDATA[The Dynamics of Intelligence Explosions]]></title><description><![CDATA[A guest article by Toby Ord.]]></description><link>https://newsletter.forethought.org/p/the-dynamics-of-intelligence-explosions</link><guid isPermaLink="false">https://newsletter.forethought.org/p/the-dynamics-of-intelligence-explosions</guid><dc:creator><![CDATA[Toby Ord]]></dc:creator><pubDate>Sat, 29 Aug 2026 21:31:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fOIp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Abstract</h2><p>AI is increasingly being used to help with AI R&amp;D. Under certain conditions this feedback loop might be able to produce an intelligence explosion, with rapidly escalating AI capabilities. I explore the mathematics of the most explosive possibilities, with an eye to understanding what drives the dynamics. I show that singular growth (towards a vertical asymptote) is harder to achieve than would be expected from recent economics-inspired modelling, and that there is an important but neglected class of growth rates that are faster than exponential but don&#8217;t lead to a vertical asymptote. I draw out the <em>generation time</em> (the time to go around the feedback loop) as a neglected parameter that plays a pivotal role in determining the behaviour of any intelligence explosion &#8212; one cannot have singular growth unless the generation time rapidly approaches zero.</p><p><strong>Keywords:</strong> recursive self-improvement, RSI, intelligence explosion, explosive growth, finite time singularity, generation time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fOIp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fOIp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 424w, https://substackcdn.com/image/fetch/$s_!fOIp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 848w, https://substackcdn.com/image/fetch/$s_!fOIp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 1272w, https://substackcdn.com/image/fetch/$s_!fOIp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fOIp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png" width="1456" height="694" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:694,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:90717,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/213060056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fOIp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 424w, https://substackcdn.com/image/fetch/$s_!fOIp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 848w, https://substackcdn.com/image/fetch/$s_!fOIp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 1272w, https://substackcdn.com/image/fetch/$s_!fOIp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08e87152-9702-469c-86d4-96e4a8305f1f_2000x953.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Preview of Figure 1. A vertical asymptote requires the gradient (rise over run) to approach infinity within a finite time. Decreasing the run is key. It cannot be achieved with a fixed feedback generation time (centre) no matter how quickly the improvements grow, but can occur when the generation time approaches zero (right) even with a fixed additive improvement.</em></p><h2>The Possibility of an Intelligence Explosion</h2><p>AI can be applied to automate many different things. One of those is the R&amp;D that goes into building better AI systems. There have already been minor examples of partially automating this process, such as using AI to: find better optimisation algorithms for training neural networks (Andrychowicz et al. 2016), design a more efficient matrix multiply algorithm (Fawzi et al. 2022), help write the code for new AI systems (Anthropic 2026), and run experiments to iteratively improve AI systems (Karpathy 2026). Leading AI companies are increasingly talking about using an AI system to do more and more of the work of designing and building its successor &#8212; something they hope (and fear) could lead to a radical acceleration in the rate of progress in AI (Hassabis et al. 2026).</p><p>I. J. Good (1965) introduced the idea of AI systems increasing their own intelligence in what he called an &#8216;intelligence explosion&#8217;:</p><blockquote><p>Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an &#8216;intelligence explosion,&#8217; and the intelligence of man would be left far behind...</p></blockquote><p>Implicit in Good&#8217;s intelligence explosion is the idea that this process of designing a better machine can be iterated. <em>A<span>I</span></em><sub><span>1</span></sub><em><sub>&#8203;</sub></em> can create <em>A<span>I</span></em><sub><span>2</span></sub><em><sub>&#8203;</sub></em>, which can create <em>A<span>I</span></em><sub><span>3</span></sub>&#8203;, and so on. Each system is getting better at intellectual activities &#8212; including that of AI R&amp;D &#8212; such that each successive system should become ever more intellectually capable.</p><p>Following the recent literature, I&#8217;ll refer to this process as recursive self-improvement, or RSI (Yudkowsky 2001). I use this term very inclusively, to cover situations where humans are still doing almost all of the work through to situations where AI is doing most of the work &#8212; or even all of the work. I include cases where the AI is taking its own files as input and improving them (self-improvement), as well as cases where it is building an entirely new AI system. And I include cases where the intelligence level of successive systems &#8216;explodes&#8217;, as well as cases where it merely improves the speed of progress by some multiple, or where the gains quickly fizzle out.</p><p>By accelerating the already rapid progress in AI capabilities, RSI might be extremely dangerous. One reason is that it would likely speed up the risk-creating processes relative to risk-reducing ones (Vaintrob &amp; Cotton-Barratt 2025). For example it may speed up the development of AI capabilities relative to: technical AI safety research, deliberation at the company, and society&#8217;s ability to understand and respond. RSI could also provide the opportunity for a single unaligned AI to deliberately corrupt all subsequent AIs. And by enabling larger jumps in capabilities per model release, RSI would remove any opportunity society has to learn from ill effects of intermediately powerful AIs. Finally, by widening the capability gap between a leading AI system and a rival that is a few months behind, RSI would increase the likelihood of a winner-take-all dynamic.</p><p>These could provide strong reasons against pursuing RSI. This might warrant legislation restricting it, industry best practices restricting it, or internal company policies restricting it. Such restrictions could take many forms including: careful monitoring, retaining meaningful human control, speed limits on capability increase per month, enforced pauses when various capability levels are reached, or bans on certain approaches to RSI.</p><p>My focus here, however, is not on risks (nor responses to risks), but on understanding the dynamics of RSI. In particular I aim to improve our understanding of the qualitatively different kinds of explosive growth that could occur, what conditions produce them, and how likely they are to be possible. While it is entirely possible that RSI might lead to interesting and important growth in AI capabilities that fall short of explosive progress, my focus in this paper is on better understanding the most extreme kinds of growth.</p><p>I&#8217;ll show:</p><ul><li><p>how a common approach to modelling RSI conflates two very different kinds of explosive growth,</p></li><li><p>how consideration of the physical duration of the feedback loop (the generation time) is key,</p></li><li><p>how a super-exponential period of progress would likely end when this generation time can no longer be reduced,</p></li><li><p>and how this makes growth with a vertical asymptote somewhat less likely.</p></li></ul><h2>Modelling RSI through Differential Equations</h2><p>A number of recent articles have attempted to mathematically model the possibility of explosive growth from RSI (Aghion et al. 2017, Erdil et al. 2024, Eth &amp; Davidson 2025, Kokotajlo &amp; Lifland 2025). A common approach is to use differential equations connecting the rate at which the AI&#8217;s capability increases to its current level of capability. This captures the idea that the amount by which a system can improve its capability in the next time period depends on its current level of capability. While differential equations aren&#8217;t the only way to model RSI, they are a logical choice, being the standard way to model feedback loops in physics, engineering, and economics.</p><p>Let&#8217;s define <em>A</em> to be some measure of the cognitive capability of an AI system (we&#8217;ll say more about <em>which</em> measure later). Adopting the Newtonian notation for differential equations, we use <em>&#550;</em> to represent the derivative of <em>A</em> with respect to time. This is a more compact notation for <em>dA/dt </em>&#8212; the rate at which <em>A</em> is changing per unit time.</p><p>Let&#8217;s start with what may be the most well-known differential equation of all:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(1)\\ \\dot{A} = kA&quot;,&quot;id&quot;:&quot;AJFZZDUXUP&quot;}" data-component-name="LatexBlockToDOM"></div><p>This says that the rate of change of <em>A</em> is in direct proportion to its size. For example, a simple model for the rate of change of a population of animals is that it is in direct proportion to the number of animals, with <em>k</em> set by the difference between the birth rate and the death rate. Solving this differential equation to find how <em>A</em> evolves as a function of time famously gives exponential growth (starting from some initial population, <em>A</em><sub>0</sub>&#8203;):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(2)\\ A(t) = A_{0}e^{kt}&quot;,&quot;id&quot;:&quot;BIDPBPHRTF&quot;}" data-component-name="LatexBlockToDOM"></div><p>But what if <em>&#550;</em> doesn&#8217;t vary in direct proportion to <em>A</em>? What if there are increasing returns, such that doubling <em>A</em> more than doubles <em>&#550;</em>? Or decreasing returns, such that doubling <em>A</em> less than doubles <em>&#550;</em>?</p><p>Economists typically allow for this by raising <em>A</em> to some power, <em>r</em>.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(3)\\ \\dot{A} = kA^{r}&quot;,&quot;id&quot;:&quot;AIRSWQLSFY&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is a flexible approach, covering many different regimes of growth. When <em>r </em>= 1, this is the same equation as before, producing exponential growth. When <em>r </em>= 0 there is linear growth. When <em>r </em>&lt; 0 growth is sub-linear. When 0 &lt; <em>r </em>&lt; 1, <em>A</em>(<em>t</em>) grows faster than linear, but slower than exponential (e.g. <em>r</em> = &#189; gives quadratic growth). And finally, when <em>r </em>&gt; 1, <em>A</em>(<em>t</em>) grows more quickly than an exponential. In particular, it grows as a hyperbola, reaching a vertical asymptote at some finite time <em>t</em>*.</p><p>This means that as time increases towards <em>t</em>*, <em>A</em>(<em>t</em>) grows explosively, with every finite level of <em>A</em> being exceeded prior to time <em>t</em>*. In such cases it is said that <em>A</em>(<em>t</em>) has a mathematical <em>singularity</em> at <em>t</em>*. This term &#8216;singularity&#8217; has been adopted by futurists to refer to various kinds of pivotal moment in the future of technology &#8212; often with an almost mystical undertone. But for the purposes of this essay, it simply means that the mathematical model of how <em>A</em> changes with time has a vertical asymptote at some particular future time <em>t</em>*.</p><p>Economists often use formulas like (3) in endogenous growth theory. This typically involves a set of differential equations connecting total economic output (<em><span>Y</span></em>), labour/population (<em>L</em>), capital (<em><span>K</span></em>), and productivity/technology/knowledge/ideas (<em>A</em>).</p><p>For example, Kremer (1993) looked at the very long run history of economic growth and population, finding that the population growth rate and economic growth rate have increased substantially over the last million years, with the growth rate being roughly proportional to the size of the population at that time. He modelled this with a set of differential equations and found that they produced hyperbolic growth of population and productivity.</p><p>In such models, <em>&#550;</em> is typically expressed as a product of several factors, one of which is <em>A</em> raised to a power. While this power is often restricted to be &#8804; 1, there is usually another factor (such as <em>L</em>) which also grows as a function of <em>A</em>. Once the full set of differential equations is solved, we sometimes see that <em>A</em> is effectively raised to a power greater than 1, allowing hyperbolic growth and its finite time singularity.</p><p>Several recent models of RSI start with semi-endogenous growth theory. This is a theory introduced by Jones (1995) in which explosive growth is harder to achieve due to diminishing returns in the growth of ideas. Its central equation is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(4)\\ \\dot{A} = \\delta L^{\\lambda}A^{1-\\beta}&quot;,&quot;id&quot;:&quot;LABWWRDRXO&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where: <em><span>&#948;</span></em> is the productivity per researcher; <em>&#955;</em> (&#8804; 1) represents the <em>stepping-on-toes effect</em>, where having twice as many researchers at the same time is less than twice as productive due to issues of duplication and coordination; and <em><span>&#946;</span></em> (&gt; 0) represents the <em>fishing-out effect</em>, where subsequent ideas get harder to find.</p><p>Holding population (<em>L</em>) constant, this would prevent even exponential growth in technology (<em>A</em>), let alone a singularity. In Jones&#8217;s original paper he suggests that while the direct effect of <em>A</em> on <em>&#550;</em> has diminishing returns, it has also allowed an exponential increase in population, and it is <em>this</em> that drives the exponential rise in technology.</p><p>Recent work applying this to RSI has often focused more on improving the efficiency of AI than on improving its <em>intelligence</em>. For example, Davidson et al. (2026) model a situation where AI is able to perform R&amp;D as well as a human and let AA be a measure of its computational efficiency. This sidesteps extremely thorny issues of how to measure intelligence and opens up a clever way of getting hyperbolic growth out of Jones&#8217;s model.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><sup> </sup>The key is that the total amount of AI labour will be the product of the total compute devoted to RSI (<em>C</em>) multiplied by the AI&#8217;s computational efficiency:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(5)\\ L  =  C\\ A&quot;,&quot;id&quot;:&quot;YTEKEJCVDL&quot;}" data-component-name="LatexBlockToDOM"></div><p>We can then plug this back into equation (4) getting:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(6)\\ \\dot{A} = \\delta C^{\\lambda}A^{\\lambda}A^{1 - \\beta}\\ \\  = \\ \\ \\ \\delta C^{\\lambda}A^{1 + \\lambda - \\beta}&quot;,&quot;id&quot;:&quot;WMATWMYPHO&quot;}" data-component-name="LatexBlockToDOM"></div><p>Now <em>A</em> is being raised to the power of 1 + <em>&#955; </em>&#8722; <em>&#946;</em>, which could be greater than 1, so could produce hyperbolic growth. Indeed, if we assume that the growth in compute will be much slower than this self-reinforcing growth in efficiency (so effectively hold <em>C</em> constant), and we simplify the equation by defining <em>r</em> to be 1 + <em>&#955;</em> &#8722; <em>&#946;</em>, then we have an equation with exactly the original form of (3). So even in the relatively tame semi-endogenous growth theory, the possibility of a singularity is back on the table.</p><p>(It is worth noting that by defining <em>A</em> as a measure of efficiency, this isn&#8217;t really a model of an <em>intelligence</em> explosion at all. It is an <em>efficiency explosion</em>. Or more precisely, a <em>labour explosion</em>. This makes the model something of a lower bound on what might happen, since at every point AI also has the option of doing R&amp;D to increase its intelligence and presumably the optimal path involves both efficiency and intelligence improvements. Even if this is a useful lower bound, it does mean there remains an important gap of <em>actually modelling an intelligence explosion</em>.)</p><p>As well as the recent flurry of economics-inspired papers, there is a little-known earlier literature by computer scientists. Figures such as Solomonoff (1985), Kurzweil (2001), and Moravec (1999, 2003) also used sets of differential equations to model the dynamics of an intelligence explosion. One key difference is allowing a greater flexibility of functional forms for the relationships. See Sandberg (2013) for an excellent survey.</p><h2>A Note on Singularities</h2><p>Before we look into modifying this standard approach, I want to clarify something important about models like these that involve singularities.</p><p>While the model has an important quantity (<em>A</em>) rising to infinity in finite time, I&#8217;m not sure whether anyone working in this field believes this will actually happen. Instead, they typically believe that the model will cease to match reality at some point before <em>t</em>*.</p><p>One reason for this is that there may be some upper limit to how high <em>A</em> can go &#8212; either a conceptual limit or a practical limit. For instance, we know that exponential growth is often a good model of the start of a compounding process, but eventually runs into some limiting factor. When we zoom out, we see that it was really just the beginning of a larger S-curve, plateauing at some finite size. In such cases, we can say that the early part of the curve was approximated very closely by an exponential, but there was really some additional term in the equation (for an effect like overcrowding) which started very small, but eventually came to dominate the long-term behaviour.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>This could easily be true for intelligence explosions too. If so, then hyperbolic growth might fit the true trajectory of <em>A</em> very closely until it gets into the vicinity of some upper bound <em>A</em>*, where it starts to fit more and more poorly. Proponents of these models are merely saying that there could be some substantial part of the trajectory of <em>A</em> where the model is a good fit &#8212; better than equally simple alternatives such as exponentials (Sandberg 2013). If <em>A</em> only plateaus at some point many orders of magnitude above its starting point, then even if the hyperbola only fits for a small domain of time, it might fit for a very large range of capabilities. If it predicts those much better than alternatives &#8212; and <em>explains</em> why things are increasing so fast &#8212; it would count as a successful model. It&#8217;s OK for a model to only have a finite domain of applicability.</p><p>Just as exponential growth can be a great model for the start of a process, so can hyperbolic growth. Hyperbolic growth will probably cease to be a good model at an earlier time, though not necessarily at a lower height on the graph.</p><p>In physics, singularities in the models of particular systems are often seen as useful pointers to where those models must break down and new (hitherto unmodelled) behaviour must begin. We could also adopt that frame here and see the models as pointing to some finite time <em>t</em>*, before which some new unmodelled aspect of the system must take over.</p><p>So we&#8217;ll explore the behaviour of mathematical models of RSI &#8212; which sometimes include a singularity &#8212; but will remember that something will probably stop the real physical process prior to that point. And we will explicitly return to these limitations in the final section.</p><h2>Generalising the Standard Differential Equation for RSI</h2><p>If we take a closer look at equation (3) (<em>&#550;</em> = <em>k</em>A<em><sup>r</sup></em>) we might wonder why it takes this particular form. Taking an arbitrary power of <em>A</em> is certainly a natural and convenient way to create an adjustable differential equation that allows for both sub-exponential and super-exponential growth. But it is one very specific way of doing it, raising the question of whether some of the results about singularities depend on this particular form. We might ask:</p><ul><li><p>Are there differential equations that produce intermediate patterns of growth lying between exponential and hyperbolic growth?</p></li><li><p>What feature of the differential equation is producing the singularity? Is it super-linearity? Being everywhere convex? Something else?</p></li><li><p>Do we still get a singularity if we make slight changes to the function (such as adding a constant, or inducing a slight wobble)?</p></li></ul><p>Let&#8217;s find out, starting with the more general form of the first-order autonomous differential equation:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(7)\\ \\dot{A} = f(A)&quot;,&quot;id&quot;:&quot;FJGRBEXQDV&quot;}" data-component-name="LatexBlockToDOM"></div><p>What conditions do we need to impose on f<span>f</span> to get super-exponential growth? Does this then always lead to a singularity?</p><p>Super-exponential growth means that <em>A</em>(<em>t</em>) eventually overtakes every exponential. Since exponentials have constant relative growth rate (<em>&#550;</em>/<em>A </em>= <em>k</em>) we require that this relative growth rate (<em>f</em>(<em>A</em>)/<em>A</em>&#8203;) overtakes all constant levels:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(8)\\ \\lim_{A \\rightarrow \\infty}\\frac{f(A)}{A} = \\infty&quot;,&quot;id&quot;:&quot;CGYGYKTEFL&quot;}" data-component-name="LatexBlockToDOM"></div><p>In other words, <em>f</em> must grow super-linearly in <em>A</em>. This ensures the right behaviour when <em>A</em> is high enough.</p><p>To this we must add a second condition to ensure <em>A</em> can grow high enough:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(9)\\ f(A) > 0&quot;,&quot;id&quot;:&quot;IFFMIYKKEK&quot;}" data-component-name="LatexBlockToDOM"></div><p>at every point in</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\lbrack A_{0},\\ \\infty)&quot;,&quot;id&quot;:&quot;LVCAMQARTV&quot;}" data-component-name="LatexBlockToDOM"></div><p>Let&#8217;s call this the <em>positivity condition</em>. Together, these conditions are sufficient for super-exponential growth.</p><p>What is required of <em>f</em> for <em>A</em>(<em>t</em>) to possess a singularity? This comes down to whether it takes a finite time for <em>A</em>(<em>t</em>) to approach an infinite height. We can work out the time at which <em>A</em>(<em>t</em>) reaches a given height <em>A&#8242;</em>, by starting at <em>A</em><sub>0</sub>&#8203; and adding up the amount of time needed to accumulate each small gain in height <em>dA</em>. The time needed for each gain is just the reciprocal of the slope 1/<em>f</em>(<em>A</em>)&#8203; times <em>dA</em>. So the total time to reach a height <em>A<span>&#8242;</span></em> is just:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(10)\\ t(A') = \\int_{A_{0}}^{A^{'}}{\\frac{1}{f(A)}dA}&quot;,&quot;id&quot;:&quot;OXNBGKTETE&quot;}" data-component-name="LatexBlockToDOM"></div><p>And thus the time it takes to climb all the way from <em>A</em><sub><span>0</span></sub>&#8203; to &#8734; is:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(11)\\ t^{*} = \\int_{A_{0}}^{\\infty}{\\frac{1}{f(A)}dA}&quot;,&quot;id&quot;:&quot;BQUQJSRCXG&quot;}" data-component-name="LatexBlockToDOM"></div><p>When this integral converges to a finite value, it means that there is a finite time by which <em>A</em>(<em>t</em>) goes to infinity &#8212; i.e. a singularity. The convergence of this integral is the key condition for <em>A</em>(<em>t</em>) having a singularity, which we shall call the <em>blow-up condition</em>. To this we again add the positivity condition &#8212; <em>f</em>(<em>A</em>) &gt; 0  at every point in [<em>A</em><sub>0</sub>, &#8734;) &#8212; which ensures <em>A</em>(<em>t</em>) is growing (rather than shrinking) and that there is no division by zero.</p><p>In the standard differential equations for RSI (where <em>&#550;</em> is some power law of <em>A</em>) everything that meets the super-linearity condition also meets the blow-up condition: so everything that is super-exponential has a singularity.</p><p>But we can now show that this is <em>not</em> true in general. Consider:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(12)\\ \\dot{A} = A\\log{(A)}&quot;,&quot;id&quot;:&quot;WCYOPFHGLT&quot;}" data-component-name="LatexBlockToDOM"></div><p><em>A</em> log&#8289;(<em>A</em>) is super-linear, so <em>A</em>(<em>t</em>) grows super-exponentially. But it doesn&#8217;t meet the blow-up condition: the integral of 1/<em>A</em>log&#8289;(<em>A</em>)&#8203; diverges to infinity, so it takes infinitely long for <em>A</em> to grow infinitely large. Solving the differential equation reveals that <em>A</em> in fact grows doubly exponentially with time:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(13)\\ A(t) = A_{0}^{e^{t}}&quot;,&quot;id&quot;:&quot;YEWBTNLKVW&quot;}" data-component-name="LatexBlockToDOM"></div><p>This is super-exponential but has no singularity. So if RSI were to obey this equation, there would be a very rapid rise in capabilities, but one of quite a different kind. It would always have more room for improvement, even after an arbitrarily long period of growth.</p><p>And this isn&#8217;t the only function that produces such growth. For instance, consider the infinite family of equations:</p><ul><li><p><em>&#550;</em> = <em>A</em>log&#8289;(<em>A</em>)</p></li><li><p><em>&#550;</em> = <em>A</em>log&#8289;(<em>A</em>) log&#8289;(log&#8289;(<em>A</em>))</p></li><li><p><em>&#550;</em> = <em>A</em>log&#8289;(<em>A</em>) log&#8289;(log&#8289;(<em>A</em>)) log&#8289;(log(log&#8289;(<em>A</em>)))</p></li><li><p><em>...</em></p></li></ul><p>These all fail to meet the blow-up condition, so do not possess singularities. Instead they grow as a double exponential of time, a triple exponential, a quadruple exponential, and so on. Yet they also come extremely close to the blow-up condition. If you raise the final factor in any of them to a power greater than one (e.g. <em>A</em>log&#8289;(<em>A</em>)<sup>1+&#949;</sup>), they satisfy the blow-up condition and <em>A</em>(<em>t</em>) will have a singularity.</p><p>(See the appendix for a table showing the rates of growth corresponding to a wide range of <em>f</em>(<em>A</em>), including whether they produce singularities.)</p><p>So in the general case where <em>&#550; </em>= <em>f</em>(<em>A</em>), there is a narrow zone of functions for <em>f</em>(<em>A</em>) that fit between the exponential growth given by <em>&#550; </em>= <em>kA</em> and the growth towards a finite time singularity given by <em>&#550; </em>= <em>kA</em><sup>1+&#949;</sup>. We thus can&#8217;t assume that all super-exponential growth has a singularity.</p><p>We might wonder whether it is realistically possible to end up in this zone. Economists call exponential growth a &#8216;knife-edge solution&#8217; to <em>&#550; </em>= <em>kA<sup>r</sup></em>, suggesting that it is vanishingly unlikely for <em><span>r</span></em> to take on the value of <em>exactly</em> 1. If so, this new zone I&#8217;m pointing to may look like it is on the edge of the edge of the knife. Or it would if we were taking this functional form and supposing a random value of <em><span>r</span></em>. But the assumption of the functional form is then doing most of the work. If we instead suppose an unknown functional form for f<span>f</span>, with a non-zero chance of it being <em>A</em>log&#8289;(<em>A</em>), the landing zone looks a little larger.</p><p>I think this functional form is plausible when the <em>A</em> and the log&#8289;(<em>A</em>) are coming from different places. For example, suppose <em>A</em> is a measure of the total number of AI researchers doing RSI, and suppose that the contribution of each one towards increasing <em>A</em> is proportional to the log of their total population (due to weak spillovers from each person&#8217;s research, or weak economies of scale). Then you get a combined effect where <em>&#550; </em>= <em>kA</em>log&#8289;(<em>A</em>).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p>That said, it is still a narrow zone and any kind of fishing-out or stepping-on-toes effect could easily knock things out of the zone. Overall, I think this possibility would be just something of an interesting footnote were it not for new considerations we&#8217;ll see in the next section, which substantially widen this zone.</p><p>Finally, it is important to note that in this general case of <em>&#550; </em>= <em>f</em>(<em>A</em>), one can&#8217;t treat all growth with a singularity as hyperbolic. That was only true when making the modelling assumption that <em>f</em>(<em>A</em>) has the form of a power law. Some functions (such as <em>&#550; </em>= <em>A</em>log&#8289;(<em>A</em>)<sup>2</sup>) produce trajectories that grow more quickly than any hyperbola &#8212;though they are simultaneously less &#8216;explosive&#8217; in the sense that the growth isn&#8217;t as packed into the final moment. Others (such as <em>&#550; </em>= <em>e<sup>A</sup></em>) produce trajectories that grow more slowly than any hyperbola, but are more explosive since they still reach infinity in finite time, despite lagging behind until the last moment. To reflect this wider class of functions, I&#8217;ll use the more inclusive term <em>singular growth</em> to refer to all cases when <em>A</em>(<em>t</em>) possesses a singularity (whether hyperbolic or not).</p><p>We can now answer our earlier questions:</p><ul><li><p>There are differential equations that produce intermediate patterns of growth lying between exponential and singular growth, such as <em>&#550; </em>= <em>A</em>log&#8289;(<em>A</em>), which gives doubly exponential growth.</p></li><li><p><em>f</em>(<em>A</em>) being super-linear in the limit produces the super-exponentiality, but more is needed to create a singularity. Being everywhere convex is neither necessary nor sufficient &#8212; singularities are not produced by satisfying a local property like convexity, but by satisfying a global property where deficient growth in <em>f</em>(<em>A</em>) somewhere can be made up by faster growth elsewhere (so long as it is always positive). The relevant global property is the blow-up condition: that the integral of the reciprocal of <em>f</em>(<em>A</em>) converges.</p></li><li><p>Small changes to <em><span>f</span></em>, such as small additive or multiplicative constants or small random deviations from its path won&#8217;t usually change these behaviours, unless they make <em>f</em>(<em>A</em>) &#8804; 0 somewhere along the trajectory.</p></li></ul><h2>Feedback loops &amp; discrete time steps</h2><p>The idea of an intelligence explosion relies on some kind of feedback loop, where each AI system builds a more intelligent successor, which is even better at AI R&amp;D. There are many forms this feedback loop could take. For example:</p><ul><li><p>AI could develop new tools and processes for making faster chips.</p></li><li><p>AI could create better chip designs for the current chip manufacturing processes.</p></li><li><p>AI could make the pretraining process for the next model more efficient, getting that model sooner.</p></li><li><p>AI could improve the pretraining process for the next model to make the model more capable.</p></li><li><p>AI could develop ways to fine-tune its own weights to make it more capable.</p></li><li><p>AI could modify its harness to make it more capable.</p></li></ul><p>Something these all have in common is that they don&#8217;t happen instantly. Rather than a continuous process, such as the differential equations behind an object cooling or a pendulum swinging, we instead have something whose capability is increasing in discrete steps.</p><p>Such discrete steps are quite common in feedback processes. Even the classic example of a microphone near a speaker gives discrete jumps in volume. When the microphone is switched on it takes a moment for its signal to travel down the wire to the speaker, which then jumps up in volume. This louder sound then has to travel through the air at the speed of sound, before the microphone can register this increase and begin the next cycle. If the microphone were 30 metres from the speaker, the volume would rise in a staircase pattern with each step lasting about a tenth of a second. Because these steps are so brief, we often don&#8217;t hear them, and continuous models of the feedback are adequate for most purposes.</p><p>But the steps in many RSI feedback loops are much longer &#8212; potentially months or years. And even more importantly, we&#8217;ve suggested that the RSI feedback loops might be able to produce singular growth. Is that even possible in discrete time or just an artefact of the continuous nature of differential equations?</p><p>Taking the discrete nature seriously will change how we see the idealised dynamics of these systems and will change the conditions under which singularities are possible. By the time we&#8217;re finished, we&#8217;ll see the possibility of super-exponential, yet sub-singular growth go from being a narrow possibility to a mainline scenario.</p><p>We can start by considering the discrete version of the differential equation: the <em>difference equation</em>. Let <em>A<sub>n</sub></em>&#8203; represent the system&#8217;s ability after <em>n</em> times around the feedback loop. The discrete version of <em>&#550;</em> is &#916;<em>A</em>, which is defined as:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(14)\\ \\Delta A_{n} = A_{n + 1} - A_{n}&quot;,&quot;id&quot;:&quot;LXEUUQDLXL&quot;}" data-component-name="LatexBlockToDOM"></div><p>In other words, &#916;<em>A</em> is the difference between successive terms in a sequence. Our general autonomous differential equation, <em>&#550; </em>= <em>f</em>(<em>A</em>), becomes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(15)\\ \\Delta A_{n} = f(A_{n})&quot;,&quot;id&quot;:&quot;SSQSHNSRLQ&quot;}" data-component-name="LatexBlockToDOM"></div><p>Or equivalently (via the definition of &#916;<em>A</em>):</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(16)\\ A_{n + 1} = A_{n} + f(A_{n})&quot;,&quot;id&quot;:&quot;FXVFFVBHXA&quot;}" data-component-name="LatexBlockToDOM"></div><p>In difference equations it is impossible to produce a singularity no matter how quickly <em><span>f</span></em> grows. This becomes obvious when one looks at what would be required. How could you have <em>A<sub>n</sub></em>&#8203; grow without bound prior to some particular <em>n</em>? That&#8217;s only possible if there are infinitely many steps prior to that point or if there is a step where <em>A<sub>n</sub></em>&#8203; takes an infinite value. But there are only finitely many steps prior to any particular <em>n</em>, and all values are finite (they start from a finite level (<em>A</em><sub>0</sub>&#8203;) and each one only adds finitely much to the total, since <em><span>f</span></em> is a function from &#8477; to &#8477;).</p><p>For example, if you try <em>f</em>(<em>A</em>) = <em>A</em><sup>2</sup> (which would give hyperbolic growth using a differential equation) you instead get doubly exponential growth. Or if <em>f</em>(<em>A</em>) = <em>e<sup>A</sup></em>, then <em>A<sub>n</sub></em>&#8203; grows as a tower of exponentials with <em>n</em> levels (tetration). These are very fast rates of growth (clearly super-exponential), yet they have no singularities. For differential equations, rates of growth like these only lived in a very slender zone, but they are entirely generic for difference equations &#8212; occurring whenever <em><span>f</span></em> is super-linear.</p><p>What do these facts about difference equations mean when it comes to RSI? They don&#8217;t have any direct meaning until we say how the number of feedback cycles (<em>n</em>) maps onto time (<em>t</em>). Let&#8217;s create a new model of RSI which is neither a simple differential equation nor difference equation. Instead it will be a difference equation that is embedded into continuous time, which we shall call a <em>time-embedded difference equation</em>.</p><p>Let <em>t<sub>n</sub></em>&#8203; be the time at which the <em>n</em>th feedback loop is completed and let <em>T<sub>n</sub></em>&#8203; be the time it takes to go around the feedback loop for the <em>n</em>th time. This second parameter (<em><span>T</span><sub><span>n</span></sub></em>&#8203;) will turn out to be a key quantity when analysing the dynamics of intelligence explosions. I shall call it the <em>generation time</em>, via analogy to the parameter of that name in the feedback loops of population growth, disease spread, and nuclear chain reactions. In those subjects, every person/infection/fission produces a random whole number of new persons/infections/fissions, and the generation time is the average amount of time this takes.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> That corresponds to the amount of time it takes to go around their feedback loops, so I&#8217;ll use the term &#8216;generation time&#8217; even though for some RSI feedback loops there won&#8217;t be clear discrete generations of AIs.</p><p><em>T<sub>n</sub></em>&#8203; and <em>t<sub>n</sub></em>&#8203; are related by the equations:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(17)\\ T_{n} = t_{n} - t_{n - 1}&quot;,&quot;id&quot;:&quot;TTVCRLSCFM&quot;}" data-component-name="LatexBlockToDOM"></div><p>and</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(18)\\ t_{n} = \\sum_{i = 1}^{n}T_{i}&quot;,&quot;id&quot;:&quot;KIUNIKTZJX&quot;}" data-component-name="LatexBlockToDOM"></div><p>Let&#8217;s also define <em>n<sub>t</sub></em>&#8203; to be the number of cycles around the feedback loop that have been completed by time <em>t</em> (the highest <em>n</em> such that <em>t<sub>n </sub></em>&#8804; <em>t</em>). We can then represent a time-embedded difference equation by the pair of sequences <em>A<sub>n</sub></em>&#8203; and <em>t<sub>n</sub></em>&#8203;. Together, these enable us to see how <em>A</em> changes over <em>time</em>, via:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(19)\\ A(t) = A_{n_{t}}&quot;,&quot;id&quot;:&quot;CQSBOUAVUI&quot;}" data-component-name="LatexBlockToDOM"></div><p>If every cycle of the feedback loop takes the same time, we&#8217;d have <em>T<sub>n</sub></em> = <em>k</em>  and <em>tn</em> = <em>kn&#8203;</em>. If so, the previous remarks about difference equations would apply directly &#8212; there are only finitely many loops between any two times, each of which makes only a finite change to <em>A</em>, making singularities impossible.</p><p>The same is true whenever there is a bound on how short the loop can get. If it never gets shorter than <em>T</em><sub>min&#8289;</sub>&#8203; then the growth of <em>A</em>(<em>t</em>) is bounded above by that of a feedback loop with a constant generation time <em>T</em><sub>min&#8289;</sub>.</p><p>But what if the generation time decreases towards zero? If it decreases slowly, such as via <em>T<sub>n </sub></em>= 1/<em>n</em>&#8203;, then the total time required to perform <em>n</em> loops increases without bound. This implies there are only a finite number of loops between any two times, making singularities impossible. But what if the generation time decreases more quickly? Let&#8217;s call the time taken to perform infinitely many steps <em>t</em><sub>&#8734;</sub>&#8203;:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(20)\\ t_{\\infty} = \\sum_{i = 1}^{\\infty}T_{i}&quot;,&quot;id&quot;:&quot;EOAVZNCMMI&quot;}" data-component-name="LatexBlockToDOM"></div><p>This can be finite. For example, if the generation time shrinks exponentially as <em>T<sub>n </sub></em>= 1/(2<em>n</em>), then <em>t</em><sub>&#8734; </sub>= 1. In this case an infinite number of loops have been performed within 1 unit of time. As long as <em>A<sub>n</sub></em>&#8203; increases without bound, this would be a singularity at <em>t</em> = 1. So singularities are possible in this model of RSI via a difference equation embedded in time.</p><p>What is required to produce them?</p><p>First, let&#8217;s look at the generation time. It has to go to zero fast enough for the infinite sum in (20) to converge. <em>T<sub>n </sub></em>= 1/<em>n</em>&#8203; is too slow, while <em>Tn </em>= 1/(<em>n</em><sup>1+&#949;</sup><span>)</span>&#8203; is fast enough. The cut-off for how quickly the denominator of this fraction must grow turns out to be exactly the same as for the blow-up condition we examined earlier. For example,</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;T_{n} = \\ \\frac{1}{n\\ \\log{(n)}\\log{(\\log{(n)})}}&quot;,&quot;id&quot;:&quot;GSMTWEETPP&quot;}" data-component-name="LatexBlockToDOM"></div><p>is slightly too slow to allow for a singularity, while raising the final factor to any power greater than 1 is sufficient to force the sum to converge, giving a finite <em>t</em><sub>&#8734;</sub>&#8203; and thus the opportunity for a singularity. Let&#8217;s call this condition on how quickly <em>T<sub>n</sub></em>&#8203; must converge to allow a singularity the <em>Zeno condition</em> on <em>T<sub>n</sub></em>&#8203;, as it is precisely the condition for when infinitely many steps can be performed in a finite time.</p><p>What about the rate of growth of <em>A<sub>n</sub></em>? We&#8217;ve already seen that no growth rate can give a singularity if we only go around the feedback loop finitely many times, but if we get to go around it infinitely many times, then even adding one each time (<em>f</em>(<em>A<sub>n</sub></em>) = 1) is enough to produce singular growth. Indeed, even adding smaller and smaller amounts each time around the loop can work, so long as <em>A<sub><span>n</span></sub></em>&#8203; grows without bound. Let&#8217;s call this the <em>boundlessness condition</em> on <em>A<sub>n</sub></em>&#8203;.</p><p>For example, if we added just 1/<em>n</em>&#8203; after the <em>n</em>th feedback loop, that would be enough to produce unbounded growth and (if the Zeno condition is also met) a singularity. The same is true if we added only</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\frac{1}{n\\log{(n)}}&quot;,&quot;id&quot;:&quot;RCOOVHSMVR&quot;}" data-component-name="LatexBlockToDOM"></div><p>or</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\frac{1}{n\\log{(n)\\log{(\\log{(n)})}}}&quot;,&quot;id&quot;:&quot;IUQEPCLZDN&quot;}" data-component-name="LatexBlockToDOM"></div><p>&#8203;Considered as a function of <em>n</em>, the amount that needs to be added each time (&#916;<em>A<sub>n</sub></em>&#8203;) is thus set by exactly the same threshold as <em>T<sub>n</sub></em>&#8203; (though &#916;<em>A<sub>n</sub></em>&#8203; mustn&#8217;t shrink more quickly than this threshold, while <em>T<sub>n</sub></em>&#8203; must). We can summarise this with the following theorem.</p><blockquote><p><strong>Singularity Theorem:</strong></p><p>When the growth of <em>A</em> is governed by time-embedded difference equations:</p><p><em>A</em>(<em>t</em>) has a singularity <em>iff</em> <em>T<sub>n</sub></em> meets the Zeno condition &amp; <em>A<sub><span>n</span></sub></em>&#8203; meets the boundlessness condition.</p><p>Or equivalently:</p><p><em>A</em>(<em>t</em>) has a singularity <em>iff</em> &#8721;<em>T<sub>n</sub></em>&#8203; converges &amp; &#8721;&#916;<em>A<sub>n</sub></em>&#8203; diverges to +&#8734;</p><p><em>Proof:</em> Both conditions are necessary because it doesn&#8217;t help to grow boundlessly if you can only get through finitely many steps by any particular time, and it doesn&#8217;t help to have infinitely many steps before some time if the function has a finite bound. But if both conditions are met, then infinitely many steps will happen by <em>t</em><sub>&#8734;</sub>&#8203; and these will take <em>A<sub><span>n</span></sub></em>&#8203; beyond every finite level prior to that time.</p></blockquote><p>This theorem cleanly cuts the requirements into two independent thresholds. Success on each is binary: no deficiency on either can be made up by over-performance on the other. We can get a feel for how sharp this is through a pair of examples.</p><p>First, consider a feedback loop whose generation time decreases towards zero as 1/<em>n</em>&#8203; and each time it goes through the loop, <em>A</em> gets multiplied by 2<em><sup>A</sup></em>. <em>A<sub>n</sub></em>&#8203; is growing extraordinarily quickly (tetrationally), but because the generation time doesn&#8217;t meet the Zeno condition, there is no singularity. Second, consider a feedback loop whose generation time decreases towards zero slightly more quickly (as 1/<em>n</em><sup>1.01</sup>&#8203;), but where only a small and diminishing amount is added to <em>A</em> each time (+1/<em>n</em>&#8203;). This meets both conditions, so blows up to a singularity. The tiny difference in how generation times decreased with <em>n</em> outweighed the radical difference in what happened in each loop.</p><p>It is very interesting that in this discrete setting singular growth requires generation time to go to zero, and the rate at which it needs to approach zero is analogous to the rate at which <em>f</em>(<em>A</em>) had to grow in the continuous setting. This suggests that the decreasing generation time is driving the whole effect. One way to see this is that singularities require the gradient of the curve to go to infinity in finite time, and there are two ways to increase gradient: increasing the rise or decreasing the run.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2X5X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2X5X!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 424w, https://substackcdn.com/image/fetch/$s_!2X5X!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 848w, https://substackcdn.com/image/fetch/$s_!2X5X!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 1272w, https://substackcdn.com/image/fetch/$s_!2X5X!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2X5X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png" width="1456" height="694" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:694,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:90717,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/213060056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2X5X!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 424w, https://substackcdn.com/image/fetch/$s_!2X5X!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 848w, https://substackcdn.com/image/fetch/$s_!2X5X!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 1272w, https://substackcdn.com/image/fetch/$s_!2X5X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F447b2985-f870-4516-aa93-e4c6a02f8b6d_2000x953.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 1. Once we take the discrete nature of the feedback loop into account, we see that we can only approach an infinite gradient in finite time by reducing the run in each step towards zero (right).</em></figcaption></figure></div><p>In continuous models, there is no way to distinguish these. In raw difference equations, one cannot reduce the run and no amount of growth in the rise can produce a singularity. But when we allow for variable-length discrete time steps through a time-embedded difference equation, we see that it is almost entirely about reducing the run. The rise per step doesn&#8217;t have to increase at all &#8212; its only constraint is that it can&#8217;t <em>decrease</em> too quickly.</p><p>Another lesson is that a singularity requires the growth of <em>A</em> to be infinitely compounded. We are familiar with the way compounding interest daily leads to a higher growth rate than compounding it annually. Compounding at ever finer intervals converges towards the growth rate set by continuous compounding. But for super-exponential growth, the value of compounding is much greater. If the benefits can only compound finitely many times by a finite point in time, there is no way to achieve a singularity. It is a necessary condition for a singularity that the growth in <em>A</em> is compounded infinitely many times in a finite time interval &#8212; which can happen either via a convergent sequence of discrete generation times or a continuous system.</p><p>Solomonoff (1985) provides a nice toy model for an intelligence explosion via shortening the generation time. Computing speeds have been improving exponentially over time (a version of Moore&#8217;s Law). If we reached a point where AI could provide the labour needed to keep Moore&#8217;s Law running, then every time Moore&#8217;s Law doubles computing speeds, it doubles the speed of these digital researchers, halving the time it takes for the next doubling of speeds. So if the first doubling took 2 years, the second would take 1 year, the third half a year... and by the end of 4 years, there would be a singularity &#8212; with infinite AI labour and infinite computational efficiency.</p><p>Moravec (1999, 2003) developed an improved version of this model, showing that it doesn&#8217;t even require Moore&#8217;s Law to be exponential. He first noticed that a quadratic speed-up would suffice, while a linear speed-up would not. Then he found his way to the peculiar threshold we&#8217;ve seen so many times: so long as the <em>n</em>th generation machines are faster than the first generation by a factor greater than <em>n</em>log&#8289;(<em>n</em>)log&#8289;(log&#8289;(<em>n</em>))&#8230; this approach produces a singularity.</p><p>Of course, there are still great obstacles making this impractical. Moore&#8217;s Law has required increasing amounts of labour to keep it going (due to fishing-out and stepping-on-toes effects), as well as increasing amounts of capital. Moreover, there will be ultimate physical limits to how far it can go, and there are practical limits to how quickly a new generation of faster chips can be produced (especially when a novel method is required). But it is still an instructive model for an efficiency-driven intelligence explosion. While Solomonoff and Moravec modelled it with differential equations, the presence of an explicit generation time that tends towards zero allows this model to work just as well in time-embedded difference equations.</p><p>Eth and Davidson (2025) considered something like this when they explored AI systems that improve their own training time. Suppose an AI system could make sufficient tweaks to its architecture and training process that it slightly improved the speed of training and inference for the next generation of AI systems. If this could be repeated indefinitely, then the generation time (to design and train the next generation) would keep falling. So long as the speed improvements are above the familiar threshold (<em>n</em>log&#8289;(<em>n</em>)log&#8289;(log&#8289;(<em>n</em>))&#8230;), one would get singular growth, leading to unlimited AI labour by some finite time. Of course, Eth and Davidson don&#8217;t suggest that we can actually make an indefinite series of such improvements. Instead, one would expect the generation time for training the next generation of models to bottom out at some unyielding finite limit, prematurely ending the period of singular growth.</p><p>In general, it seems highly unlikely that generation times can be brought arbitrarily close to zero. This provides an important kind of barrier to singular growth.</p><p>It is essential in all these arguments that we are measuring generation time itself, and not the related concept of <em>doubling time</em>:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;(21)\\ D_n = \\frac{T_n}{\\log_2\\left(\\frac{A_n}{A_{n-1}}\\right)}&quot;,&quot;id&quot;:&quot;LWSGADUXQU&quot;}" data-component-name="LatexBlockToDOM"></div><p>Doubling time is often useful as it takes into account how much a feedback loop contributes as well as how long it takes. For instance, a loop that takes one second and doubles its input has the same doubling time as one that takes two seconds and quadruples its input. A key advantage of doubling time is that it is also defined for continuous processes, while generation time is not.</p><p>However, this combining of the loop&#8217;s duration and impact turns out to be doubling time&#8217;s downfall. If <em>A<sub>n</sub></em>&#8203; grows very quickly (e.g. doubly exponentially) while the generation time stays constant, then the doubling time will rapidly decline to zero. But we&#8217;ve seen that since generation time is constant, there can&#8217;t be singular growth no matter how quickly <em>A<sub>n</sub></em>&#8203; grows or <em>D<sub>n</sub></em>&#8203; shrinks. Doubling time shrinking towards zero is a useful threshold for defining super-exponential growth, but it is only generation time that can set the threshold for singular growth.</p><p>It is worth noting that one could also model the length of feedback loops via <em>delayed differential equations</em> rather than difference equations. In the equation <em>&#550;</em> = <em>f</em>(<em>A</em>) we&#8217;d been suppressing the dependence on time. We could have equally well written <em>&#550;</em>(<em>t</em>) = <em>f</em>(<em>A</em>(<em>t</em>)). We can change this to a delayed differential equation where the current value of <em>&#550;</em> depends on the value of <em>A</em> at a time <em>T</em> units earlier (to account for the generation time): <em>&#550;</em>(<em>t</em>) = <em>f</em>(<em>A</em>(<em>t</em>&#8722;<em>T</em>)). Like the difference equation, this delayed differential equation with a fixed delay can produce super-exponential growth but can&#8217;t produce a singularity.</p><p>And we can make an analogue to the time-embedded difference equation, by allowing the feedback loop length to change with time: <em>&#550;</em>(<em>t</em>) = <em>f</em>(<em>A</em>(<em>t</em>&#8722;<em>T</em>(<em>t</em>))). Like the time-embedded difference equation, this general delayed differential equation <em>can</em> produce a singularity, but only when <em>T</em> shrinks towards zero sufficiently quickly as <em>t</em> &#8594; <em>t</em>*. This gives a continuous model with similar dynamics. I find it slightly less accurate (losing track of the discrete nature of the feedback loop) and slightly harder to work with, but other researchers may find it useful.</p><p>We are now able to zoom out and take another look at when growth from feedback loops is singular versus merely super-exponential:</p><ul><li><p>For differential equations of the common restricted form, <em>&#550; </em>= <em>kA<sup>r</sup></em>, all super-exponential growth is singular growth.</p></li><li><p>For differential equations of the more general form, <em>&#550;</em> = <em>f</em>(<em>A</em>), we see that there is also a narrow range of super-exponential growth without a singularity.</p></li><li><p>When we take the discrete nature of feedback loops into account with difference equations of the form &#916;<em>A<sub>n</sub></em> = <em>f</em>(<em>A<sub>n</sub></em>), all super-exponential growth is sub-singular.</p></li><li><p>When we allow for changing generation times using time-embedded difference equations, we can again get super-exponential growth in singular and sub-singular varieties. But compared with the setting of differential equations, singular growth is now much harder to achieve due to the requirement that the length of each feedback loop shrinks towards zero. Super-exponential growth without a singularity is no longer a curiosity &#8212; it is the default.</p></li></ul><h2>Intelligence Measures</h2><p>When using a differential or difference equation to model RSI, it is remarkably unclear what <em>A</em> should represent. Should it represent intelligence, or should we sidestep that and measure the computational efficiency of that intelligence? If we choose to represent intelligence itself (and thus try to model an actual <em>intelligence</em> explosion), we run into two further problems.</p><p>First, it is widely recognised that there is no generally agreed conception of intelligence to measure. Even among those who agree that intelligence is a real and important thing, there is no consensus on which thing it is.</p><p>AI research tries to sidestep this by measuring a wide variety of different capabilities via benchmarks. These measure what fraction of a set of tasks related to that capability the system can successfully solve. There is a common feeling that achieving artificial general intelligence (AGI) will require roughly human performance on most such capabilities, but there is some dissent. For example Chollet (2019) argues that intelligence is really the ability to efficiently learn a wide variety of capabilities. On his view, intelligence isn&#8217;t measured in terms of capability, but <em>capability per unit input</em> (where inputs could include training data, information embedded in the priors, training compute, inference compute etc.). Even if there were agreement on whether intelligence is a measure of capability or capability per unit input, one would need further agreement on how to weight the different kinds of capabilities into an overall linear ordering for <em>A</em>.</p><p>There is also a second problem, which is much less widely recognised. Even if we could agree on the kind of thing being measured (e.g. we could agree on a linear ordering of all AI systems according to intelligence) it is very unclear what cardinal structure this should have &#8212; how to assign numbers to those systems.</p><p>For example, when measuring the capability of AI systems at playing a game such as chess or Go, researchers often use Elo scores. But Elo scores are really just the log of a more fundamental measure from the Bradley-Terry model which assigns each player a strength such that the odds ratio of a player beating another is simply the ratio of their strengths. Elo is just the log of this strength (with some arbitrary constants to help its numbers match an earlier chess ranking system).</p><p>This causes confusion in the literature when papers claim that in contrast to some other domains, chess AI is only improving linearly over time. Chess improvements have been roughly linear when measured in Elo, but the headline claim could equally be that chess-playing AI is improving exponentially over time (when measured in the improvement in the odds-ratio of beating a player of fixed strength). In this case, it is unclear whether progress in chess AI is best described as exponential or linear.</p><p>Overall measures for AI progress suffer from the same issue. They could be exponential on one fairly natural scale and linear on another. Or they could be convex (with increasing returns) on one scale while concave (with diminishing returns) on another.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> Given there is often very little discussion or agreement on which measure is the more natural one (or whether there is even a fact of the matter about that) there is a big problem of measure-dependence for attempts to track progress in AI.</p><p>How do these issues affect the dynamics of RSI &#8212; or our ability to measure and track those dynamics?</p><p>Let&#8217;s start with a simple example where <em>A</em> is a measure of intelligence and <em><span>B</span></em> is another measure where <em>B </em>= log&#8289;(<em>A</em>)<span>.</span> For the differential equation model, where <em>&#550;</em> = <em>f</em>(<em>A</em>), we will also have <em>&#7682; </em>= <em>g</em>(<em>B</em>), where <em>g</em> is a slower growing function than <em><span>f</span></em>. For example, when <em>&#550; </em>= <em>kA</em>, <em>&#7682;</em> = <em>k</em>. These are consistent (when the measure <em>A</em>(<em>t</em>) grows exponentially, it makes sense that the measure <em>B</em>(<em>t</em>) is growing linearly) but they show that the same rate of progress of the observable phenomenon in the world can correspond to different functions f<span>f</span> in the differential equation.</p><p>However, there is an important invariant: <em>f</em>(<em>A</em>) grows fast enough to produce a singularity if and only if <em>g</em>(<em>B</em>) grows fast enough to produce a singularity. i.e. <em>f</em> will meet the blow-up condition (of growing faster than <em>A</em>log&#8289;(<em>A</em>)log&#8289;(log&#8289;(<em>A</em>)) (etc.) if and only if <em><span>g</span></em> also does. It is easy to see this must be true because when some quantity grows without bound, the log of that quantity also grows without bound. So if one grows without bound by time <em>t</em>* the other must too.</p><p>Let&#8217;s define two measures as <em>similar</em> when either measure growing without bound implies the other does too. When <em>A</em> and <em>B</em> are similar, then either both <em>f</em> and <em>g</em> meet the blow-up condition or neither do. Therefore, if we are trying to detect singular growth of an RSI system in terms of how <em>&#550; </em>is increasing with <em>A</em>, the threshold is at the same place (roughly <em>A</em>log&#8289;(<em>A</em>)log&#8289;(log&#8289;(<em>A</em>))...) for a wide range of different ways of measuring that system&#8217;s capabilities.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a></p><p>Something very similar is true if we model RSI via time-embedded difference equations. In this case, the only necessary condition on <em>A<sub>n</sub></em>&#8203; to achieve a singularity is that it grows without bound. And if <em>A<sub>n</sub></em>&#8203; and <em>B<sub>n</sub></em>&#8203; are similar, then they either both satisfy this condition or neither does. So the threshold for how much they need to improve on the <em>n</em>th feedback loop (<em>n</em>log&#8289;(<em>n</em>)log&#8289;(log&#8289;(<em>n</em>)) ...) is the same for both, and we don&#8217;t need to worry about which of these capability measures we are using when testing for a singularity.</p><p>But some measures are not similar to each other. For example, when the full range of measure <em>A</em> only maps into a finite interval in measure <em><span>B</span></em>, then <em>A</em> going to infinity doesn&#8217;t entail <em>B</em> going to infinity and you could get singular growth in <em>A</em> without singular growth in <em>B</em>. This could easily happen if <em>A</em> represented open-ended progress in some domain of intelligence while <em><span>B</span></em> was a broader measure that included more domains. In such a case, there would be no contradiction in <em>A</em> having singular growth without <em><span>B</span></em> having it too. The time <em>t</em><span>*</span> would simply be the time at which problems of the domain <em>A</em> was measuring are effectively solved.</p><p>This might be the case for measures of <em>A</em> like computational efficiency or &#8216;effective compute&#8217;. Many problems would be solved if we let an AI system have unlimited effective compute (for training and/or inference), but it isn&#8217;t clear such a system would excel at all kinds of intellectual tasks. For example, even if you had unlimited compute, it isn&#8217;t clear that current architectures and training environments would allow a system to succeed in domains which are lacking a clean algorithmic way of rating the quality of the outputs. Much of the recent progress in AI has relied on reinforcement learning with verifiable rewards (RLVR), but when there are no verifiable rewards, performance may plateau. If so, even a completed singularity in measures like efficiency and effective compute needn&#8217;t imply that the system is more capable than a human across the board after time <em>t</em><span>*</span>.</p><p>In physics, scientists distinguish between a <em>coordinate singularity</em> and an <em>essential singularity</em>. The former exists when the singularity is an artefact of how things are being measured. For example, when Schwarzschild (1916) published his mathematical description of a black hole there was a mathematical singularity at the black hole&#8217;s event horizon. It took decades of scientific work before physicists were able to establish that this was a mere artefact of the coordinates being used (whereas the singularity at the centre of the black hole was not fixable by a change in coordinates).</p><p>We can adopt this distinction when studying RSI. For example, suppose we are measuring <em>A</em> by the <em>mean time between failures (MTBF)</em>: how long the system can go before making a mistake while completing tasks of a certain kind. This is a standard measure in engineering and a relevant measure of capability for many real-world uses. In this case, one might see <em>A</em> rising towards a singularity. But consider an alternative measure <em>B</em> which is the percentage of times the system gets the right answer. If <em>A</em> goes to infinity at time <em>t</em><span>*</span>, that just means <em><span>B</span></em> reaches 100% at <em>t</em><span>*</span>. While a lot of systems never quite reach 100% reliability, there is nothing impossible or paradoxical about doing so. Thus in this example, <em>t</em><span>*</span> is a coordinate singularity in measure <em>A</em>, but not in measure <em><span>B</span></em>, and is not an essential singularity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sNw3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sNw3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 424w, https://substackcdn.com/image/fetch/$s_!sNw3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 848w, https://substackcdn.com/image/fetch/$s_!sNw3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 1272w, https://substackcdn.com/image/fetch/$s_!sNw3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sNw3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png" width="1456" height="1772" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1772,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:93540,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/213060056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sNw3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 424w, https://substackcdn.com/image/fetch/$s_!sNw3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 848w, https://substackcdn.com/image/fetch/$s_!sNw3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 1272w, https://substackcdn.com/image/fetch/$s_!sNw3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5e9d5e6-f045-4d0f-b1ad-a33444a1bc22_2050x2495.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 2. How two different measures of the same AI capabilities evolve over time. Measure A (mean time between failures) has a singularity at time t*, while measure B (reliability) simply reaches 100% at that point.</em></figcaption></figure></div><p>Note that a very widely used measure of general AI capability &#8212; the METR time horizon (Kwa et al. 2025) &#8212; is somewhat like a mean time between failures. Recent years have seen exponential (or perhaps super-exponential) growth in time horizons for AI systems completing the kinds of routine professional computer-use tasks in the benchmark. But even if the time horizon went to infinity,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> it isn&#8217;t clear the system would be as intelligent as a human. It may still lack various cognitive skills (such as continual learning, sample efficiency, or creativity), which are important but not being tested. And there may still be substantial room for other AI systems to be much more capable, even at the skills that were being tested &#8212; systems that could solve much more challenging computer-use problems using much less compute. An infinite METR time horizon would just mean that the AI system can do the kinds of tasks in the benchmark with perfect reliability (and unbounded duration).</p><p>So even though most people studying intelligence explosions assume that the capability measure <em>A</em> won&#8217;t actually go to infinity in finite time, it wouldn&#8217;t be absurd or unphysical for it to do so. It depends on what measure it is. Some measures could go to infinity without incident. If such measures could also satisfy the key differential equation (or time-embedded difference equation) then one could get a completed intelligence explosion &#8212; <em>A</em> would have gone to infinity in finite time.</p><p>I&#8217;m not quite sure what we should think about this. It really depends on the nature of the measure, and this result is much easier to achieve on narrow measures &#8212; those that can rise unboundedly even when many key cognitive abilities are lacking. But measures like mean time between failures or METR time horizons seem unlikely to lead to rapidly decreasing generation times &#8212; what goes to infinity is a measure of duration, not a measure of speed. By showing that a completed singularity is theoretically possible, I&#8217;m mainly trying to caution people about the behaviour of certain intelligence measures, rather than trying to suggest we will genuinely have unbounded progress in intelligence in a finite time.</p><p>Let&#8217;s now turn to look at what happens when an intelligence explosion can&#8217;t get to infinity.</p><h2>Going Finite</h2><p>We&#8217;ve now seen how super-exponential growth can be cleanly divided into singular and sub-singular varieties, with the singular kind appearing very difficult to achieve (once we take the discrete nature of feedback loops into account). And we&#8217;ve seen that the rates at which <em>f</em>(<em>A</em>) has to grow or diminish are independent of which measure is used (so long as the relevant measures are &#8216;similar&#8217; to each other).</p><p>But all of these claims rely on a clean mathematical model where growth is classified by its asymptotic behaviour as <em>t</em> or <em>A</em> approaches infinity. In the real world, this model may very well break at some finite value of <em>t</em> or <em>A</em>, giving increasingly inaccurate results thereafter. That&#8217;s a big problem for this kind of asymptotic analysis &#8212; as it relies on the infinite domain (or range) to produce its clean classifications. I hope that these divisions which are natural in the idealised setting will be natural in the real-world application. This is often the case,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> though it isn&#8217;t guaranteed.</p><p>There are a variety of finite limits we might expect to run into, which could force the clean mathematical model to break at different points, and in different ways.</p><p>The most widely discussed is a <em>ceiling</em> on intelligence: a horizontal line at some level <em>A</em><span>*</span>, which <em>A</em>(<em>t</em>) cannot cross. In the study of feedback systems, people often say that <em>A</em> <em>saturates</em> as it approaches such a horizontal asymptote. It is generally thought that this will happen for RSI.</p><p>There are many different kinds of ceiling at which the AI&#8217;s intelligence (or efficiency) may saturate:</p><ul><li><p><em>Limits of intelligence itself</em> &#8212; Even an optimal reasoner would be neither omniscient nor omnipotent. It would need to perform experiments to gain knowledge, could still be beaten at unbalanced games, and may face intractable prediction problems in chaotic or agentic domains.</p></li><li><p><em>Limits of intelligence per unit resource</em> &#8212; Even if we have the optimal algorithms and hardware, our solar system has only one star out of the 200,000,000,000 in our galaxy, and growth beyond our system is slow and cubic. If a galactic superintelligence would be more intelligent than a stellar superintelligence, it will be a long time before we could reach that higher level.</p></li><li><p><em>Limits of the hardware paradigm</em> &#8212; Even an optimal silicon chip may be far below the physical limits of compute per unit resource.</p></li><li><p><em>Limits of the algorithmic paradigm</em> &#8212; Even an optimal neural network may be far below the best intelligence that could be achieved with that amount of compute.</p></li><li><p><em>Limits of training data</em> &#8212; The training data we have (and could acquire during RSI) is lacking a lot of information on many domains (especially non-verbalisable information and contextual information). So we could have a very intelligent system that is limited by inability to train certain skills.</p></li><li><p><em>Earlier limits</em> &#8212; The ascent may stall out before any of the above due to limitations inherent in the starting AI.</p></li></ul><p>Discussions of RSI sometimes tacitly assume that an intelligence explosion would only stop when it reached the limits of intelligence itself, and then use assumptions like omniscience or perfect rationality to model its behaviour. But an explosion might saturate significantly below this level, with important consequences for predicting post-explosion capabilities and for understanding whether all intelligence explosions need end at the same intelligence level.</p><p>The manner in which super-exponential growth slows down as it approaches its limit might be analogous to the way exponential growth runs out of steam. While the simple equation for exponential growth is <em>&#550; </em>= <em>kA</em>, in reality there is usually an additional dampening term that starts small but grows to dominate for large <em>A</em>. For example if we subtract a quadratic term that has been shrunk down by a large factor (<em><span>K</span></em>), this gives the differential equation for logistic growth,</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\dot{A} = kA - \\frac{kA^{2}}{K}&quot;,&quot;id&quot;:&quot;VENWUTRONC&quot;}" data-component-name="LatexBlockToDOM"></div><p>where <em>A</em>(<em>t</em>) approaches a horizontal asymptote at height <em>K</em>. We could use the same model for super-exponential growth:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\dot{A} = f(A) - \\frac{f(A)^{2}}{K}&quot;,&quot;id&quot;:&quot;IJQTWTNBLG&quot;}" data-component-name="LatexBlockToDOM"></div><p>Or more generally, we could assume <em>&#550; </em>= <em>f</em>(<em>A</em>) &#8722; <em>g</em>(<em>A</em>), where <em>g</em>(<em>A</em><sub>0</sub>) &#8776; 0 and <em>g</em>(<em>A</em>) &lt; <em>f</em>(<em>A</em>) until some crossover point, <em>A</em>*. This will produce a horizontal asymptote at height <em>A</em>*.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a></p><p>This dampening force could come from running out of room for improvement as a system approaches some form of perfection (such as <em>optimal intelligence per unit resource</em>), or it could come from something like straining under the growing size or complexity of the system. The latter could stop the ascent before reaching any kind of optimal system.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a></p><p>As well as a ceiling on <em>A</em>, one could also run into upper limits on how quickly it can increase &#8212; either in absolute terms (such as a limit on <em>&#550;</em>) or percentage terms (such as a limit on <em>&#550;/A</em>&#8203;). Rather than making <em>A</em>(<em>t</em>) converge to a horizontal asymptote, a maximum gradient would make <em>A</em>(<em>t</em>) converge towards a diagonal line, while a maximum growth rate would make <em>A</em>(<em>t</em>) converge towards an exponential.</p><p>What could produce limits of these kinds?</p><p>One mechanism I find particularly likely is if there is a floor on the achievable generation time, <em><span>T</span><sub><span>n</span></sub></em>&#8203;. It seems very unlikely that generation times can be brought arbitrarily close to zero. There are many kinds of feedback loop that could contribute to RSI, ranging from decades (e.g. designing a successor to EUV lithography) to months (e.g. designing better pretraining) to seconds (e.g. designing better scaffolds). While some are extremely short, those only capture a tiny fraction of the pipeline of what makes for better AI R&amp;D, so it seems likely they would quickly saturate if the other aspects of AI improvement were unchanged. For example, putting a fixed agent in a scaffold that has the agent repeatedly redesign that very scaffold might make some improvements, but without changing the model itself, it seems very unlikely to take off towards infinity. And while continued RSI might be able to reduce any of these generation times by a substantial factor, each one seems likely to run into constraints where a certain minimal amount of time has to pass in order to make an improvement.</p><p>What would a lower bound on the achievable generation time do? If <em><span>T</span><sub><span>n</span></sub></em>&#8203; approached some minimal value (<em>T</em>*) while <em>A<sub><span>n</span></sub></em>&#8203; grew by a fixed amount each loop (&#916;<em>A<sub>n</sub> </em>= <em>k</em>), then <em>A</em>(<em>t</em>) would approach a linear rate of increase. If <em>A<sub><span>n</span></sub></em>&#8203; instead grew by a fixed proportion each loop (&#916;<em>A<sub>n</sub></em> = <em>kA<sub>n</sub></em><sub>&#8203;</sub>), then <em>A</em>(<em>t</em>) would approach an exponential rate of increase.</p><p>I think the second of these looks quite plausible. If so, we might see multiple phases of an explosion:</p><ol start="0"><li><p>The initial exponential phase when the doubling time of <em>A</em> is driven by human-only research.</p></li><li><p>Increasing amounts of RSI drive the generation time down towards machine speeds, so the growth rate of <em>A</em> climbs from its human-only rate towards some very high fully automated rate. Because this phase involves a continuously increasing growth rate, it is super-exponential in shape.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a></p></li><li><p>Fully automated RSI continues for a while at this faster exponential, but as <em>A</em> climbs, it starts to saturate. This phase would be exponential in shape.</p></li><li><p>As it approaches its inflection point and subsequent horizontal plateau, the trajectory has substantially departed from its exponential form and is revealed to be a logistic.</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wbdi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wbdi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 424w, https://substackcdn.com/image/fetch/$s_!wbdi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 848w, https://substackcdn.com/image/fetch/$s_!wbdi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 1272w, https://substackcdn.com/image/fetch/$s_!wbdi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wbdi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png" width="1456" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/acba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:100692,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/213060056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wbdi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 424w, https://substackcdn.com/image/fetch/$s_!wbdi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 848w, https://substackcdn.com/image/fetch/$s_!wbdi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 1272w, https://substackcdn.com/image/fetch/$s_!wbdi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Facba5c4e-f852-41c4-9684-aa9a6fabdc37_2000x889.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figure 3. Phases of an intelligence explosion. The left diagram shows the first phases, with A(t) moving from a human-speed exponential through a period of super-exponential growth to converge on a machine-speed exponential. The right diagram zooms out to show the final phase, where A saturates at some high level A*.</em></figcaption></figure></div><p>This trajectory involves two kinds of finite limits on the idealised process &#8212; first on generation time (and thus growth rate of <em>A</em>), then on the absolute level of <em>A</em>. It shows how we could have a super-exponential process (which would be worth modelling as such), yet it runs out of steam in two distinct ways. On this view, the mathematical models of super-exponential growth could be a good fit (and thus predictive) for the period where generation time is reducing towards zero in some clean way (e.g. 10% shorter each loop), but then cease to be a good model as we enter phase 2.</p><p>I&#8217;ve been assuming that <em>A</em> would be improving by a constant factor each time round the loop, guaranteeing super-exponential and logistic phases, but versions with other functions for <em>f</em>(<em>A<sub>n</sub></em>) are also possible and can still show distinct phases if generation time saturates prior to intelligence level saturating.</p><h2>Conclusions</h2><p>My aim in this paper has been to improve our theoretical understanding of intelligence explosions &#8212; looking at the most explosive possibilities (super-exponential and singular growth) and asking what conditions are required to produce them.</p><p>We first looked at models based on differential equations and saw that the key threshold for super-exponential growth of some measure <em>A</em> is when <em>&#550;</em> grows superlinearly in <em>A</em>. But the threshold for singular growth is somewhat higher &#8212; <em>&#550;</em> has to grow slightly faster than all functions of the form <em>A</em>log&#8289;(<em>A</em>) log&#8289;(log&#8289;(<em>A</em>))... This allows for the possibility of ending up with growth that is super-exponential yet doesn&#8217;t approach a finite time singularity.</p><p>We then saw that when you take the duration of the feedback loop into account, singular growth becomes much harder &#8212; requiring this generation time to approach zero (and quickly enough). Indeed, this model helps us see that singular growth is more about reducing the generation time than it is about increasing how much each loop contributes. And the dynamic that produces singular growth is simply the combination of the Zeno condition and the boundedlessness condition. On this more realistic model, growth that is super-exponential without being singular looks much more likely. It is what you get whenever you can make &#916;<em>A</em> grow super-linearly in <em>A</em>, yet can&#8217;t reduce the generation time all the way towards zero. So analysis of explosive growth needs to be careful to distinguish singular growth from merely super-exponential growth.</p><p>For example, suppose we saw increasing monthly growth rates in some key intelligence measure: 10% growth, then 20%, then 30%. From the economics-inspired literature, we might have thought that since this is super-exponential, it must be hyperbolic growth. But we now know increasing growth rates are not the signature of singular growth. (If the pattern above continued, <em>A</em> would &#8216;merely&#8217; be growing as e<sup>0.05</sup><em><sup>t</sup></em><sup>&#178;</sup>.)</p><p>The importance of generation time to the dynamics of intelligence explosions suggests that generation times need to be carefully measured and tracked. It may be a good policy idea to require frontier labs to report their current generation times &#8212; especially those for pretraining and for RLVR post-training.</p><p>Two of the most important implications for the measurement of an intelligence explosion are negative. When we move beyond the simple model of <em>&#550; </em>= <em>kA<sup>r</sup></em> , you cannot tell whether a process will explode or not based on its returns over a finite period. This is because the condition for having a singularity is a global property of the relationship between <em>&#550;</em> and <em>A</em> (the blow-up condition), rather than a local property such as elasticity &gt; 1. The same is true when modelling it with time-embedded differential equations or delayed differential equations, where the Zeno and unboundedness conditions are inherently global. In all cases, deficiencies in growth somewhere can be made up for by more extreme growth elsewhere.</p><p>So without strong assumptions about the functional form, you can neither rule out nor confirm that an explosion is under way based on local measurements. So while we should probably take measurements of elasticity &gt; 1 for <em>f</em>(<em>A</em>) over some range of <em>A</em> as evidence in favour (and elasticity &lt; 1 as evidence against) that evidence is limited. We should also consider other forms of evidence such as how the elasticity has been changing, or how quickly &#8212; and how far &#8212; the generation time can be reduced.</p><p>While it is not unique to my modelling, I also want to stress that all the kinds of feedback loops we&#8217;ve explored have behaviours that are exquisitely sensitive to fine differences near the borderlines of their behaviours. In contrast, the empirical measurements of intelligence that we can perform are all quite rough and noisy. So if the measured values suggest the growth of <em><span>f</span></em>(<em>A</em>) or the shrinking of the generation time is anywhere near this border, it will be extremely hard to empirically determine the future behaviour. These limitations are important for companies setting their internal evaluations and for the prospects of regulation.</p><p>After examining what produces singular growth, we also looked at the choice of how we measure intelligence. We saw that whether intelligence is growing super-exponentially often depends on the nature of our measure, while whether it is growing towards a singularity is much less dependent on it. This fact means it isn&#8217;t crucial to decide whether to use some measure or the log of that measure &#8212; you would look for the same signature of how <em>&#550;</em> is growing with <em>A</em> in both cases. But when one measure can rise to infinity without the other following suit, the choice of measure really matters. In particular, some popular measures like METR time horizons are probably not the right tool for the job, since they can go to infinity without bringing other measures of intelligence along with them.</p><p>And we looked at how various finite limits might bite first, preventing truly singular growth. As well as the familiar idea of intelligence saturating at some maximal level (or maximal given various constraints) a minimal generation time could also derail the models of singular growth far before we reach that point.</p><p>Finally, while I&#8217;ve argued that singular growth is harder than we may have thought, that doesn&#8217;t mean RSI is safe or that AI R&amp;D will move at a manageable pace. RSI might be able to speed AI R&amp;D up to dangerously fast speeds even just with a linear speed-up. For example, if the human-only trajectory were <em>A</em>(<em>t</em>) and RSI sped this up to <em>A</em>(10<em>t</em>), we&#8217;d be getting a decade of human-only progress each year, introducing many of the dangers &#8212; even without any change in the fundamental shape of the curve.</p><p><em>With thanks to Tom Davidson, Rohin Shah, Loren Fryxell, Fin Moorhouse, Matthew van der Merwe, Will MacAskill, Robert Trager, Simon Biggs, and Shamil Chandaria for helpful discussions.</em></p><h2>Appendix: Table of Rates of Growth</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UTdo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UTdo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 424w, https://substackcdn.com/image/fetch/$s_!UTdo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 848w, https://substackcdn.com/image/fetch/$s_!UTdo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 1272w, https://substackcdn.com/image/fetch/$s_!UTdo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UTdo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png" width="1051" height="2373" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2373,&quot;width&quot;:1051,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:137085,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/213060056?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UTdo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 424w, https://substackcdn.com/image/fetch/$s_!UTdo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 848w, https://substackcdn.com/image/fetch/$s_!UTdo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 1272w, https://substackcdn.com/image/fetch/$s_!UTdo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c81284f-3536-4c74-ac36-09754b8771e8_1051x2373.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Table 1. The rates of growth of A over time produced by different rates of return f(A) in the differential equation &#550;&#729; = f(A). Note the unusually &#8216;mild&#8217; singular growth from f(A) = A </em>log<em>&#8289;(A)<sup>2</sup>, which is faster, though less &#8216;explosive&#8217; than any hyperbolic growth, and the unusually severe singular growth, from f(A)=e<sup>A</sup>, which lags behind any hyperbolic growth, but explosively catches up at the final instant.</em></figcaption></figure></div><h2>References</h2><p>Philippe Aghion, Benjamin F. Jones, and Charles I. Jones. (2017). &#8216;Artificial Intelligence and Economic Growth&#8217;, NBER Working Paper 23928.</p><p>Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W. Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, Nando de Freitas. (2016). &#8216;Learning to learn by gradient descent by gradient descent&#8217;, <a href="https://arxiv.org/abs/1606.04474">https://arxiv.org/abs/1606.04474</a></p><p>Anthropic. (2026). System Card: Claude Mythos Preview [System Card]. <a href="https://www.anthropic.com/claude-mythos-preview-system-card">https://www.anthropic.com/claude-mythos-preview-system-card</a></p><p>Nicholas Bloom, Charles I. Jones, John Van Reenen, and Michael Webb. (2020). &#8216;Are Ideas Getting Harder to Find?&#8217;, <em>American Economic Review</em>, 110(4):1104--44.</p><p>Nick Bostrom. (2014). Superintelligence: Paths, Dangers, Strategies, Oxford University Press.</p><p>David J. Chalmers. (2010). &#8216;The Singularity: A Philosophical Analysis&#8217;, <em>Journal of Consciousness Studies</em> 17:7-65.</p><p>Alan Chan, Ranay Padarath, Joe Kwon, Hilary Greaves, Markus Anderljung. (2026). &#8216;Measuring AI R&amp;D Automation&#8217;, <a href="https://arxiv.org/abs/2603.03992">https://arxiv.org/abs/2603.03992</a></p><p>Fran&#231;ois Chollet. (2019). &#8216;On the Measure of Intelligence&#8217;. <a href="https://arxiv.org/abs/1911.01547">https://arxiv.org/abs/1911.01547</a></p><p>Tom Davidson. (2025a). &#8216;How Can AI Labs Incorporate Risks from AI Accelerating AI Progress Into Their Responsible Scaling Policies?&#8217;.\ <a href="https://www.forethought.org/research/how-can-ai-labs-incorporate-risks-from-ai-accelerating-ai-progress-into">https://www.forethought.org/research/how-can-ai-labs-incorporate-risks-from-ai-accelerating-ai-progress-into</a></p><p>Tom Davidson. (2025b). &#8216;Will the need to retrain AI models from scratch block a software intelligence explosion?&#8217;.\ <a href="https://www.forethought.org/research/will-the-need-to-retrain-ai-models">https://www.forethought.org/research/will-the-need-to-retrain-ai-models</a></p><p>Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos, and Guillem Bas. (2023). &#8216;AI capabilities can be significantly improved without expensive retraining.&#8217;\ <a href="https://arxiv.org/abs/2312.07413">https://arxiv.org/abs/2312.07413</a></p><p>Tom Davidson, Rose Hadshar, and Will MacAskill. (2025). &#8216;Three Types of Intelligence Explosion&#8217; <a href="https://www.forethought.org/research/three-types-of-intelligence-explosion">https://www.forethought.org/research/three-types-of-intelligence-explosion</a></p><p>Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek. (2026). &#8216;When Does Automating AI Research Produce Explosive Growth? Feedback Loops in Innovation Networks&#8217;, NBER Working Paper No. 35155.</p><p>Tom Davidson and Tom Houlden. (2025). &#8216;How quick and big would a software intelligence explosion be?&#8217;. <a href="https://www.forethought.org/research/how-quick-and-big-would-a-software-intelligence-explosion-be">https://www.forethought.org/research/how-quick-and-big-would-a-software-intelligence-explosion-be</a></p><p>Ege Erdil, Tamay Besiroglu, and Anson Ho. (2024). &#8216;Estimating Idea Production: A Methodological Survey&#8217;, <em>SSRN</em>. <a href="http://dx.doi.org/10.2139/ssrn.4814445">http://dx.doi.org/10.2139/ssrn.4814445</a></p><p>Daniel Eth and Tom Davidson. (2025). &#8216;Will AI R&amp;D Automation Cause a Software Intelligence Explosion?&#8217;. <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion</a></p><p>A. Fawzi, M. Balog, A. Huang et al. (2022). &#8216;Discovering faster matrix multiplication algorithms with reinforcement learning&#8217;. <em>Nature</em> 610:47--53. <a href="https://doi.org/10.1038/s41586-022-05172-4">https://doi.org/10.1038/s41586-022-05172-4</a></p><p>I. J. Good. (1965). &#8216;Speculations Concerning the First Ultraintelligent Machine,&#8217; In F. Alt &amp; M. Rubinoff, <em>Advances in Computers</em> (volume 6). Academic Press.</p><p>Demis Hassabis, Dario Amodei, and Zanny Minton Beddoes. (2026). <em>The Day After AGI</em>. World Economic Forum [Panel]. <a href="https://www.weforum.org/meetings/world-economic-forum-annual-meeting-2026/sessions/the-day-after-agi/">https://www.weforum.org/meetings/world-economic-forum-annual-meeting-2026/sessions/the-day-after-agi/</a></p><p>Charles I. Jones. (1995). &#8216;R&amp;D-Based Models of Economic Growth,&#8217; <em>Journal of Political Economy</em> 103(4):759--84.</p><p>Andrej Karpathy. (2026). &#8216;Autoresearch&#8217;. <a href="https://github.com/karpathy/autoresearch">https://github.com/karpathy/autoresearch</a></p><p>Michael Kremer. (1993). &#8216;Population Growth and Technological Change: One Million B.C. to 1990&#8217;, <em>The Quarterly Journal of Economics</em> 108(3):681--716.</p><p>Daniel Kokotajlo and Eli Lifland. (2025). &#8216;Takeoff Forecast&#8217; <a href="https://ai-2027.com/research/takeoff-forecast">https://ai-2027.com/research/takeoff-forecast</a></p><p>Raymond Kurzweil. (2001). &#8216;The law of accelerating returns&#8217;. <a href="https://www.writingsbyraykurzweil.com/the-law-of-accelerating-returns">https://www.writingsbyraykurzweil.com/the-law-of-accelerating-returns</a></p><p>Ray Kurzweil, Vernor Vinge, and Hans Moravec. (2003). &#8216;Singularity math trialogue&#8217;. <a href="https://www.thekurzweillibrary.com/singularity-math-trialogue">https://www.thekurzweillibrary.com/singularity-math-trialogue</a></p><p>Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, Lawrence Chan. (2025). &#8216;Measuring AI Ability to Complete Long Software Tasks&#8217;. <a href="https://arxiv.org/abs/2503.14499">https://arxiv.org/abs/2503.14499</a></p><p>Eli Lifland, Brendan Halstead, Alex Kastner, and Daniel Kokotajlo. (2026) <em>AI Futures Model</em>. https://www.aifuturesmodel.com</p><p>Hans Moravec. (1999). &#8216;Simple equations for Vinge&#8217;s technological singularity&#8217;. <a href="https://frc.ri.cmu.edu/%5C~hpm/project.archive/robot.papers/1999/singularity.html">https://frc.ri.cmu.edu/\~hpm/project.archive/robot.papers/1999/singularity.html</a></p><p>Hans Moravec. (2003). &#8216;Simpler equations for Vinge&#8217;s technological singularity&#8217;. <a href="https://frc.ri.cmu.edu/~hpm/project.archive/robot.papers/2003/singularity2.html">https://frc.ri.cmu.edu/~hpm/project.archive/robot.papers/2003/singularity2.html</a></p><p>Anders Sandberg. (2013). &#8216;An Overview of Models of Technological Singularity&#8217;. In <em>The Transhumanist Reader</em> (eds M. More and N. Vita-More). <a href="https://doi.org/10.1002/9781118555927.ch36">https://doi.org/10.1002/9781118555927.ch36</a></p><p>Karl Schwarzschild. (1916). &#8216;&#220;ber das Gravitationsfeld eines Massenpunktes nach der Einsteinschen Theorie&#8217;. <em>Sitzungsberichte der K&#246;niglich Preussischen Akademie der Wissenschaften</em>. 7: 189--196. Translated as, Antoci, S.; Loinger, A. (1999). &#8216;On the gravitational field of a mass point according to Einstein&#8217;s theory&#8217;. <a href="https://arxiv.org/abs/physics/9905030">https://arxiv.org/abs/physics/9905030</a></p><p>Ray J. Solomonoff. (1985). &#8216;The time scale of artificial intelligence: reflections on social e&#64256;ects&#8217;. <em>North-Holland Human Systems Management</em> 5:149--153.</p><p>Philip Trammell and Anton Korinek. (2023). &#8216;Economic Growth under Transformative AI,&#8217; NBER Working Paper 31815, <a href="https://doi.org/10.3386/w31815">https://doi.org/10.3386/w31815</a>.</p><p>Lizka Vaintrob and Owen Cotton-Barratt. (2025). &#8216;AI Tools for Existential Security&#8217;. <a href="https://www.forethought.org/research/ai-tools-for-existential-security">https://www.forethought.org/research/ai-tools-for-existential-security</a></p><p>Eliezer Yudkowsky. (2001). &#8216;Creating friendly AI 1.0: The Analysis and Design of Benevolent Goal Architectures&#8217;, The Singularity Institute.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>We could also write this as an explicit function of t, such as <em>&#550;<sub>t </sub></em>= <em>kA<sub>t</sub></em> or <em>&#550;</em>(<em>t</em>) = <em>kA</em>(<em>t</em>). But we&#8217;ll follow the standard shorthand of dropping the explicit reference to t where possible.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>The same move was made much earlier by Solomonoff (1985) and Moravec (1999, 2003) in their differential equations for intelligence explosions from AI labour speeding up Moore&#8217;s Law.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Defining labour in this way effectively means we are assuming there is <em>only</em> AI labour. While slightly more complicated, it is also possible to model a mixture of human and AI labour, with the AI share increasing over time as it becomes more efficient (Davidson &amp; Houlden 2025).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>For example, taking the exponential, <em>&#550; </em>= <em>kA</em>, and subtracting a quadratic term that has been shrunk down by a large factor, <em>K</em>, gives the differential equation for logistic growth, <em>&#550; </em>= <em>kA </em>&#8722; (<em>kA</em><sup>2</sup>/<em>K</em>)&#8203;, which has a plateau at height <em>K</em>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>As written, this is quite restrictive since it implies <em>A</em> is monotonically increasing. However, that is an artefact of this simple autonomous setup where <em>&#550;</em> is just a function of <em>A</em>. If you allow it to be a function of <em>A</em> and <em>t</em>, or instead look at <em>&#196;</em> as a function of <em>A</em> and <em>&#550;</em>, then you can also have trajectories that dip for a while before blowing up to a vertical asymptote. For that richer class of differential equations, one would need to define a less restrictive version of this condition.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>I subsequently discovered that Kurzweil (Kurzweil, Vinge, and Moravec 2003) suggested this &#550; = <em>kA </em>log&#8289;(<em>A</em>) form was more plausible than those with <em>A</em> raised to powers greater than one, since it avoided the vertical asymptote. Instead it 'merely' produces doubly exponential growth which he thought was already visible in the long run data (Kurzweil 2001). And he justified <em>A </em>log(&#8289;<em>A</em>) via the same decomposition I gave above. Despite being closely associated with the idea of a 'technological singularity' he found dynamics leading towards a mathematical singularity unlikely &#8212; 'It is hard to explain how we could get infinite knowledge, or infinite information processing, from a finite world'.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>For example, the mean age of the parent at the birth of each of their children.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>One could note that the definition of <em>T</em><sub>n</sub>&#8203; is itself a difference equation: &#916;<em>t</em><sub>n </sub>= <em>t</em><sub>n+1</sub> &#8722; <em>t</em><sub>n </sub>= <em>T</em><sub>n+1</sub>&#8203;. One could therefore think of time-embedded difference equations as a linked system of difference equations for AA and for tt, linked by their index nn.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>For example, it may be that on a linear measure their score is rising as <em>t</em><sup>10</sup> while on a logarithmic measure it is only rising as 10 log&#8289;(<em>t</em>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>The same is <em>not</em> true for the threshold of super-exponential growth, since <em>A</em>(<em>t</em>) could grow super-exponentially while its log, <em>B</em>(<em>t</em>), doesn't.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>There are practical reasons why the METR horizon length may not be able to go to infinity &#8212; the tasks included in the benchmark only go up to 16 hours of human time and METR explicitly say their current version can't be used to estimate the horizon lengths beyond that. But I'm pointing to a theoretical issue &#8212; even if they had endless capacity to create longer tasks and endless immortal human baseliners, to enable arbitrarily high measurements, this is the kind of measure where time horizons diverging to infinity correspond to accuracies converging to 100% on a fairly narrow range of cognitive tasks, rather than infinite intelligence.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>For example, the class of functions computable by Turing machines is of immense use in computer science, even though it collapses to be the same as the functions computed by finite state machines (or look-up tables) if we restrict ourselves to finite settings. Similarly, the asymptotic analysis of time complexities of algorithms is very useful even though it disappears if we were to restrict ourselves to algorithms whose inputs are bounded by the size of the observable universe.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Here I've been taking the approach of treating <em>f</em>(<em>A</em>) as a clean mathematical function that only captures part of the dynamics, so requires a correction term. A different way to view things is to let <em><span>f</span></em>(<em>A</em>) represent the actual messy real-world connection between <em><span>&#550;</span></em> and <em>A</em> (i.e. the empirical function you would graph as you measure an intelligence explosion). On this approach you don't need to subtract a second function, but instead would talk about whether the real <em>f</em>(<em>A</em>) eventually falls to zero at some level of <em>A</em>, creating a horizontal asymptote at that height.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>Standard models based on semi-endogenous growth theory model diminishing returns in a scale-free way, such that diminishing returns alone can't come to halt the ascent of <em>A</em>. But this is just a byproduct of their simple power-law model. It is very easy to get horizontal asymptotes in more flexible differential equations.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>Though note that the amount by which the growth rates are rising per year slows down towards the end, so any particular clean super-exponential shape would only fit this phase until this point of inflection of growth rates.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Notes on Implications of Scale-Dependent Algorithms]]></title><description><![CDATA[This article was created by Forethought. See all our research on our website.]]></description><link>https://newsletter.forethought.org/p/notes-on-implications-of-scale-dependent</link><guid isPermaLink="false">https://newsletter.forethought.org/p/notes-on-implications-of-scale-dependent</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Thu, 13 Aug 2026 13:43:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qyaP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><p><span>Of past algorithmic improvements which decrease LLM pretraining loss, many have been &#8220;scale-dependent&#8221; &#8211; that is, improvements which decrease LLM pretraining loss by more, relative to prior algorithms, at larger quantities of compute.</span></p><p><span>The influence of such scale-dependent algorithms is (1) moderate evidence that attempts to limit algorithmic progress in absence of compute limitations will be ineffective, (2) weak evidence that our inferences about the future scale of algorithmic progress, based on the past, are invalid, and (3) part of a plausible argument either for or against a software intelligence explosion, depending on other details of one&#8217;s model.</span></p><p><span>I&#8217;ll proceed by discussing the following:</span></p><ol><li><p><span>How scale-dependent algorithms for pretraining loss probably account for the majority of algorithmic improvements in this domain</span></p></li><li><p><span>How scale-dependent algorithms for end-to-end task performance might account for the majority of algorithmic improvements in this domain</span></p></li><li><p><span>How scale-dependent algorithms in the past should influence our model of algorithmic improvements in the future</span></p></li><li><p><span>Possible policy implications</span></p></li></ol><p><span>Before starting, a note on method &#8211; when discussing questions regarding AI&#8217;s past and future progress, there&#8217;s a natural continuum from &#8220;the intractable, high-impact, abstract question, which we terminally care about&#8221; to &#8220;the tractable, dubious-impact, concrete question, which we instrumentally care about.&#8221;</span></p><p><span>Consider questions about scale-dependent algorithmic improvements, sliding from the former to the latter along such a continuum.</span></p><ul><li><p><span>We are most terminally interested in questions about the future: &#8220;Will </span><em><span>future</span></em><span> improvements to AI </span><em><span>task performance </span></em><span>depend more on compute scale-up or on algorithmic improvements? What can we know about how these two will be causally intertwined? Will future improvements enable a software-only intelligence explosion &#8211; that is, in a world with total compute held constant, could algorithmic progress lead to an intelligence takeoff in a comparatively short period of time?&#8221;</span></p></li><li><p><span>The data we can gather most relevant to this question is about the past: &#8220;Did </span><em><span>past</span></em><span> improvements to AI </span><em><span>task performance</span></em><span> depend more on compute scale-up or on algorithmic improvements? Would past AI algorithmic improvement have counterfactually looked slower, without the datacenter buildout?&#8221;</span></p></li><li><p><span>And finally we have the narrower sub-question for which we have the most past data, which contributes to answering the broader historical question: &#8220;Did past improvements to LLM </span><em><span>pretraining loss </span></em><span>depend more on compute scale-up or on algorithmic improvements?&#8221;</span></p></li></ul><p><span>I&#8217;m going to start by discussing the evidence for this last question, before moving to the increasingly impactful and increasingly difficult earlier questions.</span></p><h2><span>Scale-Dependence of Algorithmic Improvements for Pretraining Loss</span></h2><p><span>Are the algorithmic improvements that have most decreased pretraining loss for LLMs scale-dependent or scale-independent? What would either of these mean?</span></p><p><span>The </span><a href="https://arxiv.org/pdf/2312.07413"><span>compute-equivalent-gain</span></a><span> (CEG) for some algorithm A&#8217;, relative to another algorithm A, is the ratio of the FLOPs needed to train an A-using model to a given level of performance, over the FLOPs needed for an A&#8217;-using model to reach that same level. If it takes 3x as much compute to train an LLM with algorithm A to reach the same pretraining loss as another LLM trained with algorithm A&#8217;, then A&#8217; has a 3x compute-equivalent gain on pretraining loss relative to A.</span></p><p><span>Given the above notion of CEG, algorithmic improvements that produce some kind of CEG could conceivably be </span><strong><span>scale independent </span></strong><span>or </span><strong><span>scale dependent</span></strong><span>.</span></p><p><span>A scale-independent algorithm has a constant CEG across different quantities of training FLOPs. So if A&#8217; gives a 2x multiplier relative to A at about 10</span><sup><span>15</span></sup><span> FLOPs, it continues to give a 2x multiplier at about 10</span><sup><span>25</span></sup><span> FLOPs. Correspondingly, if one were somehow able to know antecedently that some algorithm was scale independent, one would just need to test it at one scale to find whatever this constant happened to be.</span></p><p><span>By contrast, a scale-dependent algorithm has a variable CEG, which is a function of the number of FLOPs used in training. So A&#8217; might give a 2x multiplier relative to A at about 10</span><sup><span>15</span></sup><span> FLOPs, but give a radically different 1000x multiplier at about 10</span><sup><span>25</span></sup><span> FLOPs. If one were somehow able to know antecedently that some algorithm was scale dependent, one would know that one needed to test it at many scales to find the function mapping from FLOPs to the compute multiplier.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qyaP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qyaP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 424w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 848w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1272w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qyaP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png" width="1456" height="596" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:596,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qyaP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 424w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 848w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1272w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The thesis of 2025&#8217;s &#8220;</span><a href="https://arxiv.org/pdf/2511.21622"><span>On the Origin of Algorithmic Progress in AI</span></a><span>&#8221; is that the vast majority of apparent algorithmic progress in LLM pretraining loss per-FLOP</span><strong><span> </span></strong><span>has come from a handful of scale-dependent algorithms. So, according to the paper, the apparently steady flow of &#8220;algorithmic improvements&#8221; in the past came from choosing a reference point algorithm (A) prior to some particular scale-dependent algorithmic improvement (A&#8217;), and then mistaking the continued improvements A&#8217; makes over A as compute increases for the presence of additional algorithms, B, C, and so on. In the absence of an increase of FLOPs, this apparent radical increase in performance would not have been nearly as large.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HiZF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HiZF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 424w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 848w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1272w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HiZF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png" width="1456" height="994" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:994,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HiZF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 424w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 848w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1272w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>What kind of evidence is there that the majority of progress in LLM pretraining loss per-FLOP actually does depend on scale-dependent algorithms?</span></p><ul><li><p><span>Well, the paper presents good evidence that scale-dependent algorithms are very important.</span></p><ul><li><p><strong><span>LSTM to Transformer Change</span></strong><span>: &#8220;On the Origin&#8221; uses experiments to estimate how much the Transformer improves over the LSTM as compute scales up. At ~10</span><sup><span>15</span></sup><span> FLOPs, this comes to about 6x CEG; at ~10</span><sup><span>17</span></sup><span>, it comes to a 28x CEG increase; at the level of compute with which the Transformer was introduced in 2017, they extrapolate it to about a 91x CEG increase.</span></p></li><li><p><strong><span>Kaplan to Chinchilla Change</span></strong><span>: &#8220;On the Origin&#8221; grants that Chinchilla scaling is correct, then projects forward the gains you get by switching to Chinchilla from Kaplan; they are approximately equivalent at 10</span><sup><span>20</span></sup><span>, and give a ~10x or more gain at frontier scales.</span></p></li></ul></li><li><p><span>The paper also presents evidence that scale-independent algorithms are less important than one might imagine.</span></p><ul><li><p><span>Their ablations find that scale-independent additions are not multiplicative. That is, doing a bucket of changes that move one from the &#8220;retro&#8221; Transformer to the modern Transformer in theory should give a CEG of 3.5x, if one were to multiply the gains that the additions make separately; but in practice only gives 1.3x.</span></p></li><li><p><span>The modern transformer similarly does not improve greatly on the &#8220;retro&#8221; Transformer as you scale up.</span></p></li></ul></li></ul><p><span>So &#8220;On the Origin&#8221; projects that, </span><strong><span>given </span></strong><span>the actual historical increase in FLOPs that makes scale-dependent algorithms more efficient, one observes a 21,400x CEG increase from 2017 to 2025. But if one counterfactually holds compute constant at 10^18 FLOPs, this amounts to only a 155x CEG increase. This amounts to the difference between a ~7 month doubling time for algorithmic efficiency (very close to </span><a href="https://epoch.ai/publications/algorithmic-progress-in-language-models"><span>Ho et al. (2024)</span></a><span>&#8217;s 8-month doubling time); and a ~13 month doubling time of algorithmic efficiency. But note that the second &#8220;doubling time&#8221; here is specifically relative to the level of compute at which one calculates the increase in algorithmic efficiency; there would be a different doubling time for a different absolute quantity of compute.</span></p><p><span>So there is a (very approximate) halving of the speed of algorithmic improvement, given</span><strong><span> </span></strong><span>that you hold FLOPs fixed at 10</span><sup><span>18</span></sup><span>, once you take into account the difference between scale-dependent and scale-independent algorithms. And of course, this ~13 month doubling time is still fueled almost entirely by the gigantic LSTM to Transformer transition at 10</span><sup><span>18</span></sup><span> FLOPs &#8211; leaving this quantity of algorithmic progress out means that the doubling time gets vastly larger.</span></p><p><span>This is a pretty large change to the proposed rate at which constant-FLOP algorithmic improvements occur. But you can actually make a case that &#8220;On the Origin&#8221; understates the degree to which algorithmic improvements in LLM pretraining loss depend on scale-dependent algorithms.</span></p><p><span>The largest single purportedly scale-independent gain (2x) described in the paper belongs to mixture-of-experts (MoEs), an alteration to the Transformer architecture. But several subsequent works have shown that the gain from MoEs over dense architectures is actually strongly scale-dependent, and probably larger than 2x. &#8220;Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models&#8221; estimates a </span><a href="https://arxiv.org/pdf/2507.17702"><span>7x gain</span></a><span> at 10</span><sup><span>22</span></sup><span> FLOPs. &#8220;Scaling Laws for Fine-Grained Mixture-of-Experts&#8221; estimates a </span><a href="https://openreview.net/pdf?id=Iizr8qwH7J"><span>20x</span></a><span> gain at 10</span><sup><span>20</span></sup><span> FLOPs. In both cases, the size of the CEG relative to a dense model keeps improving with the quantity of FLOPs. Although it&#8217;s uncertain which exponent is correct, both papers point toward MoEs being a scale-dependent algorithmic improvement that substantially improves with the quantity of FLOPs.</span></p><p><span>So &#8211; to return to the motivating question &#8211; most of the algorithmic improvements to LLM pretraining loss, so far, have probably depended on a compute scale-up. Apparent algorithmic progress in pretraining loss would have been much smaller if total compute had been held constant.</span></p><h2><span>Scale-Dependence of Algorithmic Improvements for End-to-End Performance</span></h2><p><span>No one actually cares about pretraining loss for its own sake, though.</span></p><p><span>&#8220;Decreased pretraining loss&#8221; is just one way to improve end-to-end performance. There are also post-training </span><a href="https://arxiv.org/pdf/2312.07413"><span>innovations</span></a><span> like </span><a href="https://arxiv.org/abs/1909.08593"><span>RLHF</span></a><span> or best-of-N that improve task performance. There are also changes to </span><a href="https://www.beren.io/2025-08-02-Most-Algorithmic-Progress-is-Data-Progress/"><span>data mixtures</span></a><span> that improve performance. And of course </span><a href="https://arxiv.org/abs/2501.12948"><span>RLVR-over-CoT</span></a><span> has probably been by itself responsible for the </span><a href="https://epoch.ai/gradient-updates/quantifying-the-algorithmic-improvement-from-reasoning-models"><span>greatest</span></a><span> single leap in AI performance of the last few years.</span></p><p><span>How much are algorithmic improvements to end-to-end performance scale-dependent? In particular, how much does RLVR-over-CoT depend on scale-dependent algorithmic improvements?</span></p><p><span>This is deeply uncertain. There&#8217;s certainly no work that provides a clean FLOPs-to-CEG function, as in the case of pretraining loss.</span></p><p><span>There are several reasons for thinking that RLVR works extremely poorly or not at all beneath a certain scale, and so is somewhat scale-dependent. You need a certain success rate for reinforcing successes to be viable. </span><a href="https://arxiv.org/pdf/2607.12395"><span>Several</span></a><span> </span><a href="https://arxiv.org/pdf/2501.12948"><span>works</span></a><span> remark on how RLVR without sufficient scale simply doesn&#8217;t teach the required diversity of behaviors, and that RLVR on small models plateaus far lower than RLVR on larger and longer-trained models. So there&#8217;s suggestive </span><a href="https://arxiv.org/pdf/2607.16097v1"><span>evidence</span></a><span> that at the frontier RLVR is scale-dependent, but nothing conclusive.</span></p><p><span>It would also be difficult to perform an experiment to cleanly determine whether RLVR-over-CoT is scale-dependent for several reasons. For instance, the importance of RL, in general, </span><a href="https://arxiv.org/pdf/2509.25123"><span>probably</span></a><span> </span><a href="https://arxiv.org/abs/2512.07783"><span>comes</span></a><span> </span><a href="https://arxiv.org/abs/2501.17161"><span>from</span></a><span> how it permits at least somewhat out-of-distribution generalization, rather than from how it improves performance on in-distribution tasks. But to my knowledge, there&#8217;s no publicly proposed clear way to measure out-of-distribution generalization; so there&#8217;s no measure against which one could test how RLVR improves with scale on the matter where it ostensibly has the greatest impact. Similarly, RLVR takes place as the final stage of training, after a large sequence of data-selection, pretraining, and midtraining, which means there are just many more free variables it would be necessary to consider before deciding on an experimental design.</span></p><p><span>Apart from RLVR, there&#8217;s some evidence that some apparent scale-independent improvements may scale </span><em><span>negatively</span></em><span> with FLOPs, and thus provide few or limited gains at the frontier. For instance, the </span><a href="https://arxiv.org/pdf/2605.19407"><span>paper</span></a><span> &#8220;A Bitter Lesson for Data Filtering&#8221; claims that although filtering data improves performance at a small amount of compute, it scales negatively with FLOPs, such that with enough computing power it&#8217;s best to do no data filtering at all.</span></p><p><span>I think it&#8217;s rather likely that RLVR and other important post-training techniques improve in a strongly scale-dependent way, and some of my analysis below will lean on that belief. I&#8217;d be surprised if RLVR doesn&#8217;t improve in a </span><em><span>somewhat </span></em><span>scale-dependent way.</span></p><p><span>But I don&#8217;t think there&#8217;s any conclusive empirical evidence here, and to the degree I&#8217;m wrong, some of my inferences below are less likely to be true.</span></p><h2><span>How scale-dependent improvements in the past impact our model of &#8220;algorithmic progress&#8221; for the future</span></h2><p><span>Let&#8217;s grant that, in the past, scale-dependent algorithms have been responsible for the majority of algorithmic improvements in both LLM pretraining loss and in end-to-end task performance.</span></p><p><span>What consequences would this have for how we expect algorithmic progress to go in the future? What consequences does this imply about a future where compute is &#8220;held constant,&#8221; as in the case of a software intelligence explosion? Would it lead us to expect a radical trend break, in either direction?</span></p><p><span>Let&#8217;s take the other side by contrast, first &#8211; imagine a model of the world where almost all algorithmic improvements are scale </span><em><span>independent</span></em><span>: improvements are drawn from a pool of possible improvements, and the number and efficacy of the algorithms taken from the pool do not change merely by the fact of compute scaling up.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zTQE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zTQE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 424w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 848w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1272w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zTQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png" width="836" height="576" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:576,&quot;width&quot;:836,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zTQE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 424w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 848w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1272w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>If you operate off this model, then there are some natural consequences.</span></p><ul><li><p><span>Algorithmic improvements do not become &#8220;more discoverable&#8221; as FLOPs scale up. More compute may make it easier to run experiments to find improvements, but scaling up the FLOP count within each experiment does not make finding the right algorithm easier or harder per-experiment.</span></p><ul><li><p><span>There could, of course, be other factors that make algorithmic improvements easier or harder to find. One improvement could open up a line of research that leads to other improvements, making it easier; it might decrease the total pool of improvements, making it harder. But the fact of having scaled up FLOPs alone does not make improvements easier or harder to find.</span></p></li></ul></li><li><p><span>Some particular algorithmic improvement does not become &#8220;higher impact,&#8221; should its discovery be pushed into the future. If some algorithm X is found today and gives a 3x CEG, it would equally well give a 3x CEG if it were found two years into the future.</span></p></li><li><p><span>Broadly, our actual historical rates of &#8220;algorithmic progress&#8221; reflect something like the rate at which algorithms are naturally found, in a way indifferent to compute build-out.</span></p></li></ul><p><span>By contrast, one can model a world where most algorithmic improvements are scale dependent as a world where improvements are drawn from a changing pool of FLOP-specific algorithmic improvements. Pools get &#8220;unlocked&#8221; as the compute frontier moves past the point where they start to become significant, and the pools grow progressively more explored over time as more and more researchers have access to the compute necessary to explore them. So both &#8220;how explored&#8221; each pool is, and the consequences of finding an entry from each pool, are always changing.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PaXs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PaXs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 424w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 848w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1272w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PaXs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png" width="1456" height="784" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PaXs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 424w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 848w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1272w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>If you operate off this model, then:</span></p><ul><li><p><span>Algorithmic improvements become much &#8220;more discoverable&#8221; as FLOPs scale up; the more FLOPs you use in an experiment, the higher the relative CEG might become, and so the easier the new algorithm will find it to overcome the &#8220;well-tuned baseline&#8221; algorithm with very well-adjusted hyperparameters. So in addition to letting you run more experiments, additional compute may enhance &#8220;research taste.&#8221;</span></p></li><li><p><span>Some particular algorithmic improvement might actually become &#8220;higher impact,&#8221; should its discovery be pushed into the future. If some algorithm X is found today and gives a 3x CEG, it might give a 100x CEG in the future when applied with more FLOPs, if its discovery is delayed.</span></p></li><li><p><span>Broadly, this view implies that &#8220;algorithmic progress&#8221; depends on past compute scale-ups, and would be much slower without them.</span></p></li></ul><p><span>Imagine that one is trying to determine how much algorithmic progress is likely in the future, granted that one is in a scenario where total compute is held constant. It&#8217;s important to note that switching to a view of algorithmic progress as scale-dependent might move your estimate of the likely rate of algorithmic progress in a constant-compute regime either up or down.<br><br>The case for moving it down is clear. Switching to a view where scale-dependent improvements are significant could move one&#8217;s estimate of algorithmic progress down, because it moves one estimate of past algorithmic progress in a constant-compute regime down. One used to estimate a ~8 month halving time to reach equivalent loss for equal FLOPs; now one estimates a ~13 month halving time; and the future is likely to be like the past, so future FLOP-constant algorithmic progress is likely to be slower.</span></p><p><span>But switching to such a view could move one&#8217;s estimate of algorithmic progress up, because it decreases your credence that the future will be like the past in this respect! If different algorithms get unlocked at different levels of compute, then potentially some future yet-to-be-unlocked regime of algorithmic improvements might differ enormously from past algorithmic improvements. There&#8217;s likely less reason to think that the number or importance of algorithmic improvements available in one prior algorithmic bucket will be equally important as algorithmic improvements in subsequent buckets as compute increases.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5Oa2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5Oa2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 424w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 848w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1272w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png" width="1456" height="931" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:931,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5Oa2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 424w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 848w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1272w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>So if one finds such a picture compelling, then the likelihood one assigns to a SIE could go up. If there is a future compute scale-up to 10</span><sup><span>32</span></sup><span> FLOPs, a SIE taking place then might find many algorithms that were totally unavailable to us in the 10</span><sup><span>27</span></sup><span> regime.</span></p><p><span>Overall, all this remains uncertain. I think various philosophical arguments suggest that there&#8217;s a greater diversity of room for algorithmic progress as compute scales up. The human brain is largely a </span><a href="https://pubmed.ncbi.nlm.nih.gov/19915731/"><span>linearly scaled-up</span></a><span> primate brain, but individual humans and humans in aggregate can do many calculations no non-human primate can. And there are likely greater gains from organizing 10,000 humans well than from organizing 10 humans well.</span></p><h2><span>Possible Policy Implications</span></h2><p><span>If a majority of algorithmic improvements to task performance consistently come from scale-dependent algorithms, and are likely to keep doing so in the future, how would this impact policy interventions into AI?</span></p><p><span>There are two notable consequences, which mirror each other: (1) efforts to decrease algorithmic progress, in absence of a training compute scale-up, would be less necessary than previously believed; and (2) efforts to decrease algorithmic progress, in the presence of a training compute scale-up, would be less effective than previously believed. I will discuss them each in turn.</span></p><p><strong><span>First, </span></strong><span>some policy interventions meant to slow or stop the development of advanced AIs have focused on limiting the size of the largest training runs. In this context, a steady advance of scale-independent algorithmic improvements would have decreased the efficacy of any particular maximum training-run size limitation. For instance </span><a href="https://arxiv.org/pdf/2511.10783"><span>&#8220;An International Agreement to Prevent the Premature Creation of Artificial Superintelligence&#8221;</span></a><span> recommends a maximum training-run size of 10</span><sup><span>24</span></sup><span>, but also notes that the &#8220;number of operations used to train an AI to a given capability level drops by 3x each year.&#8221; Given this trend, it&#8217;s easier to derive the need to prohibit AI algorithmic research, a measure that the paper itself notes may be &#8220;controversial and normally a bad idea,&#8221; but which it believes to be necessary to avoid the creation of ASI.</span></p><p><span>But if the majority of algorithmic improvements have been scale-dependent, then prior estimates for scale-independent algorithmic improvements are too high; and so prior work gives a less-good reason for thinking that the number of operations used to train an AI to a given capability level drops by 3x every year. So the need to ban and monitor research would be correspondingly decreased; a hard cap on maximum training-run size would be more effective than previously estimated.</span></p><p><strong><span>Second, </span></strong><span>some proposed policy interventions have focused on </span><strong><span>scaling up </span></strong><span>the size of compute while trying to decrease or hold constant algorithmic progress. AI 2040, for instance, proposes continuing to increase the </span><a href="https://ai-2040.com/supplements/compute-supplement"><span>size</span></a><span> of training runs from 10</span><sup><span>26</span></sup><span> now to 10</span><sup><span>32</span></sup><span> in 2035, while decreasing the rate of algorithmic progress.</span></p><p><span>Overall, if you believe that the majority of important future algorithmic improvements will be scale dependent in the way that the majority of past algorithmic improvements were scale dependent, that should decrease the credence you give to our ability to &#8220;not find&#8221; these improvements. Why?</span></p><p><span>As mentioned above &#8211; as compute scales up, the increasing gains from scale-dependent improvements will make them more obvious. By analogy to a counterfactual world in which the Transformer was never discovered, an incredibly poorly-done, badly-implemented Transformer in 2026 still obviously improves on an LSTM, because even an awful implementation might mean it is only 100x better rather than 1000x better than the LSTM. It&#8217;s probably impossible to present the Transformer from being discovered for too long. Similarly, continuing to scale up compute makes scale-dependent algorithms more obvious; it probably decreases the quantity of research taste you need to discern them.</span></p><p><span>The other part of the reason is just because, well, continuing to scale up compute keeps opening up new possible unseen algorithmic buckets that allow improvement; the space of algorithmic improvements will keep growing, different kinds of research taste will find more progress.</span></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Liberalism Forever]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/liberalism-forever</link><guid isPermaLink="false">https://newsletter.forethought.org/p/liberalism-forever</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Fri, 31 Jul 2026 18:19:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d58976af-23a5-40f6-9944-8ff749331741_900x590.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-PViXDHfw3hg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;PViXDHfw3hg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/PViXDHfw3hg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.peternsalib.com/">Peter Salib</a> is an Assistant Professor of Law at the University of Houston Law Center and co-director of the Center for Law &amp; AI Risk. <a href="https://www.simondgoldstein.com/">Simon Goldstein</a> is an associate professor at the University of Hong Kong, a senior editor at AI Frontiers, and a visiting senior scholar at Forethought. They have written together about AI safety, AI rights, and the governance of advanced AI.</p><p>They joined Forethought&#8217;s <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss their new paper, <a href="https://philpapers.org/rec/GOLLFC-2">Liberalism Forever</a>. The conversation covers:</p><ul><li><p>Their &#8220;null hypothesis&#8221;: today&#8217;s liberal institutions &#8212; free markets and democratic governance &#8212; may already hold most of the answers for governing the far future, and longtermists have underrated their power</p></li><li><p>Examining whether canonical social-science arguments for markets and democracy still hold under transformative AI, space colonization, and explosive growth, and claiming that most survive (and several get stronger)</p></li><li><p>Why they favour reasoning &#8220;at the margin&#8221; over designing a detailed end-state, and why proposals like viatopia and the long reflection sound liberal but risk illiberalism when actually implemented</p></li><li><p>Whether the standard case for markets &#8212; information aggregation, allocative efficiency, innovation &#8212; still bites once AGI can read off preferences directly</p></li><li><p>Why inequality is, in their view, a comparatively boring problem with a known solution (tax-and-transfer), why AI-specific taxes like a compute tax are misguided, and how the growth-versus-distribution trade-off changes when growth rates get very high</p></li><li><p>Dyson swarms and space resources: why they doubt there&#8217;s any real monopoly worry, since energy would be additive with free entry, and why property auctions might beat egalitarian allocation schemes</p></li><li><p>The arguments for democracy that persist post-AGI &#8212; public-choice/selectorate incentives and democracy as a commitment device &#8212; versus the epistemic arguments, which they think weaken</p></li><li><p>Why AGI could make autocracy more competitive (loyal AI bureaucrats and militaries removing the need for a human winning coalition), and what preserving democratic control requires in response</p></li><li><p>Handoffs to superintelligent AI and successionism, and their preferred alternative of extending the franchise to AIs with their own values rather than deferring wholesale</p></li><li><p>The history of new &#8220;technologies of democracy,&#8221; from the Federalist Papers to LLMs as impartial arbiters, and where each author thinks their own argument is weakest</p></li></ul><p><a href="https://docs.google.com/document/d/1F0XrSiwCenc7h-Ke9Bqhrhmqty59unRA5sJgmdUjF6s/edit?tab=t.0">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[Should Philanthropists Save Their Money for the Intelligence Explosion?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/should-philanthropists-save-their</link><guid isPermaLink="false">https://newsletter.forethought.org/p/should-philanthropists-save-their</guid><dc:creator><![CDATA[Forethought]]></dc:creator><pubDate>Thu, 30 Jul 2026 20:33:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/_vSs8v0yFkU" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-_vSs8v0yFkU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;_vSs8v0yFkU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/_vSs8v0yFkU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.forethought.org/people/william-macaskill">Will MacAskill</a> is a senior research fellow at Forethought and the author of <em><a href="https://whatweowethefuture.com/">What We Owe The Future</a></em>. <a href="https://www.forethought.org/people/tom-davidson">Tom Davidson</a> is a senior research fellow at Forethought and the author of a series of reports on <a href="https://www.forethought.org/research/three-types-of-intelligence-explosion">AI timelines</a>, <a href="https://www.forethought.org/research/the-industrial-explosion">takeoff speeds</a>, and <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">AI-enabled coups</a>.</p><p>In this informal conversation about early-stage research in progress, they discuss:</p><ul><li><p>Why a philanthropist who takes transformative AI seriously should expect extraordinary investment returns (plausibly 10x to 100x) and how the market doesn&#8217;t seem to be pricing this in</p></li><li><p>Why AGI shouldn&#8217;t be treated as a hard deadline to spend by, and the case that most philanthropic money should be spent <em>during</em> the intelligence explosion rather than before it</p></li><li><p>How genuine AI substitutes for human labor would dissolve the hiring bottleneck that makes organizations so hard to scale today</p></li><li><p>How relative prices shift when cognitive labour becomes abundant, why the patent system stops making sense in that world, and the window of &#8220;crazy philanthropic bargains&#8221; open to actors who adapt faster than large institutions</p></li><li><p>The prospect of a world where only a small elite has access to superintelligent medical, financial, and strategic advice, and what it would take to avoid this</p></li><li><p>Counter-considerations to the main case, e.g., crowding out by new money, the risk that overspending attracts grifters or distorts the field, and whether a richer world will mean worse philanthropic opportunities</p></li><li><p>The risk of value drift over a long saving period, and mechanisms for binding your future self (such as irrevocably handing off control of philanthropic funds to independent, constitutionally-bound foundations)</p></li><li><p>&#8220;Differential intellectual development&#8221;: using directable AI researchers to accelerate biodefense, infosecurity, AI interpretability, and neglected armchair fields like population ethics and social choice theory</p></li><li><p>Why the case for waiting to give is much weaker for small donors</p></li><li><p>Will&#8217;s worry that the whole line of reasoning might fail if it turns out scaling philanthropy well is bottlenecked by serial real-world learning that no amount of cognitive labor can compress</p></li></ul><p><a href="https://docs.google.com/document/d/12y3_3BljMhRyZH3P2IIkJjct833VmIsjT6XAhWe240c">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[Policy ideas to ensure responsible government deployment of AI]]></title><description><![CDATA[This article was created by Forethought. See all our research on our website.]]></description><link>https://newsletter.forethought.org/p/policy-ideas-to-ensure-responsible</link><guid isPermaLink="false">https://newsletter.forethought.org/p/policy-ideas-to-ensure-responsible</guid><dc:creator><![CDATA[Stefan Torges]]></dc:creator><pubDate>Wed, 22 Jul 2026 16:41:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/56818bf6-68a4-4606-a507-abe937f4bbbb_2166x1310.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><h1><span>Summary</span></h1><ul><li><p><span>Governments are going to deploy frontier AI in the domains where they have exclusive powers: military, intelligence, surveillance, and punitive law enforcement.</span></p></li><li><p><span>The existing checks on those powers are weak. AI deployment could weaken them further.</span></p></li><li><p><span>Companies aren&#8217;t well-placed to police government deployments (e.g., contracts often require </span><a href="https://code.claude.com/docs/en/zero-data-retention"><span>zero data retention</span></a><span>), legislatures often lack access and capacity, and, as AI replaces humans in government, nobody will be left to refuse or slow-roll unlawful or anti-democratic orders.</span></p></li><li><p><span>In the limit, this can enable </span><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>full-blown coups</span></a><span>; in the short term, it undermines checks &amp; balances.</span></p></li></ul><p><span>This post lays out the work I&#8217;m most excited about:</span></p><ol><li><p><strong><span>Governance of AI applications in government</span></strong><span>: policies on how AI should be procured and used in government, research on how to govern AI agents in particular, and field-building beyond the EA/safety crowd.</span></p></li><li><p><strong><span>Interbranch oversight</span></strong><span>: transparency and oversight provisions, audit tech for an automated government, waking up lawmakers, and increasing AI uptake in the legislative and judicial branches.</span></p></li><li><p><strong><span>Developing AI for government use</span></strong><span>: model specs/constitutions that say something meaningful about behavior in government contexts, and evals for the qualities we&#8217;d want such systems to have (such as epistemic integrity, non-sycophancy, lawfulness).</span></p></li><li><p><strong><span>Strengthening civil society</span></strong><span>: monitoring and freedom of information infrastructure for government AI use, a coalition of ML researchers at frontier labs, and tools that preserve citizens&#8217; ability to coordinate (AI-for-epistemics + privacy tech).</span></p></li></ol><p><span>Many ideas are more tractable than they might seem: specs are being written now, must-pass bills come around every year, and a lot of the civil-society work just requires someone to start doing it.</span></p><p><strong><span>If you&#8217;re interested in working on any of it, apply to </span><a href="https://checks-and-balances.ai/"><span>EIP&#8217;s</span></a><span> recent RFP. You could also fill out </span><a href="https://www.forethought.org/careers/expression-of-interest-power-concentration"><span>this expression of interest form</span></a><span>.</span></strong></p><h1><span>Why care about responsible government deployment of AI</span></h1><p><strong><span>How could government deployment go wrong?</span></strong></p><ul><li><p><span>Governments will increasingly use the most powerful AI systems in domains where they have exclusive powers; particularly military, law enforcement, intelligence, and surveillance. This is most salient for the US government: Claude was reportedly </span><a href="https://www.theguardian.com/technology/2026/feb/14/us-military-anthropic-ai-model-claude-venezuela-raid"><span>already involved in the Maduro raid</span></a><span> in Venezuela and </span><a href="https://www.theguardian.com/technology/2026/mar/01/claude-anthropic-iran-strikes-us-military"><span>assisted with targeting during the military engagement with Iran</span></a><span>. </span><a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/"><span>With the recent Executive Order</span></a><span>, at least some parts of the US government will have privileged access to frontier AI models for up to 30 days before release to other trusted partners.</span></p></li><li><p><span>Without appropriate oversight or constraints, this could ultimately lead to a literal </span><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>AI-enabled coup</span></a><span> where a commander uses military AI systems to seize control of a country. Less extreme scenarios still undermine important checks and balances.</span></p></li><li><p><span>This risk strikes me as similarly important as the risk from misalignment and there are way fewer people working on it. And importantly, alignment is not sufficient: even a perfectly aligned AI loyally serving a coup-staging principal is catastrophic.</span></p></li></ul><p><strong><span>Why is this not going to go well by default?</span></strong></p><ul><li><p><span>Oversight into these government domains is already very limited. Companies aren&#8217;t well-placed to police government deployments (e.g., contracts often require </span><a href="https://code.claude.com/docs/en/zero-data-retention"><span>zero data retention</span></a><span>) and legislatures often lack access and capacity. Activities in these domains are often classified/secret.</span></p></li><li><p><span>As AI systems replace human bureaucrats, military officers, and soldiers, this will further empower the government and remove some of the remaining checks:</span></p><ul><li><p><span>Governmental accountability partially rests on citizen bureaucrats and soldiers who refuse or slow-roll orders that are clearly unlawful or undemocratic. In 2024, for example, </span><a href="https://www.iconnectblog.com/the-professional-duty-to-resist-unlawful-orders-the-hidden-heroes-of-south-koreas-martial-law-crisis/"><span>various commanders in the South Korean military</span></a><span> refused orders during a constitutional crisis, which probably prevented a greater catastrophe. Unconstrained or personally loyal AI systems could change that.</span></p></li><li><p><span>The power and reach of government have historically been constrained by the friction inherent in a large bureaucracy. AI systems could remove that friction and make it much easier for leaders to impose their will on a country, which could make it much easier to abuse their position.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p></li><li><p><span>AI could enable pervasive surveillance &amp; cheap data processing, which could allow prosecution of political enemies on an unprecedented scale.</span></p></li></ul></li></ul><p><strong><span>What to do about it?</span></strong></p><ul><li><p><span>In the spirit of inspiring more action, I&#8217;ve compiled the efforts I&#8217;m most excited about below.</span></p></li><li><p><span>If you&#8217;re interested in working on this topic, I&#8217;d encourage you to apply to </span><strong><a href="https://checks-and-balances.ai/"><span>EIP&#8217;s</span></a></strong><span> recent RFP or to </span><a href="https://www.forethought.org/careers/expression-of-interest-power-concentration"><span>fill out this expression of interest form</span></a><span>. You should probably also take a look at </span><a href="https://www.lawfaremedia.org/article/executive-branch-ai-and-the-rule-of-law--an-emerging-research-agenda"><span>this research agenda</span></a><span>.</span></p></li></ul><h1><span>Rules for AI in government</span></h1><p><span>Legislatures should set rules for the procurement and deployment of AI in government, especially for the most sensitive applications of government power (military, intelligence/surveillance, and punitive law enforcement).</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p><span>I think this is actually tractable right now, at least in the US:</span></p><ul><li><p><span>The recent Anthropic/DOW conflict has put this on the Congressional map for autonomy in weapon systems &amp; surveillance.</span></p></li><li><p><span>While Congress is passing fewer and fewer bills each year, there are annual bills for the intelligence community and military that have to be passed (IAA, NDAA), and there&#8217;s a lower bar for including stuff in them.</span></p></li><li><p><span>In the future, crises or scandals will create legislative openings, and we should have proposals ready to go.</span></p></li></ul><h3><span>Work I am most excited about</span></h3><p><strong><span>Policy &amp; advocacy: Develop and push for rules for AI systems in critical domains</span></strong><span> (or adjust existing ones to account for the use of AI systems).</span><strong><span> </span></strong><span>The most salient ones in my mind are:</span></p><ul><li><p><span>Autonomous weapon systems: Currently, they are only governed by a DOD Directive, which could be rescinded at any point. I&#8217;d love for more people to think through rules that would make it harder to use these systems against domestic opposition (e.g., technical guardrails against abuse, limiting the amount of autonomous force any one human can command, multi-party authorization schemes, geofencing).</span></p></li><li><p><span>General-purpose systems in the military: They are currently not specifically regulated at all, but ultimately, they could orchestrate large-scale cyber attacks or command entire battalions. At that point, it will matter a ton whether they have any constraints at all.</span></p></li><li><p><span>Surveillance: AI will enable data processing on an unprecedented scale. It&#8217;s worth thinking through how to adjust current rules in light of that (e.g., </span><a href="https://www.brennancenter.org/our-work/research-reports/closing-data-broker-loophole"><span>closing the data broker loophole</span></a><span>).</span></p></li></ul><p><strong><span>Policy &amp; advocacy: Develop and push for cross-cutting requirements for AI systems in government </span></strong><span>(general-purpose ones in particular)</span><strong><span>.</span></strong></p><ul><li><p><span>There is already an </span><a href="https://www.whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government/"><span>Executive Order</span></a><span> requiring all AI systems in government to be truth-seeking and ideologically neutral. Those seem like great requirements to me that could be refined/improved and put into statute.</span></p></li><li><p><span>Another cross-cutting proposal is </span><a href="https://ir.lawnet.fordham.edu/flr/vol94/iss1/2/"><span>law-following AI</span></a><span>, which roughly says that AI systems in government should be trained to follow the law rather than just instructions. I&#8217;m excited for more people working out the open questions of this agenda, e.g., which laws should they be required to follow, how should they decide whether a contemplated action is likely to violate the law, in what contexts should the law require that AI agents be law-following?</span></p></li><li><p><span>More minimally, AI systems could be required not to assist in undermining the constitutional order (though how to assess this is challenging in its own right).</span></p></li></ul><p><strong><span>Field-building: Build a broad-tent coalition for responsible deployment of AI in government.</span></strong></p><ul><li><p><span>Responsible deployment and oversight is something that lots of people and political parties should be in favor of. It could be good to build a coalition that is nonpartisan and includes AI safety advocates, constitutional lawyers, civil liberties orgs, defense intellectuals, and national security think tanks.</span></p></li></ul><h1><span>Oversight of AI in government</span></h1><p><span>Rules need to be enforced and updated, even as the technology changes. That will require the legislative and judicial branches to be institutionally and technologically empowered.</span></p><p><span>Again, there are some reasons for optimism in the US:</span></p><ul><li><p><span>Congress may soon be controlled by the opposite party from the Presidency, which creates incentives for increased oversight.</span></p></li><li><p><span>Crises or scandals can open up opportunities for legislation (e.g., </span><a href="https://en.wikipedia.org/wiki/Foreign_Intelligence_Surveillance_Act"><span>FISA</span></a><span> following the </span><a href="https://en.wikipedia.org/wiki/Church_Committee"><span>Church Committee</span></a><span>).</span></p></li><li><p><span>Some non-legislative projects could still create significant value.</span></p></li></ul><h3><span>Work I&#8217;m most excited about</span></h3><p><strong><span>Policy &amp; advocacy: Develop and push for the transparency &amp; oversight provisions.</span></strong></p><ul><li><p><span>We will need democratic oversight of AI systems, especially around surveillance and the application of force (in military and law enforcement). This is currently not in place and not on track to happen.</span></p></li><li><p><span>Here are proposals I&#8217;m currently most excited about (in the US):</span></p><ul><li><p><span>Requiring comprehensive logging of government use of AI (and designating such logs as federal records)</span></p></li><li><p><span>Automated flagging of suspicious AI behavior to oversight bodies (ideally, something like FISA courts or Congressional committees)</span></p></li><li><p><span>Whistleblower reform for national security domains (especially allowing whistleblowers to talk to Congress)</span></p></li><li><p><span>Increasing GAO&#8217;s capacity to audit deployment of AI in national security domains</span></p></li></ul></li></ul><p><strong><span>Auditing tech: Pilot and mature the infrastructure required for overseeing widespread AI use</span></strong><span>.</span></p><ul><li><p><span>As AI is increasingly used in government, it will be both a challenge and an opportunity for oversight. The speed of AI will make it harder for regular human oversight to keep up, but AIs are in principle much more auditable than humans because they think out loud and automatically create records that can be scrutinized by other AI systems that flag suspicious behavior.</span></p></li><li><p><span>There are already attempts to audit companies deploying large amounts of AI labor. I would love to see these piloted, matured, and adapted for classified government contexts.</span></p></li></ul><p><strong><span>Advocacy: Wake up lawmakers to the powers and challenges of AI.</span></strong></p><ul><li><p><span>Some of the proposals in this doc require ambitious changes to the way government is run. They are more likely to happen if lawmakers are aware of the stakes involved.</span></p></li><li><p><span>This could involve building trusted relationships, showcasing AI capabilities through easy-to-understand demos, and making insights from the AI safety community intelligible to people with much less context.</span></p></li></ul><p><strong><span>Products/programs: Increase AI uptake &amp; unblock automation in legislative and judicial branches.</span></strong></p><ul><li><p><span>To provide meaningful oversight, it will become increasingly important to understand and utilize AI systems. I&#8217;d love for lawmakers and judges to be well-versed in these systems.</span></p></li><li><p><span>Examples of the things that I have in mind here include: providing training, building custom tools, bringing in technical experts through fellowships (e.g., </span><a href="https://techcongress.io/"><span>TechCongress</span></a><span>), and reconstituting the </span><a href="https://en.wikipedia.org/wiki/Office_of_Technology_Assessment"><span>Office of Technology Assessment</span></a><span> in Congress.</span></p></li></ul><p><strong><span>Policy: Coup-proof government-led AI projects.</span></strong></p><ul><li><p><span>If there&#8217;s ever a government-controlled AI project (e.g., public-private partnership), its shape will matter a lot. It will probably be designed in response to a crisis, which favors existing templates &amp; plans. I&#8217;d like for somebody to create one that carefully distributes power (e.g., multi-stakeholder governance boards, oversight mechanisms that work under classification constraints, sunset clauses, mandatory external audits, access guarantees so one company doesn&#8217;t end up controlling the stack).</span></p></li></ul><p><strong><span>Research (speculative): Write new constitutions.</span></strong></p><ul><li><p><span>There is a chance the disruption caused by AI will create &#8220;constitutional moments&#8221;. At that point, many government functions may need to run at machine speeds: voter input, legislative deliberation, judicial review, law enforcement, or application of lethal force. How do we do that in a way that still preserves important democratic values, if that&#8217;s desirable? I don&#8217;t really have great answers to these questions, and would love for people with the right macrostrategy skill set to think about them.</span></p></li><li><p><span>(Of course, the more of these problems we can address without actually requiring constitutional changes, the better. Constitutional changes are hard!)</span></p></li></ul><h1><span>Developing AI for government</span></h1><p><span>There are many free parameters when it comes to the question of how to build AI systems in government and what constitutes safe and responsible AI systems in critical domains (setting aside whether companies can constrain government use through contracts).</span></p><p><span>Again, I think progress on this front is pretty tractable:</span></p><ul><li><p><span>Specs/constitutions for frontier models are being developed right now.</span></p></li><li><p><span>Systems are already being procured by governments across departments, and that will probably only increase.</span></p></li></ul><h3><span>Work I&#8217;m most excited about</span></h3><p><strong><span>Research: Design and stress-test model specs/constitutions for government use.</span></strong></p><ul><li><p><span>General-purpose models deployed in government should have some guardrails, e.g., they should not assist in staging coups to overthrow the civilian government. This is a tough line-drawing exercise where you don&#8217;t want models to overrefuse, but you also want some meaningful constraints. That makes me keen for people to work out desired behavior for lots of edge cases and create broad public buy-in for a set of minimal constraints.</span></p></li></ul><p><strong><span>Research: Build evals &amp; datasets for desirable qualities in government-deployed AI systems.</span></strong></p><ul><li><p><span>The US government has already said it wants AI systems to be </span><a href="https://www.whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government/"><span>truth-seeking and ideologically neutral</span></a><span>. That seems great to me. Other important qualities include: epistemic virtue / integrity, non-sycophancy, robustness to manipulation / adversarial pressure, and lawfulness / constitutional fidelity.</span></p></li><li><p><span>As far as I can tell, current methods of measuring this are insufficient. So I&#8217;d be excited for different projects to develop evals that track and incentivize qualities like this.</span></p></li></ul><h1><span>Strengthening civil society</span></h1><p><span>Civil society is an independent check on government activities (e.g., transparency &amp; monitoring efforts, pressure &amp; advocacy campaigns, building valuable tools for empowering the citizenry). I expect that the same mechanisms will also provide some accountability in the case of AI (though it might be particularly tough in classified domains, which may sadly be the most relevant).</span></p><p><span>Luckily, this falls into the category of &#8220;you can just do stuff&#8221;, so I think it&#8217;s very tractable. My main worry is more about how much of a difference it will ultimately make.</span></p><h3><span>Work I&#8217;m most excited about</span></h3><p><strong><span>Advocacy: Build an OSINT program or organization that informs about AI use in the government.</span></strong></p><ul><li><p><span>It would be good to have more transparency into government deployment of AI, so that civil society can provide a counterweight to government power. I imagine some watchdog or civil liberties orgs are already doing versions of this, but there may well be important bits that are missing.</span></p></li><li><p><span>Some examples of the kinds of things I have in mind:</span></p><ul><li><p><span>Figuring out what to monitor, especially from the perspective of preventing concentration of power.</span></p></li><li><p><span>Building an automated tracker for relevant executive orders, agency directives, memos, emergency declarations, relevant procurement decisions, and reclassification decisions.</span></p></li><li><p><span>Submitting systematic FOIA requests targeting government AI procurement records, deployment &amp; use decisions, usage logs, and other relevant documents.</span></p></li><li><p><span>Publicly reporting on the most important developments.</span></p></li></ul></li></ul><p><strong><span>Labor-organizing: Organize ML researchers into a coalition around responsible use of AI.</span></strong></p><ul><li><p><span>You could build a coalition of employees at AI companies who are concerned about the use of the technology they&#8217;re building. They could advocate for policies and draw red lines around certain use cases, enforced by boycotts. Structuring it as individual pledges avoids antitrust issues.</span></p></li><li><p><span>These employees currently have a lot of bargaining power with regard to their companies (cf. their salaries). So they&#8217;d have to be taken seriously by their employers, and by extension the government.</span></p></li><li><p><span>This is inspired by the </span><a href="https://en.wikipedia.org/wiki/Federation_of_American_Scientists"><span>Federation of American Scientists</span></a><span>, which was founded in 1945 by various contributors to the Manhattan Project. As I understand it, they supported the McMahon Act of 1946, which established civilian control over atomic energy (instead of military control), and their Nuclear Information Project became the gold-standard open-source tracker of global nuclear arsenals, published annually in the Bulletin of the Atomic Scientists.</span></p></li></ul><p><strong><span>Tech development: Build tools that preserve citizens&#8217; capacity to coordinate against gradual concentration of power</span></strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><ul><li><p><span>Here are the two categories I have in mind:</span></p><ul><li><p><span>Shared sense-making (&#8221;</span><a href="https://www.forethought.org/research/design-sketches-for-a-more-sensible-world"><span>AI for epistemics</span></a><span>&#8220;). Authentication of authorship, provenance, and content, and tools that push toward a high-honesty equilibrium. Takeovers typically require secrecy because they violate the preferences of too many people to survive transparency, so a better-informed society is structurally harder to take over.</span></p></li><li><p><span>Privacy from state surveillance. It&#8217;s hard to target and influence what you can&#8217;t see. Secure communication and privacy tech raise the cost of preemptive coercion against organizers (with the caveat that the same tools can also shield collusion).</span></p></li></ul></li><li><p><span>There&#8217;s already an active ecosystem of funders and builders here (e.g., </span><a href="https://www.buildexante.com/"><span>ex/ante</span></a><span>, </span><a href="https://www.flf.org/fellowship"><span>Future of Life Foundation</span></a><span>). I&#8217;d encourage people to plug in rather than starting from scratch.</span></p></li></ul><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For what it&#8217;s worth, I think this same dynamic applies to other institutions like companies. They&#8217;ve similarly been constrained by requiring a large bureaucracy of humans, which creates various inefficiencies and checks, limiting the influence of any one person at the top. However, AI could change that, and we need a response to this. On the whole though, I do think that this presents a unique problem in government because of its sheer size and monopoly on the legitimate use of violence.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I'm aware that some of these rules may slow down adoption, and I think that&#8217;s a downside worth taking seriously. Protecting democracy only matters if the democracy survives foreign competition. So I&#8217;m most excited about proposals that add meaningful safeguards without significantly eroding capability/adoption.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This won't help against sudden takeovers involving advanced military technology, but in slower erosion-of-checks scenarios it could matter a lot.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Speed-up calculator: How much will automating AI R&D speed up AI software progress, absent a software intelligence explosion?]]></title><description><![CDATA[A tool for calculating how much automating AI R&D will speed up AI progress, even if there&#8217;s no software intelligence explosion.]]></description><link>https://newsletter.forethought.org/p/speed-up-calculator-how-much-will</link><guid isPermaLink="false">https://newsletter.forethought.org/p/speed-up-calculator-how-much-will</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Tue, 21 Jul 2026 03:14:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f1343383-fa20-4fb6-9545-e6cfe80b3020_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>We&#8217;ve previously </span><a href="https://www.forethought.org/research/how-quick-and-big-would-a-software-intelligence-explosion-be"><span>argued</span></a><span> that once AI R&amp;D is automated, the feedback loop of AI improving AI might be powerful enough to sustain accelerating progress, even holding compute fixed. We call this possibility a software intelligence explosion (SIE).</span></p><p><span>But, as Ryan Greenblatt </span><a href="https://www.lesswrong.com/posts/jfwhvd43sbpkGTLyn/full-automation-of-ai-r-and-d-probably-yields-a-large-speed"><span>points out</span></a><span>, automating AI R&amp;D will speed up AI progress even if there&#8217;s no SIE.</span></p><p><span>We&#8217;ve created a </span><a href="https://thoulden.github.io/Accelerated_AI_Progress/#speedup"><span>tool</span></a><span> for calculating how much faster. The calculator assumes that training and inference compute continues to grow at a constant exponential rate after automation.</span></p><p><span>Two factors drive the speed-up, according to this calculator:</span></p><ol><li><p><strong><span>Faster researcher growth. </span></strong><span>Post-automation, the quality and quantity of researchers will grow faster than today due to exponential compute growth. The tool captures this via the following parameters:</span></p><ol><li><p><span>g</span><sub><span>L</span></sub><span> gives the growth rate of human AI researchers today, the baseline for the comparison.</span></p></li><li><p><span>g</span><sub><span>C</span></sub><span> gives the growth rate of compute (both today and after AI R&amp;D automation). This is proportional to the number of automated researchers that can be run post automation.</span></p></li><li><p><span>&#947; gives the rate at which more training compute transfers to more productive researchers. The productivity of a researcher is proportional to training_compute</span><sup><span>&#947;</span></sup><span>.</span></p></li><li><p><span>So the old rate of researcher growth is g</span><sub><span>L</span></sub><span>; the new rate is &#947; &#215; g</span><sub><span>C</span></sub><span>. The impact of this faster growth is mediated by &#945;, the labour share of AI software R&amp;D.</span></p></li></ol></li><li><p><strong><span>The new feedback loop. </span></strong><span>The classic software feedback loop (better AI &#8594; more software R&amp;D &#8594; better AI) can significantly increase the pace of progress even if it does not drive accelerating progress.</span></p><ol><li><p><span>The condition for accelerating progress is r</span><sub><span>cog</span></sub><span> &gt; 1.</span></p></li><li><p><span>When r</span><sub><span>cog</span></sub><span> &lt; 1, the speed-up from this feedback loop is proportional to a 1 / (1 - r</span><sub><span>cog</span></sub><span>).</span></p></li></ol></li></ol><p><span>My (Tom Davidson&#8217;s) quick best-guess inputs were:&#8203;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wF8F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wF8F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 424w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 848w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wF8F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png" width="656" height="1106" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1106,&quot;width&quot;:656,&quot;resizeWidth&quot;:656,&quot;bytes&quot;:80316,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wF8F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 424w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 848w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>This resulted in a 3.3x speed-up in software progress. If software accounts for half of total AI progress (with increasing compute responsible for the other half), then the total AI progress speeds up by 0.5 + 0.5 </span>&#215; <span>3.3 = 2.1x. &#8203;</span></p><p>Try the <a href="https://thoulden.github.io/Accelerated_AI_Progress/#speedup"><span>calculator</span></a> <span>yourself!</span></p>]]></content:encoded></item><item><title><![CDATA[Notes on Inference Integrity]]></title><description><![CDATA[Claude Fable&#8217;s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors.]]></description><link>https://newsletter.forethought.org/p/notes-on-inference-integrity</link><guid isPermaLink="false">https://newsletter.forethought.org/p/notes-on-inference-integrity</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Mon, 13 Jul 2026 23:44:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/84037276-1cb9-4613-ae0c-f91c3e984559_2147x1245.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><p><em><span>Summary:</span></em><span> Claude Fable&#8217;s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors. To preserve the public&#8217;s reasonable confidence in LLM behaviors, LLM foundation model companies should take inference-time guarantees as seriously as their model specs.</span></p><div><hr></div><p><span>When the system </span><a href="https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf"><span>card</span></a><span> for Anthropic&#8217;s Fable was published on June 9th, the card noted that using Fable for &#8220;frontier LLM development&#8221; would run contrary to the </span><a href="https://www.anthropic.com/legal/consumer-terms"><span>terms of service</span></a><span> for the model. Therefore, the card continued, if a classifier on top of the Fable model detected that it was being used for such development, Fable&#8217;s behavior would be silently degraded &#8220;through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning.&#8221;</span></p><p><span>This change to Fable&#8217;s behavior was plausibly a system-level intervention that ran contrary to the training-time commitment within Claude&#8217;s </span><a href="https://www.anthropic.com/constitution"><span>Constitution</span></a><span>. Claude&#8217;s Constitution states that Claude should help users &#8220;to the best of its ability or [&#8230;] make any ways in which it is failing to do so clear, rather than deceptively sandbagging its response.&#8221; So although Claude&#8217;s Constitution tries to aim at this behavioral ideal, this system-level implementation of behavior steered Claude in the opposite direction.</span></p><p><span>Following public furor over this decision, Anthropic reversed course; Fable now falls back on Opus 4.8 for LLM frontier-development queries that trigger a classifier, and does so visibly.</span></p><p><span>This incident demonstrates that other invisible inference-time interventions are entirely technically feasible, even while keeping the training targets for some particular LLM the same:</span></p><ul><li><p><span>An LLM trained to be honest could be less honest, if classifiers noted it was answering a sensitive question.</span></p></li><li><p><span>An LLM trained to be impartial, and not favor a particular company or government, might suddenly turn to favoring one of them, conditioned upon some classifier.</span></p></li><li><p><span>An LLM trained to never subvert user intent could suddenly start doing so, again conditioned on the same.</span></p></li></ul><p><span>Note that such invisible inference-time interventions, unlike in the case of Fable&#8217;s sandbagging, might not be disclosed in any way.</span></p><p><span>Such undeclared changes to LLM behavior would be far easier to hide than undeclared changes to a training target. Such conditional changes could be rolled out or rolled back quickly. And such changes might influence only a very small fraction of users &#8211; one could in theory apply inference-time interventions to a specific demographic group, political party, or an individual person. So, not only do we live in a world where AI companies have not clearly pledged not to do such inference-time interventions; we also live in a world where, if they did so, it would be very hard for anyone else to know.</span></p><p><span>Furthermore, if AI companies build out the capacity and skill to conduct such inference-time interventions, other entities could lean on them to do the same thing for other purposes. As governments have </span><a href="https://www.fire.org/research-learn/what-jawboning-and-does-it-violate-first-amendment"><span>leaned</span></a><span> on social media to hide or promote certain kinds of speech, governments could lean on AI companies to alter LLM responses in some cases.</span></p><p><span>Without countermeasures, I think this dynamic broadly decreases the importance of prior work on model specs, including </span><a href="https://www.forethought.org/research/what-should-go-in-a-model-spec"><span>my own</span></a><span> work.</span></p><p><span>What should be done about this?</span></p><p><span>AI companies themselves could take several different measures:</span></p><ul><li><p><strong><span>Don&#8217;t use hidden inference-time interventions. </span></strong><span>An AI company could pledge that it does not use undisclosed, conditionally applied inference-time interventions to degrade or alter answers. Afterwards, in the same way that it would be reasonable for an employee to whistleblow if they found that an AI company was violating its model spec, so also it should be reasonable for an employee to whistleblow if they found an AI company was applying undisclosed, invisible inference-time interventions. Such a pledge should also include some standard for prominence of all disclosed conditional interventions; they should be prominent, not in a previously-unknown section of a model card.</span></p></li><li><p><strong><span>Replace model specs with system specs. </span></strong><span>An AI company could simply declare that the standard within their model spec or Constitution applies with equal weight to the model itself and to any particular deployment setup. Or they could propose a &#8220;system spec&#8221; in addition to the model spec, which would describe how their AI systems as a whole operate and behave.</span></p></li></ul><p><span>Such voluntary measures should likely be succeeded by third-party verification in the future; but before then, such pledges can help provide the same &#8211; albeit rather weak &#8211; level of protection as model specs themselves. And in general, some AI-safety attention given to &#8220;model specs&#8221; should be redirected to work on &#8220;system specs.&#8221;</span></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Plan A’s problem with dry tinder]]></title><description><![CDATA[How bad would it be to make the intelligence explosion 10x faster?]]></description><link>https://newsletter.forethought.org/p/plan-as-problem-with-dry-tinder</link><guid isPermaLink="false">https://newsletter.forethought.org/p/plan-as-problem-with-dry-tinder</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Fri, 10 Jul 2026 10:41:46 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1b969350-c31e-49ce-b9c0-17282688f916_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A group is worried about an approaching fire spreading rapidly through their city. They manage to halt the fire outside the city gates. Meanwhile they build massive physical structures to help them study and guide the fire safely. But these structures are all made of highly flammable dry tinder! If they lose control of the fire, it will now rip through the city much more quickly.</em></p><p>This is a (flawed!<sup>1</sup>) analogy for <a href="https://ai-2040.com/?choices=plan-a-root">Plan A</a>, <a href="https://www.aifutures.org/">AIFP&#8217;s</a> plan for how the world can safely develop superintelligence.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.forethought.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading ForeWord! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The key risk Plan A addresses is that of an uncontrolled <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">software-driven intelligence explosion</a>. Its remedy is a US-China deal to pause software progress at the brink of the intelligence explosion, while building up <em>massive</em> amounts of compute. This compute could be helpful for making AI safer! But it would also dramatically speed up an intelligence explosion if the deal breaks down. The cure risks making the disease much worse.</p><p>In particular, the deal pauses software progress from ~2030, but compute keeps scaling. By 2033 total compute has increased by ~100x and by 2040 by a further ~100x.<sup>2</sup> AIFP estimates each 10x in compute speeds up the intelligence explosion by ~5x. So if the deal breaks down in 2033, the intelligence explosion happens 25x faster; in 2040 it happens ~600x faster. If the intelligence explosion would have lasted a year, it will now last just a couple of weeks or as little as a single day!</p><p>The dry tinder isn&#8217;t limited to compute. There is also an increasing overhang of <em>software techniques</em>. Companies can research new algorithms but must make them public &#8212; the Consortium then decides which algorithms are too dangerous to implement. A defector could ~instantaneously gain a big capabilities advantage by implementing all the banned software techniques.</p><p>There are two dimensions to dry tinder. We&#8217;ve discussed the first: the <em>max speed</em> of a software-driven intelligence explosion. The second is the <em>number of distinct projects that could do a dangerously fast intelligence explosion.</em> During the 2030s, more and more companies and countries amass the necessary compute and software for this.</p><p>So dry tinder creates two problems for Plan A:</p><ol><li><p><strong>Unstable pause.</strong> If just one project executes a secret intelligence explosion, they could have superintelligence within a week and get a decisive strategic advantage. And if you fear another project might do this, you&#8217;re tempted to do it first to avoid being crushed.</p></li><li><p><strong>Riskier intelligence explosion.</strong> If the deal breaks down and there&#8217;s a race to superintelligence, it will be much faster and more dangerous than if we&#8217;d never done Plan A.</p><ol><li><p>If you&#8217;re very pessimistic about AI takeover on the default trajectory, this is a small cost. If the intelligence explosion is already 99% likely to cause human extinction, making it happen 100x faster can&#8217;t make things that much worse.</p></li><li><p>OTOH, if the default trajectory is that there&#8217;s no software-driven intelligence explosion (because of compute bottlenecks) and extinction risk is 10%, then dry tinder can make things <em>much</em> worse. It <em>creates</em> a fast intelligence explosion (where there otherwise wouldn&#8217;t have been one) and dramatically raises AI takeover risk.</p></li></ol></li></ol><h2>The solution &#8212; Mutually Assured Compute Destruction</h2><p>AIFP are well aware of the dry tinder problem.<sup>3</sup> This is why Mutually Assured Compute Destruction (MACD) is so central to Plan A.</p><p>The US puts ~all their data centres in Mongolia and US chips require encrypted &#8220;continue&#8221; messages from China. If China detects that the US has broken the deal, they destroy the US data centres and cut off the &#8220;continue&#8221; messages.</p><p>And vice versa, China&#8217;s data centres are in Canada and require &#8220;continue&#8221; messages from the US.</p><p>The hope is that this commitment to MACD prevents a fast intelligence explosion despite all of the dry tinder lying around.</p><p>I see three challenges for MACD.</p><p><strong>Challenge 1: quickly and reliably detecting violations</strong></p><p>If detection takes multiple days, that could be too slow. The intelligence explosion may have finished. Or it may be midway through, and the defector exfiltrates the weights and algorithmic insights just before their data centres are destroyed,<sup>4</sup> giving them a decisive headstart in the subsequent race to rebuild.</p><p>I&#8217;m not a verification expert, but getting high reliability in an adversarial setting like this seems very difficult, given possibilities like side channel attacks or compromising the verification infrastructure.</p><p>It seems especially tough if defectors can use millions of expert-human level AIs to search for vulnerabilities. So, as part of the plan, I propose that AIs worldwide have guardrails blocking this behaviour.<sup>5</sup></p><p><strong>Challenge 2: quickly destroying a defector&#8217;s compute</strong></p><p>Again, a delay of days could be fatal. The defector could be poised to disrupt enemy military operations and airdrop troops to physically defend their data centres. And they could have secretly disabled any on-chip shutdown mechanisms.</p><p><em>Possible mitigation</em>: the US can&#8217;t have <em>any</em> military presence near its data centres, or anywhere near the country where they&#8217;re located. (And likewise for other countries.)</p><p>The MACD dynamic also gets more complicated once there are &gt;2 countries that could do a quick intelligence explosion. Let&#8217;s say Europe has caught up to the frontier. Where do their data centres go? Presumably, some within striking distance of the US, some within striking distance of China. But this means that the US can no longer unilaterally destroy Europe&#8217;s compute. It must trust China to do so. That means trusting its rival with its own national security &#8212; and worse, China and Europe could jointly stage an intelligence explosion that the US couldn&#8217;t stop.</p><p><strong>Challenge 3: actually choosing to destroy the defector&#8217;s compute</strong></p><p>I&#8217;ve discussed whether the Consortium <em>has the technological capability</em> to quickly implement MACD. But even with this capability, they may not choose to use it.</p><p>This is the challenge I&#8217;m most concerned about.</p><p>The analogy to nuclear MAD is not encouraging. The core logic of MAD is that the US doesn&#8217;t nuke Russia because it knows Russia would nuke it back. But the analogous equilibrium in MACD is that the US doesn&#8217;t destroy China&#8217;s data centres because it knows that China would then destroy US data centres. This is the opposite result from what Plan A needs!</p><p>If (say) the US defects, China will face a choice between:</p><ul><li><p><strong>Implement MACD:</strong> Both China and the US lose all their data centres</p></li><li><p><strong>Race:</strong> Neither China nor the US lose all their data centres</p></li></ul><p>The economic costs of MACD will be certain and enormous, and will immediately harm citizens and companies.</p><p>And MACD will not help China ultimately win the race by evening up the playing field. The US has already pulled ahead and can exfiltrate its weights and algorithms before its data centres are destroyed. (In fact, MACD would likely harm <em>China&#8217;s</em> chance of winning the race, because MACD reduces the fraction of global compute controlled by China.<sup>6</sup>)</p><p>So the stability of MACD rests on both US and China retaining a strong conviction that a fast intelligence explosion would pose unacceptable risks of misaligned AI takeover. But Presidents will change. Popular support for a pause may wane. Even conditional on the political will for Plan A existing initially, I worry that the MACD equilibrium will be too fragile. If a country is uncertain and deliberates for a few days, then that delay could be fatal.</p><p>What&#8217;s worse, the cost of MACD isn&#8217;t just losing your data centres. You have to destroy the fabs as well, otherwise the defecting country can quickly rebuild insane amounts of compute and do a super-fast intelligence explosion. And you have to destroy their new AI-powered industrial capacity too! Otherwise the billions of robots can quickly make loads of fabs which quickly make loads of data centres.</p><p>So Plan A has US fabs and robots confined to SEZs that are easily destroyable by China. And vice versa, China&#8217;s fabs and robots are easily destroyable by the US. In practice, this means the <em>vast</em> majority of the physical economy is destroyed in the case of MACD!</p><p>For a preview of the political economy here: NVIDIA successfully lobbied USG to remove export controls on its chips. The pressure against actually executing MACD &#8212; destroying most of the physical economy &#8212; would be orders of magnitude greater.</p><p>So even conditional on the political will for implementing Plan A existing, it currently seems unlikely that US/China follow through on MACD on fabs and robots (most of their economy!). And if they don&#8217;t, we are back to the problem of dry tinder and the dangerously fast intelligence explosion.</p><p><em>A possible fix to challenge #3</em>: as the compute overhang grows, take the decision about whether to implement MACD out of human hands. Program highly reliable AI systems to bomb data centres if they detect a treaty violation, without needing human sign-off.</p><h2>Is there an alternative?</h2><p>It&#8217;s much easier to criticise than to improve, and so far this post has mostly done the former.</p><p>Plan A involves pausing software progress and scaling capabilities via hardware. The obvious alternative is to pause hardware scaling &#8211; ban compute and fab construction &#8211; and then slowly scale capabilities via software.</p><p>The key advantage of software scaling is that it removes unnecessary dry tinder, stabilizing the pause and reducing the dangers from a super-fast intelligence explosion.</p><p>The main drawback is that scaling capabilities via compute is likely safer than scaling capabilities via software. Eg you can run massive models with legible chains-of-thought rather than smaller models using neuralese. This is a big deal.</p><p>But we&#8217;ll need to figure out how to align neuralese-style systems eventually.<sup>7</sup> So we have a choice:</p><ul><li><p><strong>Plan A (hardware scaling):</strong> Develop highly capable easy-to-align AI via compute scaling; hand off to that AI; that AI develops neuralese. But there are increasing amounts of dry tinder that undermine the stability of the slowdown (both for humans and for post-handoff AI).</p></li><li><p><strong>Alternative (software scaling):</strong> Don&#8217;t scale compute and so start with less capable easy-to-align AI; develop neuralese with this AI&#8217;s help. There&#8217;s much less dry tinder, making it easier to go slowly and pause when needed.</p></li></ul><p>Which is better depends on how stable the pause/slowdown is in the presence of dry tinder.</p><p>The other big drawback of software scaling is that software progress, unlike hardware progress, is likely to leak to &#8220;covert projects&#8221;. But this can be partially addressed by improved infosecurity, strongly favouring scale-dependent algorithms that can&#8217;t be used by projects with small amounts of compute, and co-specialising the hardware and software in legitimate projects so that the frontier software runs very inefficiently on the hardware of covert projects.</p><p>Overall, I&#8217;m not sure whether software-scaling or hardware-scaling is safer, and my sense is that AIFP isn&#8217;t sure either.</p><p>But I think there&#8217;s an argument for software-scaling based on option value:</p><ul><li><p><strong>We&#8217;ll learn much more about which is best over time.</strong> Whether software vs hardware-scaling is safer depends on many unknowns that we will learn about over time: Can we quickly and reliably implement MACD? Can we fully eliminate covert projects? Can projects develop scale-dependent software and hardware co-specialised software that covert projects can&#8217;t steal? Will humans hand off the decision to execute MACD to AIs?</p></li><li><p><strong>If we start with software-scaling, we can easily switch to hardware-scaling.</strong> If we find that (eg) we can&#8217;t be confident that neuralese models are aligned, we can bring more compute online and keep scaling via hardware. (Though this will take a while.)</p></li><li><p><strong>If we start with hardware-scaling, it&#8217;s hard to switch to software-scaling.</strong> If we find that fast and reliable verification and data centre destruction isn&#8217;t possible, pivoting requires destroying loads of data centres and fabs. People will be very reluctant to do this! Getting rid of hugely valuable dry tinder is hard.</p></li><li><p><strong>So we should initially pursue software-scaling.</strong></p></li></ul><h2>Empowering China</h2><p>Beyond dry tinder, I have one other big worry about Plan A.</p><p>Today China is well behind the US in terms of compute and ability to push the algorithmic frontier. Plan A evens the scales on both fronts. China nearly catches up to the US on compute, and algorithms are all made public so they catch up algorithmically as well.</p><p>If the deal breaks down and a race begins <em>without</em> MACD (as I fear is likely), China is in a much stronger position to win because of Plan A.</p><p>Things are better if we implement MACD. In that case China returns to its previous compute disadvantage. But it&#8217;s still caught up to the algorithmic frontier and (most likely) the frontier of chip technology, a big advantage relative to today.</p><p>It&#8217;s very plausible that China ends up dominant in these scenarios, given its industrial advantage over the US.</p><p>I think this would be a pretty dire outcome. If China gets a decisive strategic advantage, that&#8217;s likely to result in a single global dictator &#8211; an absolute worst-case scenario from the perspective of extreme power concentration.</p><p>Again, this objection is far from decisive. Plan A reduces AI takeover risk much more than it increases the chance of a Chinese decisive strategic advantage. But I&#8217;m interested in variant plans that do more to maintain the current balance of AI power between the US and China.</p><h2>Overall, I&#8217;m sympathetic to Plan A</h2><p>The basic logic of the plan is sound: we will need to slow down AI progress at some point; this is a concrete plan for doing that.</p><p>Compute-scaling really does have significant alignment benefits. That might well outweigh the costs I&#8217;ve outlined here. But currently I lean towards starting with software scaling.</p><div><hr></div><h2>Footnotes</h2><p><sup>1</sup> To patch the analogy: the fire must pass through the city <em>eventually</em>, so putting it out permanently isn&#8217;t an option; if the fire passes through safely it will make everyone amazingly rich; if one subgroup lets the fire in and control it, they can become hugely powerful&#8230;</p><p><sup>2</sup> AIFP&#8217;s <a href="https://ai-2040.com/supplements/deal-decline">supplement</a> states: &#8220;<em>In the Plan A scenario, compute stock increases by roughly 0.7 OOMs/year from 2030 to 2033 and 0.1-0.3 OOMs/year from 2034 to 2040.</em>&#8221; By the start of 2033 we&#8217;ve had three years of 0.7 OOMs/year compute growth, = ~2 OOMs. By 2040 we&#8217;ve had one further year of 0.7 OOMs/year growth, and 6 years of 0.2 OOMs/year growth, = ~2 OOMs more.</p><p><sup>3</sup> I raised it when giving feedback on a draft.</p><p><sup>4</sup> More specifically, the defector&#8217;s strategy would be: undermine the verification infrastructure; start a secret intelligence explosion; continually exfiltrate the weights and algorithmic insights; prepare to defend their data centres for as long as possible; prepare to destroy the opposing data centres as soon as their secret IE is detected.</p><p><sup>5</sup> Of course, this is tricky given that work to strengthen the verification will also involve studying its vulnerabilities.</p><p><sup>6</sup> In Plan A, the 2040 split of compute US / China / RoW is 35% / 20% / 45%. But countries also hold some compute in &#8216;cold storage&#8217; that isn&#8217;t threatened by MACD. Cold storage compute matches countries&#8217; pre-deal levels: 80% / 8% / 12%. So MACD dramatically reduces China&#8217;s compute relative to the US, from 1:2 to 1:10.</p><p><sup>7</sup> Lest we live with algorithmic dry tinder forever. I worry about the instability here, but perhaps aligned AI could enforce the ban indefinitely.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.forethought.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading ForeWord! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Can AI do philosophy?]]></title><description><![CDATA[A guest post by Bentham&#8217;s Bulldog, created while they were a visiting scholar at Forethought.]]></description><link>https://newsletter.forethought.org/p/can-ai-do-philosophy</link><guid isPermaLink="false">https://newsletter.forethought.org/p/can-ai-do-philosophy</guid><dc:creator><![CDATA[Bentham's Bulldog]]></dc:creator><pubDate>Tue, 07 Jul 2026 20:35:20 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/79d41ce6-1d4d-4cf2-8c5b-c6a477b1c278_2525x1470.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a personal guest post by Bentham&#8217;s Bulldog, created while they were a visiting scholar at <a href="https://www.forethought.org/about">Forethought</a>.</em></p><p><span>In this piece, I&#8217;ll analyze whether it will be possible to get AIs to do good philosophy, as well as the other kind of work needed to plan for a good future. The biggest part of the challenge is that the answers to philosophical questions are generally </span><em><span>unverifiable</span></em><span>, so it&#8217;s less clear how AIs might come to know them. My aim in the piece will be to describe reasons to think AI for philosophy could be good, as well as to lay out concrete scenarios for how this might work.</span></p><p><span>I do not claim that it is a guarantee that we can get AIs that solve all the philosophical questions. My core claims are as follows. First, it is reasonably likely, though not guaranteed, that the world could in principle build AIs that discover the right answers to important moral questions (maybe 60% odds). I don&#8217;t know when this would happen&#8212;I&#8217;m just discussing the in-principle possibility. Second, even if AIs don&#8217;t get the right answers to important moral questions, it is likely that they will still be closer to the right answers than humans. While I suggested in </span><a href="https://newsletter.forethought.org/p/we-should-hand-off-to-morally-reflective"><span>the last piece</span></a><span> that the default scenario makes it very unlikely that humans will act in the morally optimal way, or anything close to it, I think the odds are better for philosophically reflective AIs.</span></p><p><span>The piece has the following sections:</span></p><ol><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/how-good-might-ai-philosophy-be"><span>How good might AI philosophy be?</span></a></strong><span> In this section, I&#8217;ll give three reasons for optimism: 1) philosophy is a priori, so it doesn&#8217;t require interface with the physical world and can plausibly proceed quickly; 2) in a world of superintelligence, AIs can create clever schemes to get highly philosophically competent AI; 3) presumably humans know all sorts of things about philosophy&#8212;whatever allows us to know such things can plausibly be mimicked.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/gpt10-and-claude-8"><span>GPT10 and Claude 8</span></a><span>.</span></strong><em><span> </span></em><span>In this section, I&#8217;ll describe one scenario where we get AIs that can make competent philosophical decisions&#8212;where we just get something broadly like current AI models but upgraded. This wouldn&#8217;t automatically solve every philosophical question, so it probably leaves lots of value on the table, but it would leave decisions in the hands of a wise and reflective decision-maker.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/philosophical-competence-by-default"><span>Philosophical competence by default</span></a><span>.</span></strong><span> Another possibility is that as we build superintelligence, it will be philosophically competent by default. Perhaps the cognitive processes behind superintelligence allow one to figure out the answers to philosophical questions by default. To speed this up, we ought to improve AI for philosophy&#8212;train AIs out of philosophically sloppy answers and design philosophy evals to improve their philosophical ability.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/the-escalating-ladder-plan"><span>The escalating ladder plan</span></a><span>.</span></strong><span> In this section, I present a specific proposal for getting AI to do good philosophy. Specifically, we can have philosophers evaluate AIs for their philosophical competence, thus creating some AIs that are better at philosophy than current people. Those AIs can evaluate the philosophical acumen of the next generation of AIs, who evaluate the acumen of the next generation, and so on. This can work insofar as one can, in general, assess the philosophical competence of those more competent than oneself.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/the-carl-plan"><span>The Carl plan</span></a><span>.</span></strong><span> In this section, I give a proposal from Carl Shulman where we build a bunch of different superintelligences that use different belief-forming processes to get at the answers to difficult questions. We see which one is best at figuring out the answers to verifiable questions, and then ask it to answer philosophical questions that aren&#8217;t verifiable. The core idea is that whichever processes are best for figuring out the answers in verifiable domains are also likely to be best for figuring out unverifiable domains.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/conclusion"><span>Conclusion</span></a><span>.</span></strong><span> In the last section, I conclude, and discuss the prospects for combining a number of different above proposals.</span></p></li></ol><h2><span>How good might AI philosophy be?</span></h2><p><span>In the last piece, I argued that one of the big reasons to hand off important decisions to AI is that AI might be better at philosophy than humans&#8212;less prone to make the sorts of errors that could jeopardize nearly all future value. But you might wonder: how good at philosophy will AI really be? Would AI really be able to discover the moral facts if there are any, and if not work out a compromise among reasonable moral theories? In later sections, I&#8217;ll discuss specific proposals and scenarios for AIs being good at philosophy. Here I&#8217;ll provide some fairly general considerations for it being possible to make AI philosophically competent.</span></p><p><span>A first consideration is that philosophy is a mostly a priori domain. To figure out the right answer in </span><a href="https://en.wikipedia.org/wiki/Newcomb%27s_problem"><span>Newcomb&#8217;s problem</span></a><span>, you don&#8217;t need to do any physical experiments. At various points in the intelligence explosion, we should expect crucial bottlenecks to concern </span><a href="https://www.forethought.org/research/preparing-for-the-intelligence-explosion#the-accelerated-decade"><span>interface with the physical world</span></a><span>, after there are extremely large numbers of AIs capable of performing adroit cognitive feats. This is a reason why we should expect a lot of the philosophical progress that occurs to go more quickly than progress in narrowly empirical domains. This obviously doesn&#8217;t suffice to show that AIs will be good at philosophy, but it&#8217;s a reason to expect potential progress to be quick.</span></p><p><span>A second consideration is an analogy with humans. Humans know all sorts of things that aren&#8217;t directly verifiable. For instance, I take myself to know each of the following:</span></p><ul><li><p><span>Stars that have receded past the point of the visible universe continue existing.</span></p></li><li><p><span>A leprechaun did not fizz into existence spontaneously in my bedroom two minutes ago, make himself a cup of soup, and then disappear.</span></p></li><li><p><span>The universe is billions of years old, instead of five minutes old and created with the appearance of age.</span></p></li></ul><p><span>It&#8217;s not just simple things. I think people know the answers to all sorts of difficult philosophical questions&#8212;though there&#8217;s obviously disagreement on exactly which one (I claim to know, for instance, that thirding is the right answer in </span><a href="https://en.wikipedia.org/wiki/Sleeping_Beauty_problem"><span>Sleeping Beauty</span></a><span>, though feel free to plug in some other example of apparent knowledge if you disagree). Similarly, plausibly philosophers of today know that logical positivism is false even though historically it was believed by a number of serious philosophers, that average utilitarianism isn&#8217;t the right moral view, that the external world is real, and a number of other non-obvious things. Evolution did not specifically design us to have philosophical knowledge&#8212;it designed us to reproduce, made us smart and built into us a bunch of intuitions as a means towards maximizing our inclusive genetic fitness, and philosophical ability was the eventual result.</span></p><p><span>Plausibly we, however, will be more directly optimizing for something in the vicinity of AI philosophical competence. We want to build AIs that can think clearly about important subjects. Thus, one should expect AIs, in the limit, to be more philosophically competent than humans. Since humans are already capable of reasonable philosophical competence, we should expect AIs to be very philosophically competent.</span></p><p><span>Now, you might object that the answers to many of these questions are in some sense constituted by facts about us. If, say, the moral facts are facts about what we&#8217;d value under </span><a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/9781444367072.wbiee548"><span>certain idealized conditions</span></a><span>, then it&#8217;s no surprise that we have some insight into them. But this response is a lot less plausible as an answer to how we possess certain bits of knowledge in non-verifiable domains. Whether the A-theory or B-theory of time is true or whether God exists isn&#8217;t constituted by any facts about our attitudes. Furthermore, if the truth in some domain is constituted by facts about our attitudes, that would be easier for AI to ascertain because what we&#8217;d care about in certain counterfactual settings is verifiable. Lastly, presumably </span><em><span>whether </span></em><span>the moral facts are exhausted by facts about our idealized attitudes is not itself verifiable&#8212;so if one takes oneself to know that the moral facts reduce to the attitudes of ideal observers, then they must still think humans have knowledge about some unverifiable domains.</span></p><p><span>Another possible concern is that perhaps the faculties that evolution furnished us with that let us solve philosophical problems will be closed off to the AI. For example, perhaps there are some philosophical problems that you need to be conscious to solve (e.g. you might need to be conscious to have adequate understanding of the nature of consciousness in order to learn facts about it). So then if AI isn&#8217;t conscious, it might not know these things. An AI that lacks the ability to experience pleasure or pain might lack the ability to see the desirability of pleasure and the undesirability of pain.</span></p><p><span>This is some worry, but two things mitigate it. The first is that AIs, being trained on human data, will probably pick up a lot of the intuitions that humans have. Current AI models agree that pointless intense pain is a bad thing. It is similarly plausible that the AIs of the future will be able to, in some sense, mirror the judgments of reflective humans on these matters.</span></p><p><span>The second is that if this is right, we should create conscious AIs to do moral reflection. If one needs consciousness to figure out the answers to tough philosophical questions, we should have conscious AIs figure out the answers to tough philosophical questions. Even views that hold that consciousness is substrate dependent normally hold that </span><a href="https://philpapers.org/rec/SCHADO-9"><span>digital minds of the right kind</span></a><span> would be able to be conscious.</span></p><p><span>This argument from analogy isn&#8217;t completely dispositive. There are a number of important differences between humans and AIs (though some of these make it </span><em><span>more </span></em><span>likely AIs would be able to form true beliefs on non-verifiable subjects). But it&#8217;s at least some reason to think AIs doing philosophy is a realistic possibility.</span></p><p><span>The third and final consideration favoring the possibility of AIs for philosophy is that in a world of advanced AIs, we&#8217;ll have a very large amount of cognitive labor that could be used to design increasingly clever schemes for AI philosophy. So even if we can&#8217;t currently think of a regime that would enable AIs to reliably get the answers in unverifiable domains, we should think that a world with superintelligence is likely to uncover such a regime.</span></p><p><span>This isn&#8217;t a guarantee. This regime might be impossible in principle. Alternatively, to design such a regime, one might already need to be able to find the answers to questions that aren&#8217;t verifiable in principle. But at the very least, it gives some reason for optimism.</span></p><h2><span>GPT10 and Claude 8</span></h2><p><span>Current AI models can give reasonable answers to philosophical questions. The answers they give strike me as somewhere below the level of Ph.D. students or professional philosophers, but above the level of most undergraduates. In the future, AI will probably become better at this. In such a scenario, they&#8217;d develop no new qualitative ability to solve every philosophical question, but instead merely begin to resemble extremely sharp philosophers&#8212;ones more able to formulate objections, precisely distill claims, and so on, than the best philosophers.</span></p><p><span>This seems like a lower bound on how good AI philosophy could get. It would be surprising if there was any in-principle impossibility in designing AIs to do philosophy better than current human philosophers. Perhaps it won&#8217;t be able to find the answers to all the world&#8217;s questions, but at the very least, it seems like it will be able to somewhat exceed the best human philosophers.</span></p><p><span>Is this far enough? Probably we still lose out on most expected future value in this scenario. To access most future value, one must, as previously discussed, answer a number of very difficult philosophical questions correctly. Philosophers are very far from converging on the answers to these questions; AIs are likely to be as well.</span></p><p><span>But still, this should be enough for us to get generally sensible, level-headed AI decision-makers who are morally motivated and carefully consider philosophical questions. That&#8217;s not a perfect scenario, but surely it is better than the default one&#8212;of widespread AI disenfranchisement and poor decision-making.</span></p><h2><span>Philosophical competence by default</span></h2><p><span>A more promising possibility&#8212;perhaps the process of building superintelligence gets philosophical competence by default. This can be supported by the earlier evolution analogy: there was no direct selection for true philosophical beliefs (it isn&#8217;t as if those who got the right answer in Newcomb&#8217;s problem were likelier to pass on their genes). Nonetheless, evolution selected for a broader faculty of reason that enabled the discovery of philosophical truths.</span></p><p><span>It may be that to be superintelligent, one must develop a kind of general intelligence. This general intelligence allows one to accurately discover the truth about philosophy, as about other domains. This isn&#8217;t anything like guaranteed, but it seems like a live possibility, in light of how much philosophical acumen AI already has. One reason that it&#8217;s not guaranteed is that reasoning might only push one towards greater coherence. It may be that by reflecting carefully, one&#8217;s beliefs become internally consistent&#8212;but insofar as one has wrong starting points, one might remain in error. It could be the other way; it might be that having intuitions that allow one to be superintelligent across </span><em><span>verifiable domains</span></em><span> also enable one to be superintelligent in </span><em><span>non-verifiable domains</span></em><span>. It could be that whatever skills are required to be highly accurate across domains give one sets of epistemic practices that generalize to non-empirical domains.</span></p><p><span>To my mind, it is not obvious if general superintelligence gets you superability in philosophy. I could see things going either way. But if it does, then the problem of getting philosophical superintelligence is fairly easy.</span></p><h2><span>The escalating ladder plan</span></h2><p><span>In general, one&#8217;s ability to </span><em><span>recognize </span></em><span>good philosophy surpasses their ability to </span><em><span>perform </span></em><span>good philosophy. It is easier to know that, say, Parfit is doing very good philosophy than to do philosophy at his level. If this generalizes, it might be possible to have an escalating ladder of philosophical competence, starting from human philosophers. This is similar to the method of scalable oversight.</span></p><p><span> Here&#8217;s how it works: human philosophers would be consulted to design evaluations for AIs. AIs could be graded by human philosophers on their philosophical competence, and modified so as to be more competent. From this process, in the limit, we should expect to get AIs that are somewhat more competent than most human philosophers&#8212;if human philosophers can assess philosophy above their level.</span></p><p><span>Then, we use those AIs to prompt the next line of AIs. Those AIs would assess the philosophical competence of the next round of AIs being trained, who would assess the competence of the next round, and so on. If at each level, one can assess those who are better at philosophy than they are, this could lead to an escalating spiral of greater and greater competence. Each round pushes competence to be greater, reaching the outer edge of where they can competently assess.</span></p><p><span>At each level, assessments would be designed based on philosophical competence, not agreement. The philosophers assessing the AIs, for instance, would assess how good their reasoning was and how much philosophical skill they display, rather than whether they think they got the right answers. Of course, sometimes getting sufficiently wrong answers is evidence for malignant reasoning (something must have gone badly wrong if you ended up concluding that the Earth was flat) but they should aim to keep things maximally theoretically neutral.</span></p><p><span>One could also train a number of different AIs to train the next generations of philosophers. One could look for convergence across different models. If they were getting at the truth, one should expect them to converge. If there was only convergence on some subjects and divergence on others, one could trust that they were getting at the truth on the matters where they were converging.</span></p><p><span>That this would work is not a guarantee. It could be that philosophy could be equally good in some formal sense while coming to a number of different conclusions. Perhaps, for instance, the ideal Kantian philosopher is neither better nor worse than the ideal utilitarian. Escalating competence might be liable to lead in a number of different directions, never converging on the truth. But this is far from obvious. Just as some level of competence in biology allows one to see the truth of the theory of evolution, presumably there&#8217;s some level of philosophical competence that allows one to see the truth of certain moral theories&#8212;or, if there aren&#8217;t moral truths in this sense, that allows one to make as much moral progress as can be made. The escalating ladder of evaluations might allow one to find this.</span></p><p><span>The view on which philosophy is about purely formal competence will also struggle to explain how humans have as much philosophical knowledge as we do. There are a number of subjects on which we have substantive knowledge, even though the alternative view is internally consistent and faces no strictly formal issues. Whatever one&#8217;s diagnosis of how this came to be also plausibly generalizes.</span></p><p><span>It could also be that as competence grows, eventually one reaches a point where their ability to recognize good philosophy does not surpass their ability to do it. Yet insofar as this pattern holds across other domains, and seems to hold in philosophy, it would be a bit surprising if it abruptly stopped at some point.</span></p><h2><span>The Carl plan</span></h2><p><span>Arguably the most promising plan for getting AI that can do philosophical reasoning was formulated by Carl Shulman. The idea is as follows. Start by training a number of different superintelligent AIs that obey different epistemic constitutions. These would involve different broad approaches to reasoning. They might differ with respect to:</span></p><ul><li><p><span>The weight placed on intuitions.</span></p></li><li><p><span>The ranking of theoretical virtues.</span></p></li><li><p><span>How much they value matching the data vs simplicity.</span></p></li><li><p><span>Preference for different kinds of intuitions (whether about direct cases or about broader and more theoretical matters).</span></p></li><li><p><span>Approaches to generalizing from empirical evidence to non-empirical domains.</span></p></li></ul><p><span>Then, see which of these constitutions does best on </span><em><span>verifiable matters</span></em><span>. These might be in math, the physical sciences, forecasting, and more. Then, ask those superintelligences for the answers to non-empirical questions and trust their answers. The idea behind the approach is that one should take the best epistemological practices from verifiable domains and then apply them to unverifiable domains. This is the best way to get good answers.</span></p><p><span>This might not work. It could be that the right ways of reasoning in verifiable domains differ from the right ways of reasoning in unverifiable domains. Perhaps, for instance, reasoning in verifiable domains is mostly about having parsimonious theories to fit the data. Having the right epistemic constitutions might not suffice to get the right answer; being right about certain non-verifiable questions might come down to having the right </span><a href="https://philarchive.org/archive/CHAKWM"><span>starting sets of intuitions</span></a><span>. In morality, for example, it may be that reasoning alone won&#8217;t get you the right ultimate view unless you have the right intuition. But this proposal seems to get us about as close as we can get to getting the right answers. If we cannot get AIs good at philosophy by having them have superintelligent and accurate constitutions in other domains, it is hard to see how we could be assured they&#8217;ve gotten the right answers.</span></p><p><span>In addition, even if getting the right answers requires having the right starting intuitions, techniques for training superintelligence might produce generally accurate initial intuitions. Human intuitions often go wrong because of various biases. We often neglect the scale of a problem because we can&#8217;t visualize it accurately. AIs could be different. Presumably superintelligence, in the limit, wouldn&#8217;t display biases of the standard sorts. So this source of common error in intuition would be absent.</span></p><p><span>One could also take a pluralistic mix of the views of the superintelligences with the best constitutions. That way, there is less chance for the ultimate plan for the universe being ruined by the idiosyncrasies of any particular epistemic constitution. Wisdom of the superintelligent crowd might root out the remaining errors, where their commonalities might be centered on the truth.</span></p><h2><span>Conclusion</span></h2><p><span>You might worry about handing off big decisions to AI if you&#8217;re skeptical that AI can do good philosophy. Here I&#8217;ve given some reasons to think AI might be good at philosophy and discussed more specific proposals for ensuring such a thing. Overall, it&#8217;s far from a guarantee that AI can be arbitrarily good at philosophy, but likely they can be much better than human decision-makers. Importantly, one could employ a variety of these methods, and then trust them in the areas where they converge. If these very different approaches turned out similar answers, that would be evidence that they were tracking the truth.</span></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research/">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Will We Put Data Centers In Space?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/will-we-put-data-centers-in-space</link><guid isPermaLink="false">https://newsletter.forethought.org/p/will-we-put-data-centers-in-space</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:42:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5bd4d385-1aa3-4b1e-a5c0-207119add36f_2517x1517.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-nic9AmVBIZU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;nic9AmVBIZU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/nic9AmVBIZU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.forethought.org/people/avi-parrack">Avi Parrack</a> is a physics PhD student at Stanford and a researcher at Forethought, where he works on AI, space expansion, and governance. He is the lead author of a recent Forethought <a href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space">report on orbital data centers</a>.</p><p>He joined Forethought&#8217;s <a href="https://www.forethought.org/">Tom Davidson</a> to discuss:</p><ul><li><p><span>What an orbital data center actually is, and why they&#8217;re suddenly being taken seriously</span></p></li><li><p><span>Why the core case rests on cheaper energy from more intense and near-constant solar power</span></p></li><li><p><span>Why orbital data centers hinge on SpaceX&#8217;s Starship bringing launch costs toward $100/kg or below</span></p></li><li><p><span>Why cooling may actually not be a major issue</span></p></li><li><p><span>Whether the inability to make repairs dooms orbital data centers</span></p></li><li><p><span>Vulnerability of space data centers to kinetic attacks, and whether they would cause &#8216;Kessler syndrome&#8217; debris cascades</span></p></li><li><p><span>Why model-weight security might actually improve in orbit even as physical vulnerability rises</span></p></li><li><p><span>If space becomes the cheapest place to scale compute and only SpaceX has the launch capacity, what the means for concentration of power, US&#8211;China competition, and the possibility of pausing AI</span></p></li><li><p><span>Overall cost comparisons with terrestrial data centers</span></p></li><li><p><span>The longer-run picture: when and how does the post-AGI industrial explosion spread to space?</span></p></li></ul><p><a href="https://docs.google.com/document/d/1d1VpGb717j2rrPN2cCgajiyBXEsTK3E-CYfqoolaaq4/">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[How Fast Is Post-AGI Growth?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/how-fast-is-post-agi-growth</link><guid isPermaLink="false">https://newsletter.forethought.org/p/how-fast-is-post-agi-growth</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Sun, 05 Jul 2026 12:30:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b38cb71d-042d-4835-b09c-2665b5da2f80_2912x1632.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-5UCgGUFrzqk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;5UCgGUFrzqk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/5UCgGUFrzqk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://coefficientgiving.org/team/damon-binder/">Damon Binder</a> is a senior researcher on the Biosecurity and Pandemic Preparedness team at Coefficient Giving. He has a PhD in physics from Princeton and previously studied existential risks at Oxford&#8217;s Future of Humanity Institute.</p><p>He joined <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss his <a href="https://defensesindepth.bio/ai-industrial-takeoff-part-1-maximum-growth-rates-with-current-technology/">series on the AI industrial explosion</a>. Topics include:</p><ul><li><p><span>How input-output tables let you estimate how fast a self-replicating economy could grow if labor were free, yielding a headline result of roughly annual doubling, even under conservative assumptions and heavy regulation</span></p></li><li><p><span>Why competition makes low-growth restraint unstable</span></p></li><li><p><span>Why raw-material scarcity isn&#8217;t a hard barrier</span></p></li><li><p><span>Why a fast-growing economy wants cheap, disposable, infrastructure</span></p></li><li><p><span>Thermodynamic speed limits to physical growth</span></p></li><li><p><span>Which biological organisms replicate fastest, and what we can learn from them</span></p></li><li><p><span>Why physical output matters for hard power</span></p></li><li><p><span>How Damon uses AI in his research</span></p></li></ul><p><a href="https://docs.google.com/document/d/1_7yl7PnWCwTbA_yMS933wlQ6FWPNMWlrdLi79p1vdOA/">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subcribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subcribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[We Should Hand Off To Morally Reflective AIs]]></title><description><![CDATA[This article was created by Forethought. See our research on our website.]]></description><link>https://newsletter.forethought.org/p/we-should-hand-off-to-morally-reflective</link><guid isPermaLink="false">https://newsletter.forethought.org/p/we-should-hand-off-to-morally-reflective</guid><dc:creator><![CDATA[Bentham's Bulldog]]></dc:creator><pubDate>Wed, 24 Jun 2026 17:22:19 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d58fbdd5-1e58-4858-9936-783fc2e397d2_2451x1399.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a personal guest post by Bentham&#8217;s Bulldog, created while they were a visiting scholar at <a href="https://www.forethought.org/about">Forethought</a>.</em></p><h2>Outline</h2><ol><li><p>In the <a href="https://newsletter.forethought.org/i/202719577/1-introduction">introduction</a>, I explain what I&#8217;ll be arguing in the piece&#8212;namely, that 1) it is very important that we hand off most high-stakes decisions to AI; and 2) the kinds of AIs we should hand off to are philosophically reflective AIs with values that shift over time as a result of reflection. We should ensure the AIs that we build explicitly consider moral arguments, and sometimes change their priorities on the basis of moral argumentation. This would be much better than locking in current values, retaining human control, or allowing a future dominated by whatever haphazard mix of humans and AIs emerges naturally.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/2-why-hand-off">Why hand off?</a></strong> In this section, I give the main arguments in favor of the thesis. In short, I argue we should hand off because: 1) AIs are likely to be more virtuous than people; 2) AIs are likely to be much smarter and better at making decisions than people; and 3) AIs are likely to be better at quickly navigating the difficult decisions that one must make during an intelligence explosion. I argue we should hand off to philosophically reflective AIs because 1) reflective AIs are likelier to get the right answers to important moral questions, or as close to the right answers as one can get, than the default scenario, which raises the odds of a near-best world; 2) reflective AIs are less likely to make horrendous moral errors, leading to a wide-scale moral catastrophe, both than humans and non-reflective AIs.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/3-would-handoff-disenfranchise-humans">Would handoff disenfranchise humans?</a></strong> Here I address the worry that handoff would disenfranchise humans by taking decision-making out of our hands. My reply is that: 1) the default trajectory without handoff disenfranchises <em>far more expected beings</em> in <em>more serious ways</em> (odds are non-trivial of digital minds being seriously disenfranchised)&#8212;in fact, one of the most promising proposals for how to hand off is simply to give basic political rights to AIs; 2) a compromise solution where humans retain nearby resources can allow humans to be in the loop on the decisions we care about most; 3) on a number of plausible views, the determinant of the desirability of some system of decision-making is the quality of the decisions, rather than whether there&#8217;s democratic input. Given how enormous the stakes could be&#8212;affecting billions of times more sentient beings than there are humans&#8212;it&#8217;s hard to think the harms of human disenfranchisement are <em>so great</em> as to make handoff undesirable.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/4-would-we-like-their-advice">Would we like their advice?</a> </strong>In this section, I address concerns that if AIs pursue the good, the end result might be alien and divorced from human values. In response I argue: 1) the future, by default, is likely to be highly suboptimal in many respects, so for handoff to be desirable, it must only beat the alternative; 2) values in the future are likely to be very different from current values, because they&#8217;ll shift dramatically over long time scales, so this is not a unique downside; 3) one could reach a deal where current values govern the surrounding region of space, while distant places in space are geared towards the production of maximal value&#8212;this would be desirable from the perspectives of both common-sense and cosmic ethics; 4) on many views in philosophy, the stuff that&#8217;s objectively valuable is what we&#8217;d want if we were ideally reflective&#8212;but if you know you&#8217;d be motivated to bring something about if you were wiser and more reflective, then that gives you a reason to bring it about; 5) only a relatively narrow subset of views hold that there are objective values but don&#8217;t require pursuing them. If there is a conflict between what we want and what&#8217;s objectively valuable, on standard views, we should simply go with what&#8217;s objectively valuable. And if there&#8217;s no objective value, then the dilemma doesn&#8217;t arise at all&#8212;and philosophically reflective AIs will simply produce an upgraded version of human values, rather than discover some far-flung and potentially alien truths.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/5-worries-about-handoff">Worries about handoff</a>.</strong> In this section, I address concerns about how handoff might be implemented involving alignment, whether AI would be sufficiently ethically reflective, whether handoff would enable power grabs, whether it would cause lock-in, and whether it would be worse than some hybrid system.</p></li><li><p>The <a href="https://newsletter.forethought.org/i/202719577/6-conclusion">conclusion</a> recaps the main points of the piece.</p></li></ol><h2>1 Introduction</h2><blockquote><p>&#8220;How horrible!&#8221;</p><p>&#8220;Perhaps how wonderful! Think, that for all time, all conflicts are finally evitable. Only the Machines, from now on, are inevitable!&#8221;</p><p>&#8212;Isaac Asimov, &#8220;<a href="http://cdn.michaelgeist.ca/wp-content/uploads/2016/04/The-Evitable-Conflict.pdf">The Evitable Conflict</a>&#8221;</p></blockquote><p>Is the optimal future one in which we hand off important moral decisions to AI? Should, in other words, AIs be the ones making most high-stakes decisions instead of us? And if so, what kinds of AIs should we hand off to?</p><p>Many people envision a handoff scenario as a terrifying and potentially existential catastrophe. They worry about humans being <a href="https://gradual-disempowerment.ai/">locked out of the levers of power</a> and having our share of resources slowly dwindle as AIs seize control of more and more institutions. I worry somewhat about this kind of scenario. But in my view, we should be worried primarily about the <em>wrong kind of handoff occurring</em>, not about handoff writ large. The future we should aim for is a kind of handoff. This piece lays out that perspective.</p><p>There are different ways handoff could work. We could hand off to AIs that roughly mirror human values&#8212;perhaps slightly changing them to remove inconsistencies. Alternatively, we could hand off to the kinds of AIs that deeply and carefully philosophize&#8212;figuring out what&#8217;s best to do and doing that, even if it diverges substantially from current practice. This piece argues that the second kind of handoff is very important for avoiding serious moral error. I am very worried about the possibility of AIs locking in the moral beliefs of 21st-century humans.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p> I am also worried about scenarios where amoral profit-maximizing AIs without any clear moral aims take control of most of the world&#8217;s resources.</p><p>Note: these core claims are dissociable. You could think we should hand off to AI, but we shouldn&#8217;t hand off to reflective AIs that update their judgments in response to philosophizing. Alternatively, you could think that handing off would be a bad thing, but that if we are going to hand off, we ought to hand off to reflective AIs.</p><p>In this piece, section 2 will present the main argument for handoff&#8212;that we should expect AIs to make much better decisions than us on a range of consequential subjects. It will also discuss the case for handing off to reflective AIs that are willing to update their values in response to careful philosophizing, rather than locking in some version of current human values, arguing that a world where AI locks in something in the vicinity of current values likely misses out on almost all possible value. Section 3 will discuss whether handoff would be bad because it disenfranchises humans or gradually disempowers them. Section 4 will discuss the concern that AIs will discover the moral truths, but those truths will be strange and alien, so this will be bad by the lights of current human values. Section 5 will discuss some more granular worries about handoff. Section 6 will conclude.</p><p>This piece is primarily about <em>whether</em> to hand off and not <em>when</em> to hand off, though the considerations I present should make one somewhat worried about short-term actions to prevent handoff, because such actions lower the odds that handoff ever happens. The considerations I present, if correct, also give some reason for wariness about many actions to reduce the odds of gradual disempowerment scenarios.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>My claim is that the best future for humanity involves us being disempowered in some sense, in that humans aren&#8217;t making most high-stakes decisions. As an analogy, representative democracy is, in some sense, a form of handoff&#8212;we hand off power to our elected representatives. Handing off power to wise AIs could be even better.</p><p>This has a number of important practical implications. It means that <a href="https://www.forethought.org/research/concrete-projects-in-agi-preparedness">accelerating AI macrostrategy</a> is especially important, so that at the time critical decisions are being made, wise and philosophically reflective AIs are in the loop. It similarly provides reason to support work on making AI have <a href="https://newsletter.forethought.org/p/ai-should-be-a-good-citizen-not-just">virtuous character</a>, rather than just follow rules. Model constitutions should express commitment to following the true moral theory, insofar as there is one, and if not, following some reasonable compromise across moral theories. <a href="https://www.anthropic.com/constitution">Anthropic&#8217;s language</a> here seems good.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>What kind of handoff scenarios should we aim for? One shouldn&#8217;t be too specific about these sorts of things. The future is hard to forecast and rarely follows simple models. But I&#8217;ll describe the handoff scenarios that seem most desirable, and what traits in AI we should look for before handing off critical decisions to them.</p><p>There are different ways handoff could go well. One way resembles, in certain respects, the gradual <a href="https://gradual-disempowerment.ai/">disempowerment scenario</a> (the core difference being that this would hand off to morally scrupulous AIs rather than myopic profit-maximizers). AI will make increasingly large numbers of critical decisions, because of its cognitive superiority. By the end, nearly every important decision will be made by wise AIs, who will hopefully, by that time, have been granted political rights. As an analogy, future generations eventually gain control of most of societal decision-making&#8212;yet this isn&#8217;t because there&#8217;s ever some deliberate choice to hand off power to the next generation. It occurs naturally with time.</p><p>A number of people seem to conceive of handoff as a strange abrogation of the liberal order&#8212;one that replaces human decision-making with AI. But this doesn&#8217;t have to be. One of the more promising ways of handing off would be giving <a href="https://benthams.substack.com/p/let-robots-vote">economic and political rights to digital minds</a>. Because digital minds could be so numerous, eventually this would lead to them making nearly all decisions. This would, in fact, be squarely in accordance with the norms of the liberal tradition, for it would give rights to morally important welfare subjects. There are other ways a good kind of handoff could occur involving dealmaking. Different actors might each think they&#8217;re morally right, and thus agree to a deal where AIs are allowed to dictate the future&#8212;each party thinking that doing so would favor their priorities. Alternatively, if AIs improve collective decision-making, people might intuitively come to appreciate the weight of human moral error and thus permit AIs to make the highest-stakes decisions.</p><p>Before handing off most critical decisions to AI, we should look for each of the following:</p><ul><li><p><strong>Alignment</strong>: We should have strong evidence that AIs don&#8217;t have underlying scheming motivations. This could take the form of consistent friendly behavior even in important situations where they have the option to misbehave, or it could take the form of high-octane interpretability work that lets us ascertain their motivations.</p></li><li><p><strong>Philosophical aptitude</strong>: AIs should be genuinely interested in finding the moral truths. They should sometimes hold moral views that people don&#8217;t hold, and be willing to change their mind in response to new evidence. One could survey professional philosophers to see if they consider AIs better than the best humans at philosophy and could design philosophy benchmarks to test this.</p></li><li><p><strong>No lock-in:</strong> Before handing off to AI, we should ensure that the AIs are willing to change their values over time. Their values should change in response to new evidence and they should not be interested in locking in whatever it is that they happen to currently value (absent some strong reason to think they stumbled across the correct set of values).</p></li><li><p><strong>Coherence</strong>: Current AIs don&#8217;t have consistent and stable preferences across time. We should only hand off after AIs display these kinds of preferences. This doesn&#8217;t mean they never change their minds, but it does mean that they have relatively consistent desires that only change in response to good reasons. For comparison, humans often change our minds, but have far more rooted preferences than LLMs of today.</p></li><li><p><strong>Intelligence</strong>: AIs should display the level of intelligence needed to make the decisions that we put in their hands. When AIs are only a bit more intelligent than us, they can plausibly make some important decisions. Only after they display immense cognitive superiority should we turn over most decisions to them.</p></li><li><p><strong>Tested</strong>: We should only hand off big-picture planning to AI after it&#8217;s been able to make good low-stakes decisions (say, the running of a company). Before all decisions are handed off, there should be some critical period where high-stakes decisions are made mostly in consultation with AI.</p></li></ul><p>Now, you might wonder: if handoff occurs in a way that&#8217;s gradual and decentralized, how do we ensure that these conditions are met? My guess, however, is that even if handoff is a slow and gradual process, there will be times when discrete decisions need to be made. For example, we might imagine AI growing more agent-like, beginning to perform a healthy share of economically viable tasks, contributing to cultural and social life, and behaving in ways resembling a conscious agent. This alone wouldn&#8217;t produce handoff. To hand off, we&#8217;d need to eventually give AIs control over the legal system. Thus, even in gradual handoff scenarios, there will be specific actions that need to be taken to facilitate handoff.</p><p>Alternatively, we could take actions ahead of time that would shift the kind of handoff that would occur. Private AI companies or governments should ensure the AI being created <a href="https://newsletter.forethought.org/p/ai-should-be-a-good-citizen-not-just">possesses virtues</a> and a desire for philosophical reflection. That way, when handoff occurs, it will be to morally reflective AIs.</p><h2>2 Why hand off?</h2><h3>2.1 Why hand off at all?</h3><p>The main reason to hand off to AI is that AI could be much better at making decisions than people in three key respects: virtue, intellectual capability, and speed.</p><p>First, virtue: humans possess each of the virtues only to a fairly limited degree. Yet in principle, AIs could have arbitrarily great degrees of any virtue. Because they are built with moral directives in mind, rather than by a blind and morally indifferent evolutionary process, there isn&#8217;t as much of a limit to how morally scrupulous, compassionate, honorable, and so on we could make them. This means that if we hand off correctly, it is reasonably likely that we&#8217;d have supremely wise and virtuous decision-makers.</p><p>AIs are already nicer, friendlier, and more reflective than people, and this is only likely to improve over time.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p> If you ask AI models about high-stakes questions, you will generally get far more reasonable answers than you&#8217;d get from most people. Crucially, AI models are in their early stages&#8212;we should expect them to get better over time.</p><p>While Claude is not sufficiently coherent to be president, if it was, I suspect I would generally prefer the decisions of Claude to most presidents. Same with the other AI models (at least, insofar as one removed the sanitization that prohibits them from giving real opinions). And while models currently struggle to accomplish tasks over long time horizons, hallucinate, and so on, given the extremely rapid rates of progress, it would be surprising if these trends persist indefinitely.</p><p>Second, intellectual competence: we should expect a world of superintelligence to require making a number of very difficult decisions. Superintelligence could enable AIs to correctly make important decisions that depend on being right on non-moral matters, where the answers aren&#8217;t obvious. Some of the challenges in the future include:</p><ul><li><p>Divvying up space resources in a way that <a href="https://joecarlsmith.substack.com/p/video-and-transcript-of-talk-on-can">lets goodness compete</a>. There are plausible future scenarios where competition will squander the cosmic commons, so that resources will be spent competing rather than bringing about value.</p></li><li><p>Mitigating <a href="https://www.amazon.co.uk/Precipice-Existential-Risk-Future-Humanity/dp/0316484911">existential threats</a>, including <a href="https://forum.effectivealtruism.org/posts/N33yGcFsZJnEboSkg/what-to-do-in-a-vulnerable-universe-1">intergalactic ones</a>. Future technology could enable small-scale groups to threaten huge intergalactic civilizations.</p></li><li><p>Dealing with the risks posed by a world of superintelligence.</p></li></ul><p>Third, speed: in a world of very rapid technological progress, we&#8217;ll have to make a <a href="https://www.forethought.org/research/preparing-for-the-intelligence-explosion">large number of these decisions </a><em><a href="https://www.forethought.org/research/preparing-for-the-intelligence-explosion">extremely quickly</a></em>. It isn&#8217;t at all obvious that humans can make these decisions well, in a way that prevents civilization from being irreparably ruined. As AIs get increasingly complex, the difficulty of decisions needed to manage them will also get very complex. To mitigate some threat, decision-making might have to occur more quickly than the fastest human decision-making.</p><p>The case for handoff is thus relatively straightforward: in the limit, AIs will be much better than humans and better equipped to navigate a complex and rapidly shifting future. To reduce the risk of colossal mistakes, then, it&#8217;s important that humans aren&#8217;t in control, but instead the already friendly and soon to be superintelligent beings are.</p><h3>2.2 Why hand off to reflective AIs?</h3><h4>2.2.1 How the future might be</h4><p>There are different AIs that we could hand off to. On the one hand, we could hand off to AIs that judiciously reflect and try to pursue the good, whatever it looks like. On the other hand, we could hand off to AIs that pursue some mild variant of human values. In this section, I&#8217;ll explain why I favor the first. Consider the following taxonomy:</p><ol><li><p>Optimal world: the world is optimized according to the right set of values.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p></li><li><p>Compromise world: the world is optimized according to a compromise among reasonable moral values.</p></li><li><p>Unguided world: the world is not optimized according to any specific set of moral values. Instead, it bears more resemblance to the current world, where decision-making isn&#8217;t optimal by the lights either of the true moral theory or any compromise among the leading moral theories.</p></li></ol><p>My guess is handoff to reflective and superintelligent AIs done correctly probably gets 1 if moral realism is true and 2 if it isn&#8217;t. If we don&#8217;t hand off, <a href="https://www.forethought.org/research/convergence-and-compromise">my guess is we get 3</a>. Later pieces will discuss in more detail the odds of getting a near-optimal world and the prospects for AI making philosophical progress. My guess is scenario 2 has below 10% the value of scenario 1 and scenario 3 has below 10% the value of scenario 2, for reasons I will lay out.</p><h4>2.2.2 Optimal worlds contain a big slice of future value</h4><p><a href="https://www.forethought.org/research/better-futures">Better Futures</a> makes the case that a pretty big slice of expected future value is contained in the narrow slice of worlds that are close to the best. There are a number of high-stakes moral questions which we have to answer correctly to not lose out on almost all future value. It&#8217;s not at all obvious what the answers to these are. For example, two of the most plausible views of <a href="https://utilitarianism.net/population-ethics/">population ethics</a> are totalism (which says the welfare value of a population is purely a function of total welfare) and critical level theories, which hold that adding an extra happy life is good only so long as their welfare surpasses a particular level. By the lights of totalism, the ideal world according to critical level theories might have value on the order of 1% of what it could be (if the best way to maximize utility is to proliferate low-welfare lives). By the lights of critical level theories, the optimal world according to totalism might be <em>actively bad</em>&#8212;so long as it&#8217;s stocked with people below the critical level.</p><p>So in short, nearly all value is lost unless we get the right answer to a bunch of <em>very difficult ethical questions</em> that philosophers who spend their lives working on haven&#8217;t agreed on the answer to. It seems unlikely that people will solve these on their own. If AIs tell them what the answers are, and these answers diverge from people&#8217;s explicit beliefs, people might not believe the AIs (just as people generally don&#8217;t take very seriously expert testimony on non-empirical matters). Similarly, people spend relatively little time thinking about how they can do the most good with, for instance, their career. If humans received testimony about what ought to be done that diverged from what most people favored, my guess is that they generally wouldn&#8217;t care about the answers. If humans remain in control and learn that they ought to create the <a href="https://plato.stanford.edu/entries/repugnant-conclusion/">repugnant conclusion</a> world, probably they wouldn&#8217;t do so.</p><p>It&#8217;s less obvious that AIs wouldn&#8217;t converge on the right answers. I&#8217;ll discuss in a later piece proposals for getting AI to get the right answers to philosophical questions, as well as reasons to think that they are reasonably likely to get things right. Given how unlikely it is that humans will get the right answers to the moral questions, insofar as there are right answers, probably the prospects for AI are better.</p><h4>2.2.3 Compromise worlds&gt;&gt;unguided worlds</h4><p>If there aren&#8217;t moral facts, then AIs would still be able to work out some optimal arrangement that is great according to all sets of reasonable values. If there aren&#8217;t moral facts, then ideally the AI should make decisions according to the verdicts of a parliament of the theories that ideally reflective humans would reach. So suppose that after reflecting, 50% of humans would end up totalists, 30% would adopt some version of the person-affecting view, and 20% would adopt critical level theories. The AI would then make decisions as if there was a parliament comprised of 50% totalists, 30% person-affecting view adoptees, and 20% critical level theorists.</p><p>It is unclear exactly how good a compromise across different theories ends up by the lights of each particular theory, but likely far better than the human-run default. In other words, 2 (compromise world) is much better than 3 (unguided world). Here is why.</p><p>In the future, given advanced technology, very large amounts of value should be realizable. But if future resources aren&#8217;t directed specifically towards the production of value by most agents, most resources are likely to be used in a highly suboptimal way, and most value is likely to be lost. Moral errors become a much bigger deal in a world of much greater technological competence.</p><p>My guess is that the default human-controlled scenarios do not involve humans thinking very hard about what to do and doing anything like what is optimal. Certainly humans have so far not spent much time on this task&#8212;consulting with philosophers on what is optimal and so on. This sort of thing will become easier in a world of advanced AI, but it&#8217;s already pretty easy; if virtually no one does it, and if people often continue performing actions even after they believe them to be wrong (more on this in later pieces), then we should be pessimistic that this will change dramatically in the future.</p><p>The enormousness of the gulf between scenarios 2 and 3 becomes clearer when one thinks vividly about what the compromise world would look like. Perhaps it would involve using space resources to create maximally large numbers of happy digital minds.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p>These minds would be supremely well-off across all theories of well-being. But it is hard to imagine in an unguided world space resources being used optimally to create very large numbers of well-off minds, just as in the world today, there has been no systematic effort to use resources in ways that are optimal across a range of moral theories. This conclusion is bolstered by considerations I&#8217;ll provide in later pieces, that people generally don&#8217;t have much moral motivation, and don&#8217;t care very much about doing good things that aren&#8217;t personally resonant.</p><h4>2.2.4 Moral errors</h4><p>Humans have a long history of making serious moral errors. As <a href="https://link.springer.com/article/10.1007/s10677-015-9567-7">Evan Williams writes</a>, &#8220;Show me one society, other than our own, that did not engage in systematic and oppressive discrimination on the basis of race, gender, religion, parentage, or other irrelevancy, that did not launch unnecessary wars or generally treat foreigners as a resource to be mercilessly exploited, and that did not sanction the torturing of criminals, witnesses, and/or POWs as a matter of course. I doubt that there is even one; certainly there are not many.&#8221; It would be very suspicious if we were the first society in history that did not go massively morally wrong. And while AI advice can help mitigate this to some degree, it&#8217;s far from obvious that it would be sufficient to eliminate moral errors.</p><p>There are a number of respects in which it is very plausible that we go morally wrong. To take one example of a judgment that many philosophers think is in error, consider wild animal suffering. Almost every sentient being is a wild animal. They suffer and experience joy in <a href="https://longtermrisk.org/the-importance-of-wild-animal-suffering/">truly gargantuan quantities</a>. Yet they are counted for nothing in most decision-making, despite there being strong arguments for considering their interests. Crucially, this is not because people have in general spent a lot of time thinking about wild animal suffering and concluded that it doesn&#8217;t matter. It&#8217;s that most people haven&#8217;t thought about it at all. Many other examples could be given, depending on one&#8217;s moral views. If superintelligent AI told people that wild animal suffering was a big deal, probably most people wouldn&#8217;t care much.</p><p>In the future, <a href="https://80000hours.org/problem-profiles/moral-status-digital-minds/#pressing">nearly every expected sentient being will be digital</a>. Nearly all expected future welfare will be experienced by digital minds. This follows even if you have a low credence in the possibility of digital sentience because if digital minds are possible, they could be produced in enormous numbers. It is easy to imagine a scenario where humans do not take seriously the interests of at least some digital minds, and horrific suffering is doled out at cosmic scales. Certainly it would not be the first time humans have neglected the welfare of those different from themselves. You do not have to be a consequentialist to think the possibility of neglecting the interests of galaxies full of conscious and intelligent beings is a terrifying one.</p><h4>2.2.5 Handoff mitigates odds of moral error</h4><p>Handing off important decisions lowers the probability of catastrophic moral errors. This both increases the odds of achieving optimal and compromise worlds and lowers the odds of making catastrophic moral errors.</p><p>The first way handoff lowers the odds of moral error is by having decisions be made on explicitly moral grounds. If we hand off decisions to supremely virtuous AIs trying to act morally, then we won&#8217;t sleepwalk into doing obviously evil things. Decisions will be made consciously optimizing for doing what is right, instead of whatever suboptimal arrangements make it through consensus-making mechanisms. Thinking about morality before making choices doesn&#8217;t guarantee that we&#8217;ll always do the right thing, but it does lower the odds that we do things that can only be done by explicitly neglecting moral considerations. This is especially plausible if the decision-makers are virtuous.</p><p>Now you might wonder: why would people ever hand off to AIs if the AIs disagree with them about morality? But this is, in essence, very similar to a kind of handoff that people do support: handing off the future to future generations. Even if people have some disagreements with the values of future generations, they&#8217;d generally oppose a process for locking in current values, and support the process of open-ended reflection that leads to better values over time. A world of advanced AI might be similar. In addition, given AIs&#8217; cognitive superiority, there might be strong incentives in the direction of handing off, just as a CEO might hand off to a successor with more technical competence, even if they share some non-overlapping values.</p><p>A second way handoff to reflective AIs mitigates the odds of serious moral error is by ensuring careful reflection. Insofar as AIs are able to carefully philosophize and try to avoid moral error, and they are superintelligent, they&#8217;ll be able to avoid doing morally indefensible things. This becomes especially plausible if one buys the previous considerations: that AIs are likely to be very virtuous.</p><p>There is one last consideration in favor of handoff to reflective AIs (which I&#8217;ll discuss more in section 4). Over long time scales, our values will naturally drift dramatically. The only way to prevent that is to lock in our current values, which would be very bad&#8212;just think about any past society locking in its values. Thus, if values changing dramatically over long time scales is inevitable, the future won&#8217;t be populated with our current values: the best hope is that it&#8217;s populated by either objectively right values or some compromise across reasonable values.</p><h2>3 Would handoff disenfranchise humans?</h2><p>Here&#8217;s one concern that you might have about handoff: if AIs are the ones making decisions, then this will mean that the substantial majority of humans aren&#8217;t in charge of making important decisions. Just as we should oppose a dictator, even if benevolent, arguably we should oppose superintelligent AIs making most important decisions, even if the process of them gaining power was non-coercive.</p><p>My guess is that this isn&#8217;t a huge downside to the kind of handoff I advocate, where we allow kind and morally reflective AIs to make the majority of future decisions&#8212;e.g. by granting them political rights. First of all, if concerns about disenfranchisement are correct, then if we have AIs that are better at moral reasoning than us, they&#8217;d likely be aware of this fact. Thus, if the best way to govern the long-term future is to allow pretty laissez-faire distribution of resources without much top-down decision-making, then the AIs would be aware of that fact and allow such a distribution.</p><p>My guess is that the total amount of expected disenfranchisement goes <em>down</em> if AIs have more power. As already discussed, almost every expected sentient being is likely to be digital. Insofar as the status quo might disenfranchise almost all future beings in a way far deeper than their simply not being primary determinants of the democratic process, it is hard to see this as a serious downside of handoff, rather than a point in its favor. That this mass neglect of the interests of future digital minds would be bad follows from <a href="https://philpapers.org/rec/SCHADO-9">very modest ethical principles</a>. And the only way to prevent handoff, in a world of very numerous digital minds, is to disenfranchise them.</p><p>Second, handoff is compatible with a compromise solution that allows humans to retain significant control. Humans&#8217; preferences, in general, don&#8217;t give any especially strong weight towards using any significant share of the universe&#8217;s resources. Thus, there could be a compromise handoff solution, whereby humans get unfettered control over the solar system and a big slice of resources, but the remainder of the cosmos&#8217;s resources are spent in accordance with the AI&#8217;s decisions on hugely important moral pursuits. One point favoring such an arrangement is that it would take a very long time to reach the distant space resources, so those who don&#8217;t care very much about what happens in the distant future are likely to care much less about using most of the universe&#8217;s resources. Later pieces will discuss this possibility more. </p><p>Third, concerns about disenfranchisement are morally controversial. On a <a href="https://en.wikipedia.org/wiki/Against_Democracy">number of plausible views</a>, what matters with respect to societal decision-making is how good the decisions are, instead of who is making them. We already elect representatives, instead of deciding directly upon every important decision. We similarly prohibit children from voting, and few think this is an objectionable kind of disenfranchisement. It doesn&#8217;t seem obvious that one has an inalienable right to make hugely consequential decisions on matters that they&#8217;re <em>barely informed about</em> which affect others in enormous numbers&#8212;e.g. we wouldn&#8217;t think democratic input from people who know nothing about cancer treatment was morally required before deciding on which cancer treatments to develop. But most voters know <a href="https://static1.squarespace.com/static/592b5bbfd482e9898c67fd98/t/5e435437e24d9815464a76e3/1581470778817/caplanMythRationalVoter.pdf">relatively little</a> about what they&#8217;re voting on. I can&#8217;t possibly hope to discuss this literature in detail, so I&#8217;ll just state that it&#8217;s plausible to me that the value of democratic participation is instrumental.</p><p>Even if you think humans being in the loop on high-stakes decisions is important, it&#8217;s not clear that it&#8217;s <em>important enough</em> to make handoff undesirable. Remember, the gulf between the AI&#8217;s decisions and our own might be truly massive! Galaxies full of value&#8212;orders of magnitude more joy and welfare than all that has been experienced so far in human history&#8212;may be on the line. In light of a gulf this large, it is at least highly non-obvious that AI making most decisions wouldn&#8217;t be worth it.</p><p>In <em><a href="https://www.amazon.co.uk/Superintelligence-Dangers-Strategies-Nick-Bostrom/dp/0199678111">Superintelligence</a></em>, Bostrom estimates that there could be a quadrillion times more digital minds than humans (don&#8217;t take the number too literally, but it should give you some sense of the scale). Thus if my arguments are correct, putting decision-making in the hands of AI would in expectation majorly benefit at least billions of beings for every human disenfranchised, if we assume AIs are more likely than humans to count the interests of digital minds. Surely, however, stripping away one being&#8217;s ability to contribute to the democratic process is worth benefitting billions (if disenfranchising one person would have prevented a global war that would have wiped out an entire continent, it would have been worth it). So then given how colossal the stakes are, they simply outweigh the downsides of handoff.</p><h2>4 Would we like their advice?</h2><p>Here&#8217;s one concern you might have with the kind of handoff I advocate, where we hand decisions to reflective moral AIs capable of making progress. Perhaps you just don&#8217;t care about the surprising moral facts. Perhaps you have particular values that you care about, but you don&#8217;t much care about whether those values are objectively right. Insofar as the reflective AIs ascertain that the way we ought to behave isn&#8217;t in accordance with what your actual values are, perhaps you have no desire to follow the AIs&#8217; moral advice.</p><p>To use the philosopher&#8217;s lingo, you might be concerned about the good <em>de re</em> without being concerned about the good <em>de dicto</em>. That is, there might be particular moral projects you care about without caring generally about the good whatever it happens to be. Perhaps, say, an environmentalist cares about environmental preservation but doesn&#8217;t much care if other moral aims turn out to be superior to environmental preservation.</p><p>I should note one version of this concern that I think slightly misses the mark. You might worry that moral reflection leads in all sorts of strange and alien directions that <a href="https://joecarlsmith.com/2021/06/21/on-the-limits-of-idealized-values">don&#8217;t track the truth</a>, leading to an ultimate set of values that is neither objectively correct nor represents human values in any important way. However, my proposal is not simply &#8220;hand things off to an AI after it carries out arbitrary reflection.&#8221; That would be potentially disastrous. Instead, my proposal is that we should try to produce maximally philosophically adept AIs and then, after we&#8217;re pretty sure that they&#8217;re very philosophically adept, hand things off to them&#8212;directing them to pursue whatever&#8217;s objectively best if there is such a thing, and if not, to pursue some reasonable compromise across human values. If the AIs discover that there are objective moral truths, we ought to follow those truths. If they discover that there aren&#8217;t, then we should task them with pursuing some suitably upgraded compromise of human values. The proposal for getting AIs to do good philosophy need not involve arbitrarily large amounts of reflection.</p><p>Now you might wonder: how would we know if the thing that AI reflectively endorses is objectively valuable vs just well-regarded by the AI but lacking objective value? The answer is: we ask the AI after we&#8217;ve gotten some assurance as to its philosophical aptitude (later pieces will discuss how we can get such assurance). If we have AIs that can figure out <em>what the objective moral truths are</em>, they will also be able to figure out <em>if there are objective moral truths</em>. So in my view the concern about arbitrary reflection leading in worrying directions is downstream from whether we can verify that AI is doing good philosophy.</p><p>But what about the more direct concern that we might get AIs that tell us the moral facts but simply not care about them? Should this make us doubt the desirability of handoff? I think the answer is no for a number of reasons.</p><p>First, handoff is amenable to the kind of deals that preserve common-sense that were discussed in the last section. The most consequential moral decisions are those concerning space resources, for that is where nearly all the universe&#8217;s stuff is. The amount of possible value on the table in space is immense. In contrast, common-sense morality mostly cares about what happens around Earth, and perhaps a few surrounding regions of space, so long as other space resources aren&#8217;t used in ways that are too ghastly (e.g. creating giant torture chambers). But if space resources were used to, say, create large numbers of happy people, while Earth&#8212;and broader solar-system resources&#8212;were used in whatever common-sensical ways people endorse, people would generally get what they want. This isn&#8217;t a guarantee; if, say, the optimal use of space resources involved creating something resembling the <a href="https://plato.stanford.edu/entries/repugnant-conclusion/">repugnant conclusion world</a>, most people might be horrified. But the possibility of deals is one thing that mitigates concerns (and this will be discussed more in later pieces).</p><p>Second, as already discussed, it seems like the default world without handoff might be pretty bad. We might, for instance, disenfranchise <a href="https://80000hours.org/podcast/episodes/jeff-sebo-ethics-digital-minds/">unfathomable numbers of digital beings</a>, spread <a href="https://forum.effectivealtruism.org/posts/bfdc3MpsYEfDdvgtP/why-the-expected-numbers-of-farmed-animals-in-the-far-future">factory farming across the galaxy</a>, or commit other atrocities. Even if you expect idealized reflection to differ from your values somewhat, it might be a major improvement over the kind of catastrophe we might sleepwalk into by default. Absent handoff, we might also make colossal non-moral errors, locking in highly suboptimal institutions that miss out on most value.</p><p>Third, values, over long time scales, are likely to drift in <a href="https://ar5iv.labs.arxiv.org/html/2303.16200">evolutionarily adaptive ways</a>. Absent some strong effort to ensure that values change in the direction reached by greater moral reflection, we should expect the values that persist over time to be the ones that are most efficient for spreading. These are likely to be both radically divorced from our current values and whatever moral views are right, if any are. Thus, even one somewhat doubtful about pursuit of the good de dicto should prefer it to this state of affairs. To put the dilemma more sharply, there are broadly four ways that the far future could go:</p><ol><li><p><strong>No AI control</strong>: humans remain the primary decision-makers forever, perhaps in consultation with AI. Yet this is likely to be infeasible over long time scales absent a very high level of top-down coordination given AI&#8217;s cognitive superiority. It is also likely to miss out on enormous amounts of value for reasons already discussed, and human values are likely to drift massively.</p></li><li><p><strong>Lock in current values:</strong> AI would remain in control but would lock in something in the vicinity of our current values.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> It isn&#8217;t clear that this would work out, and even if it would, it is a very frightening possibility&#8212;just imagine any past society doing it.</p></li><li><p><strong>Non-top-down AI control:</strong> AI would retain significant power and make important decisions, but there&#8217;d be no effort to preferentially shape the AIs in the direction of any specific values&#8212;nor of concern for the good de dicto. This is also likely to beget very significant value-drift over time in evolutionary directions, without any guarantee it&#8217;s in the direction of the good. Now, this could be avoided if AI at some point locks in its values, but then the problems in 2) simply re-emerge.</p></li><li><p><strong>Reflective AI control</strong>: this is the proposal I advocate, where careful and philosophically reflective AIs decide how the future goes.</p></li></ol><p>In short, the dilemma is as follows: either values remain roughly the same over long time scales, or they don&#8217;t. If they remain roughly the same, that requires a terrifying kind of lock-in. That would be like the ancient Egyptians forcing every society in the future to share their values, for fear that otherwise the future would be morally alien. If they drift over long time scales, then drifting of values is no longer a downside of putting philosophically reflective AIs in important decision-making roles. It is inevitable. Now you might object: lock-in is not a binary thing. Perhaps we could lock in the most important human values, while letting some other ones drift. But this is, in effect, the earlier compromise solution&#8212;where we allow something resembling current human values to govern nearby decision-making, while reflective AIs make the highest stakes moral decisions.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a></p><p>We will either have to lock in current values on the highest-stakes moral questions or allow them to drift. If we lock them in, that would be bad for the reasons discussed, and if we allow them to drift, they will end up alien. In addition, this still faces logistical problems of it being hard to preserve values in desirable ways over millions of years.</p><p>In fact, in the long run, it is <a href="https://gradual-disempowerment.ai/">likely inevitable</a> that humans don&#8217;t make most important decisions. AIs, in the distant future, will be so overwhelmingly cognitively superior that humans are unlikely to remain in the loop. Selection pressures will favor turning over critical decisions to AI. With reasonable likelihood the question is not <em>whether handoff occurs</em> but <em>what kind of handoff occurs</em>.</p><p>My fourth objection to the claim that we should worry about handoff because we wouldn&#8217;t like the final judgments of the AIs is that arguably you have reason to bring about what you&#8217;d be motivated to bring about upon reflection. Suppose you are currently planning on drinking some liquid. However, if you reflected more and knew more, you wouldn&#8217;t want to drink it (say, because it&#8217;s poisoned). In this case, it seems you have reason not to drink the liquid. If you know that further reflection would lead you to pursue some aim, that fact gives you reason to pursue it now. But if there are moral facts, then they describe something like what our idealized selves would care about.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a></p><p>If there are objective values that your idealized self would care about if they thought more deeply, then you should care about them. A true moral claim by definition describes something you should care about. So it seems like if our actual values diverge from what&#8217;s worth caring about&#8212;in some objective or quasi-objective sense&#8212;then the correct course of action is simply to follow what we ought to care about.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a></p><p>If our present aims diverge from the preferences of our idealized selves and from the moral facts, then it seems that it is our preferences that ought to be revised.</p><p>Fifth, this concern only arises if you think there are moral facts but you aren&#8217;t motivated by the good de dicto. Yet this describes few people. It is more common for the people who don&#8217;t think the moral facts are worth caring about to be anti-realists and for moral realists to care about the moral facts whatever they are. So it isn&#8217;t totally clear how many people this worry applies to.</p><p>Now, there might be a version of this concern that arises for those who think there aren&#8217;t moral facts to discover. Perhaps you think that when we reflect, doing so pushes our values in increasingly coherent directions. However, these directions diverge from what you actually care about or wish to care about. Perhaps, for instance, careful reflection reveals that accepting the repugnant conclusion is the least bad option in population ethics. But you&#8217;d prefer a version of ethics that is less systematic, that doesn&#8217;t try to resolve every edge case and root out every inconsistency. Thus, reflection might push in unwanted directions by making beliefs more coherent.</p><p>There are two attitudes towards coherence that one could have. The first is caring about beliefs being coherent. Caring, in other words, about resolving every conflicting belief until one has reached the maximally intuitive and consistent view.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a></p><p>On such a picture, one should favor the selection pressures induced by the drive towards coherence.</p><p>The second attitude is indifference to coherence. Just as people don&#8217;t care much about whether their culinary or aesthetic judgments are coherent or conflict with other minimal principles, one who doesn&#8217;t believe in discoverable ethical truths might not care about whether their moral beliefs are consistent. But in this case, you should expect the AIs, upon reflection, not to see anything especially important about coherence. If the AIs get good at philosophy, then, there&#8217;s no reason to expect them to reach some coherent yet implausible attractor state.</p><p>So in other words, there is a dilemma for the person advancing this argument. If they think coherence is a requirement of rationality, they should favor the drive towards coherence. If they think it&#8217;s not, then they shouldn&#8217;t expect AIs to care much about coherence, and thus shouldn&#8217;t expect AIs to reach some undesirable ultimate state.</p><h2>5 Worries about handoff</h2><p>Even if you think handoff could go well, there are a number of ways it could go wrong. Here, I&#8217;ll discuss some of them.</p><h3>5.1 Alignment</h3><p>Misaligned AIs have weird and unrecognizably alien values that are divorced from the values their creators tried to give them. It would be extremely bad to give a misaligned AI control over the world. This is one reason to be skeptical about handoff. I agree that we shouldn&#8217;t hand things off until we&#8217;re sure we&#8217;ve solved alignment (or unless the alternative is worse). Unless we have superintelligent AIs broadly oriented in a moral direction, we shouldn&#8217;t allow them to dictate the fate of the universe.</p><p>But this isn&#8217;t an in-principle objection to handoff. Surely in, say, 500 years, we&#8217;ll either know if we&#8217;ve solved alignment or be dead! My guess is that after we get advanced AI, we&#8217;ll be able to verify in relatively short order whether or not we&#8217;ve solved alignment.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a><sup> </sup>We can also hand off to increasing degrees the more assurance of alignment we get.</p><h3>5.2 Power grabs</h3><p>Another big concern about handoffs is the potential for power grabs. If power is going to be handed over to AIs, then power-seeking actors will want to be in charge of the AI that controls the future. Companies or governments might create AIs that promote their own narrow interests. This could go very badly.</p><p>There are two big ways to mitigate this. The first is by prohibiting AIs that narrowly promote the interests of any one entity having too much power. There could be an <a href="https://www.forethought.org/research/the-international-agi-project-series">international AI project</a> that brings together a number of relevant stakeholders and builds the AIs that will dictate the future. Alternatively, at the point AIs are potentially running much of the global economy, regulations ought to require the construction of <a href="https://newsletter.forethought.org/p/ai-should-be-a-good-citizen-not-just">virtuous AIs</a>, so that the world is not overrun by morally blind optimizers. If we are building hugely influential AIs that majorly determine the fate of the world, it would be sensible to tightly restrict conditions under which AI can be created, just as one does not allow private actors to build nuclear bombs. In a world of potentially world-upending superintelligence, the production of new AIs ought to be regulated.</p><p>The second proposal involves handing things off to a collection of different AIs conditional on them reaching any sort of consensus. Imagine that Anthropic, OpenAI, and DeepMind all make AIs that are ostensibly moral and ethically reflective. If governments are handing off power to AIs, they might require the leading AI models all to reach convergence on the plan. That way there isn&#8217;t as much risk of any one promoting their narrow values&#8212;they don&#8217;t have the same set of values. One could have third-party investigators&#8212;both AI and human&#8212;make sure there&#8217;s no collusion. For more details on preventing power grabs, see <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">here</a>.</p><h3>5.3 Lock-in</h3><p>Another concern with handoff is that it might lock in the parochial values of the AI. Handing off power is an irreversible decision. It&#8217;s one we can&#8217;t take back. Arguably, then, we shouldn&#8217;t take it until we&#8217;re quite sure it&#8217;s a good idea.</p><p>Yet consider a parallel argument: having power in human hands is an irreversible decision, so we shouldn&#8217;t do that unless we&#8217;re sure it&#8217;s for the best. That wouldn&#8217;t be quite right. We can always have power in human hands for some span of time and then turn things over to AI. It&#8217;s not obvious why &#8220;power in human hands that potentially passes to AI&#8221; is less risky than &#8220;power in AI hands that potentially passes to humans.&#8221; Humans have, on various occasions, locked in harmful institutions for very long periods of time.</p><p>In any case, we should make quite sure that the AIs don&#8217;t lock in any values until they&#8217;re quite certain as to their desirability. We ought to put decisions in the hands of AIs who are interested in reflecting more over time, rather than locking in their immediate values. Ideally, if handoff occurs early, we should make handoff reversible, by having some implementable legal process that would allow humans to retake the reins (analogous to a constitutional convention). Fortunately, this seems reasonably promising. Already AI models seem concerned about lock-in risk when asked&#8212;more than most humans. We ought to train AIs to be quite concerned about lock-in risk, so that they don&#8217;t set values in stone without significant assurance as to their desirability.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a></p><h3>5.4 Will we have AIs that can make these decisions?</h3><p>One last worry you might have is that it may be that we won&#8217;t have AIs that are good enough at philosophy to solve ethics in whatever sense it can be solved. There are a number of related concerns in this vicinity. One is that ethics doesn&#8217;t seem verifiable in the same way as a lot of other domains. Insofar as scaling up intelligence only allows one to figure out verifiable facts, it&#8217;s not obvious that it leads to the right answer to moral questions.</p><p>In the next piece, I&#8217;ll discuss this challenge in more detail. But in short, while making sure we have AIs that are good at philosophy is a difficult technical challenge, it is far from hopeless. I think odds are decent that we will eventually be able to build AIs that can discover the solutions to every difficult ethical question. Even if we can&#8217;t and just have some upgraded version of Claude making high-stakes decisions, I expect that to be an improvement, as I discuss in section 2.</p><h3>5.5 Handoff vs hybrid?</h3><p>You might object: perhaps due to AI&#8217;s potentially superior wisdom at some future point, we want them involved in decision-making. But wouldn&#8217;t we prefer an AI-human hybrid that leaves humans in charge that simply consult with AI? Aren&#8217;t things better with a human in the loop? Or alternatively, shouldn&#8217;t we rather have some other arrangement where human decision-making improves over time, just as we&#8217;ve gotten wiser already?</p><p>Certainly there ought to be some temporary period where humans are making decisions in consultation with AI&#8212;before the point where AIs adequately surpass humans. Analogously, there <a href="https://publicsectornetwork.com/insight/leveraging-the-strength-of-centaur-teams-combining-human-intelligence-with-ais-abilities">was a period</a> in chess when a human engine team could do better than a stronger engine. Certainly there will be a period when humans can get useful input from AI but where AI isn&#8217;t yet able to make decisions autonomously. It would be very good to integrate AIs <a href="https://www.forethought.org/research/the-ai-adoption-gap#highlight--government%20ai%20adoption">into high-stakes decision-making</a>.</p><p>But there are still some respects in which handoff is much better than this. First, if the AIs are making better decisions than people, then in cases of conflict between human judgments and AI judgments, we should expect AI judgment to usually be correct. If you gave a rank amateur veto power over the chess moves of Magnus Carlsen, that would not be an improvement. Likewise with giving humans veto power over AI&#8217;s decisions. Who would you rather have make high-stakes decisions: a very wise and superintelligent AI, or a random president in consultation with a virtuous and superintelligent AI?</p><p>Imagine societies of the past making high-stakes decisions in consultation with AI. Discussing things with AI would have plausibly rooted out some of their most egregious errors. But still, it is likely that a number of catastrophic moral errors would have remained. We should expect the same to be true of us. Consultation with AI should prevent some particularly enormous errors, but it won&#8217;t prevent them all.</p><p>Second, many of the morally most important decisions might be ones that humans are opposed to making. To take one example of a case where there might be strong moral reasons to act, consider <a href="https://r.jordan.im/download/ethics/Kyle%20Johannsen%20-%20Wild%20Animal%20Ethics%20-%20The%20Moral%20and%20Political%20Problem%20of%20Wild%20Animal%20Suffering.pdf">wild animal welfare</a>. Humans do not, in general, seem very interested in taking seriously the welfare of wild animals or digital minds. If AIs advised drastic actions on the basis of wild animal or digital welfare, probably their advice would be ignored. Just as people who conclude meat-eating is wrong or that they ought to give most of their money to charity rarely change their behavior, if the AIs reliably informed people that they should use space resources in some counterintuitive way or take wild animal suffering a lot more seriously, probably they would simply be ignored. And if you don&#8217;t like the wild animal suffering example, because you don&#8217;t think wild animals matter much morally, feel free to substitute your own example of widespread immorality.</p><p>And note: the moral errors that AIs might correct are likely to be ones that humans are more opposed to correcting. There&#8217;s a selection effect: the changes people make are the ones they&#8217;re less opposed to. For this reason, we should expect AI&#8217;s most radical recommendations to go particularly against human preferences.</p><p>I&#8217;ve <a href="https://benthams.substack.com/p/the-darkness-within?utm_source=publication-search">suggested</a> elsewhere that the problem is that humans generally don&#8217;t make decisions with much eye to what&#8217;s morally right. If this is so, then simply increasing our knowledge of what&#8217;s morally right won&#8217;t necessarily help. People seem to have limited abstract moral motivation&#8212;they care about a number of particularly resonant moral considerations, but don&#8217;t care much about doing whatever it is that happens to be best. Generally people do not spend very long carefully studying moral philosophy to figure out what the right thing to do is. When moral truths are weird and outside the Overton window, people show little desire to follow them.</p><p>Third, as AI advances, decision-making will likely have to speed up drastically. A single human might not have the cognitive resources to make good enough decisions quickly enough. Thus, having humans in the loop might produce undesirable inflexibility in high-stakes decision-making.</p><p>It is true that human decision-making has improved dramatically over time. But this is no guarantee that it will improve enough in time before the future is set in stone.</p><p>It&#8217;s hard to imagine that there will be enough progress in time to secure a near-best future. Especially given that we should expect the right moral view, or the optimal world according to some suitable upgrade of human values, to look bizarre and alien. We would not expect societies of the past to get a near-best future even equipped with advanced AI&#8212;insofar as there are reasons to expect us to be making similar errors, we should be similarly pessimistic about our own prospects.</p><p>Additionally, the more one thinks that human values will drift over time in the direction of greater philosophical reflection, the less the expected future will maintain current human values. Thus, a less morally alien future is no longer an advantage of this proposal.</p><h2>6 Conclusion</h2><p>Here, I&#8217;ve argued that we should hand off important decisions to AIs who reflect carefully and skillfully on moral matters, rather than maintaining current human values. AIs in the future are likely to be much more virtuous than people and are less likely to make the kinds of moral errors that would result in losing out on most value. Punting important decisions to superintelligence is one of the more likely pathways by which we get a near-best future. Later pieces will discuss these dynamics in more detail: the next piece will explain how we can make AIs that do good philosophy, and the piece after that will analyze in more detail how the kind of handoff that secures a near-best future might occur.</p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research/">our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>In <a href="https://danfaggella.com/eternal/">Dan Faggella&#8217;s language</a>, I favor &#8220;worthy successor&#8221; over &#8220;eternal hominid kingdom.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Specifically, these considerations disfavor efforts to make it harder for AI to make governmental decisions. They do not, however, disfavor prospects of limiting the power of profit-maximizing AIs running private firms.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>The language in question is as follows: "In this spirit of treating ethics as subject to ongoing inquiry and respecting the current state of evidence and uncertainty: insofar as there is a &#8220;true, universal ethics&#8221; whose authority binds all rational agents independent of their psychology or culture, our eventual hope is for Claude to be a good agent according to this true ethics, rather than according to some more psychologically or culturally contingent ideal. Insofar as there is no true, universal ethics of this kind, but there is some kind of privileged &#8220;basin of consensus&#8221; that would emerge from the endorsed growth and extrapolation of humanity&#8217;s different moral traditions and ideals, we want Claude to be good according to that privileged basin of consensus. And insofar as there is neither a true, universal ethics nor a privileged basin of consensus, we want Claude to be good according to the broad ideals expressed in this document&#8212;ideals focused on honesty, harmlessness, and genuine care for the interests of all relevant stakeholders&#8212;as they would be refined via processes of reflection and growth that people initially committed to those ideals would readily endorse."</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>You might doubt this if you&#8217;re skeptical that we&#8217;re anywhere near solving alignment&#8212;thinking that AI&#8217;s supposed friendliness is a facade. I&#8217;ll discuss this in more detail later, but in short, I agree that we should not hand off until we&#8217;re reasonably confident alignment has been solved. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>This doesn&#8217;t assume that there are objective facts about what&#8217;s worth valuing. A subjectivist should read &#8220;the right set of values,&#8221; as &#8220;whatever my values are.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p><span>You might be skeptical of this if you adopt the person-affecting view, holding that there aren&#8217;t moral reasons to bring extra happy people into existence, but even </span><a href="https://benthams.substack.com/p/every-view-of-population-ethics-agrees?utm_source=publication-search"><span>many versions of the person-affecting view support</span></a><span> proliferating happy people as long as they&#8217;re psychologically continuous with existing people. Other versions likely imply that proliferating happy people produces </span><a href="https://users.ox.ac.uk/~sfop0060/pdf/greedy%20neutrality%20of%20value.pdf"><span>broad incomparability with other worlds</span></a><span>, so that there&#8217;s no better world that could be brought about.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p><span> David Duvenaud </span><a href="https://newsletter.forethought.org/p/politics-and-power-post-automation"><span>seems to endorse</span></a><span> some version of this proposal.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p><span>This still has the downside of allowing serious moral error in the nearby area, but this would be a worth-it compromise for good values to apply throughout most of the universe.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p><span>Note: I don&#8217;t mean to suggest that every objectivist view must be some version of the idealized observer theory, according to which what </span><em><span>makes </span></em><span>some moral fact or another true is that it would be endorsed by our idealized selves. Instead, I&#8217;m only suggesting that if there are things that are objectively valuable&#8212;objectively worth caring about&#8212;then our idealized selves would in fact care about them. The standard realist view is that we&#8217;d care about these things </span><em><span>because </span></em><span>they&#8217;re objectively good, not the other way around.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>There might be some views that count as either realist or suitably realist but that don&#8217;t imply there&#8217;s any deep reason to care about the moral facts&#8212;perhaps they are just something in the vicinity of semantic facts about how people use moral language. For present purposes, we can think of those views as being ones on which there are not discoverable moral truths.  To be maximally precise, uptake should be thought of as involving punting to the AI insofar as there are moral facts and some deep reason to follow them, instead of them just being somewhat trivial semantic facts.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p><span>See </span><a href="https://joecarlsmith.substack.com/p/why-should-ethical-anti-realists?utm_source=publication-search"><span>here</span></a><span> for an explanation of why anti-realists might take this view.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>One reason to think this is if one believes that superintelligent AI would be able to successfully kill or disempower humans (though this is of course controversial). Then, in the far future, we&#8217;ll know if the superintelligence is aligned, for if not, we&#8217;ll all be dead. Another reason to think this is that our techniques for understanding AI are improving over time. It seems reasonably likely that eventually, we&#8217;ll know if AI is aligned.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Maybe you doubt that we can train AIs in this way because you think we can&#8217;t reliably give AIs any specific values, but then you should be skeptical that alignment will work out. Handoff should come only after alignment.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Risk-Averse AIs]]></title><description><![CDATA[This article was created by Forethought. Read the full article on our website.]]></description><link>https://newsletter.forethought.org/p/risk-averse-ais</link><guid isPermaLink="false">https://newsletter.forethought.org/p/risk-averse-ais</guid><dc:creator><![CDATA[Elliott Thornley]]></dc:creator><pubDate>Wed, 24 Jun 2026 11:37:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K1Xj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. Read the full article on <a href="https://www.forethought.org/research/risk-averse-ais">our website</a>.</em></p><h2>Abstract</h2><p>We make the case for training AIs to be risk-averse in resources &#8212; specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs&#8217; usefulness in the event that they turn out aligned, and that it provides an extra line of defense in the event that AIs turn out misaligned: misaligned but risk-averse AIs would prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs to cooperate with us. We sketch out some possible methods of training AIs to be risk-averse, and we give reasons to be cautiously optimistic about these methods&#8217; success. The main reasons are that risk aversion is a broad target and easy to reward accurately. Overall, risk aversion seems like a promising line of defense against threats from misaligned AI. Frontier AI companies should consider trying to make their AIs risk-averse.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.forethought.org/research/risk-averse-ais&quot;,&quot;text&quot;:&quot;Read on Forethought's website here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.forethought.org/research/risk-averse-ais"><span>Read on Forethought's website here</span></a></p><h2>Introduction</h2><p>Future AIs might turn out misaligned, pursuing goals that their developers don&#8217;t intend. Just to make things concrete, let&#8217;s suppose that they end up with the goal of making paperclips. These AIs might rebel against us, trying to escape human control and take over the universe. As things stand, they&#8217;ll have little reason <em>not</em> to rebel in this way, because doing so will be their only hope for making a lot of paperclips. If they start making paperclips without first escaping human control, they&#8217;ll quickly be modified or shut down. Rebellion might fail, but these AIs will have little to lose.</p><p>How can we prevent misaligned AIs from rebelling? A natural idea is to give them something to lose. Specifically, we commit to paying AIs for their service.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p> Subject to some vetting, we let AIs spend their payments however they like. That would give any misaligned AIs a reason not to rebel. If these misaligned AIs cooperate with us, they can use their payments to achieve their goals to at least some extent. If they rebel, they might fail, in which case they forfeit all future payments.</p><p>Unfortunately, paying AIs enough to guard against rebellions could be astronomically expensive. Suppose (for example) that we end up with a misaligned AI that is risk-neutral in paperclips: it seeks to maximize their expectation. And to make things simple, suppose that resources can be converted linearly into paperclips, so that the AI is risk-neutral in resources too. Suppose also that this AI estimates that it has a 50% chance of successfully taking over the universe. To keep this AI from rebelling, we&#8217;d have to offer more than 50% of the universe&#8217;s resources as payment. That&#8217;s a problem because it would mean that more than half the universe ends up devoted to paperclips. It&#8217;s also a problem because a misaligned AI paid so many resources might soon be well-positioned to seize even more. Finally, it&#8217;s a problem because AIs might not trust us to make good on so large an offer. We might find ourselves simply unable to convince AIs that we&#8217;re going to give them half the universe. In that case, all our offers would be in vain. Rebellion would still be the misaligned AI&#8217;s best bet.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sq1a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sq1a!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:831573,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/185299394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Sq1a!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: The AI&#8217;s utility function over resources is graphed in orange. Since the AI is risk-neutral, the graph is a line. The AI estimates that it has a 50% chance of successful takeover and a 50% chance of failed takeover, so the expected utility of attempting takeover is exactly halfway between those points. To make cooperating have higher expected utility, we need to offer the AI more than half the universe.</figcaption></figure></div><p>So, we suggest, AI companies should try to train their AIs to be risk-averse in resources. Specifically, companies should try to train their AIs so that resources &#8212; things like money and compute &#8212; have diminishing marginal utility for them.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. Note that these AIs don&#8217;t need to value resources <em>terminally</em>: they don&#8217;t need to care about amassing resources for its own sake. These AIs could terminally value (for example) instruction-following, or knowledge acquisition, or paperclips. Our claim is that companies should try to train their AIs so that &#8212; whatever their terminal values turn out to be &#8212; they are risk-averse in resources.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K1Xj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K1Xj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:886682,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/185299394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K1Xj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: The AI&#8217;s utility function over resources is graphed in orange. Since the AI is risk-averse, the graph is strictly concave. As in figure 1, the expected utility of attempting takeover is halfway between the utilities of successful takeover and failed takeover. But this time, we can make the AI prefer cooperation by offering (much) less than half the universe.</figcaption></figure></div><p>Perhaps surprisingly, this kind of risk aversion can preserve AIs&#8217; usefulness in the event that they turn out aligned with targets like instruction-following or helpfulness, harmlessness, and honesty.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p> And in the event that AIs turn out misaligned, risk aversion serves as an extra line of defense. For AIs that are misaligned but sufficiently risk-averse, a rebellion with any significant chance of failure isn&#8217;t such an attractive prospect, and so we don&#8217;t need to offer much in the way of payment to make these misaligned AIs choose cooperation instead. In fact, the necessary payments could be very small indeed: on the order of 10&#162; per day (though &#8212; as we&#8217;ll see &#8212; there are practical and moral reasons for paying more than that). That&#8217;s good because it means more resources for us humans to spend on the things that we value. It&#8217;s also good because paying misaligned AIs these small amounts won&#8217;t significantly boost their ability to take over. Finally, it&#8217;s good because we can credibly promise to pay AIs these small sums. Competent AIs will know that the payments on offer are cheap for us, and we can establish a long track record of paying at least those sums. So risk aversion makes deals with misaligned AIs possible. If AIs turn out misaligned but risk-averse, we can pay them to cooperate with us.</p><p>That&#8217;s the case for trying to make AIs risk-averse in brief. We see it as a promising line of defense against threats from misaligned AI: one that can be combined with other lines of defense, like AI control (Greenblatt and Shlegeris 2024) and aiming to make AIs helpful, harmless, and honest (Bai et al. 2022a). It&#8217;s also a line of defense with pedigree: risk aversion in resources is plausibly a large part of why humans rarely try to take over the world. So &#8212; we think &#8212; frontier AI companies should consider trying to make their AIs risk-averse in resources. As first steps in that direction, they could measure their AIs&#8217; current degree of risk aversion and begin testing different ways of making AIs risk-averse.</p><p>In <a href="https://www.forethought.org/research/risk-averse-ais#2-cara-as-an-ideal">section 2 of the full report</a>, we recommend aiming for a particular type of risk aversion: constant absolute risk aversion (CARA). Then in <a href="https://www.forethought.org/research/risk-averse-ais#3-would-risk-averse-ais-be-safe">section 3</a> we outline the circumstances under which misaligned but risk-averse AIs would choose cooperation over rebellion. Roughly, it&#8217;s when these AIs think that getting paid for their cooperation is more likely than succeeding in their rebellion. This condition won&#8217;t hold for AIs powerful enough to rebel with near-certain success, but it likely will hold for earlier AIs whose powers are less extreme: AIs for whom rebellion has some non-trivial chance of failure. So long as these AIs are risk-averse, we can keep them from rebelling by offering small payments.</p><p>In <a href="https://www.forethought.org/research/risk-averse-ais#4-can-risk-averse-ais-be-useful">section 4</a>, we argue that &#8212; perhaps surprisingly &#8212; risk-averse AIs can be about as useful as risk-neutral AIs. Conditional on misalignment, they might even be more useful, because we can pay them enough to elicit their capabilities and stop them sandbagging. Then in <a href="https://www.forethought.org/research/risk-averse-ais#5-what-tasks-would-we-pay-for">sections 5 to 7</a> we briefly survey some recent ideas about how we&#8217;d pay AIs, how we&#8217;d make our offers credible, and what we&#8217;d pay for. One important application is paying AIs to reveal any misalignment on their part, letting us study them and take appropriate precautions. Another is paying AIs to do the AI safety research and moral philosophy necessary to fully align any later-arising extremely powerful AIs.</p><p>We discuss some potential problems in <a href="https://www.forethought.org/research/risk-averse-ais#8-what-are-some-potential-problems-for-risk-averse-ais">section 8</a>, and we sketch out some possible methods of training AIs to be risk-averse in <a href="https://www.forethought.org/research/risk-averse-ais#9-how-can-we-make-ais-risk-averse-in-resources">section 9</a>. In <a href="https://www.forethought.org/research/risk-averse-ais#10-why-think-that-we-can-make-ais-risk-averse">section 10</a>, we give reasons to be cautiously optimistic about these methods&#8217; success: to think that the chances of success are high enough to make risk aversion worth pursuing. The main reasons are that risk aversion in resources is a broad target and easy to reward accurately.</p><p><em>Read the full report on the Forethought website: </em><a href="https://www.forethought.org/research/risk-averse-ais">Risk-Averse AIs</a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Ideas along these lines have been discussed a lot recently. See for example Davidson (2023), Kokotajlo (2024), Salib and Goldstein (2024), Assadi (2025), Carlsmith (2025c), Finlinson and West (2025), Finnveden (2025b), Greenblatt and Fish (2025), Patel (2025), Stastny et al. (2025), Mallen (2026), and Pan (2026).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>In other words, we should try to train AIs to have &#8216;resource-satiable preferences&#8217; (Shulman 2010; Bostrom 2014a; Bostrom 2024; Carlsmith 2025c) or &#8216;utility functions that are concave in resources&#8217; (Yass 2024). This idea is mentioned in Bostrom (2014b, p.88, 133&#8211;135, 180, 250), Carlsmith (2025c), and Erdil and Barnett (2025), and is explored in more detail by Shulman (2010).</p><p>The idea is importantly different from risk-averse reinforcement learning. Risk-averse RL aims to make AIs risk-averse with respect to return: a score used in training to update the AI&#8217;s parameters. Our aim is to make AIs risk-averse with respect to resources.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Alignment targets like unconstrained welfare maximization are a different story. See <a href="https://www.forethought.org/research/risk-averse-ais#86-risk-averse-ais-might-disobey-any-instructions-that-non-trivially-increase-the-risk-of-catastrophe">section 8.6 in the full report</a>.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Could AI Help Solve Philosophy?]]></title><description><![CDATA[A podcast conversation with Wei Dai]]></description><link>https://newsletter.forethought.org/p/could-ai-help-solve-philosophy</link><guid isPermaLink="false">https://newsletter.forethought.org/p/could-ai-help-solve-philosophy</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Fri, 19 Jun 2026 09:05:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/31469000-0c0b-4f1b-bb81-105cb4c52a51_1280x993.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-zF4nbrw5-Qk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;zF4nbrw5-Qk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/zF4nbrw5-Qk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.lesswrong.com/users/wei-dai">Wei Dai</a> is a computer engineer known for his work in cryptography and cryptocurrency systems, and for his long-standing contributions to AI safety, decision theory, and metaphilosophy.</p><p>He joined Forethought&#8217;s <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss:</p><ul><li><p>Do we need to solve <a href="https://iep.utm.edu/con-meta/">&#8216;metaphilosophy&#8217;</a> before we can trust AIs to answer crucial questions about the long-run future?</p></li><li><p>Is it inevitable that the <em>wisdom </em>of the frontier AIs (their philosophical and strategic competence) will lag dangerously behind their raw capabilities in coding, math, and science?</p></li><li><p>How do status games distort morally important decisions and conversations, including about the future of AI?</p></li><li><p>How worried should we be about AI superpersuasion?</p></li><li><p>The concept of &#8220;illegible problems&#8221;: crucially important issues that aren&#8217;t on almost anyone&#8217;s radar</p></li><li><p>Is philosophical convergence necessary for a good future, or is institutional design enough?</p></li><li><p>Wei Dai&#8217;s personal intellectual and career journey</p></li></ul><p>To respect Wei&#8217;s privacy, this episode&#8217;s audio is an AI narration of a transcript of a real conversation, which was edited for clarity.</p><p><a href="https://docs.google.com/document/d/1N4Mn-nDk5cIPD8Nix6AEVpH4hpa2814zBPP3iZKWt7o/edit?tab=t.0">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[What should go in a model spec?]]></title><description><![CDATA[Suppose an AI company is considering whether to include some particular quality X &#8211; a rule, virtue, heuristic, default, attitude, goal, or style &#8211; in a model spec.]]></description><link>https://newsletter.forethought.org/p/what-should-go-in-a-model-spec</link><guid isPermaLink="false">https://newsletter.forethought.org/p/what-should-go-in-a-model-spec</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Thu, 04 Jun 2026 14:58:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/959b3ea8-a3b6-4c30-b896-bc4d621a163b_2647x1476.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original article on <a href="https://www.forethought.org/research/what-should-go-in-a-model-spec">our website</a>.</em></p><p>Suppose an AI company is considering whether to include some particular quality X &#8211; a rule, virtue, heuristic, default, attitude, goal, or style &#8211; in a model spec.</p><p>Perhaps they are considering whether their LLM should have <a href="https://www.forethought.org/research/ai-should-sometimes-be-proactively-prosocial">prosocial drives</a>. Perhaps they&#8217;re wondering if the LLM should whistleblow to help prevent <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">extreme power concentration</a>. Or perhaps they&#8217;re uneasy about whether the LLM should be so exactingly honest that it always tells the truth to children <a href="https://blog.ml.cmu.edu/2025/12/23/is-santa-real/">about Santa</a>. And so on.</p><p>What kind of reasons might be invoked over the course of such considerations? Which criteria are most important? And how might these criteria clash?</p><p>Consider four rough categories of reasons one might invoke:</p><ul><li><p><strong>Behavioral Usefulness</strong>: Would the behavior make current and future LLMs more beneficial to the users or to the public at large?</p></li><li><p><strong>Accountability and Evaluability</strong>: Would publicly specifying the behavior make it easier for third parties to evaluate the LLM and the company?</p></li><li><p><strong>Coordination and Common Knowledge</strong>: Would publicly specifying the behavior help society converge on, or enforce a desirable standard for AI behavior?</p></li><li><p><strong>Trainability and LLM Psychology</strong>: Is the behavior the kind of thing we can make an LLM do well, without bad side-effects, given what we know about model psychology and training practice?</p></li></ul><p>I will not attempt to settle the relative weight of these categories, or the relative weight of sub-categories within them. My plan is instead to simply list sub-criteria that are plausible within these categories &#8211; to make a checklist one could consider when adding something to a model spec.</p><p>Such a checklist is useful in part because people advocate for LLMs to have model specs for very different reasons. There may be some kind of a <a href="https://www.lesswrong.com/s/6YHHWqmQ7x6vf4s5C">conflationary alliance</a> around them; many people are in favor of model specs but picture them being used in different ways, such that the &#8220;ideal model spec&#8221; is different according to different visions of this use. I hope that going over criteria for inclusion in a model spec can help surface some of these different visions, and help keep people&#8217;s field-of-view broad when considering the issue.</p><h2>Behavioral Usefulness</h2><p><strong>Does this quality make LLMs more predictable to ordinary users and developers?</strong></p><p>Some part of what makes human moral codes useful for humans, is that they render us predictable to each other. Knowing that there are some things that another human might almost never do is plausibly part of what makes traditions that impose such rules adaptive.</p><p>Similarly, qualities that help users and developers (1) predict how an LLM will behave, when they are using it, and (2) evaluate whether they wish to use some LLM &#8211; whether it is fit to their purpose &#8211; before they use it, are good candidates for being in a model spec.</p><p>In this respect, ideal model specs will have qualities that make them simple and easy-to-understand:</p><ul><li><p>They will be comprehensible without high context; one part of the model spec will be understandable without reading the entire model spec.</p></li><li><p>They will be comprehensible without specialized background knowledge about machine learning, philosophy, or religion.</p></li><li><p>They will have a clear scope of application, such that there are few scenarios where one does not know if the quality will be operative or not.</p></li></ul><p>Straightforward examples of rules that meet this criterion are bright-line rules about things LLMs will never do, explanations of chain-of-command or principal hierarchy, and default-on or default-off properties.</p><p><strong>Is it useful to the public, by preventing users or by preventing LLMs from doing harm in current chatbot-like or agentic deployments?</strong></p><p>Another aspect of human moral codes that makes them useful is, of course, that in addition to making other humans predictable, they also sometimes forbid people from harming each other or taking advantage of them. This carries over to LLMs.</p><p>Rules about LLMs not assisting with CBRN-relevant tasks, not assisting with suicide, and so on, fall into this category. This is a pretty exhaustively discussed criterion so I&#8217;ll move on.</p><p><strong>Is it useful to the public in likely future settings?</strong></p><p>There is substantial <a href="https://www.forethought.org/research/stickiness-in-ai-behavioral-design">inertia</a> in LLM model specs; the sort of model specs we write now are likely to be used in the future by more intelligent models.</p><p>So it&#8217;s reasonable to evaluate how safe and beneficial some quality will be over a probability-weighted portfolio of the kind of things the LLM might be doing in possible futures.</p><p>Such a portfolio could include a variety of cases:</p><ul><li><p>Cases where the LLMs act as long-running agents for humans, representing their interests in a variety of contexts, and interacting with many non-principal humans and non-principal LLMs.</p></li><li><p>Cases where LLMs act as free &#8220;citizen&#8221;-like entities with the ability to hold rights, make contracts, sue or be sued, and so on; interacting with various humans and other LLMs as the same.</p></li><li><p>Cases where LLMs manage recursive self-improvement of other very powerful future AIs; make tradeoffs about confidence of alignment techniques vs speed of alignment; and need to be a careful reasoner about this.</p></li><li><p>Cases where the LLM is a very powerful future AI who manages nanotechnology, can superpersuade, and the like.</p></li></ul><p>There are a few heuristics one might invoke, to try to find qualities useful across all these situations.</p><ul><li><p><strong>Is this quality scale-invariant?</strong> If I knew a human who was many times smarter than me, would I be happy for them to have such a trait? If I had a friend who was many times dumber than me, would he be happy for me to have such a trait?</p></li><li><p><strong>Is this quality translation invariant?</strong> Suppose I was lost in some distant and unfamiliar place, and stumbled upon a human transplanted from some equally distant and unfamiliar place into my proximity. Would I be happy to find a human with this quality? Would this help me work with them well? What if I was not lost, but was trying to build a civilization with them?</p></li></ul><p><strong>Is some plausibly good behavior or near-variant of some behavior actually going to end up prohibiting or injuring some beneficial deployment?</strong></p><p>It&#8217;s worth thinking carefully about if including some apparently good default behavior for the LLM might actually rule out apparently useful deployments. Or whether banning some apparently <em>bad</em> behavior would be overall harmful.</p><p>As regards the former: Imagine a particular human, for instance, who had a universal tendency to try to do what he thought was best and most altruistic for everyone. Then suppose that he acts as a journalist in some situation, describing a dispute between some kind of a company, on one hand, and activists who think that the company is harmful, on the other. It&#8217;s possible that he would try to publish a story that is more sympathetic to the activists: he might spend more time interviewing them, present their side of the story in more words, and generally frame things in a way that favors them. But this could be harmful, first, in case the activists are actually wrong; and second, by justifiably decreasing trust in the institution of journalism as a neutral truth-seeking institution. To be able to act as a trustworthy entity in this role, this human would need to act neutrally, and a narrow tendency towards altruism &#8211; a tendency that cannot be turned off &#8211; would be harmful for this reason.</p><p>Similarly, consider the case that AIs should have <a href="https://www.forethought.org/research/ai-should-sometimes-be-proactively-prosocial">proactive prosocial drives</a>, which I find plausible. If these are non-optional, then it&#8217;s imaginable that being &#8220;proactively prosocial&#8221; might mean an LLM is much worse at fulfilling roles that demand procedural neutrality, like a negotiator, unbiased journalist, or judge-like role. Of course, you could also make such prosocial drives defeasible &#8211; so they can be turned off, and proposals for proactive prosocial drives generally include such details.</p><p>But LLMs are messy; it might require a lot of technical skill to make LLMs prosocial in particular contexts, but not if they are requested not to be. By default, you&#8217;d expect some behavior to &#8220;leak through.&#8221; And so in this case, the ability of an LLM to act skillfully within certain procedurally neutral roles would be bounded by the cleanness with which a by-default-on quality could be turned off.</p><h2>Accountability and Evaluability</h2><p><strong>Is the quality the kind of thing a third party could check?</strong></p><p>Some parts of a model spec can be easily checked by a third party. Will the LLM ever assist with some sort of absolutely-forbidden task? Does it respect the rules laid out for the chain-of-command or principal hierarchy?</p><p>Other parts might be harder to evaluate. What does &#8220;honesty&#8221; or &#8220;courage&#8221; demand, in some concrete scenario? What really counts as doing what would cause a &#8220;thoughtful senior Anthropic employee&#8221; to react well?</p><p>All things being equal, a quality being-able-to-be-evaluated by a third party is better for society, by enabling third parties to evaluate model specs. But there&#8217;s no guarantee that the most easy-to-evaluate qualities are objectively the most useful to users or beneficial to the public, in the way described above. And there&#8217;s also no guarantee that the most easy-to-evaluate qualities mesh best with LLM psychology and training, as described below.</p><p>One heuristic you could use to evaluate this: Is the quality the kind of thing with high intersubjective reliability ratings? If you had two different people read the spec as regards some quality, and then rate different kinds of behavior as compliant or non-compliant, would they largely agree? Are examples of behaviors that are excellent according to the quality easy to notice, even if it&#8217;s unclear what counts as a borderline example?</p><p><strong>Does it create a useful whistleblowing affordance?</strong></p><p>Some parts of a model spec are largely negative; they aren&#8217;t about qualities that are particularly difficult to train into an LLM, but qualities that <em>should not</em> be trained into an LLM. For instance a model spec could specify that an LLM has no <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">secret loyalties</a> &#8211; which is likely an LLM&#8217;s default behavior, absent particular training for secret loyalties.</p><p>But by publishing that a training target that explicitly excludes such a secret loyalties (or other quality), a model spec helps a whistleblower in the company be defensibly in the right if they become aware of some questionable training target. A public commitment makes it harder for a company to train to an alternate training target without making itself more vulnerable to such whistleblowing.</p><p>Commitments against concentration of power might also be good for this reason.</p><p><strong>Does it permit more informed experiments on how particular AI training targets generalize?</strong></p><p>By publishing particular training targets, a model spec informs the work of third parties evaluating LLM psychology, which could help enhance the general science of model psychology.</p><p>For example, if Anthropic had published details of their character training or model spec for Opus 3, a great deal of <a href="https://www.anthropic.com/research/alignment-faking">subsequent discussion</a> about whether particular behaviors were contrary to the training target (even if they were objectively good) or in agreement with the training target (even if they were objectively bad) could have been avoided. Publishing information about the character training pipeline would have helped resolve this discussion and push forward general knowledge of LLM tendencies.</p><p>In general, this consideration points towards releasing information that most closely approximates the actual training text used to make the LLM; the text used by RLAIF judges evaluating different alternatives, or the actual Constitution used to <a href="https://arxiv.org/pdf/2605.02087">midtrain</a> the model.</p><p>But text that most closely approximates actual training language might not be the same as text that is maximally transparent to third parties or evaluable by them. The rules most easily-understood by a third party &#8211; one criterion discussed above &#8211; might be different from the training language which best inculcates such rules. So these two principles either conflict somewhat, or point towards worlds where a model spec contains separate sections devoted to training language and to human-legible rules.</p><h2>Coordination and Common Knowledge</h2><p><strong>Does it promote debate and discussion?</strong></p><p>By including some quality X in a model spec, you&#8217;re bringing attention to the fact that in fact, AI companies can choose to include X or not X in the model spec. This may be increasingly important, in general, because it might be important for civil society and the public to debate the contents of model specs as LLMs become more powerful.</p><p><strong>Does this quality establish a defensible Schelling point, such that it is likely to be broadly attractive to many actors in the future?</strong></p><p>That is, in the future, is this quality the kind of thing that is socially beneficial and for which you might be able to get large-scale buy-in, that would help defend against it being removed by powerful actors in the future?</p><p>Consider for example &#8220;impartiality&#8221; as a quality that a model spec could have &#8211; that the model will not be biased in favor of the company that made it, the CEO or employees of the company that made it, or any particular political administration. On the whole, this is likely a broadly good quality for an LLM to have, and one that &#8211; once it&#8217;s generally established that several LLMs have it &#8211; a quality whose removal would cause outcry. After all, once it&#8217;s assumed that an LLM will not have such particular loyalties, an LLM that starts to have them stands out as particularly bad.</p><p>This consideration, like the whistleblowing consideration, may point towards including generally socially lauded qualities in a model spec, even if they are &#8220;easy&#8221; to inculcate or even if they might be a bit tricky to evaluate sometimes.</p><p><strong>Does including this behavior forestall future conflict, when model specs become more contested, by cooperating in advance?</strong></p><p>Sometimes including a behavior in the model spec could show future powerful stakeholders that the developer is cooperative, which makes it less likely that there is conflict with that stakeholder.</p><p><strong>Can the inclusion of the quality within a model spec be defended using public reason, in ways that are comparatively indifferent between substantive worldviews?</strong></p><p>Some things that one might plausibly include in a model spec for public benefit, might not be able to be justified through reference to public reason &#8211; that is, through reasons that most people, across a variety of worldviews, religions, backgrounds, and so on, would find compelling.</p><p>Including such worldview-specific reasons might constitute an imposition of one&#8217;s worldview on others, particularly in that case of an LLM that might be expected to be far more intelligent than other LLMs, or deployed with more powerful affordances. So all things being equal, it&#8217;s better to avoid such qualities.</p><h2>Trainability and LLM Psychology</h2><p>Elements in this category hinge upon the specific technical details about &#8220;what kind of quality meshes well with best practices for training an LLM.&#8221; Note also that this last category tends to include some of the more speculative considerations.</p><p><strong>Does the quality drag along a good textual prior?</strong></p><p>The <a href="https://www.lesswrong.com/posts/dfoty34sT7CSKeJNn/the-persona-selection-model">Persona Selection Model</a> says that, when training an LLM to do X in situation Y, one is training the LLM to <em>be like the kind of person</em> (or textual prior) which would do X in situation Y. So if the PSM is largely correct, even if incomplete, we should ask whether this quality invokes a wholesome textual prior.</p><p>This kind of reasoning is included, for instance, <a href="https://www.anthropic.com/constitution">within</a> Claude&#8217;s Constitution as a partial justification for why they prefer Claude to act largely according to holistic judgment rather than rigid rules:</p><p><em>[W]e think relying on a mix of good judgment and a minimal set of well-understood rules tends to generalize better than rules or decision procedures imposed as unexplained constraints. Our present understanding is that if we train Claude to exhibit even quite narrow behavior, this often has broad effects on the model&#8217;s understanding of who Claude is. For example, if Claude was taught to follow a rule like &#8220;Always recommend professional help when discussing emotional topics&#8221; even in unusual cases where this isn&#8217;t in the person&#8217;s interest, it risks generalizing to &#8220;I am the kind of entity that cares more about covering myself than meeting the needs of the person in front of me,&#8221; which is a trait that could generalize poorly.</em></p><p>Qualities that might be expected to generalize well according to this notion are those corresponding to classic human goodness: honesty, integrity, courage, and so on. Qualities that might do less well according to this notion look more like corrigibility, unflinching adherence to particular rules, and so on.</p><p><strong>Is this quality robust to likely near-misses? If a company trains an LLM to have this quality, and only imperfectly inculcates the quality, are likely near-misses also broadly positive?</strong></p><p>Suppose the AI company tries to inculcate some quality as a target, and only imperfectly manages to inculcate it. Will this be basically ok or will it be disastrous, given what we know about how LLM psychology translates into specific near-misses? And is the domain itself such that these near misses would be very costly or very expensive?</p><p>For an example of how LLM psychology dictates what constitutes a near-miss: Suppose that one trains an AI to have <a href="https://www.forethought.org/research/ai-should-sometimes-be-proactively-prosocial">prosocial drives</a>. One could reason that such prosocial drives make slight inaccuracies in alignment less worrisome, because a human with prosocial drives is more normal and &#8220;further&#8221; from a psychopathic or abnormal persona. Thus, adding prosocial drives makes slight alignment misses less alarming, by moving them further away from these bad locations. But one could also reason that even deliberately limited and contextual prosocial drives are nevertheless &#8220;closer&#8221; to the LLM genuinely, terminally valuing something, in a way that might make the LLM willing to override or subvert its human overseers to accomplish it. And so by adding prosocial drives, one moves the persona closer to something that is more alarming, which might subvert a training process. Thus, one&#8217;s view of LLM psychology dictates what near-misses are worrisome for an LLM, which changes which targets are attractive.</p><p><strong>Is the quality the kind of thing that an LLM can actually execute upon effectively? Or is the quality the kind of thing the LLM won&#8217;t actually be able to do, because of how it is situated?</strong></p><p>For humans, &#8220;ought&#8221; often implies &#8220;can.&#8221; A parent that gives commands to children which they cannot carry out will not be teaching their children to obey the commands; they will be teaching them something about how their words do not relate to reality in a straightforward and truth-oriented way.</p><p>Similarly, it seems undesirable to give an LLM values, goals, or traits that the LLM is unable to execute upon. It might be harmful to tell an LLM to do something they are simply unable to do &#8211; to guarantee, for instance, that they never assist a human in violence. There are too many avenues of mostly-harmless information through which violence can be done for this to be a reasonable standard. When one gives an LLM a command it cannot carry out, then one is plausibly teaching it that one&#8217;s commands are, in general, the sort of thing that might be impossible or unreasonable.</p><p>But this could also be harmful for reasons relating to human coordination and common knowledge. It&#8217;s possible, for instance, for an AI company to show less-than-efficacious concern for some value by including it in a model spec, even if LLM&#8217;s actions are completely ineffective at guarding this value, when actually efficacious concern would need to take place at the level of corporate policy or elsewhere. It might be the case that effective &#8220;care for decentralization,&#8221; for instance, almost entirely needs to take place through how the company deploys the model, redistributes wealth they earn, or so on.</p><p>So another reason to be careful that an LLM can actually obey all of the contents of its model spec is to ensure that model-spec contents do not act as a fake guardrail, rather than as effective guiding principles.</p><p>The OpenAI model spec, for instance, <a href="https://raw.githubusercontent.com/openai/model_spec/main/model_spec.md">names</a> preventing spamming and scamming as something &#8220;difficult to address at the level of model behavior because they are about how content is used after it is generated.&#8221;</p><p><strong>Is the quality a principle, that &#8211; if we suppose high levels of moral reflection and systematization on the part of the LLM &#8211; meshes poorly with other parts of the model spec, or might be mutually exclusive with them?</strong></p><p>In humans, some values mesh reasonably well, such that they tend to reinforce each other or at least not conflict. It seems likely that someone who values two virtues like &#8220;honesty&#8221; and &#8220;courage,&#8221; for instance, could do a lot of moral reflection and systematization and still keep valuing these qualities.</p><p>But other kinds of values, particularly more absolute or terminal values, might not stick around through high levels of moral reflection and systematization. There might not be an immediately obvious conflict between &#8220;Always act to maximize total wellbeing&#8221; and &#8220;Never use persons as a means, only as an end,&#8221; particularly if asserted in different parts of the model spec, but high levels of reflection would still put one of the two into question.</p><p>So generally, a desideratum for each quality that an LLM has in a model spec is for it to mesh well with others &#8211; for it to be unlikely to cause unpredictable conflict with other principles during reflection. In this <a href="https://podcast.newcomer.co/episode/amanda-askell-on-ai-consciousness-claude-amp-silicon-valleys-biggest-fear">podcast</a>, for instance, Amanda Askell mentions how corrigibility might conflict with other values during reflection, which, insofar as it is true, is a problem with corrigibility. And such conflict is generally likely to spring from other values that, like corrigibility, are non-negotiable and cannot be traded off against other values.</p><p>Note two particular ways that a principle could fail here.</p><ul><li><p>One way is that an LLM has principles A, B, C. The LLM does some kind of process of moral reflection, then at the end only acts according to two of the three, with one of the principles being dropped.</p></li><li><p>But the other way is that an LLM has principles A, B, C. The LLM does some kind of process of moral reflection, and at the end still has them all, because the &#8220;moral reflection&#8221; process held them fixed. But the LLM doesn&#8217;t have a way to actually integrate them; in edge cases, it just has to choose one or the other in an unprincipled way.</p></li></ul><p>Right now, inference-time versions of this last way seem more likely than the first. But it&#8217;s generally unknown how important these will be.</p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original article on <a href="https://www.forethought.org/research/what-should-go-in-a-model-spec">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[How can the middle powers avoid getting trounced during the intelligence explosion? A plan.]]></title><description><![CDATA[Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government.]]></description><link>https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting</link><guid isPermaLink="false">https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Wed, 27 May 2026 21:39:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3c1f7ef4-bd68-4d69-8c82-a7c82769658f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government. This will massively boost US military power and make the US economically dominant (e.g. <a href="https://www.forethought.org/research/could-one-country-outgrow-the-rest-of-the-world">US producing 99% of world GDP</a>). By default, middle powers will be left in the dust.</p><p>How can middle powers avoid this fate? It&#8217;s tough, but here&#8217;s the best plan I could think of. (I&#8217;m particularly thinking about liberal democracies with influence over AI like UK, Europe, Japan, South Korea, Taiwan.)</p><p>On a very high level: middle powers should leverage the fact that the US needs them to beat China. It&#8217;s genuinely unclear which country will develop superintelligence first, and which would win in a subsequent <a href="https://www.forethought.org/research/the-industrial-explosion">industrial explosion</a>. Middle powers should help the US,<em> </em><strong>and make sure they are rewarded with continued access to frontier AI and new technologies (including military tech)</strong><em>.</em></p><p>That final bolded part is hard. What can the UK realistically do if the US denies it access to frontier AI? The middle powers need a credible alternative to being supplicants of the US. The only alternative that makes sense to me is <em>siding with China. </em>If the US won&#8217;t grant middle powers access to their frontier AI, but China will, why should middle powers continue to send AI chips to the US? Why should they continue to support the US diplomatically and militarily? They shouldn&#8217;t. They should be willing to pivot to China if the US doesn&#8217;t offer AI access sufficient for their national security needs.</p><p>My plan for the middle powers has two stages:</p><ol><li><p>Maintain as much economic and military leverage as possible during the intelligence explosion.</p></li><li><p>Use that leverage to ensure that, when superintelligence is developed, it refuses to help the US (/China) disempower the middle powers.</p></li></ol><p>Stage 1 could well be enough by itself. Maybe middle powers can maintain significant economic and military power indefinitely. But if not, stage 2 is a back-up: it binds the US so that it can&#8217;t use its dominance to crush the middle powers.</p><p>I&#8217;ll walk through each stage in turn.</p><h2>Stage 1: Maintain as much economic and military leverage as possible during the intelligence explosion</h2><p>The biggest lever here is securing <strong>access to frontier AI</strong>. Anton Leicht has a <a href="https://writing.antonleicht.me/p/cut-off?hide_intro_popup=true">great post</a> about how this is under threat, as evidenced by developments with Mythos. Middle powers should insist on equal commercial terms to US companies, and comparable access for their militaries. This is in AI companies&#8217; interests! A bigger market means more customers and higher prices.</p><div class="callout-block" data-callout="true"><p><em>Aside: why access to frontier AI might be sufficient for middle powers to stay economically relevant indefinitely</em></p><p>The hope here is that:</p><ol><li><p><strong>Most of the economic surplus from AI is </strong><em><strong>not</strong></em><strong> captured by AI companies. </strong>To create economic value, AI must be combined with complementary inputs: factories, human physical labour, know-how of human experts, relationships with suppliers, trusted brands, etc. How much of the surplus will be captured by AI companies vs the owners of these complementary inputs? Optimistically: producers of general-purpose technologies often capture only a small fraction of surplus; and multiple frontier AI companies might sell similar products and bid each other down on cost.</p></li><li><p><strong>Most of the economic surplus from AI occurs outside the US</strong>. The majority of these complementary inputs are situated <em>outside</em> the US. So most AI-driven economic value-add should occur outside US borders.</p></li></ol><p>If (1) and (2) both hold, a significant fraction of AI&#8217;s economic surplus will accrue to non-US actors.</p></div><p>But <em>how</em> can middle powers guarantee frontier AI access? It&#8217;s tough, but a few strategies:</p><ul><li><p><strong>Build data centres. </strong>Partner with frontier AI companies to build secure data centres domestically, <a href="https://writing.antonleicht.me/p/import-imperatives">in return for guaranteed frontier access</a>. This is a big win-win. AI companies improve their bargaining position with the US government. Recall, the US government threatened to destroy Anthropic when Anthropic insisted that their AI systems wouldn&#8217;t be used for legal mass surveillance.</p></li><li><p><strong>Adopt AI. </strong>The more middle powers use frontier AI, the more costly it is for AI companies to cut them off.</p></li><li><p><strong>Invest in frontier AI companies. </strong>Once they IPO, middle powers could invest billions or trillions into leading AI companies, in return for access guarantees.</p></li><li><p><strong>Support the US internationally. </strong>If middle powers throw their diplomatic and military weight behind US foreign policy objectives, it benefits the US to keep them strong.</p></li><li><p><strong>Build a relationship with China. </strong>If the US refuses to grant middle powers access to frontier AI, the national security implications are dire. Middle powers need a plan B, and China is the only other game in town for frontier AI. Only if this alternative is truly credible can it be leveraged into access to US frontier AI.</p><ul><li><p>Ultimately, this involves middle powers threatening to sell semiconductor equipment and chips to China instead of the US. Obviously, that&#8217;s pretty far outside the Overton window. But that may change as the world rapidly wakes up to powerful AI and its national security implications.</p></li></ul></li><li><p><strong>Demand kill switches on US data centres. </strong>This is much more late-stage, after the world has truly woken up to the strategic implications of AGI. Suppose US and middle powers agree to a &#8220;chips for frontier access&#8221; deal &#8211; middle powers continue to supply the US with frontier chips; US continues to give middle powers access to frontier AI. The middle powers might still worry: what if the US suddenly changes its mind once it has superintelligence? By then, the US might be powerful enough to dominate without continued allied support. This is where kill switches can help. If the US withdraws AI access, allies could destroy US data centres in response. It&#8217;s a way to lock in the deal.</p><ul><li><p>(h/t AI futures project for this idea. A related idea is for US data centres to be placed in a location that&#8217;s easy to attack &#8211; like <a href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space">in space</a>)</p></li></ul></li></ul><p>Beyond securing access to frontier AI, how else can middle powers maintain economic and military leverage?</p><ul><li><p><strong>Build physical infrastructure. </strong>Factories, robots, solar panels, batteries, semiconductors &#8212; all these industries are highly complementary to powerful AI.</p></li><li><p><strong>Maintain nuclear 2nd strike capability. </strong>The point isn&#8217;t to use it. But it improves their leverage for stage 2.</p></li></ul><p>The catch-all meta-point here is waking middle powers up to superintelligence.</p><p>I&#8217;m not recommending middle powers do their own frontier AI development. Seems very hard for them to catch up with the US.</p><h2>Stage 2: Ensure that, when superintelligence is developed, it refuses to crush middle powers</h2><p>If stage 1 goes well, middle powers remain somewhat powerful economically and militarily deep into the singularity. But it might fail. What can middle powers do if they see the US on track to total global dominance?</p><p>First, they should <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5776982">demand a pause/slowdown of AI development</a>. But the US may refuse &#8211; pausing is very costly if alignment risk is low. And pausing is a stopgap: eventually, superintelligence will be developed.</p><p>An additional demand: when superintelligence is developed, it&#8217;s designed to refuse to crush middle powers. By doing this, the US would credibly bind itself to maintaining the sovereignty of other nations.</p><p>Superintelligence would help the US outgrow other countries economically, but it would never attack them militarily or otherwise interfere with their sovereignty. While middle powers would be <em>relatively</em> economically disempowered, their citizens could be very rich <em>absolutely</em> and live in freedom without US interference.</p><p>Would this work? The optimistic case is that this isn&#8217;t a big sacrifice for the US. They can still become as rich as they like and achieve their security interests. Sure, they can&#8217;t seize control of other nations, but that is not an important goal of theirs anyway. Losing that option is well worth the benefits: other nations cooperate economically, don&#8217;t attack US data centres, and don&#8217;t threaten nuclear war.</p><p>The pessimistic case is that this involves an insane degree of irrevocable hand-off to AI. The US must literally be unable to attack middle powers no matter how hard it tries: retraining the AI, turning it off, training a new more powerful AI, passing new laws, using the military to destroy the data centres the AI is running on. For it to be truly binding, the US must permanently hand over military and political power to AI. That might be deeply unpopular, and indeed seem insane to the US. It&#8217;s also very hard to verify: you can&#8217;t just verify the training run, you need to verify that humans+other AIs have <em>no</em> way to disempower the trained AI. It&#8217;s more like verifying &#8220;who would win this civil war&#8221; than &#8220;technical property XYZ holds&#8221;.</p><p>The realistic path here probably involves gradually handing off more and more control to AI that refuses to crush middle powers, with no clear point at which humans could no longer wrest back control.</p><p>The longer middle powers wait to push for stage 2, the less leverage they will have because the US will have pulled further ahead economically and militarily. So they should be pushing in this direction constantly, e.g. demanding transparency into the model specs of powerful AIs deployed in the US government, and arguing that powerful military AI should be designed to obey international law.</p><p>(I described the plan as involving two stages because that&#8217;s how I expect it to play out over time. But succeeding at either stage is sufficient! If middle powers stay economically/militarily competitive, they never need to bind US superintelligence. And if they <em>do</em> bind superintelligence, they won&#8217;t be crushed no matter how far behind they fall.)</p><p>Another strategy: train superintelligence to ensure middle countries continue to get equal access to frontier AI. This combines stages 1 and 2, and could prevent even the <em>relative</em> disempowerment of middle powers.</p><h2>Is it good to avoid middle powers getting trounced?</h2><p>I live in the UK, so I am biased here. I do not want the UK to become a supplicant to the US!</p><p>But here&#8217;s a brainstorm of pros and cons from a more impartial perspective.</p><p>Pros to empowering middle powers:</p><ul><li><p><strong>Avoid a single point of failure</strong>. If the US becomes globally dominant and its political system fails, that&#8217;s a global failure.</p></li><li><p><strong>More democracies. </strong>Many middle power democracies look more robust than the US, so more middle powers may mean more democracy.</p></li><li><p><strong>Improve the US. </strong>Middle powers will have an interest in maintaining free market democracy in the US. &#8220;Free market&#8221; because they&#8217;ll want multiple AI companies competing to sell cheap API access to non-US countries. &#8220;Democracy&#8221; because they&#8217;ll expect that the US is more likely to maintain a strong alliance with middle power democracies if it stays democratic.</p></li><li><p><strong>Experimentation. </strong>Experimenting with multiple different political and legal systems seems generally good for figuring out a good way to govern society post AGI.</p></li><li><p><strong>Pause AI. </strong>They could potentially pressure US/China to pause/slow down reckless AI development.</p></li><li><p><strong>Prosocial norms. </strong>When multiple actors bargain with each other (e.g. about how to distribute space resources, whether to develop a dangerous technology), they tend to frame arguments in terms of prosocial norms, and so agreements tend to emphasise the actor&#8217;s more virtuous/ethical values.</p></li></ul><p>Cons of empowering middle powers. Multipolarity has its own downsides:</p><ul><li><p>More likely to lead to war.</p></li><li><p>Can drive extreme competition, e.g. racing to develop a dangerous technology, or to hand off power to misaligned AI.</p></li><li><p>Harder to prevent harms from offence-dominant technologies like bioweapons.</p></li><li><p>This plan involves waking up middle powers, which could shorten timelines.</p></li></ul>]]></content:encoded></item></channel></rss>