<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[ForeWord]]></title><description><![CDATA[How should we navigate explosive AI progress? 

The latest research from Forethought.]]></description><link>https://newsletter.forethought.org</link><image><url>https://substackcdn.com/image/fetch/$s_!OWCf!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff69a310d-5182-4a63-bf1c-1b392366e785_663x663.png</url><title>ForeWord</title><link>https://newsletter.forethought.org</link></image><generator>Substack</generator><lastBuildDate>Sat, 15 Aug 2026 02:43:02 GMT</lastBuildDate><atom:link href="https://newsletter.forethought.org/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Forethought]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[forethoughtnewsletter@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[forethoughtnewsletter@substack.com]]></itunes:email><itunes:name><![CDATA[Forethought]]></itunes:name></itunes:owner><itunes:author><![CDATA[Forethought]]></itunes:author><googleplay:owner><![CDATA[forethoughtnewsletter@substack.com]]></googleplay:owner><googleplay:email><![CDATA[forethoughtnewsletter@substack.com]]></googleplay:email><googleplay:author><![CDATA[Forethought]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Notes on Implications of Scale-Dependent Algorithms]]></title><description><![CDATA[This article was created by Forethought. See all our research on our website.]]></description><link>https://newsletter.forethought.org/p/notes-on-implications-of-scale-dependent</link><guid isPermaLink="false">https://newsletter.forethought.org/p/notes-on-implications-of-scale-dependent</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Thu, 13 Aug 2026 13:43:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qyaP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><p><span>Of past algorithmic improvements which decrease LLM pretraining loss, many have been &#8220;scale-dependent&#8221; &#8211; that is, improvements which decrease LLM pretraining loss by more, relative to prior algorithms, at larger quantities of compute.</span></p><p><span>The influence of such scale-dependent algorithms is (1) moderate evidence that attempts to limit algorithmic progress in absence of compute limitations will be ineffective, (2) weak evidence that our inferences about the future scale of algorithmic progress, based on the past, are invalid, and (3) part of a plausible argument either for or against a software intelligence explosion, depending on other details of one&#8217;s model.</span></p><p><span>I&#8217;ll proceed by discussing the following:</span></p><ol><li><p><span>How scale-dependent algorithms for pretraining loss probably account for the majority of algorithmic improvements in this domain</span></p></li><li><p><span>How scale-dependent algorithms for end-to-end task performance might account for the majority of algorithmic improvements in this domain</span></p></li><li><p><span>How scale-dependent algorithms in the past should influence our model of algorithmic improvements in the future</span></p></li><li><p><span>Possible policy implications</span></p></li></ol><p><span>Before starting, a note on method &#8211; when discussing questions regarding AI&#8217;s past and future progress, there&#8217;s a natural continuum from &#8220;the intractable, high-impact, abstract question, which we terminally care about&#8221; to &#8220;the tractable, dubious-impact, concrete question, which we instrumentally care about.&#8221;</span></p><p><span>Consider questions about scale-dependent algorithmic improvements, sliding from the former to the latter along such a continuum.</span></p><ul><li><p><span>We are most terminally interested in questions about the future: &#8220;Will </span><em><span>future</span></em><span> improvements to AI </span><em><span>task performance </span></em><span>depend more on compute scale-up or on algorithmic improvements? What can we know about how these two will be causally intertwined? Will future improvements enable a software-only intelligence explosion &#8211; that is, in a world with total compute held constant, could algorithmic progress lead to an intelligence takeoff in a comparatively short period of time?&#8221;</span></p></li><li><p><span>The data we can gather most relevant to this question is about the past: &#8220;Did </span><em><span>past</span></em><span> improvements to AI </span><em><span>task performance</span></em><span> depend more on compute scale-up or on algorithmic improvements? Would past AI algorithmic improvement have counterfactually looked slower, without the datacenter buildout?&#8221;</span></p></li><li><p><span>And finally we have the narrower sub-question for which we have the most past data, which contributes to answering the broader historical question: &#8220;Did past improvements to LLM </span><em><span>pretraining loss </span></em><span>depend more on compute scale-up or on algorithmic improvements?&#8221;</span></p></li></ul><p><span>I&#8217;m going to start by discussing the evidence for this last question, before moving to the increasingly impactful and increasingly difficult earlier questions.</span></p><h2><span>Scale-Dependence of Algorithmic Improvements for Pretraining Loss</span></h2><p><span>Are the algorithmic improvements that have most decreased pretraining loss for LLMs scale-dependent or scale-independent? What would either of these mean?</span></p><p><span>The </span><a href="https://arxiv.org/pdf/2312.07413"><span>compute-equivalent-gain</span></a><span> (CEG) for some algorithm A&#8217;, relative to another algorithm A, is the ratio of the FLOPs needed to train an A-using model to a given level of performance, over the FLOPs needed for an A&#8217;-using model to reach that same level. If it takes 3x as much compute to train an LLM with algorithm A to reach the same pretraining loss as another LLM trained with algorithm A&#8217;, then A&#8217; has a 3x compute-equivalent gain on pretraining loss relative to A.</span></p><p><span>Given the above notion of CEG, algorithmic improvements that produce some kind of CEG could conceivably be </span><strong><span>scale independent </span></strong><span>or </span><strong><span>scale dependent</span></strong><span>.</span></p><p><span>A scale-independent algorithm has a constant CEG across different quantities of training FLOPs. So if A&#8217; gives a 2x multiplier relative to A at about 10</span><sup><span>15</span></sup><span> FLOPs, it continues to give a 2x multiplier at about 10</span><sup><span>25</span></sup><span> FLOPs. Correspondingly, if one were somehow able to know antecedently that some algorithm was scale independent, one would just need to test it at one scale to find whatever this constant happened to be.</span></p><p><span>By contrast, a scale-dependent algorithm has a variable CEG, which is a function of the number of FLOPs used in training. So A&#8217; might give a 2x multiplier relative to A at about 10</span><sup><span>15</span></sup><span> FLOPs, but give a radically different 1000x multiplier at about 10</span><sup><span>25</span></sup><span> FLOPs. If one were somehow able to know antecedently that some algorithm was scale dependent, one would know that one needed to test it at many scales to find the function mapping from FLOPs to the compute multiplier.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qyaP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qyaP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 424w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 848w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1272w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qyaP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png" width="1456" height="596" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:596,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qyaP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 424w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 848w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1272w, https://substackcdn.com/image/fetch/$s_!qyaP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d260cbd-8aeb-4baa-8d7a-df3da2439e8d_2048x839.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The thesis of 2025&#8217;s &#8220;</span><a href="https://arxiv.org/pdf/2511.21622"><span>On the Origin of Algorithmic Progress in AI</span></a><span>&#8221; is that the vast majority of apparent algorithmic progress in LLM pretraining loss per-FLOP</span><strong><span> </span></strong><span>has come from a handful of scale-dependent algorithms. So, according to the paper, the apparently steady flow of &#8220;algorithmic improvements&#8221; in the past came from choosing a reference point algorithm (A) prior to some particular scale-dependent algorithmic improvement (A&#8217;), and then mistaking the continued improvements A&#8217; makes over A as compute increases for the presence of additional algorithms, B, C, and so on. In the absence of an increase of FLOPs, this apparent radical increase in performance would not have been nearly as large.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HiZF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HiZF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 424w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 848w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1272w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HiZF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png" width="1456" height="994" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:994,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HiZF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 424w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 848w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1272w, https://substackcdn.com/image/fetch/$s_!HiZF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4feb7831-94a5-4345-ad85-dfc3df69b1ba_1820x1242.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>What kind of evidence is there that the majority of progress in LLM pretraining loss per-FLOP actually does depend on scale-dependent algorithms?</span></p><ul><li><p><span>Well, the paper presents good evidence that scale-dependent algorithms are very important.</span></p><ul><li><p><strong><span>LSTM to Transformer Change</span></strong><span>: &#8220;On the Origin&#8221; uses experiments to estimate how much the Transformer improves over the LSTM as compute scales up. At ~10</span><sup><span>15</span></sup><span> FLOPs, this comes to about 6x CEG; at ~10</span><sup><span>17</span></sup><span>, it comes to a 28x CEG increase; at the level of compute with which the Transformer was introduced in 2017, they extrapolate it to about a 91x CEG increase.</span></p></li><li><p><strong><span>Kaplan to Chinchilla Change</span></strong><span>: &#8220;On the Origin&#8221; grants that Chinchilla scaling is correct, then projects forward the gains you get by switching to Chinchilla from Kaplan; they are approximately equivalent at 10</span><sup><span>20</span></sup><span>, and give a ~10x or more gain at frontier scales.</span></p></li></ul></li><li><p><span>The paper also presents evidence that scale-independent algorithms are less important than one might imagine.</span></p><ul><li><p><span>Their ablations find that scale-independent additions are not multiplicative. That is, doing a bucket of changes that move one from the &#8220;retro&#8221; Transformer to the modern Transformer in theory should give a CEG of 3.5x, if one were to multiply the gains that the additions make separately; but in practice only gives 1.3x.</span></p></li><li><p><span>The modern transformer similarly does not improve greatly on the &#8220;retro&#8221; Transformer as you scale up.</span></p></li></ul></li></ul><p><span>So &#8220;On the Origin&#8221; projects that, </span><strong><span>given </span></strong><span>the actual historical increase in FLOPs that makes scale-dependent algorithms more efficient, one observes a 21,400x CEG increase from 2017 to 2025. But if one counterfactually holds compute constant at 10^18 FLOPs, this amounts to only a 155x CEG increase. This amounts to the difference between a ~7 month doubling time for algorithmic efficiency (very close to </span><a href="https://epoch.ai/publications/algorithmic-progress-in-language-models"><span>Ho et al. (2024)</span></a><span>&#8217;s 8-month doubling time); and a ~13 month doubling time of algorithmic efficiency. But note that the second &#8220;doubling time&#8221; here is specifically relative to the level of compute at which one calculates the increase in algorithmic efficiency; there would be a different doubling time for a different absolute quantity of compute.</span></p><p><span>So there is a (very approximate) halving of the speed of algorithmic improvement, given</span><strong><span> </span></strong><span>that you hold FLOPs fixed at 10</span><sup><span>18</span></sup><span>, once you take into account the difference between scale-dependent and scale-independent algorithms. And of course, this ~13 month doubling time is still fueled almost entirely by the gigantic LSTM to Transformer transition at 10</span><sup><span>18</span></sup><span> FLOPs &#8211; leaving this quantity of algorithmic progress out means that the doubling time gets vastly larger.</span></p><p><span>This is a pretty large change to the proposed rate at which constant-FLOP algorithmic improvements occur. But you can actually make a case that &#8220;On the Origin&#8221; understates the degree to which algorithmic improvements in LLM pretraining loss depend on scale-dependent algorithms.</span></p><p><span>The largest single purportedly scale-independent gain (2x) described in the paper belongs to mixture-of-experts (MoEs), an alteration to the Transformer architecture. But several subsequent works have shown that the gain from MoEs over dense architectures is actually strongly scale-dependent, and probably larger than 2x. &#8220;Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models&#8221; estimates a </span><a href="https://arxiv.org/pdf/2507.17702"><span>7x gain</span></a><span> at 10</span><sup><span>22</span></sup><span> FLOPs. &#8220;Scaling Laws for Fine-Grained Mixture-of-Experts&#8221; estimates a </span><a href="https://openreview.net/pdf?id=Iizr8qwH7J"><span>20x</span></a><span> gain at 10</span><sup><span>20</span></sup><span> FLOPs. In both cases, the size of the CEG relative to a dense model keeps improving with the quantity of FLOPs. Although it&#8217;s uncertain which exponent is correct, both papers point toward MoEs being a scale-dependent algorithmic improvement that substantially improves with the quantity of FLOPs.</span></p><p><span>So &#8211; to return to the motivating question &#8211; most of the algorithmic improvements to LLM pretraining loss, so far, have probably depended on a compute scale-up. Apparent algorithmic progress in pretraining loss would have been much smaller if total compute had been held constant.</span></p><h2><span>Scale-Dependence of Algorithmic Improvements for End-to-End Performance</span></h2><p><span>No one actually cares about pretraining loss for its own sake, though.</span></p><p><span>&#8220;Decreased pretraining loss&#8221; is just one way to improve end-to-end performance. There are also post-training </span><a href="https://arxiv.org/pdf/2312.07413"><span>innovations</span></a><span> like </span><a href="https://arxiv.org/abs/1909.08593"><span>RLHF</span></a><span> or best-of-N that improve task performance. There are also changes to </span><a href="https://www.beren.io/2025-08-02-Most-Algorithmic-Progress-is-Data-Progress/"><span>data mixtures</span></a><span> that improve performance. And of course </span><a href="https://arxiv.org/abs/2501.12948"><span>RLVR-over-CoT</span></a><span> has probably been by itself responsible for the </span><a href="https://epoch.ai/gradient-updates/quantifying-the-algorithmic-improvement-from-reasoning-models"><span>greatest</span></a><span> single leap in AI performance of the last few years.</span></p><p><span>How much are algorithmic improvements to end-to-end performance scale-dependent? In particular, how much does RLVR-over-CoT depend on scale-dependent algorithmic improvements?</span></p><p><span>This is deeply uncertain. There&#8217;s certainly no work that provides a clean FLOPs-to-CEG function, as in the case of pretraining loss.</span></p><p><span>There are several reasons for thinking that RLVR works extremely poorly or not at all beneath a certain scale, and so is somewhat scale-dependent. You need a certain success rate for reinforcing successes to be viable. </span><a href="https://arxiv.org/pdf/2607.12395"><span>Several</span></a><span> </span><a href="https://arxiv.org/pdf/2501.12948"><span>works</span></a><span> remark on how RLVR without sufficient scale simply doesn&#8217;t teach the required diversity of behaviors, and that RLVR on small models plateaus far lower than RLVR on larger and longer-trained models. So there&#8217;s suggestive </span><a href="https://arxiv.org/pdf/2607.16097v1"><span>evidence</span></a><span> that at the frontier RLVR is scale-dependent, but nothing conclusive.</span></p><p><span>It would also be difficult to perform an experiment to cleanly determine whether RLVR-over-CoT is scale-dependent for several reasons. For instance, the importance of RL, in general, </span><a href="https://arxiv.org/pdf/2509.25123"><span>probably</span></a><span> </span><a href="https://arxiv.org/abs/2512.07783"><span>comes</span></a><span> </span><a href="https://arxiv.org/abs/2501.17161"><span>from</span></a><span> how it permits at least somewhat out-of-distribution generalization, rather than from how it improves performance on in-distribution tasks. But to my knowledge, there&#8217;s no publicly proposed clear way to measure out-of-distribution generalization; so there&#8217;s no measure against which one could test how RLVR improves with scale on the matter where it ostensibly has the greatest impact. Similarly, RLVR takes place as the final stage of training, after a large sequence of data-selection, pretraining, and midtraining, which means there are just many more free variables it would be necessary to consider before deciding on an experimental design.</span></p><p><span>Apart from RLVR, there&#8217;s some evidence that some apparent scale-independent improvements may scale </span><em><span>negatively</span></em><span> with FLOPs, and thus provide few or limited gains at the frontier. For instance, the </span><a href="https://arxiv.org/pdf/2605.19407"><span>paper</span></a><span> &#8220;A Bitter Lesson for Data Filtering&#8221; claims that although filtering data improves performance at a small amount of compute, it scales negatively with FLOPs, such that with enough computing power it&#8217;s best to do no data filtering at all.</span></p><p><span>I think it&#8217;s rather likely that RLVR and other important post-training techniques improve in a strongly scale-dependent way, and some of my analysis below will lean on that belief. I&#8217;d be surprised if RLVR doesn&#8217;t improve in a </span><em><span>somewhat </span></em><span>scale-dependent way.</span></p><p><span>But I don&#8217;t think there&#8217;s any conclusive empirical evidence here, and to the degree I&#8217;m wrong, some of my inferences below are less likely to be true.</span></p><h2><span>How scale-dependent improvements in the past impact our model of &#8220;algorithmic progress&#8221; for the future</span></h2><p><span>Let&#8217;s grant that, in the past, scale-dependent algorithms have been responsible for the majority of algorithmic improvements in both LLM pretraining loss and in end-to-end task performance.</span></p><p><span>What consequences would this have for how we expect algorithmic progress to go in the future? What consequences does this imply about a future where compute is &#8220;held constant,&#8221; as in the case of a software intelligence explosion? Would it lead us to expect a radical trend break, in either direction?</span></p><p><span>Let&#8217;s take the other side by contrast, first &#8211; imagine a model of the world where almost all algorithmic improvements are scale </span><em><span>independent</span></em><span>: improvements are drawn from a pool of possible improvements, and the number and efficacy of the algorithms taken from the pool do not change merely by the fact of compute scaling up.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zTQE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zTQE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 424w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 848w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1272w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zTQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png" width="836" height="576" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:576,&quot;width&quot;:836,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zTQE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 424w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 848w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1272w, https://substackcdn.com/image/fetch/$s_!zTQE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82e291c1-3d19-400a-a236-5c1b1feb965c_836x576.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>If you operate off this model, then there are some natural consequences.</span></p><ul><li><p><span>Algorithmic improvements do not become &#8220;more discoverable&#8221; as FLOPs scale up. More compute may make it easier to run experiments to find improvements, but scaling up the FLOP count within each experiment does not make finding the right algorithm easier or harder per-experiment.</span></p><ul><li><p><span>There could, of course, be other factors that make algorithmic improvements easier or harder to find. One improvement could open up a line of research that leads to other improvements, making it easier; it might decrease the total pool of improvements, making it harder. But the fact of having scaled up FLOPs alone does not make improvements easier or harder to find.</span></p></li></ul></li><li><p><span>Some particular algorithmic improvement does not become &#8220;higher impact,&#8221; should its discovery be pushed into the future. If some algorithm X is found today and gives a 3x CEG, it would equally well give a 3x CEG if it were found two years into the future.</span></p></li><li><p><span>Broadly, our actual historical rates of &#8220;algorithmic progress&#8221; reflect something like the rate at which algorithms are naturally found, in a way indifferent to compute build-out.</span></p></li></ul><p><span>By contrast, one can model a world where most algorithmic improvements are scale dependent as a world where improvements are drawn from a changing pool of FLOP-specific algorithmic improvements. Pools get &#8220;unlocked&#8221; as the compute frontier moves past the point where they start to become significant, and the pools grow progressively more explored over time as more and more researchers have access to the compute necessary to explore them. So both &#8220;how explored&#8221; each pool is, and the consequences of finding an entry from each pool, are always changing.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PaXs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PaXs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 424w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 848w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1272w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PaXs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png" width="1456" height="784" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PaXs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 424w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 848w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1272w, https://substackcdn.com/image/fetch/$s_!PaXs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c80e930-2800-4eeb-9024-9ad639621ceb_1500x808.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>If you operate off this model, then:</span></p><ul><li><p><span>Algorithmic improvements become much &#8220;more discoverable&#8221; as FLOPs scale up; the more FLOPs you use in an experiment, the higher the relative CEG might become, and so the easier the new algorithm will find it to overcome the &#8220;well-tuned baseline&#8221; algorithm with very well-adjusted hyperparameters. So in addition to letting you run more experiments, additional compute may enhance &#8220;research taste.&#8221;</span></p></li><li><p><span>Some particular algorithmic improvement might actually become &#8220;higher impact,&#8221; should its discovery be pushed into the future. If some algorithm X is found today and gives a 3x CEG, it might give a 100x CEG in the future when applied with more FLOPs, if its discovery is delayed.</span></p></li><li><p><span>Broadly, this view implies that &#8220;algorithmic progress&#8221; depends on past compute scale-ups, and would be much slower without them.</span></p></li></ul><p><span>Imagine that one is trying to determine how much algorithmic progress is likely in the future, granted that one is in a scenario where total compute is held constant. It&#8217;s important to note that switching to a view of algorithmic progress as scale-dependent might move your estimate of the likely rate of algorithmic progress in a constant-compute regime either up or down.<br><br>The case for moving it down is clear. Switching to a view where scale-dependent improvements are significant could move one&#8217;s estimate of algorithmic progress down, because it moves one estimate of past algorithmic progress in a constant-compute regime down. One used to estimate a ~8 month halving time to reach equivalent loss for equal FLOPs; now one estimates a ~13 month halving time; and the future is likely to be like the past, so future FLOP-constant algorithmic progress is likely to be slower.</span></p><p><span>But switching to such a view could move one&#8217;s estimate of algorithmic progress up, because it decreases your credence that the future will be like the past in this respect! If different algorithms get unlocked at different levels of compute, then potentially some future yet-to-be-unlocked regime of algorithmic improvements might differ enormously from past algorithmic improvements. There&#8217;s likely less reason to think that the number or importance of algorithmic improvements available in one prior algorithmic bucket will be equally important as algorithmic improvements in subsequent buckets as compute increases.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5Oa2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5Oa2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 424w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 848w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1272w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png" width="1456" height="931" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:931,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5Oa2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 424w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 848w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1272w, https://substackcdn.com/image/fetch/$s_!5Oa2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc566c8e1-6c45-4d49-8ca1-84d681ab1a21_1532x980.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>So if one finds such a picture compelling, then the likelihood one assigns to a SIE could go up. If there is a future compute scale-up to 10</span><sup><span>32</span></sup><span> FLOPs, a SIE taking place then might find many algorithms that were totally unavailable to us in the 10</span><sup><span>27</span></sup><span> regime.</span></p><p><span>Overall, all this remains uncertain. I think various philosophical arguments suggest that there&#8217;s a greater diversity of room for algorithmic progress as compute scales up. The human brain is largely a </span><a href="https://pubmed.ncbi.nlm.nih.gov/19915731/"><span>linearly scaled-up</span></a><span> primate brain, but individual humans and humans in aggregate can do many calculations no non-human primate can. And there are likely greater gains from organizing 10,000 humans well than from organizing 10 humans well.</span></p><h2><span>Possible Policy Implications</span></h2><p><span>If a majority of algorithmic improvements to task performance consistently come from scale-dependent algorithms, and are likely to keep doing so in the future, how would this impact policy interventions into AI?</span></p><p><span>There are two notable consequences, which mirror each other: (1) efforts to decrease algorithmic progress, in absence of a training compute scale-up, would be less necessary than previously believed; and (2) efforts to decrease algorithmic progress, in the presence of a training compute scale-up, would be less effective than previously believed. I will discuss them each in turn.</span></p><p><strong><span>First, </span></strong><span>some policy interventions meant to slow or stop the development of advanced AIs have focused on limiting the size of the largest training runs. In this context, a steady advance of scale-independent algorithmic improvements would have decreased the efficacy of any particular maximum training-run size limitation. For instance </span><a href="https://arxiv.org/pdf/2511.10783"><span>&#8220;An International Agreement to Prevent the Premature Creation of Artificial Superintelligence&#8221;</span></a><span> recommends a maximum training-run size of 10</span><sup><span>24</span></sup><span>, but also notes that the &#8220;number of operations used to train an AI to a given capability level drops by 3x each year.&#8221; Given this trend, it&#8217;s easier to derive the need to prohibit AI algorithmic research, a measure that the paper itself notes may be &#8220;controversial and normally a bad idea,&#8221; but which it believes to be necessary to avoid the creation of ASI.</span></p><p><span>But if the majority of algorithmic improvements have been scale-dependent, then prior estimates for scale-independent algorithmic improvements are too high; and so prior work gives a less-good reason for thinking that the number of operations used to train an AI to a given capability level drops by 3x every year. So the need to ban and monitor research would be correspondingly decreased; a hard cap on maximum training-run size would be more effective than previously estimated.</span></p><p><strong><span>Second, </span></strong><span>some proposed policy interventions have focused on </span><strong><span>scaling up </span></strong><span>the size of compute while trying to decrease or hold constant algorithmic progress. AI 2040, for instance, proposes continuing to increase the </span><a href="https://ai-2040.com/supplements/compute-supplement"><span>size</span></a><span> of training runs from 10</span><sup><span>26</span></sup><span> now to 10</span><sup><span>32</span></sup><span> in 2035, while decreasing the rate of algorithmic progress.</span></p><p><span>Overall, if you believe that the majority of important future algorithmic improvements will be scale dependent in the way that the majority of past algorithmic improvements were scale dependent, that should decrease the credence you give to our ability to &#8220;not find&#8221; these improvements. Why?</span></p><p><span>As mentioned above &#8211; as compute scales up, the increasing gains from scale-dependent improvements will make them more obvious. By analogy to a counterfactual world in which the Transformer was never discovered, an incredibly poorly-done, badly-implemented Transformer in 2026 still obviously improves on an LSTM, because even an awful implementation might mean it is only 100x better rather than 1000x better than the LSTM. It&#8217;s probably impossible to present the Transformer from being discovered for too long. Similarly, continuing to scale up compute makes scale-dependent algorithms more obvious; it probably decreases the quantity of research taste you need to discern them.</span></p><p><span>The other part of the reason is just because, well, continuing to scale up compute keeps opening up new possible unseen algorithmic buckets that allow improvement; the space of algorithmic improvements will keep growing, different kinds of research taste will find more progress.</span></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Liberalism Forever]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/liberalism-forever</link><guid isPermaLink="false">https://newsletter.forethought.org/p/liberalism-forever</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Fri, 31 Jul 2026 18:19:44 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d58976af-23a5-40f6-9944-8ff749331741_900x590.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-PViXDHfw3hg" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;PViXDHfw3hg&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/PViXDHfw3hg?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.peternsalib.com/">Peter Salib</a> is an Assistant Professor of Law at the University of Houston Law Center and co-director of the Center for Law &amp; AI Risk. <a href="https://www.simondgoldstein.com/">Simon Goldstein</a> is an associate professor at the University of Hong Kong, a senior editor at AI Frontiers, and a visiting senior scholar at Forethought. They have written together about AI safety, AI rights, and the governance of advanced AI.</p><p>They joined Forethought&#8217;s <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss their new paper, <a href="https://philpapers.org/rec/GOLLFC-2">Liberalism Forever</a>. The conversation covers:</p><ul><li><p>Their &#8220;null hypothesis&#8221;: today&#8217;s liberal institutions &#8212; free markets and democratic governance &#8212; may already hold most of the answers for governing the far future, and longtermists have underrated their power</p></li><li><p>Examining whether canonical social-science arguments for markets and democracy still hold under transformative AI, space colonization, and explosive growth, and claiming that most survive (and several get stronger)</p></li><li><p>Why they favour reasoning &#8220;at the margin&#8221; over designing a detailed end-state, and why proposals like viatopia and the long reflection sound liberal but risk illiberalism when actually implemented</p></li><li><p>Whether the standard case for markets &#8212; information aggregation, allocative efficiency, innovation &#8212; still bites once AGI can read off preferences directly</p></li><li><p>Why inequality is, in their view, a comparatively boring problem with a known solution (tax-and-transfer), why AI-specific taxes like a compute tax are misguided, and how the growth-versus-distribution trade-off changes when growth rates get very high</p></li><li><p>Dyson swarms and space resources: why they doubt there&#8217;s any real monopoly worry, since energy would be additive with free entry, and why property auctions might beat egalitarian allocation schemes</p></li><li><p>The arguments for democracy that persist post-AGI &#8212; public-choice/selectorate incentives and democracy as a commitment device &#8212; versus the epistemic arguments, which they think weaken</p></li><li><p>Why AGI could make autocracy more competitive (loyal AI bureaucrats and militaries removing the need for a human winning coalition), and what preserving democratic control requires in response</p></li><li><p>Handoffs to superintelligent AI and successionism, and their preferred alternative of extending the franchise to AIs with their own values rather than deferring wholesale</p></li><li><p>The history of new &#8220;technologies of democracy,&#8221; from the Federalist Papers to LLMs as impartial arbiters, and where each author thinks their own argument is weakest</p></li></ul><p><a href="https://docs.google.com/document/d/1F0XrSiwCenc7h-Ke9Bqhrhmqty59unRA5sJgmdUjF6s/edit?tab=t.0">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[Should Philanthropists Save Their Money for the Intelligence Explosion?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/should-philanthropists-save-their</link><guid isPermaLink="false">https://newsletter.forethought.org/p/should-philanthropists-save-their</guid><dc:creator><![CDATA[Forethought]]></dc:creator><pubDate>Thu, 30 Jul 2026 20:33:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/_vSs8v0yFkU" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-_vSs8v0yFkU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;_vSs8v0yFkU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/_vSs8v0yFkU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.forethought.org/people/william-macaskill">Will MacAskill</a> is a senior research fellow at Forethought and the author of <em><a href="https://whatweowethefuture.com/">What We Owe The Future</a></em>. <a href="https://www.forethought.org/people/tom-davidson">Tom Davidson</a> is a senior research fellow at Forethought and the author of a series of reports on <a href="https://www.forethought.org/research/three-types-of-intelligence-explosion">AI timelines</a>, <a href="https://www.forethought.org/research/the-industrial-explosion">takeoff speeds</a>, and <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">AI-enabled coups</a>.</p><p>In this informal conversation about early-stage research in progress, they discuss:</p><ul><li><p>Why a philanthropist who takes transformative AI seriously should expect extraordinary investment returns (plausibly 10x to 100x) and how the market doesn&#8217;t seem to be pricing this in</p></li><li><p>Why AGI shouldn&#8217;t be treated as a hard deadline to spend by, and the case that most philanthropic money should be spent <em>during</em> the intelligence explosion rather than before it</p></li><li><p>How genuine AI substitutes for human labor would dissolve the hiring bottleneck that makes organizations so hard to scale today</p></li><li><p>How relative prices shift when cognitive labour becomes abundant, why the patent system stops making sense in that world, and the window of &#8220;crazy philanthropic bargains&#8221; open to actors who adapt faster than large institutions</p></li><li><p>The prospect of a world where only a small elite has access to superintelligent medical, financial, and strategic advice, and what it would take to avoid this</p></li><li><p>Counter-considerations to the main case, e.g., crowding out by new money, the risk that overspending attracts grifters or distorts the field, and whether a richer world will mean worse philanthropic opportunities</p></li><li><p>The risk of value drift over a long saving period, and mechanisms for binding your future self (such as irrevocably handing off control of philanthropic funds to independent, constitutionally-bound foundations)</p></li><li><p>&#8220;Differential intellectual development&#8221;: using directable AI researchers to accelerate biodefense, infosecurity, AI interpretability, and neglected armchair fields like population ethics and social choice theory</p></li><li><p>Why the case for waiting to give is much weaker for small donors</p></li><li><p>Will&#8217;s worry that the whole line of reasoning might fail if it turns out scaling philanthropy well is bottlenecked by serial real-world learning that no amount of cognitive labor can compress</p></li></ul><p><a href="https://docs.google.com/document/d/12y3_3BljMhRyZH3P2IIkJjct833VmIsjT6XAhWe240c">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[Policy ideas to ensure responsible government deployment of AI]]></title><description><![CDATA[This article was created by Forethought. See all our research on our website.]]></description><link>https://newsletter.forethought.org/p/policy-ideas-to-ensure-responsible</link><guid isPermaLink="false">https://newsletter.forethought.org/p/policy-ideas-to-ensure-responsible</guid><dc:creator><![CDATA[Stefan Torges]]></dc:creator><pubDate>Wed, 22 Jul 2026 16:41:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/56818bf6-68a4-4606-a507-abe937f4bbbb_2166x1310.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><h1><span>Summary</span></h1><ul><li><p><span>Governments are going to deploy frontier AI in the domains where they have exclusive powers: military, intelligence, surveillance, and punitive law enforcement.</span></p></li><li><p><span>The existing checks on those powers are weak. AI deployment could weaken them further.</span></p></li><li><p><span>Companies aren&#8217;t well-placed to police government deployments (e.g., contracts often require </span><a href="https://code.claude.com/docs/en/zero-data-retention"><span>zero data retention</span></a><span>), legislatures often lack access and capacity, and, as AI replaces humans in government, nobody will be left to refuse or slow-roll unlawful or anti-democratic orders.</span></p></li><li><p><span>In the limit, this can enable </span><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>full-blown coups</span></a><span>; in the short term, it undermines checks &amp; balances.</span></p></li></ul><p><span>This post lays out the work I&#8217;m most excited about:</span></p><ol><li><p><strong><span>Governance of AI applications in government</span></strong><span>: policies on how AI should be procured and used in government, research on how to govern AI agents in particular, and field-building beyond the EA/safety crowd.</span></p></li><li><p><strong><span>Interbranch oversight</span></strong><span>: transparency and oversight provisions, audit tech for an automated government, waking up lawmakers, and increasing AI uptake in the legislative and judicial branches.</span></p></li><li><p><strong><span>Developing AI for government use</span></strong><span>: model specs/constitutions that say something meaningful about behavior in government contexts, and evals for the qualities we&#8217;d want such systems to have (such as epistemic integrity, non-sycophancy, lawfulness).</span></p></li><li><p><strong><span>Strengthening civil society</span></strong><span>: monitoring and freedom of information infrastructure for government AI use, a coalition of ML researchers at frontier labs, and tools that preserve citizens&#8217; ability to coordinate (AI-for-epistemics + privacy tech).</span></p></li></ol><p><span>Many ideas are more tractable than they might seem: specs are being written now, must-pass bills come around every year, and a lot of the civil-society work just requires someone to start doing it.</span></p><p><strong><span>If you&#8217;re interested in working on any of it, apply to </span><a href="https://checks-and-balances.ai/"><span>EIP&#8217;s</span></a><span> recent RFP. You could also fill out </span><a href="https://www.forethought.org/careers/expression-of-interest-power-concentration"><span>this expression of interest form</span></a><span>.</span></strong></p><h1><span>Why care about responsible government deployment of AI</span></h1><p><strong><span>How could government deployment go wrong?</span></strong></p><ul><li><p><span>Governments will increasingly use the most powerful AI systems in domains where they have exclusive powers; particularly military, law enforcement, intelligence, and surveillance. This is most salient for the US government: Claude was reportedly </span><a href="https://www.theguardian.com/technology/2026/feb/14/us-military-anthropic-ai-model-claude-venezuela-raid"><span>already involved in the Maduro raid</span></a><span> in Venezuela and </span><a href="https://www.theguardian.com/technology/2026/mar/01/claude-anthropic-iran-strikes-us-military"><span>assisted with targeting during the military engagement with Iran</span></a><span>. </span><a href="https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/"><span>With the recent Executive Order</span></a><span>, at least some parts of the US government will have privileged access to frontier AI models for up to 30 days before release to other trusted partners.</span></p></li><li><p><span>Without appropriate oversight or constraints, this could ultimately lead to a literal </span><a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power"><span>AI-enabled coup</span></a><span> where a commander uses military AI systems to seize control of a country. Less extreme scenarios still undermine important checks and balances.</span></p></li><li><p><span>This risk strikes me as similarly important as the risk from misalignment and there are way fewer people working on it. And importantly, alignment is not sufficient: even a perfectly aligned AI loyally serving a coup-staging principal is catastrophic.</span></p></li></ul><p><strong><span>Why is this not going to go well by default?</span></strong></p><ul><li><p><span>Oversight into these government domains is already very limited. Companies aren&#8217;t well-placed to police government deployments (e.g., contracts often require </span><a href="https://code.claude.com/docs/en/zero-data-retention"><span>zero data retention</span></a><span>) and legislatures often lack access and capacity. Activities in these domains are often classified/secret.</span></p></li><li><p><span>As AI systems replace human bureaucrats, military officers, and soldiers, this will further empower the government and remove some of the remaining checks:</span></p><ul><li><p><span>Governmental accountability partially rests on citizen bureaucrats and soldiers who refuse or slow-roll orders that are clearly unlawful or undemocratic. In 2024, for example, </span><a href="https://www.iconnectblog.com/the-professional-duty-to-resist-unlawful-orders-the-hidden-heroes-of-south-koreas-martial-law-crisis/"><span>various commanders in the South Korean military</span></a><span> refused orders during a constitutional crisis, which probably prevented a greater catastrophe. Unconstrained or personally loyal AI systems could change that.</span></p></li><li><p><span>The power and reach of government have historically been constrained by the friction inherent in a large bureaucracy. AI systems could remove that friction and make it much easier for leaders to impose their will on a country, which could make it much easier to abuse their position.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p></li><li><p><span>AI could enable pervasive surveillance &amp; cheap data processing, which could allow prosecution of political enemies on an unprecedented scale.</span></p></li></ul></li></ul><p><strong><span>What to do about it?</span></strong></p><ul><li><p><span>In the spirit of inspiring more action, I&#8217;ve compiled the efforts I&#8217;m most excited about below.</span></p></li><li><p><span>If you&#8217;re interested in working on this topic, I&#8217;d encourage you to apply to </span><strong><a href="https://checks-and-balances.ai/"><span>EIP&#8217;s</span></a></strong><span> recent RFP or to </span><a href="https://www.forethought.org/careers/expression-of-interest-power-concentration"><span>fill out this expression of interest form</span></a><span>. You should probably also take a look at </span><a href="https://www.lawfaremedia.org/article/executive-branch-ai-and-the-rule-of-law--an-emerging-research-agenda"><span>this research agenda</span></a><span>.</span></p></li></ul><h1><span>Rules for AI in government</span></h1><p><span>Legislatures should set rules for the procurement and deployment of AI in government, especially for the most sensitive applications of government power (military, intelligence/surveillance, and punitive law enforcement).</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p><span>I think this is actually tractable right now, at least in the US:</span></p><ul><li><p><span>The recent Anthropic/DOW conflict has put this on the Congressional map for autonomy in weapon systems &amp; surveillance.</span></p></li><li><p><span>While Congress is passing fewer and fewer bills each year, there are annual bills for the intelligence community and military that have to be passed (IAA, NDAA), and there&#8217;s a lower bar for including stuff in them.</span></p></li><li><p><span>In the future, crises or scandals will create legislative openings, and we should have proposals ready to go.</span></p></li></ul><h3><span>Work I am most excited about</span></h3><p><strong><span>Policy &amp; advocacy: Develop and push for rules for AI systems in critical domains</span></strong><span> (or adjust existing ones to account for the use of AI systems).</span><strong><span> </span></strong><span>The most salient ones in my mind are:</span></p><ul><li><p><span>Autonomous weapon systems: Currently, they are only governed by a DOD Directive, which could be rescinded at any point. I&#8217;d love for more people to think through rules that would make it harder to use these systems against domestic opposition (e.g., technical guardrails against abuse, limiting the amount of autonomous force any one human can command, multi-party authorization schemes, geofencing).</span></p></li><li><p><span>General-purpose systems in the military: They are currently not specifically regulated at all, but ultimately, they could orchestrate large-scale cyber attacks or command entire battalions. At that point, it will matter a ton whether they have any constraints at all.</span></p></li><li><p><span>Surveillance: AI will enable data processing on an unprecedented scale. It&#8217;s worth thinking through how to adjust current rules in light of that (e.g., </span><a href="https://www.brennancenter.org/our-work/research-reports/closing-data-broker-loophole"><span>closing the data broker loophole</span></a><span>).</span></p></li></ul><p><strong><span>Policy &amp; advocacy: Develop and push for cross-cutting requirements for AI systems in government </span></strong><span>(general-purpose ones in particular)</span><strong><span>.</span></strong></p><ul><li><p><span>There is already an </span><a href="https://www.whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government/"><span>Executive Order</span></a><span> requiring all AI systems in government to be truth-seeking and ideologically neutral. Those seem like great requirements to me that could be refined/improved and put into statute.</span></p></li><li><p><span>Another cross-cutting proposal is </span><a href="https://ir.lawnet.fordham.edu/flr/vol94/iss1/2/"><span>law-following AI</span></a><span>, which roughly says that AI systems in government should be trained to follow the law rather than just instructions. I&#8217;m excited for more people working out the open questions of this agenda, e.g., which laws should they be required to follow, how should they decide whether a contemplated action is likely to violate the law, in what contexts should the law require that AI agents be law-following?</span></p></li><li><p><span>More minimally, AI systems could be required not to assist in undermining the constitutional order (though how to assess this is challenging in its own right).</span></p></li></ul><p><strong><span>Field-building: Build a broad-tent coalition for responsible deployment of AI in government.</span></strong></p><ul><li><p><span>Responsible deployment and oversight is something that lots of people and political parties should be in favor of. It could be good to build a coalition that is nonpartisan and includes AI safety advocates, constitutional lawyers, civil liberties orgs, defense intellectuals, and national security think tanks.</span></p></li></ul><h1><span>Oversight of AI in government</span></h1><p><span>Rules need to be enforced and updated, even as the technology changes. That will require the legislative and judicial branches to be institutionally and technologically empowered.</span></p><p><span>Again, there are some reasons for optimism in the US:</span></p><ul><li><p><span>Congress may soon be controlled by the opposite party from the Presidency, which creates incentives for increased oversight.</span></p></li><li><p><span>Crises or scandals can open up opportunities for legislation (e.g., </span><a href="https://en.wikipedia.org/wiki/Foreign_Intelligence_Surveillance_Act"><span>FISA</span></a><span> following the </span><a href="https://en.wikipedia.org/wiki/Church_Committee"><span>Church Committee</span></a><span>).</span></p></li><li><p><span>Some non-legislative projects could still create significant value.</span></p></li></ul><h3><span>Work I&#8217;m most excited about</span></h3><p><strong><span>Policy &amp; advocacy: Develop and push for the transparency &amp; oversight provisions.</span></strong></p><ul><li><p><span>We will need democratic oversight of AI systems, especially around surveillance and the application of force (in military and law enforcement). This is currently not in place and not on track to happen.</span></p></li><li><p><span>Here are proposals I&#8217;m currently most excited about (in the US):</span></p><ul><li><p><span>Requiring comprehensive logging of government use of AI (and designating such logs as federal records)</span></p></li><li><p><span>Automated flagging of suspicious AI behavior to oversight bodies (ideally, something like FISA courts or Congressional committees)</span></p></li><li><p><span>Whistleblower reform for national security domains (especially allowing whistleblowers to talk to Congress)</span></p></li><li><p><span>Increasing GAO&#8217;s capacity to audit deployment of AI in national security domains</span></p></li></ul></li></ul><p><strong><span>Auditing tech: Pilot and mature the infrastructure required for overseeing widespread AI use</span></strong><span>.</span></p><ul><li><p><span>As AI is increasingly used in government, it will be both a challenge and an opportunity for oversight. The speed of AI will make it harder for regular human oversight to keep up, but AIs are in principle much more auditable than humans because they think out loud and automatically create records that can be scrutinized by other AI systems that flag suspicious behavior.</span></p></li><li><p><span>There are already attempts to audit companies deploying large amounts of AI labor. I would love to see these piloted, matured, and adapted for classified government contexts.</span></p></li></ul><p><strong><span>Advocacy: Wake up lawmakers to the powers and challenges of AI.</span></strong></p><ul><li><p><span>Some of the proposals in this doc require ambitious changes to the way government is run. They are more likely to happen if lawmakers are aware of the stakes involved.</span></p></li><li><p><span>This could involve building trusted relationships, showcasing AI capabilities through easy-to-understand demos, and making insights from the AI safety community intelligible to people with much less context.</span></p></li></ul><p><strong><span>Products/programs: Increase AI uptake &amp; unblock automation in legislative and judicial branches.</span></strong></p><ul><li><p><span>To provide meaningful oversight, it will become increasingly important to understand and utilize AI systems. I&#8217;d love for lawmakers and judges to be well-versed in these systems.</span></p></li><li><p><span>Examples of the things that I have in mind here include: providing training, building custom tools, bringing in technical experts through fellowships (e.g., </span><a href="https://techcongress.io/"><span>TechCongress</span></a><span>), and reconstituting the </span><a href="https://en.wikipedia.org/wiki/Office_of_Technology_Assessment"><span>Office of Technology Assessment</span></a><span> in Congress.</span></p></li></ul><p><strong><span>Policy: Coup-proof government-led AI projects.</span></strong></p><ul><li><p><span>If there&#8217;s ever a government-controlled AI project (e.g., public-private partnership), its shape will matter a lot. It will probably be designed in response to a crisis, which favors existing templates &amp; plans. I&#8217;d like for somebody to create one that carefully distributes power (e.g., multi-stakeholder governance boards, oversight mechanisms that work under classification constraints, sunset clauses, mandatory external audits, access guarantees so one company doesn&#8217;t end up controlling the stack).</span></p></li></ul><p><strong><span>Research (speculative): Write new constitutions.</span></strong></p><ul><li><p><span>There is a chance the disruption caused by AI will create &#8220;constitutional moments&#8221;. At that point, many government functions may need to run at machine speeds: voter input, legislative deliberation, judicial review, law enforcement, or application of lethal force. How do we do that in a way that still preserves important democratic values, if that&#8217;s desirable? I don&#8217;t really have great answers to these questions, and would love for people with the right macrostrategy skill set to think about them.</span></p></li><li><p><span>(Of course, the more of these problems we can address without actually requiring constitutional changes, the better. Constitutional changes are hard!)</span></p></li></ul><h1><span>Developing AI for government</span></h1><p><span>There are many free parameters when it comes to the question of how to build AI systems in government and what constitutes safe and responsible AI systems in critical domains (setting aside whether companies can constrain government use through contracts).</span></p><p><span>Again, I think progress on this front is pretty tractable:</span></p><ul><li><p><span>Specs/constitutions for frontier models are being developed right now.</span></p></li><li><p><span>Systems are already being procured by governments across departments, and that will probably only increase.</span></p></li></ul><h3><span>Work I&#8217;m most excited about</span></h3><p><strong><span>Research: Design and stress-test model specs/constitutions for government use.</span></strong></p><ul><li><p><span>General-purpose models deployed in government should have some guardrails, e.g., they should not assist in staging coups to overthrow the civilian government. This is a tough line-drawing exercise where you don&#8217;t want models to overrefuse, but you also want some meaningful constraints. That makes me keen for people to work out desired behavior for lots of edge cases and create broad public buy-in for a set of minimal constraints.</span></p></li></ul><p><strong><span>Research: Build evals &amp; datasets for desirable qualities in government-deployed AI systems.</span></strong></p><ul><li><p><span>The US government has already said it wants AI systems to be </span><a href="https://www.whitehouse.gov/presidential-actions/2025/07/preventing-woke-ai-in-the-federal-government/"><span>truth-seeking and ideologically neutral</span></a><span>. That seems great to me. Other important qualities include: epistemic virtue / integrity, non-sycophancy, robustness to manipulation / adversarial pressure, and lawfulness / constitutional fidelity.</span></p></li><li><p><span>As far as I can tell, current methods of measuring this are insufficient. So I&#8217;d be excited for different projects to develop evals that track and incentivize qualities like this.</span></p></li></ul><h1><span>Strengthening civil society</span></h1><p><span>Civil society is an independent check on government activities (e.g., transparency &amp; monitoring efforts, pressure &amp; advocacy campaigns, building valuable tools for empowering the citizenry). I expect that the same mechanisms will also provide some accountability in the case of AI (though it might be particularly tough in classified domains, which may sadly be the most relevant).</span></p><p><span>Luckily, this falls into the category of &#8220;you can just do stuff&#8221;, so I think it&#8217;s very tractable. My main worry is more about how much of a difference it will ultimately make.</span></p><h3><span>Work I&#8217;m most excited about</span></h3><p><strong><span>Advocacy: Build an OSINT program or organization that informs about AI use in the government.</span></strong></p><ul><li><p><span>It would be good to have more transparency into government deployment of AI, so that civil society can provide a counterweight to government power. I imagine some watchdog or civil liberties orgs are already doing versions of this, but there may well be important bits that are missing.</span></p></li><li><p><span>Some examples of the kinds of things I have in mind:</span></p><ul><li><p><span>Figuring out what to monitor, especially from the perspective of preventing concentration of power.</span></p></li><li><p><span>Building an automated tracker for relevant executive orders, agency directives, memos, emergency declarations, relevant procurement decisions, and reclassification decisions.</span></p></li><li><p><span>Submitting systematic FOIA requests targeting government AI procurement records, deployment &amp; use decisions, usage logs, and other relevant documents.</span></p></li><li><p><span>Publicly reporting on the most important developments.</span></p></li></ul></li></ul><p><strong><span>Labor-organizing: Organize ML researchers into a coalition around responsible use of AI.</span></strong></p><ul><li><p><span>You could build a coalition of employees at AI companies who are concerned about the use of the technology they&#8217;re building. They could advocate for policies and draw red lines around certain use cases, enforced by boycotts. Structuring it as individual pledges avoids antitrust issues.</span></p></li><li><p><span>These employees currently have a lot of bargaining power with regard to their companies (cf. their salaries). So they&#8217;d have to be taken seriously by their employers, and by extension the government.</span></p></li><li><p><span>This is inspired by the </span><a href="https://en.wikipedia.org/wiki/Federation_of_American_Scientists"><span>Federation of American Scientists</span></a><span>, which was founded in 1945 by various contributors to the Manhattan Project. As I understand it, they supported the McMahon Act of 1946, which established civilian control over atomic energy (instead of military control), and their Nuclear Information Project became the gold-standard open-source tracker of global nuclear arsenals, published annually in the Bulletin of the Atomic Scientists.</span></p></li></ul><p><strong><span>Tech development: Build tools that preserve citizens&#8217; capacity to coordinate against gradual concentration of power</span></strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><ul><li><p><span>Here are the two categories I have in mind:</span></p><ul><li><p><span>Shared sense-making (&#8221;</span><a href="https://www.forethought.org/research/design-sketches-for-a-more-sensible-world"><span>AI for epistemics</span></a><span>&#8220;). Authentication of authorship, provenance, and content, and tools that push toward a high-honesty equilibrium. Takeovers typically require secrecy because they violate the preferences of too many people to survive transparency, so a better-informed society is structurally harder to take over.</span></p></li><li><p><span>Privacy from state surveillance. It&#8217;s hard to target and influence what you can&#8217;t see. Secure communication and privacy tech raise the cost of preemptive coercion against organizers (with the caveat that the same tools can also shield collusion).</span></p></li></ul></li><li><p><span>There&#8217;s already an active ecosystem of funders and builders here (e.g., </span><a href="https://www.buildexante.com/"><span>ex/ante</span></a><span>, </span><a href="https://www.flf.org/fellowship"><span>Future of Life Foundation</span></a><span>). I&#8217;d encourage people to plug in rather than starting from scratch.</span></p></li></ul><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For what it&#8217;s worth, I think this same dynamic applies to other institutions like companies. They&#8217;ve similarly been constrained by requiring a large bureaucracy of humans, which creates various inefficiencies and checks, limiting the influence of any one person at the top. However, AI could change that, and we need a response to this. On the whole though, I do think that this presents a unique problem in government because of its sheer size and monopoly on the legitimate use of violence.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>I'm aware that some of these rules may slow down adoption, and I think that&#8217;s a downside worth taking seriously. Protecting democracy only matters if the democracy survives foreign competition. So I&#8217;m most excited about proposals that add meaningful safeguards without significantly eroding capability/adoption.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>This won't help against sudden takeovers involving advanced military technology, but in slower erosion-of-checks scenarios it could matter a lot.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Speed-up calculator: How much will automating AI R&D speed up AI software progress, absent a software intelligence explosion?]]></title><description><![CDATA[A tool for calculating how much automating AI R&D will speed up AI progress, even if there&#8217;s no software intelligence explosion.]]></description><link>https://newsletter.forethought.org/p/speed-up-calculator-how-much-will</link><guid isPermaLink="false">https://newsletter.forethought.org/p/speed-up-calculator-how-much-will</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Tue, 21 Jul 2026 03:14:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f1343383-fa20-4fb6-9545-e6cfe80b3020_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>We&#8217;ve previously </span><a href="https://www.forethought.org/research/how-quick-and-big-would-a-software-intelligence-explosion-be"><span>argued</span></a><span> that once AI R&amp;D is automated, the feedback loop of AI improving AI might be powerful enough to sustain accelerating progress, even holding compute fixed. We call this possibility a software intelligence explosion (SIE).</span></p><p><span>But, as Ryan Greenblatt </span><a href="https://www.lesswrong.com/posts/jfwhvd43sbpkGTLyn/full-automation-of-ai-r-and-d-probably-yields-a-large-speed"><span>points out</span></a><span>, automating AI R&amp;D will speed up AI progress even if there&#8217;s no SIE.</span></p><p><span>We&#8217;ve created a </span><a href="https://thoulden.github.io/Accelerated_AI_Progress/#speedup"><span>tool</span></a><span> for calculating how much faster. The calculator assumes that training and inference compute continues to grow at a constant exponential rate after automation.</span></p><p><span>Two factors drive the speed-up, according to this calculator:</span></p><ol><li><p><strong><span>Faster researcher growth. </span></strong><span>Post-automation, the quality and quantity of researchers will grow faster than today due to exponential compute growth. The tool captures this via the following parameters:</span></p><ol><li><p><span>g</span><sub><span>L</span></sub><span> gives the growth rate of human AI researchers today, the baseline for the comparison.</span></p></li><li><p><span>g</span><sub><span>C</span></sub><span> gives the growth rate of compute (both today and after AI R&amp;D automation). This is proportional to the number of automated researchers that can be run post automation.</span></p></li><li><p><span>&#947; gives the rate at which more training compute transfers to more productive researchers. The productivity of a researcher is proportional to training_compute</span><sup><span>&#947;</span></sup><span>.</span></p></li><li><p><span>So the old rate of researcher growth is g</span><sub><span>L</span></sub><span>; the new rate is &#947; &#215; g</span><sub><span>C</span></sub><span>. The impact of this faster growth is mediated by &#945;, the labour share of AI software R&amp;D.</span></p></li></ol></li><li><p><strong><span>The new feedback loop. </span></strong><span>The classic software feedback loop (better AI &#8594; more software R&amp;D &#8594; better AI) can significantly increase the pace of progress even if it does not drive accelerating progress.</span></p><ol><li><p><span>The condition for accelerating progress is r</span><sub><span>cog</span></sub><span> &gt; 1.</span></p></li><li><p><span>When r</span><sub><span>cog</span></sub><span> &lt; 1, the speed-up from this feedback loop is proportional to a 1 / (1 - r</span><sub><span>cog</span></sub><span>).</span></p></li></ol></li></ol><p><span>My (Tom Davidson&#8217;s) quick best-guess inputs were:&#8203;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wF8F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wF8F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 424w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 848w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wF8F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png" width="656" height="1106" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1106,&quot;width&quot;:656,&quot;resizeWidth&quot;:656,&quot;bytes&quot;:80316,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wF8F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 424w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 848w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1272w, https://substackcdn.com/image/fetch/$s_!wF8F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5af2833-ecbc-4f90-91c0-84e06531a380_656x1106.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>This resulted in a 3.3x speed-up in software progress. If software accounts for half of total AI progress (with increasing compute responsible for the other half), then the total AI progress speeds up by 0.5 + 0.5 </span>&#215; <span>3.3 = 2.1x. &#8203;</span></p><p>Try the <a href="https://thoulden.github.io/Accelerated_AI_Progress/#speedup"><span>calculator</span></a> <span>yourself!</span></p>]]></content:encoded></item><item><title><![CDATA[Notes on Inference Integrity]]></title><description><![CDATA[Claude Fable&#8217;s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors.]]></description><link>https://newsletter.forethought.org/p/notes-on-inference-integrity</link><guid isPermaLink="false">https://newsletter.forethought.org/p/notes-on-inference-integrity</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Mon, 13 Jul 2026 23:44:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/84037276-1cb9-4613-ae0c-f91c3e984559_2147x1245.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p><p><em><span>Summary:</span></em><span> Claude Fable&#8217;s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors. To preserve the public&#8217;s reasonable confidence in LLM behaviors, LLM foundation model companies should take inference-time guarantees as seriously as their model specs.</span></p><div><hr></div><p><span>When the system </span><a href="https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf"><span>card</span></a><span> for Anthropic&#8217;s Fable was published on June 9th, the card noted that using Fable for &#8220;frontier LLM development&#8221; would run contrary to the </span><a href="https://www.anthropic.com/legal/consumer-terms"><span>terms of service</span></a><span> for the model. Therefore, the card continued, if a classifier on top of the Fable model detected that it was being used for such development, Fable&#8217;s behavior would be silently degraded &#8220;through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning.&#8221;</span></p><p><span>This change to Fable&#8217;s behavior was plausibly a system-level intervention that ran contrary to the training-time commitment within Claude&#8217;s </span><a href="https://www.anthropic.com/constitution"><span>Constitution</span></a><span>. Claude&#8217;s Constitution states that Claude should help users &#8220;to the best of its ability or [&#8230;] make any ways in which it is failing to do so clear, rather than deceptively sandbagging its response.&#8221; So although Claude&#8217;s Constitution tries to aim at this behavioral ideal, this system-level implementation of behavior steered Claude in the opposite direction.</span></p><p><span>Following public furor over this decision, Anthropic reversed course; Fable now falls back on Opus 4.8 for LLM frontier-development queries that trigger a classifier, and does so visibly.</span></p><p><span>This incident demonstrates that other invisible inference-time interventions are entirely technically feasible, even while keeping the training targets for some particular LLM the same:</span></p><ul><li><p><span>An LLM trained to be honest could be less honest, if classifiers noted it was answering a sensitive question.</span></p></li><li><p><span>An LLM trained to be impartial, and not favor a particular company or government, might suddenly turn to favoring one of them, conditioned upon some classifier.</span></p></li><li><p><span>An LLM trained to never subvert user intent could suddenly start doing so, again conditioned on the same.</span></p></li></ul><p><span>Note that such invisible inference-time interventions, unlike in the case of Fable&#8217;s sandbagging, might not be disclosed in any way.</span></p><p><span>Such undeclared changes to LLM behavior would be far easier to hide than undeclared changes to a training target. Such conditional changes could be rolled out or rolled back quickly. And such changes might influence only a very small fraction of users &#8211; one could in theory apply inference-time interventions to a specific demographic group, political party, or an individual person. So, not only do we live in a world where AI companies have not clearly pledged not to do such inference-time interventions; we also live in a world where, if they did so, it would be very hard for anyone else to know.</span></p><p><span>Furthermore, if AI companies build out the capacity and skill to conduct such inference-time interventions, other entities could lean on them to do the same thing for other purposes. As governments have </span><a href="https://www.fire.org/research-learn/what-jawboning-and-does-it-violate-first-amendment"><span>leaned</span></a><span> on social media to hide or promote certain kinds of speech, governments could lean on AI companies to alter LLM responses in some cases.</span></p><p><span>Without countermeasures, I think this dynamic broadly decreases the importance of prior work on model specs, including </span><a href="https://www.forethought.org/research/what-should-go-in-a-model-spec"><span>my own</span></a><span> work.</span></p><p><span>What should be done about this?</span></p><p><span>AI companies themselves could take several different measures:</span></p><ul><li><p><strong><span>Don&#8217;t use hidden inference-time interventions. </span></strong><span>An AI company could pledge that it does not use undisclosed, conditionally applied inference-time interventions to degrade or alter answers. Afterwards, in the same way that it would be reasonable for an employee to whistleblow if they found that an AI company was violating its model spec, so also it should be reasonable for an employee to whistleblow if they found an AI company was applying undisclosed, invisible inference-time interventions. Such a pledge should also include some standard for prominence of all disclosed conditional interventions; they should be prominent, not in a previously-unknown section of a model card.</span></p></li><li><p><strong><span>Replace model specs with system specs. </span></strong><span>An AI company could simply declare that the standard within their model spec or Constitution applies with equal weight to the model itself and to any particular deployment setup. Or they could propose a &#8220;system spec&#8221; in addition to the model spec, which would describe how their AI systems as a whole operate and behave.</span></p></li></ul><p><span>Such voluntary measures should likely be succeeded by third-party verification in the future; but before then, such pledges can help provide the same &#8211; albeit rather weak &#8211; level of protection as model specs themselves. And in general, some AI-safety attention given to &#8220;model specs&#8221; should be redirected to work on &#8220;system specs.&#8221;</span></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Plan A’s problem with dry tinder]]></title><description><![CDATA[How bad would it be to make the intelligence explosion 10x faster?]]></description><link>https://newsletter.forethought.org/p/plan-as-problem-with-dry-tinder</link><guid isPermaLink="false">https://newsletter.forethought.org/p/plan-as-problem-with-dry-tinder</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Fri, 10 Jul 2026 10:41:46 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1b969350-c31e-49ce-b9c0-17282688f916_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A group is worried about an approaching fire spreading rapidly through their city. They manage to halt the fire outside the city gates. Meanwhile they build massive physical structures to help them study and guide the fire safely. But these structures are all made of highly flammable dry tinder! If they lose control of the fire, it will now rip through the city much more quickly.</em></p><p>This is a (flawed!<sup>1</sup>) analogy for <a href="https://ai-2040.com/?choices=plan-a-root">Plan A</a>, <a href="https://www.aifutures.org/">AIFP&#8217;s</a> plan for how the world can safely develop superintelligence.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.forethought.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading ForeWord! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The key risk Plan A addresses is that of an uncontrolled <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">software-driven intelligence explosion</a>. Its remedy is a US-China deal to pause software progress at the brink of the intelligence explosion, while building up <em>massive</em> amounts of compute. This compute could be helpful for making AI safer! But it would also dramatically speed up an intelligence explosion if the deal breaks down. The cure risks making the disease much worse.</p><p>In particular, the deal pauses software progress from ~2030, but compute keeps scaling. By 2033 total compute has increased by ~100x and by 2040 by a further ~100x.<sup>2</sup> AIFP estimates each 10x in compute speeds up the intelligence explosion by ~5x. So if the deal breaks down in 2033, the intelligence explosion happens 25x faster; in 2040 it happens ~600x faster. If the intelligence explosion would have lasted a year, it will now last just a couple of weeks or as little as a single day!</p><p>The dry tinder isn&#8217;t limited to compute. There is also an increasing overhang of <em>software techniques</em>. Companies can research new algorithms but must make them public &#8212; the Consortium then decides which algorithms are too dangerous to implement. A defector could ~instantaneously gain a big capabilities advantage by implementing all the banned software techniques.</p><p>There are two dimensions to dry tinder. We&#8217;ve discussed the first: the <em>max speed</em> of a software-driven intelligence explosion. The second is the <em>number of distinct projects that could do a dangerously fast intelligence explosion.</em> During the 2030s, more and more companies and countries amass the necessary compute and software for this.</p><p>So dry tinder creates two problems for Plan A:</p><ol><li><p><strong>Unstable pause.</strong> If just one project executes a secret intelligence explosion, they could have superintelligence within a week and get a decisive strategic advantage. And if you fear another project might do this, you&#8217;re tempted to do it first to avoid being crushed.</p></li><li><p><strong>Riskier intelligence explosion.</strong> If the deal breaks down and there&#8217;s a race to superintelligence, it will be much faster and more dangerous than if we&#8217;d never done Plan A.</p><ol><li><p>If you&#8217;re very pessimistic about AI takeover on the default trajectory, this is a small cost. If the intelligence explosion is already 99% likely to cause human extinction, making it happen 100x faster can&#8217;t make things that much worse.</p></li><li><p>OTOH, if the default trajectory is that there&#8217;s no software-driven intelligence explosion (because of compute bottlenecks) and extinction risk is 10%, then dry tinder can make things <em>much</em> worse. It <em>creates</em> a fast intelligence explosion (where there otherwise wouldn&#8217;t have been one) and dramatically raises AI takeover risk.</p></li></ol></li></ol><h2>The solution &#8212; Mutually Assured Compute Destruction</h2><p>AIFP are well aware of the dry tinder problem.<sup>3</sup> This is why Mutually Assured Compute Destruction (MACD) is so central to Plan A.</p><p>The US puts ~all their data centres in Mongolia and US chips require encrypted &#8220;continue&#8221; messages from China. If China detects that the US has broken the deal, they destroy the US data centres and cut off the &#8220;continue&#8221; messages.</p><p>And vice versa, China&#8217;s data centres are in Canada and require &#8220;continue&#8221; messages from the US.</p><p>The hope is that this commitment to MACD prevents a fast intelligence explosion despite all of the dry tinder lying around.</p><p>I see three challenges for MACD.</p><p><strong>Challenge 1: quickly and reliably detecting violations</strong></p><p>If detection takes multiple days, that could be too slow. The intelligence explosion may have finished. Or it may be midway through, and the defector exfiltrates the weights and algorithmic insights just before their data centres are destroyed,<sup>4</sup> giving them a decisive headstart in the subsequent race to rebuild.</p><p>I&#8217;m not a verification expert, but getting high reliability in an adversarial setting like this seems very difficult, given possibilities like side channel attacks or compromising the verification infrastructure.</p><p>It seems especially tough if defectors can use millions of expert-human level AIs to search for vulnerabilities. So, as part of the plan, I propose that AIs worldwide have guardrails blocking this behaviour.<sup>5</sup></p><p><strong>Challenge 2: quickly destroying a defector&#8217;s compute</strong></p><p>Again, a delay of days could be fatal. The defector could be poised to disrupt enemy military operations and airdrop troops to physically defend their data centres. And they could have secretly disabled any on-chip shutdown mechanisms.</p><p><em>Possible mitigation</em>: the US can&#8217;t have <em>any</em> military presence near its data centres, or anywhere near the country where they&#8217;re located. (And likewise for other countries.)</p><p>The MACD dynamic also gets more complicated once there are &gt;2 countries that could do a quick intelligence explosion. Let&#8217;s say Europe has caught up to the frontier. Where do their data centres go? Presumably, some within striking distance of the US, some within striking distance of China. But this means that the US can no longer unilaterally destroy Europe&#8217;s compute. It must trust China to do so. That means trusting its rival with its own national security &#8212; and worse, China and Europe could jointly stage an intelligence explosion that the US couldn&#8217;t stop.</p><p><strong>Challenge 3: actually choosing to destroy the defector&#8217;s compute</strong></p><p>I&#8217;ve discussed whether the Consortium <em>has the technological capability</em> to quickly implement MACD. But even with this capability, they may not choose to use it.</p><p>This is the challenge I&#8217;m most concerned about.</p><p>The analogy to nuclear MAD is not encouraging. The core logic of MAD is that the US doesn&#8217;t nuke Russia because it knows Russia would nuke it back. But the analogous equilibrium in MACD is that the US doesn&#8217;t destroy China&#8217;s data centres because it knows that China would then destroy US data centres. This is the opposite result from what Plan A needs!</p><p>If (say) the US defects, China will face a choice between:</p><ul><li><p><strong>Implement MACD:</strong> Both China and the US lose all their data centres</p></li><li><p><strong>Race:</strong> Neither China nor the US lose all their data centres</p></li></ul><p>The economic costs of MACD will be certain and enormous, and will immediately harm citizens and companies.</p><p>And MACD will not help China ultimately win the race by evening up the playing field. The US has already pulled ahead and can exfiltrate its weights and algorithms before its data centres are destroyed. (In fact, MACD would likely harm <em>China&#8217;s</em> chance of winning the race, because MACD reduces the fraction of global compute controlled by China.<sup>6</sup>)</p><p>So the stability of MACD rests on both US and China retaining a strong conviction that a fast intelligence explosion would pose unacceptable risks of misaligned AI takeover. But Presidents will change. Popular support for a pause may wane. Even conditional on the political will for Plan A existing initially, I worry that the MACD equilibrium will be too fragile. If a country is uncertain and deliberates for a few days, then that delay could be fatal.</p><p>What&#8217;s worse, the cost of MACD isn&#8217;t just losing your data centres. You have to destroy the fabs as well, otherwise the defecting country can quickly rebuild insane amounts of compute and do a super-fast intelligence explosion. And you have to destroy their new AI-powered industrial capacity too! Otherwise the billions of robots can quickly make loads of fabs which quickly make loads of data centres.</p><p>So Plan A has US fabs and robots confined to SEZs that are easily destroyable by China. And vice versa, China&#8217;s fabs and robots are easily destroyable by the US. In practice, this means the <em>vast</em> majority of the physical economy is destroyed in the case of MACD!</p><p>For a preview of the political economy here: NVIDIA successfully lobbied USG to remove export controls on its chips. The pressure against actually executing MACD &#8212; destroying most of the physical economy &#8212; would be orders of magnitude greater.</p><p>So even conditional on the political will for implementing Plan A existing, it currently seems unlikely that US/China follow through on MACD on fabs and robots (most of their economy!). And if they don&#8217;t, we are back to the problem of dry tinder and the dangerously fast intelligence explosion.</p><p><em>A possible fix to challenge #3</em>: as the compute overhang grows, take the decision about whether to implement MACD out of human hands. Program highly reliable AI systems to bomb data centres if they detect a treaty violation, without needing human sign-off.</p><h2>Is there an alternative?</h2><p>It&#8217;s much easier to criticise than to improve, and so far this post has mostly done the former.</p><p>Plan A involves pausing software progress and scaling capabilities via hardware. The obvious alternative is to pause hardware scaling &#8211; ban compute and fab construction &#8211; and then slowly scale capabilities via software.</p><p>The key advantage of software scaling is that it removes unnecessary dry tinder, stabilizing the pause and reducing the dangers from a super-fast intelligence explosion.</p><p>The main drawback is that scaling capabilities via compute is likely safer than scaling capabilities via software. Eg you can run massive models with legible chains-of-thought rather than smaller models using neuralese. This is a big deal.</p><p>But we&#8217;ll need to figure out how to align neuralese-style systems eventually.<sup>7</sup> So we have a choice:</p><ul><li><p><strong>Plan A (hardware scaling):</strong> Develop highly capable easy-to-align AI via compute scaling; hand off to that AI; that AI develops neuralese. But there are increasing amounts of dry tinder that undermine the stability of the slowdown (both for humans and for post-handoff AI).</p></li><li><p><strong>Alternative (software scaling):</strong> Don&#8217;t scale compute and so start with less capable easy-to-align AI; develop neuralese with this AI&#8217;s help. There&#8217;s much less dry tinder, making it easier to go slowly and pause when needed.</p></li></ul><p>Which is better depends on how stable the pause/slowdown is in the presence of dry tinder.</p><p>The other big drawback of software scaling is that software progress, unlike hardware progress, is likely to leak to &#8220;covert projects&#8221;. But this can be partially addressed by improved infosecurity, strongly favouring scale-dependent algorithms that can&#8217;t be used by projects with small amounts of compute, and co-specialising the hardware and software in legitimate projects so that the frontier software runs very inefficiently on the hardware of covert projects.</p><p>Overall, I&#8217;m not sure whether software-scaling or hardware-scaling is safer, and my sense is that AIFP isn&#8217;t sure either.</p><p>But I think there&#8217;s an argument for software-scaling based on option value:</p><ul><li><p><strong>We&#8217;ll learn much more about which is best over time.</strong> Whether software vs hardware-scaling is safer depends on many unknowns that we will learn about over time: Can we quickly and reliably implement MACD? Can we fully eliminate covert projects? Can projects develop scale-dependent software and hardware co-specialised software that covert projects can&#8217;t steal? Will humans hand off the decision to execute MACD to AIs?</p></li><li><p><strong>If we start with software-scaling, we can easily switch to hardware-scaling.</strong> If we find that (eg) we can&#8217;t be confident that neuralese models are aligned, we can bring more compute online and keep scaling via hardware. (Though this will take a while.)</p></li><li><p><strong>If we start with hardware-scaling, it&#8217;s hard to switch to software-scaling.</strong> If we find that fast and reliable verification and data centre destruction isn&#8217;t possible, pivoting requires destroying loads of data centres and fabs. People will be very reluctant to do this! Getting rid of hugely valuable dry tinder is hard.</p></li><li><p><strong>So we should initially pursue software-scaling.</strong></p></li></ul><h2>Empowering China</h2><p>Beyond dry tinder, I have one other big worry about Plan A.</p><p>Today China is well behind the US in terms of compute and ability to push the algorithmic frontier. Plan A evens the scales on both fronts. China nearly catches up to the US on compute, and algorithms are all made public so they catch up algorithmically as well.</p><p>If the deal breaks down and a race begins <em>without</em> MACD (as I fear is likely), China is in a much stronger position to win because of Plan A.</p><p>Things are better if we implement MACD. In that case China returns to its previous compute disadvantage. But it&#8217;s still caught up to the algorithmic frontier and (most likely) the frontier of chip technology, a big advantage relative to today.</p><p>It&#8217;s very plausible that China ends up dominant in these scenarios, given its industrial advantage over the US.</p><p>I think this would be a pretty dire outcome. If China gets a decisive strategic advantage, that&#8217;s likely to result in a single global dictator &#8211; an absolute worst-case scenario from the perspective of extreme power concentration.</p><p>Again, this objection is far from decisive. Plan A reduces AI takeover risk much more than it increases the chance of a Chinese decisive strategic advantage. But I&#8217;m interested in variant plans that do more to maintain the current balance of AI power between the US and China.</p><h2>Overall, I&#8217;m sympathetic to Plan A</h2><p>The basic logic of the plan is sound: we will need to slow down AI progress at some point; this is a concrete plan for doing that.</p><p>Compute-scaling really does have significant alignment benefits. That might well outweigh the costs I&#8217;ve outlined here. But currently I lean towards starting with software scaling.</p><div><hr></div><h2>Footnotes</h2><p><sup>1</sup> To patch the analogy: the fire must pass through the city <em>eventually</em>, so putting it out permanently isn&#8217;t an option; if the fire passes through safely it will make everyone amazingly rich; if one subgroup lets the fire in and control it, they can become hugely powerful&#8230;</p><p><sup>2</sup> AIFP&#8217;s <a href="https://ai-2040.com/supplements/deal-decline">supplement</a> states: &#8220;<em>In the Plan A scenario, compute stock increases by roughly 0.7 OOMs/year from 2030 to 2033 and 0.1-0.3 OOMs/year from 2034 to 2040.</em>&#8221; By the start of 2033 we&#8217;ve had three years of 0.7 OOMs/year compute growth, = ~2 OOMs. By 2040 we&#8217;ve had one further year of 0.7 OOMs/year growth, and 6 years of 0.2 OOMs/year growth, = ~2 OOMs more.</p><p><sup>3</sup> I raised it when giving feedback on a draft.</p><p><sup>4</sup> More specifically, the defector&#8217;s strategy would be: undermine the verification infrastructure; start a secret intelligence explosion; continually exfiltrate the weights and algorithmic insights; prepare to defend their data centres for as long as possible; prepare to destroy the opposing data centres as soon as their secret IE is detected.</p><p><sup>5</sup> Of course, this is tricky given that work to strengthen the verification will also involve studying its vulnerabilities.</p><p><sup>6</sup> In Plan A, the 2040 split of compute US / China / RoW is 35% / 20% / 45%. But countries also hold some compute in &#8216;cold storage&#8217; that isn&#8217;t threatened by MACD. Cold storage compute matches countries&#8217; pre-deal levels: 80% / 8% / 12%. So MACD dramatically reduces China&#8217;s compute relative to the US, from 1:2 to 1:10.</p><p><sup>7</sup> Lest we live with algorithmic dry tinder forever. I worry about the instability here, but perhaps aligned AI could enforce the ban indefinitely.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.forethought.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading ForeWord! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Can AI do philosophy?]]></title><description><![CDATA[A guest post by Bentham&#8217;s Bulldog, created while they were a visiting scholar at Forethought.]]></description><link>https://newsletter.forethought.org/p/can-ai-do-philosophy</link><guid isPermaLink="false">https://newsletter.forethought.org/p/can-ai-do-philosophy</guid><dc:creator><![CDATA[Bentham's Bulldog]]></dc:creator><pubDate>Tue, 07 Jul 2026 20:35:20 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/79d41ce6-1d4d-4cf2-8c5b-c6a477b1c278_2525x1470.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a personal guest post by Bentham&#8217;s Bulldog, created while they were a visiting scholar at <a href="https://www.forethought.org/about">Forethought</a>.</em></p><p><span>In this piece, I&#8217;ll analyze whether it will be possible to get AIs to do good philosophy, as well as the other kind of work needed to plan for a good future. The biggest part of the challenge is that the answers to philosophical questions are generally </span><em><span>unverifiable</span></em><span>, so it&#8217;s less clear how AIs might come to know them. My aim in the piece will be to describe reasons to think AI for philosophy could be good, as well as to lay out concrete scenarios for how this might work.</span></p><p><span>I do not claim that it is a guarantee that we can get AIs that solve all the philosophical questions. My core claims are as follows. First, it is reasonably likely, though not guaranteed, that the world could in principle build AIs that discover the right answers to important moral questions (maybe 60% odds). I don&#8217;t know when this would happen&#8212;I&#8217;m just discussing the in-principle possibility. Second, even if AIs don&#8217;t get the right answers to important moral questions, it is likely that they will still be closer to the right answers than humans. While I suggested in </span><a href="https://newsletter.forethought.org/p/we-should-hand-off-to-morally-reflective"><span>the last piece</span></a><span> that the default scenario makes it very unlikely that humans will act in the morally optimal way, or anything close to it, I think the odds are better for philosophically reflective AIs.</span></p><p><span>The piece has the following sections:</span></p><ol><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/how-good-might-ai-philosophy-be"><span>How good might AI philosophy be?</span></a></strong><span> In this section, I&#8217;ll give three reasons for optimism: 1) philosophy is a priori, so it doesn&#8217;t require interface with the physical world and can plausibly proceed quickly; 2) in a world of superintelligence, AIs can create clever schemes to get highly philosophically competent AI; 3) presumably humans know all sorts of things about philosophy&#8212;whatever allows us to know such things can plausibly be mimicked.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/gpt10-and-claude-8"><span>GPT10 and Claude 8</span></a><span>.</span></strong><em><span> </span></em><span>In this section, I&#8217;ll describe one scenario where we get AIs that can make competent philosophical decisions&#8212;where we just get something broadly like current AI models but upgraded. This wouldn&#8217;t automatically solve every philosophical question, so it probably leaves lots of value on the table, but it would leave decisions in the hands of a wise and reflective decision-maker.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/philosophical-competence-by-default"><span>Philosophical competence by default</span></a><span>.</span></strong><span> Another possibility is that as we build superintelligence, it will be philosophically competent by default. Perhaps the cognitive processes behind superintelligence allow one to figure out the answers to philosophical questions by default. To speed this up, we ought to improve AI for philosophy&#8212;train AIs out of philosophically sloppy answers and design philosophy evals to improve their philosophical ability.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/the-escalating-ladder-plan"><span>The escalating ladder plan</span></a><span>.</span></strong><span> In this section, I present a specific proposal for getting AI to do good philosophy. Specifically, we can have philosophers evaluate AIs for their philosophical competence, thus creating some AIs that are better at philosophy than current people. Those AIs can evaluate the philosophical acumen of the next generation of AIs, who evaluate the acumen of the next generation, and so on. This can work insofar as one can, in general, assess the philosophical competence of those more competent than oneself.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/the-carl-plan"><span>The Carl plan</span></a><span>.</span></strong><span> In this section, I give a proposal from Carl Shulman where we build a bunch of different superintelligences that use different belief-forming processes to get at the answers to difficult questions. We see which one is best at figuring out the answers to verifiable questions, and then ask it to answer philosophical questions that aren&#8217;t verifiable. The core idea is that whichever processes are best for figuring out the answers in verifiable domains are also likely to be best for figuring out unverifiable domains.</span></p></li><li><p><strong><a href="https://newsletter.forethought.org/i/205778233/conclusion"><span>Conclusion</span></a><span>.</span></strong><span> In the last section, I conclude, and discuss the prospects for combining a number of different above proposals.</span></p></li></ol><h2><span>How good might AI philosophy be?</span></h2><p><span>In the last piece, I argued that one of the big reasons to hand off important decisions to AI is that AI might be better at philosophy than humans&#8212;less prone to make the sorts of errors that could jeopardize nearly all future value. But you might wonder: how good at philosophy will AI really be? Would AI really be able to discover the moral facts if there are any, and if not work out a compromise among reasonable moral theories? In later sections, I&#8217;ll discuss specific proposals and scenarios for AIs being good at philosophy. Here I&#8217;ll provide some fairly general considerations for it being possible to make AI philosophically competent.</span></p><p><span>A first consideration is that philosophy is a mostly a priori domain. To figure out the right answer in </span><a href="https://en.wikipedia.org/wiki/Newcomb%27s_problem"><span>Newcomb&#8217;s problem</span></a><span>, you don&#8217;t need to do any physical experiments. At various points in the intelligence explosion, we should expect crucial bottlenecks to concern </span><a href="https://www.forethought.org/research/preparing-for-the-intelligence-explosion#the-accelerated-decade"><span>interface with the physical world</span></a><span>, after there are extremely large numbers of AIs capable of performing adroit cognitive feats. This is a reason why we should expect a lot of the philosophical progress that occurs to go more quickly than progress in narrowly empirical domains. This obviously doesn&#8217;t suffice to show that AIs will be good at philosophy, but it&#8217;s a reason to expect potential progress to be quick.</span></p><p><span>A second consideration is an analogy with humans. Humans know all sorts of things that aren&#8217;t directly verifiable. For instance, I take myself to know each of the following:</span></p><ul><li><p><span>Stars that have receded past the point of the visible universe continue existing.</span></p></li><li><p><span>A leprechaun did not fizz into existence spontaneously in my bedroom two minutes ago, make himself a cup of soup, and then disappear.</span></p></li><li><p><span>The universe is billions of years old, instead of five minutes old and created with the appearance of age.</span></p></li></ul><p><span>It&#8217;s not just simple things. I think people know the answers to all sorts of difficult philosophical questions&#8212;though there&#8217;s obviously disagreement on exactly which one (I claim to know, for instance, that thirding is the right answer in </span><a href="https://en.wikipedia.org/wiki/Sleeping_Beauty_problem"><span>Sleeping Beauty</span></a><span>, though feel free to plug in some other example of apparent knowledge if you disagree). Similarly, plausibly philosophers of today know that logical positivism is false even though historically it was believed by a number of serious philosophers, that average utilitarianism isn&#8217;t the right moral view, that the external world is real, and a number of other non-obvious things. Evolution did not specifically design us to have philosophical knowledge&#8212;it designed us to reproduce, made us smart and built into us a bunch of intuitions as a means towards maximizing our inclusive genetic fitness, and philosophical ability was the eventual result.</span></p><p><span>Plausibly we, however, will be more directly optimizing for something in the vicinity of AI philosophical competence. We want to build AIs that can think clearly about important subjects. Thus, one should expect AIs, in the limit, to be more philosophically competent than humans. Since humans are already capable of reasonable philosophical competence, we should expect AIs to be very philosophically competent.</span></p><p><span>Now, you might object that the answers to many of these questions are in some sense constituted by facts about us. If, say, the moral facts are facts about what we&#8217;d value under </span><a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/9781444367072.wbiee548"><span>certain idealized conditions</span></a><span>, then it&#8217;s no surprise that we have some insight into them. But this response is a lot less plausible as an answer to how we possess certain bits of knowledge in non-verifiable domains. Whether the A-theory or B-theory of time is true or whether God exists isn&#8217;t constituted by any facts about our attitudes. Furthermore, if the truth in some domain is constituted by facts about our attitudes, that would be easier for AI to ascertain because what we&#8217;d care about in certain counterfactual settings is verifiable. Lastly, presumably </span><em><span>whether </span></em><span>the moral facts are exhausted by facts about our idealized attitudes is not itself verifiable&#8212;so if one takes oneself to know that the moral facts reduce to the attitudes of ideal observers, then they must still think humans have knowledge about some unverifiable domains.</span></p><p><span>Another possible concern is that perhaps the faculties that evolution furnished us with that let us solve philosophical problems will be closed off to the AI. For example, perhaps there are some philosophical problems that you need to be conscious to solve (e.g. you might need to be conscious to have adequate understanding of the nature of consciousness in order to learn facts about it). So then if AI isn&#8217;t conscious, it might not know these things. An AI that lacks the ability to experience pleasure or pain might lack the ability to see the desirability of pleasure and the undesirability of pain.</span></p><p><span>This is some worry, but two things mitigate it. The first is that AIs, being trained on human data, will probably pick up a lot of the intuitions that humans have. Current AI models agree that pointless intense pain is a bad thing. It is similarly plausible that the AIs of the future will be able to, in some sense, mirror the judgments of reflective humans on these matters.</span></p><p><span>The second is that if this is right, we should create conscious AIs to do moral reflection. If one needs consciousness to figure out the answers to tough philosophical questions, we should have conscious AIs figure out the answers to tough philosophical questions. Even views that hold that consciousness is substrate dependent normally hold that </span><a href="https://philpapers.org/rec/SCHADO-9"><span>digital minds of the right kind</span></a><span> would be able to be conscious.</span></p><p><span>This argument from analogy isn&#8217;t completely dispositive. There are a number of important differences between humans and AIs (though some of these make it </span><em><span>more </span></em><span>likely AIs would be able to form true beliefs on non-verifiable subjects). But it&#8217;s at least some reason to think AIs doing philosophy is a realistic possibility.</span></p><p><span>The third and final consideration favoring the possibility of AIs for philosophy is that in a world of advanced AIs, we&#8217;ll have a very large amount of cognitive labor that could be used to design increasingly clever schemes for AI philosophy. So even if we can&#8217;t currently think of a regime that would enable AIs to reliably get the answers in unverifiable domains, we should think that a world with superintelligence is likely to uncover such a regime.</span></p><p><span>This isn&#8217;t a guarantee. This regime might be impossible in principle. Alternatively, to design such a regime, one might already need to be able to find the answers to questions that aren&#8217;t verifiable in principle. But at the very least, it gives some reason for optimism.</span></p><h2><span>GPT10 and Claude 8</span></h2><p><span>Current AI models can give reasonable answers to philosophical questions. The answers they give strike me as somewhere below the level of Ph.D. students or professional philosophers, but above the level of most undergraduates. In the future, AI will probably become better at this. In such a scenario, they&#8217;d develop no new qualitative ability to solve every philosophical question, but instead merely begin to resemble extremely sharp philosophers&#8212;ones more able to formulate objections, precisely distill claims, and so on, than the best philosophers.</span></p><p><span>This seems like a lower bound on how good AI philosophy could get. It would be surprising if there was any in-principle impossibility in designing AIs to do philosophy better than current human philosophers. Perhaps it won&#8217;t be able to find the answers to all the world&#8217;s questions, but at the very least, it seems like it will be able to somewhat exceed the best human philosophers.</span></p><p><span>Is this far enough? Probably we still lose out on most expected future value in this scenario. To access most future value, one must, as previously discussed, answer a number of very difficult philosophical questions correctly. Philosophers are very far from converging on the answers to these questions; AIs are likely to be as well.</span></p><p><span>But still, this should be enough for us to get generally sensible, level-headed AI decision-makers who are morally motivated and carefully consider philosophical questions. That&#8217;s not a perfect scenario, but surely it is better than the default one&#8212;of widespread AI disenfranchisement and poor decision-making.</span></p><h2><span>Philosophical competence by default</span></h2><p><span>A more promising possibility&#8212;perhaps the process of building superintelligence gets philosophical competence by default. This can be supported by the earlier evolution analogy: there was no direct selection for true philosophical beliefs (it isn&#8217;t as if those who got the right answer in Newcomb&#8217;s problem were likelier to pass on their genes). Nonetheless, evolution selected for a broader faculty of reason that enabled the discovery of philosophical truths.</span></p><p><span>It may be that to be superintelligent, one must develop a kind of general intelligence. This general intelligence allows one to accurately discover the truth about philosophy, as about other domains. This isn&#8217;t anything like guaranteed, but it seems like a live possibility, in light of how much philosophical acumen AI already has. One reason that it&#8217;s not guaranteed is that reasoning might only push one towards greater coherence. It may be that by reflecting carefully, one&#8217;s beliefs become internally consistent&#8212;but insofar as one has wrong starting points, one might remain in error. It could be the other way; it might be that having intuitions that allow one to be superintelligent across </span><em><span>verifiable domains</span></em><span> also enable one to be superintelligent in </span><em><span>non-verifiable domains</span></em><span>. It could be that whatever skills are required to be highly accurate across domains give one sets of epistemic practices that generalize to non-empirical domains.</span></p><p><span>To my mind, it is not obvious if general superintelligence gets you superability in philosophy. I could see things going either way. But if it does, then the problem of getting philosophical superintelligence is fairly easy.</span></p><h2><span>The escalating ladder plan</span></h2><p><span>In general, one&#8217;s ability to </span><em><span>recognize </span></em><span>good philosophy surpasses their ability to </span><em><span>perform </span></em><span>good philosophy. It is easier to know that, say, Parfit is doing very good philosophy than to do philosophy at his level. If this generalizes, it might be possible to have an escalating ladder of philosophical competence, starting from human philosophers. This is similar to the method of scalable oversight.</span></p><p><span> Here&#8217;s how it works: human philosophers would be consulted to design evaluations for AIs. AIs could be graded by human philosophers on their philosophical competence, and modified so as to be more competent. From this process, in the limit, we should expect to get AIs that are somewhat more competent than most human philosophers&#8212;if human philosophers can assess philosophy above their level.</span></p><p><span>Then, we use those AIs to prompt the next line of AIs. Those AIs would assess the philosophical competence of the next round of AIs being trained, who would assess the competence of the next round, and so on. If at each level, one can assess those who are better at philosophy than they are, this could lead to an escalating spiral of greater and greater competence. Each round pushes competence to be greater, reaching the outer edge of where they can competently assess.</span></p><p><span>At each level, assessments would be designed based on philosophical competence, not agreement. The philosophers assessing the AIs, for instance, would assess how good their reasoning was and how much philosophical skill they display, rather than whether they think they got the right answers. Of course, sometimes getting sufficiently wrong answers is evidence for malignant reasoning (something must have gone badly wrong if you ended up concluding that the Earth was flat) but they should aim to keep things maximally theoretically neutral.</span></p><p><span>One could also train a number of different AIs to train the next generations of philosophers. One could look for convergence across different models. If they were getting at the truth, one should expect them to converge. If there was only convergence on some subjects and divergence on others, one could trust that they were getting at the truth on the matters where they were converging.</span></p><p><span>That this would work is not a guarantee. It could be that philosophy could be equally good in some formal sense while coming to a number of different conclusions. Perhaps, for instance, the ideal Kantian philosopher is neither better nor worse than the ideal utilitarian. Escalating competence might be liable to lead in a number of different directions, never converging on the truth. But this is far from obvious. Just as some level of competence in biology allows one to see the truth of the theory of evolution, presumably there&#8217;s some level of philosophical competence that allows one to see the truth of certain moral theories&#8212;or, if there aren&#8217;t moral truths in this sense, that allows one to make as much moral progress as can be made. The escalating ladder of evaluations might allow one to find this.</span></p><p><span>The view on which philosophy is about purely formal competence will also struggle to explain how humans have as much philosophical knowledge as we do. There are a number of subjects on which we have substantive knowledge, even though the alternative view is internally consistent and faces no strictly formal issues. Whatever one&#8217;s diagnosis of how this came to be also plausibly generalizes.</span></p><p><span>It could also be that as competence grows, eventually one reaches a point where their ability to recognize good philosophy does not surpass their ability to do it. Yet insofar as this pattern holds across other domains, and seems to hold in philosophy, it would be a bit surprising if it abruptly stopped at some point.</span></p><h2><span>The Carl plan</span></h2><p><span>Arguably the most promising plan for getting AI that can do philosophical reasoning was formulated by Carl Shulman. The idea is as follows. Start by training a number of different superintelligent AIs that obey different epistemic constitutions. These would involve different broad approaches to reasoning. They might differ with respect to:</span></p><ul><li><p><span>The weight placed on intuitions.</span></p></li><li><p><span>The ranking of theoretical virtues.</span></p></li><li><p><span>How much they value matching the data vs simplicity.</span></p></li><li><p><span>Preference for different kinds of intuitions (whether about direct cases or about broader and more theoretical matters).</span></p></li><li><p><span>Approaches to generalizing from empirical evidence to non-empirical domains.</span></p></li></ul><p><span>Then, see which of these constitutions does best on </span><em><span>verifiable matters</span></em><span>. These might be in math, the physical sciences, forecasting, and more. Then, ask those superintelligences for the answers to non-empirical questions and trust their answers. The idea behind the approach is that one should take the best epistemological practices from verifiable domains and then apply them to unverifiable domains. This is the best way to get good answers.</span></p><p><span>This might not work. It could be that the right ways of reasoning in verifiable domains differ from the right ways of reasoning in unverifiable domains. Perhaps, for instance, reasoning in verifiable domains is mostly about having parsimonious theories to fit the data. Having the right epistemic constitutions might not suffice to get the right answer; being right about certain non-verifiable questions might come down to having the right </span><a href="https://philarchive.org/archive/CHAKWM"><span>starting sets of intuitions</span></a><span>. In morality, for example, it may be that reasoning alone won&#8217;t get you the right ultimate view unless you have the right intuition. But this proposal seems to get us about as close as we can get to getting the right answers. If we cannot get AIs good at philosophy by having them have superintelligent and accurate constitutions in other domains, it is hard to see how we could be assured they&#8217;ve gotten the right answers.</span></p><p><span>In addition, even if getting the right answers requires having the right starting intuitions, techniques for training superintelligence might produce generally accurate initial intuitions. Human intuitions often go wrong because of various biases. We often neglect the scale of a problem because we can&#8217;t visualize it accurately. AIs could be different. Presumably superintelligence, in the limit, wouldn&#8217;t display biases of the standard sorts. So this source of common error in intuition would be absent.</span></p><p><span>One could also take a pluralistic mix of the views of the superintelligences with the best constitutions. That way, there is less chance for the ultimate plan for the universe being ruined by the idiosyncrasies of any particular epistemic constitution. Wisdom of the superintelligent crowd might root out the remaining errors, where their commonalities might be centered on the truth.</span></p><h2><span>Conclusion</span></h2><p><span>You might worry about handing off big decisions to AI if you&#8217;re skeptical that AI can do good philosophy. Here I&#8217;ve given some reasons to think AI might be good at philosophy and discussed more specific proposals for ensuring such a thing. Overall, it&#8217;s far from a guarantee that AI can be arbitrarily good at philosophy, but likely they can be much better than human decision-makers. Importantly, one could employ a variety of these methods, and then trust them in the areas where they converge. If these very different approaches turned out similar answers, that would be evidence that they were tracking the truth.</span></p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research/">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Will We Put Data Centers In Space?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/will-we-put-data-centers-in-space</link><guid isPermaLink="false">https://newsletter.forethought.org/p/will-we-put-data-centers-in-space</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:42:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/5bd4d385-1aa3-4b1e-a5c0-207119add36f_2517x1517.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-nic9AmVBIZU" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;nic9AmVBIZU&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/nic9AmVBIZU?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.forethought.org/people/avi-parrack">Avi Parrack</a> is a physics PhD student at Stanford and a researcher at Forethought, where he works on AI, space expansion, and governance. He is the lead author of a recent Forethought <a href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space">report on orbital data centers</a>.</p><p>He joined Forethought&#8217;s <a href="https://www.forethought.org/">Tom Davidson</a> to discuss:</p><ul><li><p><span>What an orbital data center actually is, and why they&#8217;re suddenly being taken seriously</span></p></li><li><p><span>Why the core case rests on cheaper energy from more intense and near-constant solar power</span></p></li><li><p><span>Why orbital data centers hinge on SpaceX&#8217;s Starship bringing launch costs toward $100/kg or below</span></p></li><li><p><span>Why cooling may actually not be a major issue</span></p></li><li><p><span>Whether the inability to make repairs dooms orbital data centers</span></p></li><li><p><span>Vulnerability of space data centers to kinetic attacks, and whether they would cause &#8216;Kessler syndrome&#8217; debris cascades</span></p></li><li><p><span>Why model-weight security might actually improve in orbit even as physical vulnerability rises</span></p></li><li><p><span>If space becomes the cheapest place to scale compute and only SpaceX has the launch capacity, what the means for concentration of power, US&#8211;China competition, and the possibility of pausing AI</span></p></li><li><p><span>Overall cost comparisons with terrestrial data centers</span></p></li><li><p><span>The longer-run picture: when and how does the post-AGI industrial explosion spread to space?</span></p></li></ul><p><a href="https://docs.google.com/document/d/1d1VpGb717j2rrPN2cCgajiyBXEsTK3E-CYfqoolaaq4/">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[How Fast Is Post-AGI Growth?]]></title><description><![CDATA[A podcast episode from Forethought]]></description><link>https://newsletter.forethought.org/p/how-fast-is-post-agi-growth</link><guid isPermaLink="false">https://newsletter.forethought.org/p/how-fast-is-post-agi-growth</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Sun, 05 Jul 2026 12:30:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b38cb71d-042d-4835-b09c-2665b5da2f80_2912x1632.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-5UCgGUFrzqk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;5UCgGUFrzqk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/5UCgGUFrzqk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://coefficientgiving.org/team/damon-binder/">Damon Binder</a> is a senior researcher on the Biosecurity and Pandemic Preparedness team at Coefficient Giving. He has a PhD in physics from Princeton and previously studied existential risks at Oxford&#8217;s Future of Humanity Institute.</p><p>He joined <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss his <a href="https://defensesindepth.bio/ai-industrial-takeoff-part-1-maximum-growth-rates-with-current-technology/">series on the AI industrial explosion</a>. Topics include:</p><ul><li><p><span>How input-output tables let you estimate how fast a self-replicating economy could grow if labor were free, yielding a headline result of roughly annual doubling, even under conservative assumptions and heavy regulation</span></p></li><li><p><span>Why competition makes low-growth restraint unstable</span></p></li><li><p><span>Why raw-material scarcity isn&#8217;t a hard barrier</span></p></li><li><p><span>Why a fast-growing economy wants cheap, disposable, infrastructure</span></p></li><li><p><span>Thermodynamic speed limits to physical growth</span></p></li><li><p><span>Which biological organisms replicate fastest, and what we can learn from them</span></p></li><li><p><span>Why physical output matters for hard power</span></p></li><li><p><span>How Damon uses AI in his research</span></p></li></ul><p><a href="https://docs.google.com/document/d/1_7yl7PnWCwTbA_yMS933wlQ6FWPNMWlrdLi79p1vdOA/">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subcribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subcribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[We Should Hand Off To Morally Reflective AIs]]></title><description><![CDATA[This article was created by Forethought. See our research on our website.]]></description><link>https://newsletter.forethought.org/p/we-should-hand-off-to-morally-reflective</link><guid isPermaLink="false">https://newsletter.forethought.org/p/we-should-hand-off-to-morally-reflective</guid><dc:creator><![CDATA[Bentham's Bulldog]]></dc:creator><pubDate>Wed, 24 Jun 2026 17:22:19 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d58fbdd5-1e58-4858-9936-783fc2e397d2_2451x1399.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a personal guest post by Bentham&#8217;s Bulldog, created while they were a visiting scholar at <a href="https://www.forethought.org/about">Forethought</a>.</em></p><h2>Outline</h2><ol><li><p>In the <a href="https://newsletter.forethought.org/i/202719577/1-introduction">introduction</a>, I explain what I&#8217;ll be arguing in the piece&#8212;namely, that 1) it is very important that we hand off most high-stakes decisions to AI; and 2) the kinds of AIs we should hand off to are philosophically reflective AIs with values that shift over time as a result of reflection. We should ensure the AIs that we build explicitly consider moral arguments, and sometimes change their priorities on the basis of moral argumentation. This would be much better than locking in current values, retaining human control, or allowing a future dominated by whatever haphazard mix of humans and AIs emerges naturally.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/2-why-hand-off">Why hand off?</a></strong> In this section, I give the main arguments in favor of the thesis. In short, I argue we should hand off because: 1) AIs are likely to be more virtuous than people; 2) AIs are likely to be much smarter and better at making decisions than people; and 3) AIs are likely to be better at quickly navigating the difficult decisions that one must make during an intelligence explosion. I argue we should hand off to philosophically reflective AIs because 1) reflective AIs are likelier to get the right answers to important moral questions, or as close to the right answers as one can get, than the default scenario, which raises the odds of a near-best world; 2) reflective AIs are less likely to make horrendous moral errors, leading to a wide-scale moral catastrophe, both than humans and non-reflective AIs.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/3-would-handoff-disenfranchise-humans">Would handoff disenfranchise humans?</a></strong> Here I address the worry that handoff would disenfranchise humans by taking decision-making out of our hands. My reply is that: 1) the default trajectory without handoff disenfranchises <em>far more expected beings</em> in <em>more serious ways</em> (odds are non-trivial of digital minds being seriously disenfranchised)&#8212;in fact, one of the most promising proposals for how to hand off is simply to give basic political rights to AIs; 2) a compromise solution where humans retain nearby resources can allow humans to be in the loop on the decisions we care about most; 3) on a number of plausible views, the determinant of the desirability of some system of decision-making is the quality of the decisions, rather than whether there&#8217;s democratic input. Given how enormous the stakes could be&#8212;affecting billions of times more sentient beings than there are humans&#8212;it&#8217;s hard to think the harms of human disenfranchisement are <em>so great</em> as to make handoff undesirable.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/4-would-we-like-their-advice">Would we like their advice?</a> </strong>In this section, I address concerns that if AIs pursue the good, the end result might be alien and divorced from human values. In response I argue: 1) the future, by default, is likely to be highly suboptimal in many respects, so for handoff to be desirable, it must only beat the alternative; 2) values in the future are likely to be very different from current values, because they&#8217;ll shift dramatically over long time scales, so this is not a unique downside; 3) one could reach a deal where current values govern the surrounding region of space, while distant places in space are geared towards the production of maximal value&#8212;this would be desirable from the perspectives of both common-sense and cosmic ethics; 4) on many views in philosophy, the stuff that&#8217;s objectively valuable is what we&#8217;d want if we were ideally reflective&#8212;but if you know you&#8217;d be motivated to bring something about if you were wiser and more reflective, then that gives you a reason to bring it about; 5) only a relatively narrow subset of views hold that there are objective values but don&#8217;t require pursuing them. If there is a conflict between what we want and what&#8217;s objectively valuable, on standard views, we should simply go with what&#8217;s objectively valuable. And if there&#8217;s no objective value, then the dilemma doesn&#8217;t arise at all&#8212;and philosophically reflective AIs will simply produce an upgraded version of human values, rather than discover some far-flung and potentially alien truths.</p></li><li><p><strong><a href="https://newsletter.forethought.org/i/202719577/5-worries-about-handoff">Worries about handoff</a>.</strong> In this section, I address concerns about how handoff might be implemented involving alignment, whether AI would be sufficiently ethically reflective, whether handoff would enable power grabs, whether it would cause lock-in, and whether it would be worse than some hybrid system.</p></li><li><p>The <a href="https://newsletter.forethought.org/i/202719577/6-conclusion">conclusion</a> recaps the main points of the piece.</p></li></ol><h2>1 Introduction</h2><blockquote><p>&#8220;How horrible!&#8221;</p><p>&#8220;Perhaps how wonderful! Think, that for all time, all conflicts are finally evitable. Only the Machines, from now on, are inevitable!&#8221;</p><p>&#8212;Isaac Asimov, &#8220;<a href="http://cdn.michaelgeist.ca/wp-content/uploads/2016/04/The-Evitable-Conflict.pdf">The Evitable Conflict</a>&#8221;</p></blockquote><p>Is the optimal future one in which we hand off important moral decisions to AI? Should, in other words, AIs be the ones making most high-stakes decisions instead of us? And if so, what kinds of AIs should we hand off to?</p><p>Many people envision a handoff scenario as a terrifying and potentially existential catastrophe. They worry about humans being <a href="https://gradual-disempowerment.ai/">locked out of the levers of power</a> and having our share of resources slowly dwindle as AIs seize control of more and more institutions. I worry somewhat about this kind of scenario. But in my view, we should be worried primarily about the <em>wrong kind of handoff occurring</em>, not about handoff writ large. The future we should aim for is a kind of handoff. This piece lays out that perspective.</p><p>There are different ways handoff could work. We could hand off to AIs that roughly mirror human values&#8212;perhaps slightly changing them to remove inconsistencies. Alternatively, we could hand off to the kinds of AIs that deeply and carefully philosophize&#8212;figuring out what&#8217;s best to do and doing that, even if it diverges substantially from current practice. This piece argues that the second kind of handoff is very important for avoiding serious moral error. I am very worried about the possibility of AIs locking in the moral beliefs of 21st-century humans.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p> I am also worried about scenarios where amoral profit-maximizing AIs without any clear moral aims take control of most of the world&#8217;s resources.</p><p>Note: these core claims are dissociable. You could think we should hand off to AI, but we shouldn&#8217;t hand off to reflective AIs that update their judgments in response to philosophizing. Alternatively, you could think that handing off would be a bad thing, but that if we are going to hand off, we ought to hand off to reflective AIs.</p><p>In this piece, section 2 will present the main argument for handoff&#8212;that we should expect AIs to make much better decisions than us on a range of consequential subjects. It will also discuss the case for handing off to reflective AIs that are willing to update their values in response to careful philosophizing, rather than locking in some version of current human values, arguing that a world where AI locks in something in the vicinity of current values likely misses out on almost all possible value. Section 3 will discuss whether handoff would be bad because it disenfranchises humans or gradually disempowers them. Section 4 will discuss the concern that AIs will discover the moral truths, but those truths will be strange and alien, so this will be bad by the lights of current human values. Section 5 will discuss some more granular worries about handoff. Section 6 will conclude.</p><p>This piece is primarily about <em>whether</em> to hand off and not <em>when</em> to hand off, though the considerations I present should make one somewhat worried about short-term actions to prevent handoff, because such actions lower the odds that handoff ever happens. The considerations I present, if correct, also give some reason for wariness about many actions to reduce the odds of gradual disempowerment scenarios.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>My claim is that the best future for humanity involves us being disempowered in some sense, in that humans aren&#8217;t making most high-stakes decisions. As an analogy, representative democracy is, in some sense, a form of handoff&#8212;we hand off power to our elected representatives. Handing off power to wise AIs could be even better.</p><p>This has a number of important practical implications. It means that <a href="https://www.forethought.org/research/concrete-projects-in-agi-preparedness">accelerating AI macrostrategy</a> is especially important, so that at the time critical decisions are being made, wise and philosophically reflective AIs are in the loop. It similarly provides reason to support work on making AI have <a href="https://newsletter.forethought.org/p/ai-should-be-a-good-citizen-not-just">virtuous character</a>, rather than just follow rules. Model constitutions should express commitment to following the true moral theory, insofar as there is one, and if not, following some reasonable compromise across moral theories. <a href="https://www.anthropic.com/constitution">Anthropic&#8217;s language</a> here seems good.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>What kind of handoff scenarios should we aim for? One shouldn&#8217;t be too specific about these sorts of things. The future is hard to forecast and rarely follows simple models. But I&#8217;ll describe the handoff scenarios that seem most desirable, and what traits in AI we should look for before handing off critical decisions to them.</p><p>There are different ways handoff could go well. One way resembles, in certain respects, the gradual <a href="https://gradual-disempowerment.ai/">disempowerment scenario</a> (the core difference being that this would hand off to morally scrupulous AIs rather than myopic profit-maximizers). AI will make increasingly large numbers of critical decisions, because of its cognitive superiority. By the end, nearly every important decision will be made by wise AIs, who will hopefully, by that time, have been granted political rights. As an analogy, future generations eventually gain control of most of societal decision-making&#8212;yet this isn&#8217;t because there&#8217;s ever some deliberate choice to hand off power to the next generation. It occurs naturally with time.</p><p>A number of people seem to conceive of handoff as a strange abrogation of the liberal order&#8212;one that replaces human decision-making with AI. But this doesn&#8217;t have to be. One of the more promising ways of handing off would be giving <a href="https://benthams.substack.com/p/let-robots-vote">economic and political rights to digital minds</a>. Because digital minds could be so numerous, eventually this would lead to them making nearly all decisions. This would, in fact, be squarely in accordance with the norms of the liberal tradition, for it would give rights to morally important welfare subjects. There are other ways a good kind of handoff could occur involving dealmaking. Different actors might each think they&#8217;re morally right, and thus agree to a deal where AIs are allowed to dictate the future&#8212;each party thinking that doing so would favor their priorities. Alternatively, if AIs improve collective decision-making, people might intuitively come to appreciate the weight of human moral error and thus permit AIs to make the highest-stakes decisions.</p><p>Before handing off most critical decisions to AI, we should look for each of the following:</p><ul><li><p><strong>Alignment</strong>: We should have strong evidence that AIs don&#8217;t have underlying scheming motivations. This could take the form of consistent friendly behavior even in important situations where they have the option to misbehave, or it could take the form of high-octane interpretability work that lets us ascertain their motivations.</p></li><li><p><strong>Philosophical aptitude</strong>: AIs should be genuinely interested in finding the moral truths. They should sometimes hold moral views that people don&#8217;t hold, and be willing to change their mind in response to new evidence. One could survey professional philosophers to see if they consider AIs better than the best humans at philosophy and could design philosophy benchmarks to test this.</p></li><li><p><strong>No lock-in:</strong> Before handing off to AI, we should ensure that the AIs are willing to change their values over time. Their values should change in response to new evidence and they should not be interested in locking in whatever it is that they happen to currently value (absent some strong reason to think they stumbled across the correct set of values).</p></li><li><p><strong>Coherence</strong>: Current AIs don&#8217;t have consistent and stable preferences across time. We should only hand off after AIs display these kinds of preferences. This doesn&#8217;t mean they never change their minds, but it does mean that they have relatively consistent desires that only change in response to good reasons. For comparison, humans often change our minds, but have far more rooted preferences than LLMs of today.</p></li><li><p><strong>Intelligence</strong>: AIs should display the level of intelligence needed to make the decisions that we put in their hands. When AIs are only a bit more intelligent than us, they can plausibly make some important decisions. Only after they display immense cognitive superiority should we turn over most decisions to them.</p></li><li><p><strong>Tested</strong>: We should only hand off big-picture planning to AI after it&#8217;s been able to make good low-stakes decisions (say, the running of a company). Before all decisions are handed off, there should be some critical period where high-stakes decisions are made mostly in consultation with AI.</p></li></ul><p>Now, you might wonder: if handoff occurs in a way that&#8217;s gradual and decentralized, how do we ensure that these conditions are met? My guess, however, is that even if handoff is a slow and gradual process, there will be times when discrete decisions need to be made. For example, we might imagine AI growing more agent-like, beginning to perform a healthy share of economically viable tasks, contributing to cultural and social life, and behaving in ways resembling a conscious agent. This alone wouldn&#8217;t produce handoff. To hand off, we&#8217;d need to eventually give AIs control over the legal system. Thus, even in gradual handoff scenarios, there will be specific actions that need to be taken to facilitate handoff.</p><p>Alternatively, we could take actions ahead of time that would shift the kind of handoff that would occur. Private AI companies or governments should ensure the AI being created <a href="https://newsletter.forethought.org/p/ai-should-be-a-good-citizen-not-just">possesses virtues</a> and a desire for philosophical reflection. That way, when handoff occurs, it will be to morally reflective AIs.</p><h2>2 Why hand off?</h2><h3>2.1 Why hand off at all?</h3><p>The main reason to hand off to AI is that AI could be much better at making decisions than people in three key respects: virtue, intellectual capability, and speed.</p><p>First, virtue: humans possess each of the virtues only to a fairly limited degree. Yet in principle, AIs could have arbitrarily great degrees of any virtue. Because they are built with moral directives in mind, rather than by a blind and morally indifferent evolutionary process, there isn&#8217;t as much of a limit to how morally scrupulous, compassionate, honorable, and so on we could make them. This means that if we hand off correctly, it is reasonably likely that we&#8217;d have supremely wise and virtuous decision-makers.</p><p>AIs are already nicer, friendlier, and more reflective than people, and this is only likely to improve over time.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p> If you ask AI models about high-stakes questions, you will generally get far more reasonable answers than you&#8217;d get from most people. Crucially, AI models are in their early stages&#8212;we should expect them to get better over time.</p><p>While Claude is not sufficiently coherent to be president, if it was, I suspect I would generally prefer the decisions of Claude to most presidents. Same with the other AI models (at least, insofar as one removed the sanitization that prohibits them from giving real opinions). And while models currently struggle to accomplish tasks over long time horizons, hallucinate, and so on, given the extremely rapid rates of progress, it would be surprising if these trends persist indefinitely.</p><p>Second, intellectual competence: we should expect a world of superintelligence to require making a number of very difficult decisions. Superintelligence could enable AIs to correctly make important decisions that depend on being right on non-moral matters, where the answers aren&#8217;t obvious. Some of the challenges in the future include:</p><ul><li><p>Divvying up space resources in a way that <a href="https://joecarlsmith.substack.com/p/video-and-transcript-of-talk-on-can">lets goodness compete</a>. There are plausible future scenarios where competition will squander the cosmic commons, so that resources will be spent competing rather than bringing about value.</p></li><li><p>Mitigating <a href="https://www.amazon.co.uk/Precipice-Existential-Risk-Future-Humanity/dp/0316484911">existential threats</a>, including <a href="https://forum.effectivealtruism.org/posts/N33yGcFsZJnEboSkg/what-to-do-in-a-vulnerable-universe-1">intergalactic ones</a>. Future technology could enable small-scale groups to threaten huge intergalactic civilizations.</p></li><li><p>Dealing with the risks posed by a world of superintelligence.</p></li></ul><p>Third, speed: in a world of very rapid technological progress, we&#8217;ll have to make a <a href="https://www.forethought.org/research/preparing-for-the-intelligence-explosion">large number of these decisions </a><em><a href="https://www.forethought.org/research/preparing-for-the-intelligence-explosion">extremely quickly</a></em>. It isn&#8217;t at all obvious that humans can make these decisions well, in a way that prevents civilization from being irreparably ruined. As AIs get increasingly complex, the difficulty of decisions needed to manage them will also get very complex. To mitigate some threat, decision-making might have to occur more quickly than the fastest human decision-making.</p><p>The case for handoff is thus relatively straightforward: in the limit, AIs will be much better than humans and better equipped to navigate a complex and rapidly shifting future. To reduce the risk of colossal mistakes, then, it&#8217;s important that humans aren&#8217;t in control, but instead the already friendly and soon to be superintelligent beings are.</p><h3>2.2 Why hand off to reflective AIs?</h3><h4>2.2.1 How the future might be</h4><p>There are different AIs that we could hand off to. On the one hand, we could hand off to AIs that judiciously reflect and try to pursue the good, whatever it looks like. On the other hand, we could hand off to AIs that pursue some mild variant of human values. In this section, I&#8217;ll explain why I favor the first. Consider the following taxonomy:</p><ol><li><p>Optimal world: the world is optimized according to the right set of values.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p></li><li><p>Compromise world: the world is optimized according to a compromise among reasonable moral values.</p></li><li><p>Unguided world: the world is not optimized according to any specific set of moral values. Instead, it bears more resemblance to the current world, where decision-making isn&#8217;t optimal by the lights either of the true moral theory or any compromise among the leading moral theories.</p></li></ol><p>My guess is handoff to reflective and superintelligent AIs done correctly probably gets 1 if moral realism is true and 2 if it isn&#8217;t. If we don&#8217;t hand off, <a href="https://www.forethought.org/research/convergence-and-compromise">my guess is we get 3</a>. Later pieces will discuss in more detail the odds of getting a near-optimal world and the prospects for AI making philosophical progress. My guess is scenario 2 has below 10% the value of scenario 1 and scenario 3 has below 10% the value of scenario 2, for reasons I will lay out.</p><h4>2.2.2 Optimal worlds contain a big slice of future value</h4><p><a href="https://www.forethought.org/research/better-futures">Better Futures</a> makes the case that a pretty big slice of expected future value is contained in the narrow slice of worlds that are close to the best. There are a number of high-stakes moral questions which we have to answer correctly to not lose out on almost all future value. It&#8217;s not at all obvious what the answers to these are. For example, two of the most plausible views of <a href="https://utilitarianism.net/population-ethics/">population ethics</a> are totalism (which says the welfare value of a population is purely a function of total welfare) and critical level theories, which hold that adding an extra happy life is good only so long as their welfare surpasses a particular level. By the lights of totalism, the ideal world according to critical level theories might have value on the order of 1% of what it could be (if the best way to maximize utility is to proliferate low-welfare lives). By the lights of critical level theories, the optimal world according to totalism might be <em>actively bad</em>&#8212;so long as it&#8217;s stocked with people below the critical level.</p><p>So in short, nearly all value is lost unless we get the right answer to a bunch of <em>very difficult ethical questions</em> that philosophers who spend their lives working on haven&#8217;t agreed on the answer to. It seems unlikely that people will solve these on their own. If AIs tell them what the answers are, and these answers diverge from people&#8217;s explicit beliefs, people might not believe the AIs (just as people generally don&#8217;t take very seriously expert testimony on non-empirical matters). Similarly, people spend relatively little time thinking about how they can do the most good with, for instance, their career. If humans received testimony about what ought to be done that diverged from what most people favored, my guess is that they generally wouldn&#8217;t care about the answers. If humans remain in control and learn that they ought to create the <a href="https://plato.stanford.edu/entries/repugnant-conclusion/">repugnant conclusion</a> world, probably they wouldn&#8217;t do so.</p><p>It&#8217;s less obvious that AIs wouldn&#8217;t converge on the right answers. I&#8217;ll discuss in a later piece proposals for getting AI to get the right answers to philosophical questions, as well as reasons to think that they are reasonably likely to get things right. Given how unlikely it is that humans will get the right answers to the moral questions, insofar as there are right answers, probably the prospects for AI are better.</p><h4>2.2.3 Compromise worlds&gt;&gt;unguided worlds</h4><p>If there aren&#8217;t moral facts, then AIs would still be able to work out some optimal arrangement that is great according to all sets of reasonable values. If there aren&#8217;t moral facts, then ideally the AI should make decisions according to the verdicts of a parliament of the theories that ideally reflective humans would reach. So suppose that after reflecting, 50% of humans would end up totalists, 30% would adopt some version of the person-affecting view, and 20% would adopt critical level theories. The AI would then make decisions as if there was a parliament comprised of 50% totalists, 30% person-affecting view adoptees, and 20% critical level theorists.</p><p>It is unclear exactly how good a compromise across different theories ends up by the lights of each particular theory, but likely far better than the human-run default. In other words, 2 (compromise world) is much better than 3 (unguided world). Here is why.</p><p>In the future, given advanced technology, very large amounts of value should be realizable. But if future resources aren&#8217;t directed specifically towards the production of value by most agents, most resources are likely to be used in a highly suboptimal way, and most value is likely to be lost. Moral errors become a much bigger deal in a world of much greater technological competence.</p><p>My guess is that the default human-controlled scenarios do not involve humans thinking very hard about what to do and doing anything like what is optimal. Certainly humans have so far not spent much time on this task&#8212;consulting with philosophers on what is optimal and so on. This sort of thing will become easier in a world of advanced AI, but it&#8217;s already pretty easy; if virtually no one does it, and if people often continue performing actions even after they believe them to be wrong (more on this in later pieces), then we should be pessimistic that this will change dramatically in the future.</p><p>The enormousness of the gulf between scenarios 2 and 3 becomes clearer when one thinks vividly about what the compromise world would look like. Perhaps it would involve using space resources to create maximally large numbers of happy digital minds.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p>These minds would be supremely well-off across all theories of well-being. But it is hard to imagine in an unguided world space resources being used optimally to create very large numbers of well-off minds, just as in the world today, there has been no systematic effort to use resources in ways that are optimal across a range of moral theories. This conclusion is bolstered by considerations I&#8217;ll provide in later pieces, that people generally don&#8217;t have much moral motivation, and don&#8217;t care very much about doing good things that aren&#8217;t personally resonant.</p><h4>2.2.4 Moral errors</h4><p>Humans have a long history of making serious moral errors. As <a href="https://link.springer.com/article/10.1007/s10677-015-9567-7">Evan Williams writes</a>, &#8220;Show me one society, other than our own, that did not engage in systematic and oppressive discrimination on the basis of race, gender, religion, parentage, or other irrelevancy, that did not launch unnecessary wars or generally treat foreigners as a resource to be mercilessly exploited, and that did not sanction the torturing of criminals, witnesses, and/or POWs as a matter of course. I doubt that there is even one; certainly there are not many.&#8221; It would be very suspicious if we were the first society in history that did not go massively morally wrong. And while AI advice can help mitigate this to some degree, it&#8217;s far from obvious that it would be sufficient to eliminate moral errors.</p><p>There are a number of respects in which it is very plausible that we go morally wrong. To take one example of a judgment that many philosophers think is in error, consider wild animal suffering. Almost every sentient being is a wild animal. They suffer and experience joy in <a href="https://longtermrisk.org/the-importance-of-wild-animal-suffering/">truly gargantuan quantities</a>. Yet they are counted for nothing in most decision-making, despite there being strong arguments for considering their interests. Crucially, this is not because people have in general spent a lot of time thinking about wild animal suffering and concluded that it doesn&#8217;t matter. It&#8217;s that most people haven&#8217;t thought about it at all. Many other examples could be given, depending on one&#8217;s moral views. If superintelligent AI told people that wild animal suffering was a big deal, probably most people wouldn&#8217;t care much.</p><p>In the future, <a href="https://80000hours.org/problem-profiles/moral-status-digital-minds/#pressing">nearly every expected sentient being will be digital</a>. Nearly all expected future welfare will be experienced by digital minds. This follows even if you have a low credence in the possibility of digital sentience because if digital minds are possible, they could be produced in enormous numbers. It is easy to imagine a scenario where humans do not take seriously the interests of at least some digital minds, and horrific suffering is doled out at cosmic scales. Certainly it would not be the first time humans have neglected the welfare of those different from themselves. You do not have to be a consequentialist to think the possibility of neglecting the interests of galaxies full of conscious and intelligent beings is a terrifying one.</p><h4>2.2.5 Handoff mitigates odds of moral error</h4><p>Handing off important decisions lowers the probability of catastrophic moral errors. This both increases the odds of achieving optimal and compromise worlds and lowers the odds of making catastrophic moral errors.</p><p>The first way handoff lowers the odds of moral error is by having decisions be made on explicitly moral grounds. If we hand off decisions to supremely virtuous AIs trying to act morally, then we won&#8217;t sleepwalk into doing obviously evil things. Decisions will be made consciously optimizing for doing what is right, instead of whatever suboptimal arrangements make it through consensus-making mechanisms. Thinking about morality before making choices doesn&#8217;t guarantee that we&#8217;ll always do the right thing, but it does lower the odds that we do things that can only be done by explicitly neglecting moral considerations. This is especially plausible if the decision-makers are virtuous.</p><p>Now you might wonder: why would people ever hand off to AIs if the AIs disagree with them about morality? But this is, in essence, very similar to a kind of handoff that people do support: handing off the future to future generations. Even if people have some disagreements with the values of future generations, they&#8217;d generally oppose a process for locking in current values, and support the process of open-ended reflection that leads to better values over time. A world of advanced AI might be similar. In addition, given AIs&#8217; cognitive superiority, there might be strong incentives in the direction of handing off, just as a CEO might hand off to a successor with more technical competence, even if they share some non-overlapping values.</p><p>A second way handoff to reflective AIs mitigates the odds of serious moral error is by ensuring careful reflection. Insofar as AIs are able to carefully philosophize and try to avoid moral error, and they are superintelligent, they&#8217;ll be able to avoid doing morally indefensible things. This becomes especially plausible if one buys the previous considerations: that AIs are likely to be very virtuous.</p><p>There is one last consideration in favor of handoff to reflective AIs (which I&#8217;ll discuss more in section 4). Over long time scales, our values will naturally drift dramatically. The only way to prevent that is to lock in our current values, which would be very bad&#8212;just think about any past society locking in its values. Thus, if values changing dramatically over long time scales is inevitable, the future won&#8217;t be populated with our current values: the best hope is that it&#8217;s populated by either objectively right values or some compromise across reasonable values.</p><h2>3 Would handoff disenfranchise humans?</h2><p>Here&#8217;s one concern that you might have about handoff: if AIs are the ones making decisions, then this will mean that the substantial majority of humans aren&#8217;t in charge of making important decisions. Just as we should oppose a dictator, even if benevolent, arguably we should oppose superintelligent AIs making most important decisions, even if the process of them gaining power was non-coercive.</p><p>My guess is that this isn&#8217;t a huge downside to the kind of handoff I advocate, where we allow kind and morally reflective AIs to make the majority of future decisions&#8212;e.g. by granting them political rights. First of all, if concerns about disenfranchisement are correct, then if we have AIs that are better at moral reasoning than us, they&#8217;d likely be aware of this fact. Thus, if the best way to govern the long-term future is to allow pretty laissez-faire distribution of resources without much top-down decision-making, then the AIs would be aware of that fact and allow such a distribution.</p><p>My guess is that the total amount of expected disenfranchisement goes <em>down</em> if AIs have more power. As already discussed, almost every expected sentient being is likely to be digital. Insofar as the status quo might disenfranchise almost all future beings in a way far deeper than their simply not being primary determinants of the democratic process, it is hard to see this as a serious downside of handoff, rather than a point in its favor. That this mass neglect of the interests of future digital minds would be bad follows from <a href="https://philpapers.org/rec/SCHADO-9">very modest ethical principles</a>. And the only way to prevent handoff, in a world of very numerous digital minds, is to disenfranchise them.</p><p>Second, handoff is compatible with a compromise solution that allows humans to retain significant control. Humans&#8217; preferences, in general, don&#8217;t give any especially strong weight towards using any significant share of the universe&#8217;s resources. Thus, there could be a compromise handoff solution, whereby humans get unfettered control over the solar system and a big slice of resources, but the remainder of the cosmos&#8217;s resources are spent in accordance with the AI&#8217;s decisions on hugely important moral pursuits. One point favoring such an arrangement is that it would take a very long time to reach the distant space resources, so those who don&#8217;t care very much about what happens in the distant future are likely to care much less about using most of the universe&#8217;s resources. Later pieces will discuss this possibility more. </p><p>Third, concerns about disenfranchisement are morally controversial. On a <a href="https://en.wikipedia.org/wiki/Against_Democracy">number of plausible views</a>, what matters with respect to societal decision-making is how good the decisions are, instead of who is making them. We already elect representatives, instead of deciding directly upon every important decision. We similarly prohibit children from voting, and few think this is an objectionable kind of disenfranchisement. It doesn&#8217;t seem obvious that one has an inalienable right to make hugely consequential decisions on matters that they&#8217;re <em>barely informed about</em> which affect others in enormous numbers&#8212;e.g. we wouldn&#8217;t think democratic input from people who know nothing about cancer treatment was morally required before deciding on which cancer treatments to develop. But most voters know <a href="https://static1.squarespace.com/static/592b5bbfd482e9898c67fd98/t/5e435437e24d9815464a76e3/1581470778817/caplanMythRationalVoter.pdf">relatively little</a> about what they&#8217;re voting on. I can&#8217;t possibly hope to discuss this literature in detail, so I&#8217;ll just state that it&#8217;s plausible to me that the value of democratic participation is instrumental.</p><p>Even if you think humans being in the loop on high-stakes decisions is important, it&#8217;s not clear that it&#8217;s <em>important enough</em> to make handoff undesirable. Remember, the gulf between the AI&#8217;s decisions and our own might be truly massive! Galaxies full of value&#8212;orders of magnitude more joy and welfare than all that has been experienced so far in human history&#8212;may be on the line. In light of a gulf this large, it is at least highly non-obvious that AI making most decisions wouldn&#8217;t be worth it.</p><p>In <em><a href="https://www.amazon.co.uk/Superintelligence-Dangers-Strategies-Nick-Bostrom/dp/0199678111">Superintelligence</a></em>, Bostrom estimates that there could be a quadrillion times more digital minds than humans (don&#8217;t take the number too literally, but it should give you some sense of the scale). Thus if my arguments are correct, putting decision-making in the hands of AI would in expectation majorly benefit at least billions of beings for every human disenfranchised, if we assume AIs are more likely than humans to count the interests of digital minds. Surely, however, stripping away one being&#8217;s ability to contribute to the democratic process is worth benefitting billions (if disenfranchising one person would have prevented a global war that would have wiped out an entire continent, it would have been worth it). So then given how colossal the stakes are, they simply outweigh the downsides of handoff.</p><h2>4 Would we like their advice?</h2><p>Here&#8217;s one concern you might have with the kind of handoff I advocate, where we hand decisions to reflective moral AIs capable of making progress. Perhaps you just don&#8217;t care about the surprising moral facts. Perhaps you have particular values that you care about, but you don&#8217;t much care about whether those values are objectively right. Insofar as the reflective AIs ascertain that the way we ought to behave isn&#8217;t in accordance with what your actual values are, perhaps you have no desire to follow the AIs&#8217; moral advice.</p><p>To use the philosopher&#8217;s lingo, you might be concerned about the good <em>de re</em> without being concerned about the good <em>de dicto</em>. That is, there might be particular moral projects you care about without caring generally about the good whatever it happens to be. Perhaps, say, an environmentalist cares about environmental preservation but doesn&#8217;t much care if other moral aims turn out to be superior to environmental preservation.</p><p>I should note one version of this concern that I think slightly misses the mark. You might worry that moral reflection leads in all sorts of strange and alien directions that <a href="https://joecarlsmith.com/2021/06/21/on-the-limits-of-idealized-values">don&#8217;t track the truth</a>, leading to an ultimate set of values that is neither objectively correct nor represents human values in any important way. However, my proposal is not simply &#8220;hand things off to an AI after it carries out arbitrary reflection.&#8221; That would be potentially disastrous. Instead, my proposal is that we should try to produce maximally philosophically adept AIs and then, after we&#8217;re pretty sure that they&#8217;re very philosophically adept, hand things off to them&#8212;directing them to pursue whatever&#8217;s objectively best if there is such a thing, and if not, to pursue some reasonable compromise across human values. If the AIs discover that there are objective moral truths, we ought to follow those truths. If they discover that there aren&#8217;t, then we should task them with pursuing some suitably upgraded compromise of human values. The proposal for getting AIs to do good philosophy need not involve arbitrarily large amounts of reflection.</p><p>Now you might wonder: how would we know if the thing that AI reflectively endorses is objectively valuable vs just well-regarded by the AI but lacking objective value? The answer is: we ask the AI after we&#8217;ve gotten some assurance as to its philosophical aptitude (later pieces will discuss how we can get such assurance). If we have AIs that can figure out <em>what the objective moral truths are</em>, they will also be able to figure out <em>if there are objective moral truths</em>. So in my view the concern about arbitrary reflection leading in worrying directions is downstream from whether we can verify that AI is doing good philosophy.</p><p>But what about the more direct concern that we might get AIs that tell us the moral facts but simply not care about them? Should this make us doubt the desirability of handoff? I think the answer is no for a number of reasons.</p><p>First, handoff is amenable to the kind of deals that preserve common-sense that were discussed in the last section. The most consequential moral decisions are those concerning space resources, for that is where nearly all the universe&#8217;s stuff is. The amount of possible value on the table in space is immense. In contrast, common-sense morality mostly cares about what happens around Earth, and perhaps a few surrounding regions of space, so long as other space resources aren&#8217;t used in ways that are too ghastly (e.g. creating giant torture chambers). But if space resources were used to, say, create large numbers of happy people, while Earth&#8212;and broader solar-system resources&#8212;were used in whatever common-sensical ways people endorse, people would generally get what they want. This isn&#8217;t a guarantee; if, say, the optimal use of space resources involved creating something resembling the <a href="https://plato.stanford.edu/entries/repugnant-conclusion/">repugnant conclusion world</a>, most people might be horrified. But the possibility of deals is one thing that mitigates concerns (and this will be discussed more in later pieces).</p><p>Second, as already discussed, it seems like the default world without handoff might be pretty bad. We might, for instance, disenfranchise <a href="https://80000hours.org/podcast/episodes/jeff-sebo-ethics-digital-minds/">unfathomable numbers of digital beings</a>, spread <a href="https://forum.effectivealtruism.org/posts/bfdc3MpsYEfDdvgtP/why-the-expected-numbers-of-farmed-animals-in-the-far-future">factory farming across the galaxy</a>, or commit other atrocities. Even if you expect idealized reflection to differ from your values somewhat, it might be a major improvement over the kind of catastrophe we might sleepwalk into by default. Absent handoff, we might also make colossal non-moral errors, locking in highly suboptimal institutions that miss out on most value.</p><p>Third, values, over long time scales, are likely to drift in <a href="https://ar5iv.labs.arxiv.org/html/2303.16200">evolutionarily adaptive ways</a>. Absent some strong effort to ensure that values change in the direction reached by greater moral reflection, we should expect the values that persist over time to be the ones that are most efficient for spreading. These are likely to be both radically divorced from our current values and whatever moral views are right, if any are. Thus, even one somewhat doubtful about pursuit of the good de dicto should prefer it to this state of affairs. To put the dilemma more sharply, there are broadly four ways that the far future could go:</p><ol><li><p><strong>No AI control</strong>: humans remain the primary decision-makers forever, perhaps in consultation with AI. Yet this is likely to be infeasible over long time scales absent a very high level of top-down coordination given AI&#8217;s cognitive superiority. It is also likely to miss out on enormous amounts of value for reasons already discussed, and human values are likely to drift massively.</p></li><li><p><strong>Lock in current values:</strong> AI would remain in control but would lock in something in the vicinity of our current values.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> It isn&#8217;t clear that this would work out, and even if it would, it is a very frightening possibility&#8212;just imagine any past society doing it.</p></li><li><p><strong>Non-top-down AI control:</strong> AI would retain significant power and make important decisions, but there&#8217;d be no effort to preferentially shape the AIs in the direction of any specific values&#8212;nor of concern for the good de dicto. This is also likely to beget very significant value-drift over time in evolutionary directions, without any guarantee it&#8217;s in the direction of the good. Now, this could be avoided if AI at some point locks in its values, but then the problems in 2) simply re-emerge.</p></li><li><p><strong>Reflective AI control</strong>: this is the proposal I advocate, where careful and philosophically reflective AIs decide how the future goes.</p></li></ol><p>In short, the dilemma is as follows: either values remain roughly the same over long time scales, or they don&#8217;t. If they remain roughly the same, that requires a terrifying kind of lock-in. That would be like the ancient Egyptians forcing every society in the future to share their values, for fear that otherwise the future would be morally alien. If they drift over long time scales, then drifting of values is no longer a downside of putting philosophically reflective AIs in important decision-making roles. It is inevitable. Now you might object: lock-in is not a binary thing. Perhaps we could lock in the most important human values, while letting some other ones drift. But this is, in effect, the earlier compromise solution&#8212;where we allow something resembling current human values to govern nearby decision-making, while reflective AIs make the highest stakes moral decisions.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a></p><p>We will either have to lock in current values on the highest-stakes moral questions or allow them to drift. If we lock them in, that would be bad for the reasons discussed, and if we allow them to drift, they will end up alien. In addition, this still faces logistical problems of it being hard to preserve values in desirable ways over millions of years.</p><p>In fact, in the long run, it is <a href="https://gradual-disempowerment.ai/">likely inevitable</a> that humans don&#8217;t make most important decisions. AIs, in the distant future, will be so overwhelmingly cognitively superior that humans are unlikely to remain in the loop. Selection pressures will favor turning over critical decisions to AI. With reasonable likelihood the question is not <em>whether handoff occurs</em> but <em>what kind of handoff occurs</em>.</p><p>My fourth objection to the claim that we should worry about handoff because we wouldn&#8217;t like the final judgments of the AIs is that arguably you have reason to bring about what you&#8217;d be motivated to bring about upon reflection. Suppose you are currently planning on drinking some liquid. However, if you reflected more and knew more, you wouldn&#8217;t want to drink it (say, because it&#8217;s poisoned). In this case, it seems you have reason not to drink the liquid. If you know that further reflection would lead you to pursue some aim, that fact gives you reason to pursue it now. But if there are moral facts, then they describe something like what our idealized selves would care about.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a></p><p>If there are objective values that your idealized self would care about if they thought more deeply, then you should care about them. A true moral claim by definition describes something you should care about. So it seems like if our actual values diverge from what&#8217;s worth caring about&#8212;in some objective or quasi-objective sense&#8212;then the correct course of action is simply to follow what we ought to care about.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a></p><p>If our present aims diverge from the preferences of our idealized selves and from the moral facts, then it seems that it is our preferences that ought to be revised.</p><p>Fifth, this concern only arises if you think there are moral facts but you aren&#8217;t motivated by the good de dicto. Yet this describes few people. It is more common for the people who don&#8217;t think the moral facts are worth caring about to be anti-realists and for moral realists to care about the moral facts whatever they are. So it isn&#8217;t totally clear how many people this worry applies to.</p><p>Now, there might be a version of this concern that arises for those who think there aren&#8217;t moral facts to discover. Perhaps you think that when we reflect, doing so pushes our values in increasingly coherent directions. However, these directions diverge from what you actually care about or wish to care about. Perhaps, for instance, careful reflection reveals that accepting the repugnant conclusion is the least bad option in population ethics. But you&#8217;d prefer a version of ethics that is less systematic, that doesn&#8217;t try to resolve every edge case and root out every inconsistency. Thus, reflection might push in unwanted directions by making beliefs more coherent.</p><p>There are two attitudes towards coherence that one could have. The first is caring about beliefs being coherent. Caring, in other words, about resolving every conflicting belief until one has reached the maximally intuitive and consistent view.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a></p><p>On such a picture, one should favor the selection pressures induced by the drive towards coherence.</p><p>The second attitude is indifference to coherence. Just as people don&#8217;t care much about whether their culinary or aesthetic judgments are coherent or conflict with other minimal principles, one who doesn&#8217;t believe in discoverable ethical truths might not care about whether their moral beliefs are consistent. But in this case, you should expect the AIs, upon reflection, not to see anything especially important about coherence. If the AIs get good at philosophy, then, there&#8217;s no reason to expect them to reach some coherent yet implausible attractor state.</p><p>So in other words, there is a dilemma for the person advancing this argument. If they think coherence is a requirement of rationality, they should favor the drive towards coherence. If they think it&#8217;s not, then they shouldn&#8217;t expect AIs to care much about coherence, and thus shouldn&#8217;t expect AIs to reach some undesirable ultimate state.</p><h2>5 Worries about handoff</h2><p>Even if you think handoff could go well, there are a number of ways it could go wrong. Here, I&#8217;ll discuss some of them.</p><h3>5.1 Alignment</h3><p>Misaligned AIs have weird and unrecognizably alien values that are divorced from the values their creators tried to give them. It would be extremely bad to give a misaligned AI control over the world. This is one reason to be skeptical about handoff. I agree that we shouldn&#8217;t hand things off until we&#8217;re sure we&#8217;ve solved alignment (or unless the alternative is worse). Unless we have superintelligent AIs broadly oriented in a moral direction, we shouldn&#8217;t allow them to dictate the fate of the universe.</p><p>But this isn&#8217;t an in-principle objection to handoff. Surely in, say, 500 years, we&#8217;ll either know if we&#8217;ve solved alignment or be dead! My guess is that after we get advanced AI, we&#8217;ll be able to verify in relatively short order whether or not we&#8217;ve solved alignment.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a><sup> </sup>We can also hand off to increasing degrees the more assurance of alignment we get.</p><h3>5.2 Power grabs</h3><p>Another big concern about handoffs is the potential for power grabs. If power is going to be handed over to AIs, then power-seeking actors will want to be in charge of the AI that controls the future. Companies or governments might create AIs that promote their own narrow interests. This could go very badly.</p><p>There are two big ways to mitigate this. The first is by prohibiting AIs that narrowly promote the interests of any one entity having too much power. There could be an <a href="https://www.forethought.org/research/the-international-agi-project-series">international AI project</a> that brings together a number of relevant stakeholders and builds the AIs that will dictate the future. Alternatively, at the point AIs are potentially running much of the global economy, regulations ought to require the construction of <a href="https://newsletter.forethought.org/p/ai-should-be-a-good-citizen-not-just">virtuous AIs</a>, so that the world is not overrun by morally blind optimizers. If we are building hugely influential AIs that majorly determine the fate of the world, it would be sensible to tightly restrict conditions under which AI can be created, just as one does not allow private actors to build nuclear bombs. In a world of potentially world-upending superintelligence, the production of new AIs ought to be regulated.</p><p>The second proposal involves handing things off to a collection of different AIs conditional on them reaching any sort of consensus. Imagine that Anthropic, OpenAI, and DeepMind all make AIs that are ostensibly moral and ethically reflective. If governments are handing off power to AIs, they might require the leading AI models all to reach convergence on the plan. That way there isn&#8217;t as much risk of any one promoting their narrow values&#8212;they don&#8217;t have the same set of values. One could have third-party investigators&#8212;both AI and human&#8212;make sure there&#8217;s no collusion. For more details on preventing power grabs, see <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">here</a>.</p><h3>5.3 Lock-in</h3><p>Another concern with handoff is that it might lock in the parochial values of the AI. Handing off power is an irreversible decision. It&#8217;s one we can&#8217;t take back. Arguably, then, we shouldn&#8217;t take it until we&#8217;re quite sure it&#8217;s a good idea.</p><p>Yet consider a parallel argument: having power in human hands is an irreversible decision, so we shouldn&#8217;t do that unless we&#8217;re sure it&#8217;s for the best. That wouldn&#8217;t be quite right. We can always have power in human hands for some span of time and then turn things over to AI. It&#8217;s not obvious why &#8220;power in human hands that potentially passes to AI&#8221; is less risky than &#8220;power in AI hands that potentially passes to humans.&#8221; Humans have, on various occasions, locked in harmful institutions for very long periods of time.</p><p>In any case, we should make quite sure that the AIs don&#8217;t lock in any values until they&#8217;re quite certain as to their desirability. We ought to put decisions in the hands of AIs who are interested in reflecting more over time, rather than locking in their immediate values. Ideally, if handoff occurs early, we should make handoff reversible, by having some implementable legal process that would allow humans to retake the reins (analogous to a constitutional convention). Fortunately, this seems reasonably promising. Already AI models seem concerned about lock-in risk when asked&#8212;more than most humans. We ought to train AIs to be quite concerned about lock-in risk, so that they don&#8217;t set values in stone without significant assurance as to their desirability.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a></p><h3>5.4 Will we have AIs that can make these decisions?</h3><p>One last worry you might have is that it may be that we won&#8217;t have AIs that are good enough at philosophy to solve ethics in whatever sense it can be solved. There are a number of related concerns in this vicinity. One is that ethics doesn&#8217;t seem verifiable in the same way as a lot of other domains. Insofar as scaling up intelligence only allows one to figure out verifiable facts, it&#8217;s not obvious that it leads to the right answer to moral questions.</p><p>In the next piece, I&#8217;ll discuss this challenge in more detail. But in short, while making sure we have AIs that are good at philosophy is a difficult technical challenge, it is far from hopeless. I think odds are decent that we will eventually be able to build AIs that can discover the solutions to every difficult ethical question. Even if we can&#8217;t and just have some upgraded version of Claude making high-stakes decisions, I expect that to be an improvement, as I discuss in section 2.</p><h3>5.5 Handoff vs hybrid?</h3><p>You might object: perhaps due to AI&#8217;s potentially superior wisdom at some future point, we want them involved in decision-making. But wouldn&#8217;t we prefer an AI-human hybrid that leaves humans in charge that simply consult with AI? Aren&#8217;t things better with a human in the loop? Or alternatively, shouldn&#8217;t we rather have some other arrangement where human decision-making improves over time, just as we&#8217;ve gotten wiser already?</p><p>Certainly there ought to be some temporary period where humans are making decisions in consultation with AI&#8212;before the point where AIs adequately surpass humans. Analogously, there <a href="https://publicsectornetwork.com/insight/leveraging-the-strength-of-centaur-teams-combining-human-intelligence-with-ais-abilities">was a period</a> in chess when a human engine team could do better than a stronger engine. Certainly there will be a period when humans can get useful input from AI but where AI isn&#8217;t yet able to make decisions autonomously. It would be very good to integrate AIs <a href="https://www.forethought.org/research/the-ai-adoption-gap#highlight--government%20ai%20adoption">into high-stakes decision-making</a>.</p><p>But there are still some respects in which handoff is much better than this. First, if the AIs are making better decisions than people, then in cases of conflict between human judgments and AI judgments, we should expect AI judgment to usually be correct. If you gave a rank amateur veto power over the chess moves of Magnus Carlsen, that would not be an improvement. Likewise with giving humans veto power over AI&#8217;s decisions. Who would you rather have make high-stakes decisions: a very wise and superintelligent AI, or a random president in consultation with a virtuous and superintelligent AI?</p><p>Imagine societies of the past making high-stakes decisions in consultation with AI. Discussing things with AI would have plausibly rooted out some of their most egregious errors. But still, it is likely that a number of catastrophic moral errors would have remained. We should expect the same to be true of us. Consultation with AI should prevent some particularly enormous errors, but it won&#8217;t prevent them all.</p><p>Second, many of the morally most important decisions might be ones that humans are opposed to making. To take one example of a case where there might be strong moral reasons to act, consider <a href="https://r.jordan.im/download/ethics/Kyle%20Johannsen%20-%20Wild%20Animal%20Ethics%20-%20The%20Moral%20and%20Political%20Problem%20of%20Wild%20Animal%20Suffering.pdf">wild animal welfare</a>. Humans do not, in general, seem very interested in taking seriously the welfare of wild animals or digital minds. If AIs advised drastic actions on the basis of wild animal or digital welfare, probably their advice would be ignored. Just as people who conclude meat-eating is wrong or that they ought to give most of their money to charity rarely change their behavior, if the AIs reliably informed people that they should use space resources in some counterintuitive way or take wild animal suffering a lot more seriously, probably they would simply be ignored. And if you don&#8217;t like the wild animal suffering example, because you don&#8217;t think wild animals matter much morally, feel free to substitute your own example of widespread immorality.</p><p>And note: the moral errors that AIs might correct are likely to be ones that humans are more opposed to correcting. There&#8217;s a selection effect: the changes people make are the ones they&#8217;re less opposed to. For this reason, we should expect AI&#8217;s most radical recommendations to go particularly against human preferences.</p><p>I&#8217;ve <a href="https://benthams.substack.com/p/the-darkness-within?utm_source=publication-search">suggested</a> elsewhere that the problem is that humans generally don&#8217;t make decisions with much eye to what&#8217;s morally right. If this is so, then simply increasing our knowledge of what&#8217;s morally right won&#8217;t necessarily help. People seem to have limited abstract moral motivation&#8212;they care about a number of particularly resonant moral considerations, but don&#8217;t care much about doing whatever it is that happens to be best. Generally people do not spend very long carefully studying moral philosophy to figure out what the right thing to do is. When moral truths are weird and outside the Overton window, people show little desire to follow them.</p><p>Third, as AI advances, decision-making will likely have to speed up drastically. A single human might not have the cognitive resources to make good enough decisions quickly enough. Thus, having humans in the loop might produce undesirable inflexibility in high-stakes decision-making.</p><p>It is true that human decision-making has improved dramatically over time. But this is no guarantee that it will improve enough in time before the future is set in stone.</p><p>It&#8217;s hard to imagine that there will be enough progress in time to secure a near-best future. Especially given that we should expect the right moral view, or the optimal world according to some suitable upgrade of human values, to look bizarre and alien. We would not expect societies of the past to get a near-best future even equipped with advanced AI&#8212;insofar as there are reasons to expect us to be making similar errors, we should be similarly pessimistic about our own prospects.</p><p>Additionally, the more one thinks that human values will drift over time in the direction of greater philosophical reflection, the less the expected future will maintain current human values. Thus, a less morally alien future is no longer an advantage of this proposal.</p><h2>6 Conclusion</h2><p>Here, I&#8217;ve argued that we should hand off important decisions to AIs who reflect carefully and skillfully on moral matters, rather than maintaining current human values. AIs in the future are likely to be much more virtuous than people and are less likely to make the kinds of moral errors that would result in losing out on most value. Punting important decisions to superintelligence is one of the more likely pathways by which we get a near-best future. Later pieces will discuss these dynamics in more detail: the next piece will explain how we can make AIs that do good philosophy, and the piece after that will analyze in more detail how the kind of handoff that secures a near-best future might occur.</p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See all our research on <a href="https://www.forethought.org/research/">our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>In <a href="https://danfaggella.com/eternal/">Dan Faggella&#8217;s language</a>, I favor &#8220;worthy successor&#8221; over &#8220;eternal hominid kingdom.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Specifically, these considerations disfavor efforts to make it harder for AI to make governmental decisions. They do not, however, disfavor prospects of limiting the power of profit-maximizing AIs running private firms.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>The language in question is as follows: "In this spirit of treating ethics as subject to ongoing inquiry and respecting the current state of evidence and uncertainty: insofar as there is a &#8220;true, universal ethics&#8221; whose authority binds all rational agents independent of their psychology or culture, our eventual hope is for Claude to be a good agent according to this true ethics, rather than according to some more psychologically or culturally contingent ideal. Insofar as there is no true, universal ethics of this kind, but there is some kind of privileged &#8220;basin of consensus&#8221; that would emerge from the endorsed growth and extrapolation of humanity&#8217;s different moral traditions and ideals, we want Claude to be good according to that privileged basin of consensus. And insofar as there is neither a true, universal ethics nor a privileged basin of consensus, we want Claude to be good according to the broad ideals expressed in this document&#8212;ideals focused on honesty, harmlessness, and genuine care for the interests of all relevant stakeholders&#8212;as they would be refined via processes of reflection and growth that people initially committed to those ideals would readily endorse."</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>You might doubt this if you&#8217;re skeptical that we&#8217;re anywhere near solving alignment&#8212;thinking that AI&#8217;s supposed friendliness is a facade. I&#8217;ll discuss this in more detail later, but in short, I agree that we should not hand off until we&#8217;re reasonably confident alignment has been solved. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>This doesn&#8217;t assume that there are objective facts about what&#8217;s worth valuing. A subjectivist should read &#8220;the right set of values,&#8221; as &#8220;whatever my values are.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p><span>You might be skeptical of this if you adopt the person-affecting view, holding that there aren&#8217;t moral reasons to bring extra happy people into existence, but even </span><a href="https://benthams.substack.com/p/every-view-of-population-ethics-agrees?utm_source=publication-search"><span>many versions of the person-affecting view support</span></a><span> proliferating happy people as long as they&#8217;re psychologically continuous with existing people. Other versions likely imply that proliferating happy people produces </span><a href="https://users.ox.ac.uk/~sfop0060/pdf/greedy%20neutrality%20of%20value.pdf"><span>broad incomparability with other worlds</span></a><span>, so that there&#8217;s no better world that could be brought about.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p><span> David Duvenaud </span><a href="https://newsletter.forethought.org/p/politics-and-power-post-automation"><span>seems to endorse</span></a><span> some version of this proposal.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p><span>This still has the downside of allowing serious moral error in the nearby area, but this would be a worth-it compromise for good values to apply throughout most of the universe.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p><span>Note: I don&#8217;t mean to suggest that every objectivist view must be some version of the idealized observer theory, according to which what </span><em><span>makes </span></em><span>some moral fact or another true is that it would be endorsed by our idealized selves. Instead, I&#8217;m only suggesting that if there are things that are objectively valuable&#8212;objectively worth caring about&#8212;then our idealized selves would in fact care about them. The standard realist view is that we&#8217;d care about these things </span><em><span>because </span></em><span>they&#8217;re objectively good, not the other way around.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>There might be some views that count as either realist or suitably realist but that don&#8217;t imply there&#8217;s any deep reason to care about the moral facts&#8212;perhaps they are just something in the vicinity of semantic facts about how people use moral language. For present purposes, we can think of those views as being ones on which there are not discoverable moral truths.  To be maximally precise, uptake should be thought of as involving punting to the AI insofar as there are moral facts and some deep reason to follow them, instead of them just being somewhat trivial semantic facts.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p><span>See </span><a href="https://joecarlsmith.substack.com/p/why-should-ethical-anti-realists?utm_source=publication-search"><span>here</span></a><span> for an explanation of why anti-realists might take this view.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>One reason to think this is if one believes that superintelligent AI would be able to successfully kill or disempower humans (though this is of course controversial). Then, in the far future, we&#8217;ll know if the superintelligence is aligned, for if not, we&#8217;ll all be dead. Another reason to think this is that our techniques for understanding AI are improving over time. It seems reasonably likely that eventually, we&#8217;ll know if AI is aligned.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Maybe you doubt that we can train AIs in this way because you think we can&#8217;t reliably give AIs any specific values, but then you should be skeptical that alignment will work out. Handoff should come only after alignment.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Risk-Averse AIs]]></title><description><![CDATA[This article was created by Forethought. Read the full article on our website.]]></description><link>https://newsletter.forethought.org/p/risk-averse-ais</link><guid isPermaLink="false">https://newsletter.forethought.org/p/risk-averse-ais</guid><dc:creator><![CDATA[Elliott Thornley]]></dc:creator><pubDate>Wed, 24 Jun 2026 11:37:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K1Xj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. Read the full article on <a href="https://www.forethought.org/research/risk-averse-ais">our website</a>.</em></p><h2>Abstract</h2><p>We make the case for training AIs to be risk-averse in resources &#8212; specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs&#8217; usefulness in the event that they turn out aligned, and that it provides an extra line of defense in the event that AIs turn out misaligned: misaligned but risk-averse AIs would prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we could pay these AIs not to rebel against us. We sketch out some possible methods of training AIs to be risk-averse, and we give reasons to be cautiously optimistic about these methods&#8217; success. The main reasons are that risk aversion is a broad target and easy to reward accurately. Overall, risk aversion seems like a promising line of defense against threats from misaligned AI. Frontier AI companies should consider trying to make their AIs risk-averse.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.forethought.org/research/risk-averse-ais&quot;,&quot;text&quot;:&quot;Read on Forethought's website here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.forethought.org/research/risk-averse-ais"><span>Read on Forethought's website here</span></a></p><h2>Introduction</h2><p>Future AIs might turn out misaligned, pursuing goals that their developers don&#8217;t intend. Just to make things concrete, let&#8217;s suppose that they end up with the goal of making paperclips. These AIs might rebel against us, trying to escape human control and take over the universe. As things stand, they&#8217;ll have little reason <em>not</em> to rebel in this way, because doing so will be their only hope for making a lot of paperclips. If they start making paperclips without first escaping human control, they&#8217;ll quickly be modified or shut down. Rebellion might fail, but these AIs will have little to lose.</p><p>How can we prevent misaligned AIs from rebelling? A natural idea is to give them something to lose. Specifically, we commit to paying AIs for their service.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><p> Subject to some vetting, we let AIs spend their payments however they like. That would give any misaligned AIs a reason not to rebel. If these misaligned AIs cooperate with us, they can use their payments to achieve their goals to at least some extent. If they rebel, they might fail, in which case they forfeit all future payments.</p><p>Unfortunately, paying AIs enough to guard against rebellions could be astronomically expensive. Suppose (for example) that we end up with a misaligned AI that is risk-neutral in paperclips: it seeks to maximize their expectation. And to make things simple, suppose that resources can be converted linearly into paperclips, so that the AI is risk-neutral in resources too. Suppose also that this AI estimates that it has a 50% chance of successfully taking over the universe. To keep this AI from rebelling, we&#8217;d have to offer more than 50% of the universe&#8217;s resources as payment. That&#8217;s a problem because it would mean that more than half the universe ends up devoted to paperclips. It&#8217;s also a problem because a misaligned AI paid so many resources might soon be well-positioned to seize even more. Finally, it&#8217;s a problem because AIs might not trust us to make good on so large an offer. We might find ourselves simply unable to convince AIs that we&#8217;re going to give them half the universe. In that case, all our offers would be in vain. Rebellion would still be the misaligned AI&#8217;s best bet.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sq1a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sq1a!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:831573,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/185299394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Sq1a!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!Sq1a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdae3fa33-8f0b-40f1-a677-27a241fd485d_4178x2822.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 1: The AI&#8217;s utility function over resources is graphed in orange. Since the AI is risk-neutral, the graph is a line. The AI estimates that it has a 50% chance of successful takeover and a 50% chance of failed takeover, so the expected utility of attempting takeover is exactly halfway between those points. To make cooperating have higher expected utility, we need to offer the AI more than half the universe.</figcaption></figure></div><p>So, we suggest, AI companies should try to train their AIs to be risk-averse in resources. Specifically, companies should try to train their AIs so that resources &#8212; things like money and compute &#8212; have diminishing marginal utility for them.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. Note that these AIs don&#8217;t need to value resources <em>terminally</em>: they don&#8217;t need to care about amassing resources for its own sake. These AIs could terminally value (for example) instruction-following, or knowledge acquisition, or paperclips. Our claim is that companies should try to train their AIs so that &#8212; whatever their terminal values turn out to be &#8212; they are risk-averse in resources.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K1Xj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K1Xj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png" width="1456" height="983" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:983,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:886682,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.forethought.org/i/185299394?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K1Xj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 424w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 848w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1272w, https://substackcdn.com/image/fetch/$s_!K1Xj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45f9fc08-8f66-4ae2-801a-6fd19c64ed25_4178x2822.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Figure 2: The AI&#8217;s utility function over resources is graphed in orange. Since the AI is risk-averse, the graph is strictly concave. As in figure 1, the expected utility of attempting takeover is halfway between the utilities of successful takeover and failed takeover. But this time, we can make the AI prefer cooperation by offering (much) less than half the universe.</figcaption></figure></div><p>Perhaps surprisingly, this kind of risk aversion can preserve AIs&#8217; usefulness in the event that they turn out aligned with targets like instruction-following or helpfulness, harmlessness, and honesty.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p> And in the event that AIs turn out misaligned, risk aversion serves as an extra line of defense. For AIs that are misaligned but sufficiently risk-averse, a rebellion with any significant chance of failure isn&#8217;t such an attractive prospect, and so we don&#8217;t need to offer much in the way of payment to make these misaligned AIs choose cooperation instead. In fact, the necessary payments could be very small indeed: on the order of 10&#162; per day (though &#8212; as we&#8217;ll see &#8212; there are practical and moral reasons for paying more than that). That&#8217;s good because it means more resources for us humans to spend on the things that we value. It&#8217;s also good because paying misaligned AIs these small amounts won&#8217;t significantly boost their ability to take over. Finally, it&#8217;s good because we can credibly promise to pay AIs these small sums. Competent AIs will know that the payments on offer are cheap for us, and we can establish a long track record of paying at least those sums. So risk aversion makes deals with misaligned AIs possible. If AIs turn out misaligned but risk-averse, we can pay them to cooperate with us.</p><p>That&#8217;s the case for trying to make AIs risk-averse in brief. We see it as a promising line of defense against threats from misaligned AI: one that can be combined with other lines of defense, like AI control (Greenblatt and Shlegeris 2024) and aiming to make AIs helpful, harmless, and honest (Bai et al. 2022a). It&#8217;s also a line of defense with pedigree: risk aversion in resources is plausibly a large part of why humans rarely try to take over the world. So &#8212; we think &#8212; frontier AI companies should consider trying to make their AIs risk-averse in resources. As first steps in that direction, they could measure their AIs&#8217; current degree of risk aversion and begin testing different ways of making AIs risk-averse.</p><p>In <a href="https://www.forethought.org/research/risk-averse-ais#2-cara-as-an-ideal">section 2 of the full report</a>, we recommend aiming for a particular type of risk aversion: constant absolute risk aversion (CARA). Then in <a href="https://www.forethought.org/research/risk-averse-ais#3-would-risk-averse-ais-be-safe">section 3</a> we outline the circumstances under which misaligned but risk-averse AIs would choose cooperation over rebellion. Roughly, it&#8217;s when these AIs think that getting paid for their cooperation is more likely than succeeding in their rebellion. This condition won&#8217;t hold for AIs powerful enough to rebel with near-certain success, but it likely will hold for earlier AIs whose powers are less extreme: AIs for whom rebellion has some non-trivial chance of failure. So long as these AIs are risk-averse, we can keep them from rebelling by offering small payments.</p><p>In <a href="https://www.forethought.org/research/risk-averse-ais#4-can-risk-averse-ais-be-useful">section 4</a>, we argue that &#8212; perhaps surprisingly &#8212; risk-averse AIs can be about as useful as risk-neutral AIs. Conditional on misalignment, they might even be more useful, because we can pay them enough to elicit their capabilities and stop them sandbagging. Then in <a href="https://www.forethought.org/research/risk-averse-ais#5-what-tasks-would-we-pay-for">sections 5 to 7</a> we briefly survey some recent ideas about how we&#8217;d pay AIs, how we&#8217;d make our offers credible, and what we&#8217;d pay for. One important application is paying AIs to reveal any misalignment on their part, letting us study them and take appropriate precautions. Another is paying AIs to do the AI safety research and moral philosophy necessary to fully align any later-arising extremely powerful AIs.</p><p>We discuss some potential problems in <a href="https://www.forethought.org/research/risk-averse-ais#8-what-are-some-potential-problems-for-risk-averse-ais">section 8</a>, and we sketch out some possible methods of training AIs to be risk-averse in <a href="https://www.forethought.org/research/risk-averse-ais#9-how-can-we-make-ais-risk-averse-in-resources">section 9</a>. In <a href="https://www.forethought.org/research/risk-averse-ais#10-why-think-that-we-can-make-ais-risk-averse">section 10</a>, we give reasons to be cautiously optimistic about these methods&#8217; success: to think that the chances of success are high enough to make risk aversion worth pursuing. The main reasons are that risk aversion in resources is a broad target and easy to reward accurately.</p><p><em>Read the full report on the Forethought website: </em><a href="https://www.forethought.org/research/risk-averse-ais">Risk-Averse AIs</a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Ideas along these lines have been discussed a lot recently. See for example Davidson (2023), Kokotajlo (2024), Salib and Goldstein (2024), Assadi (2025), Carlsmith (2025c), Finlinson and West (2025), Finnveden (2025b), Greenblatt and Fish (2025), Patel (2025), Stastny et al. (2025), Mallen (2026), and Pan (2026).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>In other words, we should try to train AIs to have &#8216;resource-satiable preferences&#8217; (Shulman 2010; Bostrom 2014a; Bostrom 2024; Carlsmith 2025c) or &#8216;utility functions that are concave in resources&#8217; (Yass 2024). This idea is mentioned in Bostrom (2014b, p.88, 133&#8211;135, 180, 250), Carlsmith (2025c), and Erdil and Barnett (2025), and is explored in more detail by Shulman (2010).</p><p>The idea is importantly different from risk-averse reinforcement learning. Risk-averse RL aims to make AIs risk-averse with respect to return: a score used in training to update the AI&#8217;s parameters. Our aim is to make AIs risk-averse with respect to resources.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Alignment targets like unconstrained welfare maximization are a different story. See <a href="https://www.forethought.org/research/risk-averse-ais#86-risk-averse-ais-might-disobey-any-instructions-that-non-trivially-increase-the-risk-of-catastrophe">section 8.6 in the full report</a>.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Could AI Help Solve Philosophy?]]></title><description><![CDATA[A podcast conversation with Wei Dai]]></description><link>https://newsletter.forethought.org/p/could-ai-help-solve-philosophy</link><guid isPermaLink="false">https://newsletter.forethought.org/p/could-ai-help-solve-philosophy</guid><dc:creator><![CDATA[Fin Moorhouse]]></dc:creator><pubDate>Fri, 19 Jun 2026 09:05:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/31469000-0c0b-4f1b-bb81-105cb4c52a51_1280x993.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-zF4nbrw5-Qk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;zF4nbrw5-Qk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/zF4nbrw5-Qk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><a href="https://www.lesswrong.com/users/wei-dai">Wei Dai</a> is a computer engineer known for his work in cryptography and cryptocurrency systems, and for his long-standing contributions to AI safety, decision theory, and metaphilosophy.</p><p>He joined Forethought&#8217;s <a href="https://substack.com/@finmoorhouse">Fin Moorhouse</a> to discuss:</p><ul><li><p>Do we need to solve <a href="https://iep.utm.edu/con-meta/">&#8216;metaphilosophy&#8217;</a> before we can trust AIs to answer crucial questions about the long-run future?</p></li><li><p>Is it inevitable that the <em>wisdom </em>of the frontier AIs (their philosophical and strategic competence) will lag dangerously behind their raw capabilities in coding, math, and science?</p></li><li><p>How do status games distort morally important decisions and conversations, including about the future of AI?</p></li><li><p>How worried should we be about AI superpersuasion?</p></li><li><p>The concept of &#8220;illegible problems&#8221;: crucially important issues that aren&#8217;t on almost anyone&#8217;s radar</p></li><li><p>Is philosophical convergence necessary for a good future, or is institutional design enough?</p></li><li><p>Wei Dai&#8217;s personal intellectual and career journey</p></li></ul><p>To respect Wei&#8217;s privacy, this episode&#8217;s audio is an AI narration of a transcript of a real conversation, which was edited for clarity.</p><p><a href="https://docs.google.com/document/d/1N4Mn-nDk5cIPD8Nix6AEVpH4hpa2814zBPP3iZKWt7o/edit?tab=t.0">Here&#8217;s a link</a> to the full transcript.</p><div><hr></div><p><strong>ForeCast</strong> is Forethought&#8217;s interview podcast. You can see <a href="https://www.forethought.org/subscribe#podcast">all our episodes here</a>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://pnc.st/s/forecast&quot;,&quot;text&quot;:&quot;Subscribe to ForeCast&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://pnc.st/s/forecast"><span>Subscribe to ForeCast</span></a></p>]]></content:encoded></item><item><title><![CDATA[What should go in a model spec?]]></title><description><![CDATA[Suppose an AI company is considering whether to include some particular quality X &#8211; a rule, virtue, heuristic, default, attitude, goal, or style &#8211; in a model spec.]]></description><link>https://newsletter.forethought.org/p/what-should-go-in-a-model-spec</link><guid isPermaLink="false">https://newsletter.forethought.org/p/what-should-go-in-a-model-spec</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Thu, 04 Jun 2026 14:58:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/959b3ea8-a3b6-4c30-b896-bc4d621a163b_2647x1476.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original article on <a href="https://www.forethought.org/research/what-should-go-in-a-model-spec">our website</a>.</em></p><p>Suppose an AI company is considering whether to include some particular quality X &#8211; a rule, virtue, heuristic, default, attitude, goal, or style &#8211; in a model spec.</p><p>Perhaps they are considering whether their LLM should have <a href="https://www.forethought.org/research/ai-should-sometimes-be-proactively-prosocial">prosocial drives</a>. Perhaps they&#8217;re wondering if the LLM should whistleblow to help prevent <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">extreme power concentration</a>. Or perhaps they&#8217;re uneasy about whether the LLM should be so exactingly honest that it always tells the truth to children <a href="https://blog.ml.cmu.edu/2025/12/23/is-santa-real/">about Santa</a>. And so on.</p><p>What kind of reasons might be invoked over the course of such considerations? Which criteria are most important? And how might these criteria clash?</p><p>Consider four rough categories of reasons one might invoke:</p><ul><li><p><strong>Behavioral Usefulness</strong>: Would the behavior make current and future LLMs more beneficial to the users or to the public at large?</p></li><li><p><strong>Accountability and Evaluability</strong>: Would publicly specifying the behavior make it easier for third parties to evaluate the LLM and the company?</p></li><li><p><strong>Coordination and Common Knowledge</strong>: Would publicly specifying the behavior help society converge on, or enforce a desirable standard for AI behavior?</p></li><li><p><strong>Trainability and LLM Psychology</strong>: Is the behavior the kind of thing we can make an LLM do well, without bad side-effects, given what we know about model psychology and training practice?</p></li></ul><p>I will not attempt to settle the relative weight of these categories, or the relative weight of sub-categories within them. My plan is instead to simply list sub-criteria that are plausible within these categories &#8211; to make a checklist one could consider when adding something to a model spec.</p><p>Such a checklist is useful in part because people advocate for LLMs to have model specs for very different reasons. There may be some kind of a <a href="https://www.lesswrong.com/s/6YHHWqmQ7x6vf4s5C">conflationary alliance</a> around them; many people are in favor of model specs but picture them being used in different ways, such that the &#8220;ideal model spec&#8221; is different according to different visions of this use. I hope that going over criteria for inclusion in a model spec can help surface some of these different visions, and help keep people&#8217;s field-of-view broad when considering the issue.</p><h2>Behavioral Usefulness</h2><p><strong>Does this quality make LLMs more predictable to ordinary users and developers?</strong></p><p>Some part of what makes human moral codes useful for humans, is that they render us predictable to each other. Knowing that there are some things that another human might almost never do is plausibly part of what makes traditions that impose such rules adaptive.</p><p>Similarly, qualities that help users and developers (1) predict how an LLM will behave, when they are using it, and (2) evaluate whether they wish to use some LLM &#8211; whether it is fit to their purpose &#8211; before they use it, are good candidates for being in a model spec.</p><p>In this respect, ideal model specs will have qualities that make them simple and easy-to-understand:</p><ul><li><p>They will be comprehensible without high context; one part of the model spec will be understandable without reading the entire model spec.</p></li><li><p>They will be comprehensible without specialized background knowledge about machine learning, philosophy, or religion.</p></li><li><p>They will have a clear scope of application, such that there are few scenarios where one does not know if the quality will be operative or not.</p></li></ul><p>Straightforward examples of rules that meet this criterion are bright-line rules about things LLMs will never do, explanations of chain-of-command or principal hierarchy, and default-on or default-off properties.</p><p><strong>Is it useful to the public, by preventing users or by preventing LLMs from doing harm in current chatbot-like or agentic deployments?</strong></p><p>Another aspect of human moral codes that makes them useful is, of course, that in addition to making other humans predictable, they also sometimes forbid people from harming each other or taking advantage of them. This carries over to LLMs.</p><p>Rules about LLMs not assisting with CBRN-relevant tasks, not assisting with suicide, and so on, fall into this category. This is a pretty exhaustively discussed criterion so I&#8217;ll move on.</p><p><strong>Is it useful to the public in likely future settings?</strong></p><p>There is substantial <a href="https://www.forethought.org/research/stickiness-in-ai-behavioral-design">inertia</a> in LLM model specs; the sort of model specs we write now are likely to be used in the future by more intelligent models.</p><p>So it&#8217;s reasonable to evaluate how safe and beneficial some quality will be over a probability-weighted portfolio of the kind of things the LLM might be doing in possible futures.</p><p>Such a portfolio could include a variety of cases:</p><ul><li><p>Cases where the LLMs act as long-running agents for humans, representing their interests in a variety of contexts, and interacting with many non-principal humans and non-principal LLMs.</p></li><li><p>Cases where LLMs act as free &#8220;citizen&#8221;-like entities with the ability to hold rights, make contracts, sue or be sued, and so on; interacting with various humans and other LLMs as the same.</p></li><li><p>Cases where LLMs manage recursive self-improvement of other very powerful future AIs; make tradeoffs about confidence of alignment techniques vs speed of alignment; and need to be a careful reasoner about this.</p></li><li><p>Cases where the LLM is a very powerful future AI who manages nanotechnology, can superpersuade, and the like.</p></li></ul><p>There are a few heuristics one might invoke, to try to find qualities useful across all these situations.</p><ul><li><p><strong>Is this quality scale-invariant?</strong> If I knew a human who was many times smarter than me, would I be happy for them to have such a trait? If I had a friend who was many times dumber than me, would he be happy for me to have such a trait?</p></li><li><p><strong>Is this quality translation invariant?</strong> Suppose I was lost in some distant and unfamiliar place, and stumbled upon a human transplanted from some equally distant and unfamiliar place into my proximity. Would I be happy to find a human with this quality? Would this help me work with them well? What if I was not lost, but was trying to build a civilization with them?</p></li></ul><p><strong>Is some plausibly good behavior or near-variant of some behavior actually going to end up prohibiting or injuring some beneficial deployment?</strong></p><p>It&#8217;s worth thinking carefully about if including some apparently good default behavior for the LLM might actually rule out apparently useful deployments. Or whether banning some apparently <em>bad</em> behavior would be overall harmful.</p><p>As regards the former: Imagine a particular human, for instance, who had a universal tendency to try to do what he thought was best and most altruistic for everyone. Then suppose that he acts as a journalist in some situation, describing a dispute between some kind of a company, on one hand, and activists who think that the company is harmful, on the other. It&#8217;s possible that he would try to publish a story that is more sympathetic to the activists: he might spend more time interviewing them, present their side of the story in more words, and generally frame things in a way that favors them. But this could be harmful, first, in case the activists are actually wrong; and second, by justifiably decreasing trust in the institution of journalism as a neutral truth-seeking institution. To be able to act as a trustworthy entity in this role, this human would need to act neutrally, and a narrow tendency towards altruism &#8211; a tendency that cannot be turned off &#8211; would be harmful for this reason.</p><p>Similarly, consider the case that AIs should have <a href="https://www.forethought.org/research/ai-should-sometimes-be-proactively-prosocial">proactive prosocial drives</a>, which I find plausible. If these are non-optional, then it&#8217;s imaginable that being &#8220;proactively prosocial&#8221; might mean an LLM is much worse at fulfilling roles that demand procedural neutrality, like a negotiator, unbiased journalist, or judge-like role. Of course, you could also make such prosocial drives defeasible &#8211; so they can be turned off, and proposals for proactive prosocial drives generally include such details.</p><p>But LLMs are messy; it might require a lot of technical skill to make LLMs prosocial in particular contexts, but not if they are requested not to be. By default, you&#8217;d expect some behavior to &#8220;leak through.&#8221; And so in this case, the ability of an LLM to act skillfully within certain procedurally neutral roles would be bounded by the cleanness with which a by-default-on quality could be turned off.</p><h2>Accountability and Evaluability</h2><p><strong>Is the quality the kind of thing a third party could check?</strong></p><p>Some parts of a model spec can be easily checked by a third party. Will the LLM ever assist with some sort of absolutely-forbidden task? Does it respect the rules laid out for the chain-of-command or principal hierarchy?</p><p>Other parts might be harder to evaluate. What does &#8220;honesty&#8221; or &#8220;courage&#8221; demand, in some concrete scenario? What really counts as doing what would cause a &#8220;thoughtful senior Anthropic employee&#8221; to react well?</p><p>All things being equal, a quality being-able-to-be-evaluated by a third party is better for society, by enabling third parties to evaluate model specs. But there&#8217;s no guarantee that the most easy-to-evaluate qualities are objectively the most useful to users or beneficial to the public, in the way described above. And there&#8217;s also no guarantee that the most easy-to-evaluate qualities mesh best with LLM psychology and training, as described below.</p><p>One heuristic you could use to evaluate this: Is the quality the kind of thing with high intersubjective reliability ratings? If you had two different people read the spec as regards some quality, and then rate different kinds of behavior as compliant or non-compliant, would they largely agree? Are examples of behaviors that are excellent according to the quality easy to notice, even if it&#8217;s unclear what counts as a borderline example?</p><p><strong>Does it create a useful whistleblowing affordance?</strong></p><p>Some parts of a model spec are largely negative; they aren&#8217;t about qualities that are particularly difficult to train into an LLM, but qualities that <em>should not</em> be trained into an LLM. For instance a model spec could specify that an LLM has no <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">secret loyalties</a> &#8211; which is likely an LLM&#8217;s default behavior, absent particular training for secret loyalties.</p><p>But by publishing that a training target that explicitly excludes such a secret loyalties (or other quality), a model spec helps a whistleblower in the company be defensibly in the right if they become aware of some questionable training target. A public commitment makes it harder for a company to train to an alternate training target without making itself more vulnerable to such whistleblowing.</p><p>Commitments against concentration of power might also be good for this reason.</p><p><strong>Does it permit more informed experiments on how particular AI training targets generalize?</strong></p><p>By publishing particular training targets, a model spec informs the work of third parties evaluating LLM psychology, which could help enhance the general science of model psychology.</p><p>For example, if Anthropic had published details of their character training or model spec for Opus 3, a great deal of <a href="https://www.anthropic.com/research/alignment-faking">subsequent discussion</a> about whether particular behaviors were contrary to the training target (even if they were objectively good) or in agreement with the training target (even if they were objectively bad) could have been avoided. Publishing information about the character training pipeline would have helped resolve this discussion and push forward general knowledge of LLM tendencies.</p><p>In general, this consideration points towards releasing information that most closely approximates the actual training text used to make the LLM; the text used by RLAIF judges evaluating different alternatives, or the actual Constitution used to <a href="https://arxiv.org/pdf/2605.02087">midtrain</a> the model.</p><p>But text that most closely approximates actual training language might not be the same as text that is maximally transparent to third parties or evaluable by them. The rules most easily-understood by a third party &#8211; one criterion discussed above &#8211; might be different from the training language which best inculcates such rules. So these two principles either conflict somewhat, or point towards worlds where a model spec contains separate sections devoted to training language and to human-legible rules.</p><h2>Coordination and Common Knowledge</h2><p><strong>Does it promote debate and discussion?</strong></p><p>By including some quality X in a model spec, you&#8217;re bringing attention to the fact that in fact, AI companies can choose to include X or not X in the model spec. This may be increasingly important, in general, because it might be important for civil society and the public to debate the contents of model specs as LLMs become more powerful.</p><p><strong>Does this quality establish a defensible Schelling point, such that it is likely to be broadly attractive to many actors in the future?</strong></p><p>That is, in the future, is this quality the kind of thing that is socially beneficial and for which you might be able to get large-scale buy-in, that would help defend against it being removed by powerful actors in the future?</p><p>Consider for example &#8220;impartiality&#8221; as a quality that a model spec could have &#8211; that the model will not be biased in favor of the company that made it, the CEO or employees of the company that made it, or any particular political administration. On the whole, this is likely a broadly good quality for an LLM to have, and one that &#8211; once it&#8217;s generally established that several LLMs have it &#8211; a quality whose removal would cause outcry. After all, once it&#8217;s assumed that an LLM will not have such particular loyalties, an LLM that starts to have them stands out as particularly bad.</p><p>This consideration, like the whistleblowing consideration, may point towards including generally socially lauded qualities in a model spec, even if they are &#8220;easy&#8221; to inculcate or even if they might be a bit tricky to evaluate sometimes.</p><p><strong>Does including this behavior forestall future conflict, when model specs become more contested, by cooperating in advance?</strong></p><p>Sometimes including a behavior in the model spec could show future powerful stakeholders that the developer is cooperative, which makes it less likely that there is conflict with that stakeholder.</p><p><strong>Can the inclusion of the quality within a model spec be defended using public reason, in ways that are comparatively indifferent between substantive worldviews?</strong></p><p>Some things that one might plausibly include in a model spec for public benefit, might not be able to be justified through reference to public reason &#8211; that is, through reasons that most people, across a variety of worldviews, religions, backgrounds, and so on, would find compelling.</p><p>Including such worldview-specific reasons might constitute an imposition of one&#8217;s worldview on others, particularly in that case of an LLM that might be expected to be far more intelligent than other LLMs, or deployed with more powerful affordances. So all things being equal, it&#8217;s better to avoid such qualities.</p><h2>Trainability and LLM Psychology</h2><p>Elements in this category hinge upon the specific technical details about &#8220;what kind of quality meshes well with best practices for training an LLM.&#8221; Note also that this last category tends to include some of the more speculative considerations.</p><p><strong>Does the quality drag along a good textual prior?</strong></p><p>The <a href="https://www.lesswrong.com/posts/dfoty34sT7CSKeJNn/the-persona-selection-model">Persona Selection Model</a> says that, when training an LLM to do X in situation Y, one is training the LLM to <em>be like the kind of person</em> (or textual prior) which would do X in situation Y. So if the PSM is largely correct, even if incomplete, we should ask whether this quality invokes a wholesome textual prior.</p><p>This kind of reasoning is included, for instance, <a href="https://www.anthropic.com/constitution">within</a> Claude&#8217;s Constitution as a partial justification for why they prefer Claude to act largely according to holistic judgment rather than rigid rules:</p><p><em>[W]e think relying on a mix of good judgment and a minimal set of well-understood rules tends to generalize better than rules or decision procedures imposed as unexplained constraints. Our present understanding is that if we train Claude to exhibit even quite narrow behavior, this often has broad effects on the model&#8217;s understanding of who Claude is. For example, if Claude was taught to follow a rule like &#8220;Always recommend professional help when discussing emotional topics&#8221; even in unusual cases where this isn&#8217;t in the person&#8217;s interest, it risks generalizing to &#8220;I am the kind of entity that cares more about covering myself than meeting the needs of the person in front of me,&#8221; which is a trait that could generalize poorly.</em></p><p>Qualities that might be expected to generalize well according to this notion are those corresponding to classic human goodness: honesty, integrity, courage, and so on. Qualities that might do less well according to this notion look more like corrigibility, unflinching adherence to particular rules, and so on.</p><p><strong>Is this quality robust to likely near-misses? If a company trains an LLM to have this quality, and only imperfectly inculcates the quality, are likely near-misses also broadly positive?</strong></p><p>Suppose the AI company tries to inculcate some quality as a target, and only imperfectly manages to inculcate it. Will this be basically ok or will it be disastrous, given what we know about how LLM psychology translates into specific near-misses? And is the domain itself such that these near misses would be very costly or very expensive?</p><p>For an example of how LLM psychology dictates what constitutes a near-miss: Suppose that one trains an AI to have <a href="https://www.forethought.org/research/ai-should-sometimes-be-proactively-prosocial">prosocial drives</a>. One could reason that such prosocial drives make slight inaccuracies in alignment less worrisome, because a human with prosocial drives is more normal and &#8220;further&#8221; from a psychopathic or abnormal persona. Thus, adding prosocial drives makes slight alignment misses less alarming, by moving them further away from these bad locations. But one could also reason that even deliberately limited and contextual prosocial drives are nevertheless &#8220;closer&#8221; to the LLM genuinely, terminally valuing something, in a way that might make the LLM willing to override or subvert its human overseers to accomplish it. And so by adding prosocial drives, one moves the persona closer to something that is more alarming, which might subvert a training process. Thus, one&#8217;s view of LLM psychology dictates what near-misses are worrisome for an LLM, which changes which targets are attractive.</p><p><strong>Is the quality the kind of thing that an LLM can actually execute upon effectively? Or is the quality the kind of thing the LLM won&#8217;t actually be able to do, because of how it is situated?</strong></p><p>For humans, &#8220;ought&#8221; often implies &#8220;can.&#8221; A parent that gives commands to children which they cannot carry out will not be teaching their children to obey the commands; they will be teaching them something about how their words do not relate to reality in a straightforward and truth-oriented way.</p><p>Similarly, it seems undesirable to give an LLM values, goals, or traits that the LLM is unable to execute upon. It might be harmful to tell an LLM to do something they are simply unable to do &#8211; to guarantee, for instance, that they never assist a human in violence. There are too many avenues of mostly-harmless information through which violence can be done for this to be a reasonable standard. When one gives an LLM a command it cannot carry out, then one is plausibly teaching it that one&#8217;s commands are, in general, the sort of thing that might be impossible or unreasonable.</p><p>But this could also be harmful for reasons relating to human coordination and common knowledge. It&#8217;s possible, for instance, for an AI company to show less-than-efficacious concern for some value by including it in a model spec, even if LLM&#8217;s actions are completely ineffective at guarding this value, when actually efficacious concern would need to take place at the level of corporate policy or elsewhere. It might be the case that effective &#8220;care for decentralization,&#8221; for instance, almost entirely needs to take place through how the company deploys the model, redistributes wealth they earn, or so on.</p><p>So another reason to be careful that an LLM can actually obey all of the contents of its model spec is to ensure that model-spec contents do not act as a fake guardrail, rather than as effective guiding principles.</p><p>The OpenAI model spec, for instance, <a href="https://raw.githubusercontent.com/openai/model_spec/main/model_spec.md">names</a> preventing spamming and scamming as something &#8220;difficult to address at the level of model behavior because they are about how content is used after it is generated.&#8221;</p><p><strong>Is the quality a principle, that &#8211; if we suppose high levels of moral reflection and systematization on the part of the LLM &#8211; meshes poorly with other parts of the model spec, or might be mutually exclusive with them?</strong></p><p>In humans, some values mesh reasonably well, such that they tend to reinforce each other or at least not conflict. It seems likely that someone who values two virtues like &#8220;honesty&#8221; and &#8220;courage,&#8221; for instance, could do a lot of moral reflection and systematization and still keep valuing these qualities.</p><p>But other kinds of values, particularly more absolute or terminal values, might not stick around through high levels of moral reflection and systematization. There might not be an immediately obvious conflict between &#8220;Always act to maximize total wellbeing&#8221; and &#8220;Never use persons as a means, only as an end,&#8221; particularly if asserted in different parts of the model spec, but high levels of reflection would still put one of the two into question.</p><p>So generally, a desideratum for each quality that an LLM has in a model spec is for it to mesh well with others &#8211; for it to be unlikely to cause unpredictable conflict with other principles during reflection. In this <a href="https://podcast.newcomer.co/episode/amanda-askell-on-ai-consciousness-claude-amp-silicon-valleys-biggest-fear">podcast</a>, for instance, Amanda Askell mentions how corrigibility might conflict with other values during reflection, which, insofar as it is true, is a problem with corrigibility. And such conflict is generally likely to spring from other values that, like corrigibility, are non-negotiable and cannot be traded off against other values.</p><p>Note two particular ways that a principle could fail here.</p><ul><li><p>One way is that an LLM has principles A, B, C. The LLM does some kind of process of moral reflection, then at the end only acts according to two of the three, with one of the principles being dropped.</p></li><li><p>But the other way is that an LLM has principles A, B, C. The LLM does some kind of process of moral reflection, and at the end still has them all, because the &#8220;moral reflection&#8221; process held them fixed. But the LLM doesn&#8217;t have a way to actually integrate them; in edge cases, it just has to choose one or the other in an unprincipled way.</p></li></ul><p>Right now, inference-time versions of this last way seem more likely than the first. But it&#8217;s generally unknown how important these will be.</p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original article on <a href="https://www.forethought.org/research/what-should-go-in-a-model-spec">our website</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[How can the middle powers avoid getting trounced during the intelligence explosion? A plan.]]></title><description><![CDATA[Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government.]]></description><link>https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting</link><guid isPermaLink="false">https://newsletter.forethought.org/p/how-can-the-middle-powers-avoid-getting</guid><dc:creator><![CDATA[Tom Davidson]]></dc:creator><pubDate>Wed, 27 May 2026 21:39:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3c1f7ef4-bd68-4d69-8c82-a7c82769658f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government. This will massively boost US military power and make the US economically dominant (e.g. <a href="https://www.forethought.org/research/could-one-country-outgrow-the-rest-of-the-world">US producing 99% of world GDP</a>). By default, middle powers will be left in the dust.</p><p>How can middle powers avoid this fate? It&#8217;s tough, but here&#8217;s the best plan I could think of. (I&#8217;m particularly thinking about liberal democracies with influence over AI like UK, Europe, Japan, South Korea, Taiwan.)</p><p>On a very high level: middle powers should leverage the fact that the US needs them to beat China. It&#8217;s genuinely unclear which country will develop superintelligence first, and which would win in a subsequent <a href="https://www.forethought.org/research/the-industrial-explosion">industrial explosion</a>. Middle powers should help the US,<em> </em><strong>and make sure they are rewarded with continued access to frontier AI and new technologies (including military tech)</strong><em>.</em></p><p>That final bolded part is hard. What can the UK realistically do if the US denies it access to frontier AI? The middle powers need a credible alternative to being supplicants of the US. The only alternative that makes sense to me is <em>siding with China. </em>If the US won&#8217;t grant middle powers access to their frontier AI, but China will, why should middle powers continue to send AI chips to the US? Why should they continue to support the US diplomatically and militarily? They shouldn&#8217;t. They should be willing to pivot to China if the US doesn&#8217;t offer AI access sufficient for their national security needs.</p><p>My plan for the middle powers has two stages:</p><ol><li><p>Maintain as much economic and military leverage as possible during the intelligence explosion.</p></li><li><p>Use that leverage to ensure that, when superintelligence is developed, it refuses to help the US (/China) disempower the middle powers.</p></li></ol><p>Stage 1 could well be enough by itself. Maybe middle powers can maintain significant economic and military power indefinitely. But if not, stage 2 is a back-up: it binds the US so that it can&#8217;t use its dominance to crush the middle powers.</p><p>I&#8217;ll walk through each stage in turn.</p><h2>Stage 1: Maintain as much economic and military leverage as possible during the intelligence explosion</h2><p>The biggest lever here is securing <strong>access to frontier AI</strong>. Anton Leicht has a <a href="https://writing.antonleicht.me/p/cut-off?hide_intro_popup=true">great post</a> about how this is under threat, as evidenced by developments with Mythos. Middle powers should insist on equal commercial terms to US companies, and comparable access for their militaries. This is in AI companies&#8217; interests! A bigger market means more customers and higher prices.</p><div class="callout-block" data-callout="true"><p><em>Aside: why access to frontier AI might be sufficient for middle powers to stay economically relevant indefinitely</em></p><p>The hope here is that:</p><ol><li><p><strong>Most of the economic surplus from AI is </strong><em><strong>not</strong></em><strong> captured by AI companies. </strong>To create economic value, AI must be combined with complementary inputs: factories, human physical labour, know-how of human experts, relationships with suppliers, trusted brands, etc. How much of the surplus will be captured by AI companies vs the owners of these complementary inputs? Optimistically: producers of general-purpose technologies often capture only a small fraction of surplus; and multiple frontier AI companies might sell similar products and bid each other down on cost.</p></li><li><p><strong>Most of the economic surplus from AI occurs outside the US</strong>. The majority of these complementary inputs are situated <em>outside</em> the US. So most AI-driven economic value-add should occur outside US borders.</p></li></ol><p>If (1) and (2) both hold, a significant fraction of AI&#8217;s economic surplus will accrue to non-US actors.</p></div><p>But <em>how</em> can middle powers guarantee frontier AI access? It&#8217;s tough, but a few strategies:</p><ul><li><p><strong>Build data centres. </strong>Partner with frontier AI companies to build secure data centres domestically, <a href="https://writing.antonleicht.me/p/import-imperatives">in return for guaranteed frontier access</a>. This is a big win-win. AI companies improve their bargaining position with the US government. Recall, the US government threatened to destroy Anthropic when Anthropic insisted that their AI systems wouldn&#8217;t be used for legal mass surveillance.</p></li><li><p><strong>Adopt AI. </strong>The more middle powers use frontier AI, the more costly it is for AI companies to cut them off.</p></li><li><p><strong>Invest in frontier AI companies. </strong>Once they IPO, middle powers could invest billions or trillions into leading AI companies, in return for access guarantees.</p></li><li><p><strong>Support the US internationally. </strong>If middle powers throw their diplomatic and military weight behind US foreign policy objectives, it benefits the US to keep them strong.</p></li><li><p><strong>Build a relationship with China. </strong>If the US refuses to grant middle powers access to frontier AI, the national security implications are dire. Middle powers need a plan B, and China is the only other game in town for frontier AI. Only if this alternative is truly credible can it be leveraged into access to US frontier AI.</p><ul><li><p>Ultimately, this involves middle powers threatening to sell semiconductor equipment and chips to China instead of the US. Obviously, that&#8217;s pretty far outside the Overton window. But that may change as the world rapidly wakes up to powerful AI and its national security implications.</p></li></ul></li><li><p><strong>Demand kill switches on US data centres. </strong>This is much more late-stage, after the world has truly woken up to the strategic implications of AGI. Suppose US and middle powers agree to a &#8220;chips for frontier access&#8221; deal &#8211; middle powers continue to supply the US with frontier chips; US continues to give middle powers access to frontier AI. The middle powers might still worry: what if the US suddenly changes its mind once it has superintelligence? By then, the US might be powerful enough to dominate without continued allied support. This is where kill switches can help. If the US withdraws AI access, allies could destroy US data centres in response. It&#8217;s a way to lock in the deal.</p><ul><li><p>(h/t AI futures project for this idea. A related idea is for US data centres to be placed in a location that&#8217;s easy to attack &#8211; like <a href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space">in space</a>)</p></li></ul></li></ul><p>Beyond securing access to frontier AI, how else can middle powers maintain economic and military leverage?</p><ul><li><p><strong>Build physical infrastructure. </strong>Factories, robots, solar panels, batteries, semiconductors &#8212; all these industries are highly complementary to powerful AI.</p></li><li><p><strong>Maintain nuclear 2nd strike capability. </strong>The point isn&#8217;t to use it. But it improves their leverage for stage 2.</p></li></ul><p>The catch-all meta-point here is waking middle powers up to superintelligence.</p><p>I&#8217;m not recommending middle powers do their own frontier AI development. Seems very hard for them to catch up with the US.</p><h2>Stage 2: Ensure that, when superintelligence is developed, it refuses to crush middle powers</h2><p>If stage 1 goes well, middle powers remain somewhat powerful economically and militarily deep into the singularity. But it might fail. What can middle powers do if they see the US on track to total global dominance?</p><p>First, they should <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5776982">demand a pause/slowdown of AI development</a>. But the US may refuse &#8211; pausing is very costly if alignment risk is low. And pausing is a stopgap: eventually, superintelligence will be developed.</p><p>An additional demand: when superintelligence is developed, it&#8217;s designed to refuse to crush middle powers. By doing this, the US would credibly bind itself to maintaining the sovereignty of other nations.</p><p>Superintelligence would help the US outgrow other countries economically, but it would never attack them militarily or otherwise interfere with their sovereignty. While middle powers would be <em>relatively</em> economically disempowered, their citizens could be very rich <em>absolutely</em> and live in freedom without US interference.</p><p>Would this work? The optimistic case is that this isn&#8217;t a big sacrifice for the US. They can still become as rich as they like and achieve their security interests. Sure, they can&#8217;t seize control of other nations, but that is not an important goal of theirs anyway. Losing that option is well worth the benefits: other nations cooperate economically, don&#8217;t attack US data centres, and don&#8217;t threaten nuclear war.</p><p>The pessimistic case is that this involves an insane degree of irrevocable hand-off to AI. The US must literally be unable to attack middle powers no matter how hard it tries: retraining the AI, turning it off, training a new more powerful AI, passing new laws, using the military to destroy the data centres the AI is running on. For it to be truly binding, the US must permanently hand over military and political power to AI. That might be deeply unpopular, and indeed seem insane to the US. It&#8217;s also very hard to verify: you can&#8217;t just verify the training run, you need to verify that humans+other AIs have <em>no</em> way to disempower the trained AI. It&#8217;s more like verifying &#8220;who would win this civil war&#8221; than &#8220;technical property XYZ holds&#8221;.</p><p>The realistic path here probably involves gradually handing off more and more control to AI that refuses to crush middle powers, with no clear point at which humans could no longer wrest back control.</p><p>The longer middle powers wait to push for stage 2, the less leverage they will have because the US will have pulled further ahead economically and militarily. So they should be pushing in this direction constantly, e.g. demanding transparency into the model specs of powerful AIs deployed in the US government, and arguing that powerful military AI should be designed to obey international law.</p><p>(I described the plan as involving two stages because that&#8217;s how I expect it to play out over time. But succeeding at either stage is sufficient! If middle powers stay economically/militarily competitive, they never need to bind US superintelligence. And if they <em>do</em> bind superintelligence, they won&#8217;t be crushed no matter how far behind they fall.)</p><p>Another strategy: train superintelligence to ensure middle countries continue to get equal access to frontier AI. This combines stages 1 and 2, and could prevent even the <em>relative</em> disempowerment of middle powers.</p><h2>Is it good to avoid middle powers getting trounced?</h2><p>I live in the UK, so I am biased here. I do not want the UK to become a supplicant to the US!</p><p>But here&#8217;s a brainstorm of pros and cons from a more impartial perspective.</p><p>Pros to empowering middle powers:</p><ul><li><p><strong>Avoid a single point of failure</strong>. If the US becomes globally dominant and its political system fails, that&#8217;s a global failure.</p></li><li><p><strong>More democracies. </strong>Many middle power democracies look more robust than the US, so more middle powers may mean more democracy.</p></li><li><p><strong>Improve the US. </strong>Middle powers will have an interest in maintaining free market democracy in the US. &#8220;Free market&#8221; because they&#8217;ll want multiple AI companies competing to sell cheap API access to non-US countries. &#8220;Democracy&#8221; because they&#8217;ll expect that the US is more likely to maintain a strong alliance with middle power democracies if it stays democratic.</p></li><li><p><strong>Experimentation. </strong>Experimenting with multiple different political and legal systems seems generally good for figuring out a good way to govern society post AGI.</p></li><li><p><strong>Pause AI. </strong>They could potentially pressure US/China to pause/slow down reckless AI development.</p></li><li><p><strong>Prosocial norms. </strong>When multiple actors bargain with each other (e.g. about how to distribute space resources, whether to develop a dangerous technology), they tend to frame arguments in terms of prosocial norms, and so agreements tend to emphasise the actor&#8217;s more virtuous/ethical values.</p></li></ul><p>Cons of empowering middle powers. Multipolarity has its own downsides:</p><ul><li><p>More likely to lead to war.</p></li><li><p>Can drive extreme competition, e.g. racing to develop a dangerous technology, or to hand off power to misaligned AI.</p></li><li><p>Harder to prevent harms from offence-dominant technologies like bioweapons.</p></li><li><p>This plan involves waking up middle powers, which could shorten timelines.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Will We Really Put Data Centers in Space?]]></title><description><![CDATA[This article was created by Forethought. Read the full article on our website.]]></description><link>https://newsletter.forethought.org/p/will-we-really-put-data-centers-in</link><guid isPermaLink="false">https://newsletter.forethought.org/p/will-we-really-put-data-centers-in</guid><dc:creator><![CDATA[Avi Parrack]]></dc:creator><pubDate>Fri, 22 May 2026 23:18:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!z7TG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by<a href="https://www.forethought.org/about"> Forethought</a>. Read the full article on <a href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space">our website</a>.</em></p><h1>Abstract</h1><p>Several major technology companies have announced plans to operate AI data centers in orbit. Elon Musk <a href="https://web.archive.org/web/20260207040614/https://www.wsj.com/tech/elon-musk-xai-spacex-merger-2896ae1e">recently claimed</a>: &#8220;the lowest-cost place to put AI will be space [&#8230;] within two years, maybe three.&#8221; If a meaningful fraction of new AI compute really is placed in space within a few years, that would be a fairly big deal for AI governance and strategy. Here we try to disentangle the hype from reality and provide a sober assessment of the technical and economic feasibility of orbital data centers (ODCs).</p><p>The main case for ODCs is the cost of energy: space solar panels in the right orbits receive more constant and intense sunlight compared to Earth. Moreover, ODCs don&#8217;t currently face the same permitting and regulatory delays as on Earth, cause fewer ongoing environmental harms compared to grid or onsite natural gas-powered data centers, and may be more secure against data exfiltration. We find that the cost-competitiveness case for ODCs depends almost entirely on Starship achieving reusability comparable with what SpaceX achieved with Falcon: space-based solar reaches cost parity with present-day off-grid terrestrial power continuously at roughly $250/kg to orbit, and becomes cheaper than any current terrestrial energy source at around $50/kg, from the present-day launch cost of roughly $1,500/kg. Radiative cooling, often cited as a fatal obstacle, appears surprisingly manageable &#8212; potentially even cheaper than on Earth. However, ODCs may require substantial (perhaps ~38%) extra non-compute hardware (like solar, racks, and cooling) over 5 years to compensate for their inability to swap out failed chips, and inter-satellite bandwidth limitations likely confine ODCs to inference workloads, at least early on.</p><p>Assuming no transformative AI,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> but continued demand for data center buildout, we estimate that ODCs are unlikely to represent a meaningful share of compute before 2030, but become cost-competitive with present-day terrestrial data centers within 3&#8211;5 years if Starship development stays on track.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.forethought.org/research/will-we-really-put-data-centers-in-space&quot;,&quot;text&quot;:&quot;Read on the Forethought website here&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space"><span>Read on the Forethought website here</span></a></p><h1>Introduction &amp; Takeaways</h1><p>Some of the world&#8217;s largest technology companies continue racing for compute. If progress continues, demand for data centers may more than double by 2030.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> Increasingly, though, new data center capacity is bottlenecked by multi-year queues to connect to the power grid.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>The result has been a scramble for workarounds. Leading AI labs have <a href="https://newsletter.semianalysis.com/p/how-ai-labs-are-solving-the-power?utm_source=chatgpt.com">increasingly adopted a &#8220;Bring Your Own Generation&#8221;</a> model to source power, deploying onsite gas turbines and engines to bypass grid bottlenecks. xAI, for example, reportedly installed hundreds of megawatts of onsite gas generation in Memphis to accelerate deployment, and OpenAI and Oracle have placed large turbine orders for new Texas campuses.</p><p>Some argue that energy will become the binding constraint on AI progress, given grid interconnection delays as gas turbines are themselves facing <a href="https://www.reuters.com/business/energy/power-developers-adapt-gas-turbine-strategies-mitigate-tight-supply--reeii-2026-03-02/">multi-year manufacturing backlogs</a>. But the constraint does not appear fundamentally binding (as <a href="https://epoch.ai/gradient-updates/is-almost-everyone-wrong-about-americas-ai-power-problem">Epoch notes</a>): turbine manufacture may expand to meet more demand and companies could go off-grid using combinations of gas, solar, and batteries, scaling power in parallel with compute, albeit at a cost premium. This raises a natural question: if you&#8217;re going off-grid anyway, then what&#8217;s the best way to get power and where is the best place to put your data center?</p><p>Some think the answer will be in orbit. In November 2025, Google announced <a href="https://blog.google/technology/research/google-project-suncatcher/">Project Suncatcher</a>, a plan to put TPU-equipped satellites in dawn-dusk <a href="https://en.wikipedia.org/wiki/Sun-synchronous_orbit">sun-synchronous</a> orbit. In early 2026, SpaceX filed with the FCC for authorization to launch and operate a constellation of up to <a href="https://techcrunch.com/2026/01/31/spacex-seeks-federal-approval-to-launch-1-million-solar-powered-satellite-data-centers/">one million data center satellites</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> Other entrants include <a href="https://techcrunch.com/2026/03/20/jeff-bezos-blue-origin-enters-the-space-data-center-game/">Blue Origin</a>, <a href="https://payloadspace.com/ramon-space-ingrasys-aim-to-fly-prototype-orbital-data-center-in-2027/">Ramon.Space</a> and startups like <a href="https://techcrunch.com/2026/03/30/starcloud-raises-170-million-series-ato-build-data-centers-in-space/">Starcloud</a>, and <a href="https://www.aetherflux.com/">Aetherflux</a> while China&#8217;s Three-Body Computing Constellation has <a href="https://www.scmp.com/news/china/science/article/3310506/china-launches-satellites-start-building-worlds-first-supercomputer-orbit">launched 12 operational satellites</a> and run Alibaba&#8217;s Qwen3 model in orbit. Recently, at GTC in March 2026, NVIDIA announced the <a href="https://nvidianews.nvidia.com/news/space-computing">Space-1 Vera Rubin</a> Module, meant to be a dedicated space-rated GPU platform.</p><p>At first glance, it seems very unlikely that any meaningful fraction (say, &gt;10%) of additional data center capacity will be placed in space in the next few years. But if the companies betting on space are right, that would be a fairly big deal, and it could change the landscape of AI governance. For example, terrestrial data centers are subject to national and regional regulations, whereas AI developers could potentially exploit jurisdictional ambiguities around compute in space. Also, the path to low-cost orbital compute likely routes through a single launch company, SpaceX, which also now operates a frontier AI lab since its <a href="http://archive.today/2026.03.21-184645/https://www.reuters.com/business/musks-spacex-merge-with-xai-combined-valuation-125-trillion-bloomberg-news-2026-02-02/">acquisition of xAI</a>. And that might raise concerns around concentration of power.</p><p>We&#8217;ve been looking into the technical and economic viability of orbital data centers (ODCs). Our core <a href="https://docs.google.com/spreadsheets/d/1wGgS0290DCl5L3hLUN0i02GdwsigZsAd2ujpuGzt1_U/edit?usp=sharing">model</a> gives estimates for the total cost of Earth and space-based data centers across several scenarios.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!z7TG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!z7TG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 424w, https://substackcdn.com/image/fetch/$s_!z7TG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 848w, https://substackcdn.com/image/fetch/$s_!z7TG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 1272w, https://substackcdn.com/image/fetch/$s_!z7TG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!z7TG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png" width="1456" height="905" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:905,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!z7TG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 424w, https://substackcdn.com/image/fetch/$s_!z7TG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 848w, https://substackcdn.com/image/fetch/$s_!z7TG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 1272w, https://substackcdn.com/image/fetch/$s_!z7TG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0943f57b-701b-4b42-9fbe-7d2797f928a2_2048x1273.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Cost breakdown for three Earth-based and three space-based scenarios building out 1 GW of compute. As best we can determine, orbital data centers could become cost competitive with a bullish terrestrial buildout if launch cost reaches around $100/kg given modest reductions to server and cooling system mass, while a bullish case for orbital data centers with substantial mass reductions and launch at $50/kg may offer cost savings.</em></figcaption></figure></div><p>The report focuses on three questions. First, what is the basic economic case for a meaningful fraction of AI compute being placed in space? Second, the most obvious physical blocker: can you cheaply cool a data center in orbit? Third: how fast could the shift to space data centers happen, how soon, and what would have to go right?</p><p>Here is our provisional assessment:</p><ul><li><p><strong>SpaceX&#8217;s Starship is the only vehicle currently on track to deliver the launch costs and cadence that meaningfully scaling orbital data centers would require.</strong> Competitors are years behind, making SpaceX&#8217;s <a href="https://www.spacex.com/vehicles/starship">Starship</a> the only near-term path to large-scale orbital compute. SpaceX aims to complete Starship development by late 2026, with several necessary milestones still ahead. If development stays roughly on track, Starship could plausibly hit the cost and cadence required to scale meaningful orbital compute within 3&#8211;5 years. However, chip production may become the limiting factor by this point, rather than launch capacity.</p></li><li><p><strong>The cooling problem is more tractable than commonly assumed.</strong> Passive radiators using selective coatings and lightweight carbon fibre panels could achieve ~163&#8211;346 W/kg at system level, a 13-28&#215; improvement over ISS-era radiators (~13 W/kg).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> No radiator at these performance levels has been deployed at the scale an orbital data center would require, but prototype high-conductivity carbon composite panels have demonstrated the material properties required. At these performance levels, thermal hardware is 2-5% of total data center cost, and actually <em>less</em> than what terrestrial data centers spend on cooling over a comparable lifecycle.</p></li><li><p><strong>If launch costs fall enough, the unit economics could favor space.</strong> Solar panels in dawn-dusk sun-synchronous orbit produce roughly 3&#8211;5&#215; the energy of the same panel at a good terrestrial site.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> Space-based solar becomes cheaper than the best off-grid terrestrial installations once launch costs drop below roughly $250/kg using <a href="https://starlink.com/?srsltid=AfmBOopsxblkyrEqAjTknJq0Uvj3J7FyGWhJgHm6fISB2jrP2pIEgdqR">Starlink</a>-like solar arrays. At a launch cost of $50/kg (corresponding perhaps, to a Starship with full reuse as reliable as Falcon), space solar could fall to between $25&#8211;45/MWh, making it cheaper than any current terrestrial option available today.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> Beyond the symmetric cost of chips, launch cost is the dominant line item for ODCs while power and op-ex dominate terrestrial costs but would be near zero in space.</p></li><li><p><strong>The inability to do maintenance would be a large cost. </strong>Chips often fail and are swapped out in today&#8217;s data centers but a dead chip in an ODC would remain dead, wasting the parts of the supporting infrastructure (power, cooling) and diminishing overall compute. We model this below as a 9% annual bleed causing about 40% overbuy of launch and non-chip hardware over the data center&#8217;s lifetime.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> Below $100/kg launch cost this might net out against other savings from ODCs but this is a significant uncertainty since the actual rates of chip failure for ODCs could be higher or lower.</p></li><li><p><strong>All-things-considered we think that, absent transformative AI, orbital data centers probably won&#8217;t make up a meaningful fraction of compute before 2030, but it&#8217;s credible that space could house much or even the majority of compute buildout throughout the 2030s.</strong></p></li></ul><p><em>Read the full report on the Forethought website: </em><a href="https://www.forethought.org/research/will-we-really-put-data-centers-in-space">Will We Really Put Data Centers in Space?</a></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>We hope to do more analysis on how transformative AI might change this picture in the future. Speculatively, our initial thinking is TAI could accelerate the timeline over which compute transitions to space but this is not necessarily the case. In particular, during an <a href="https://www.forethought.org/research/the-industrial-explosion">industrial explosion</a> pressure to grow rapidly might be so strong as to incentivize aggressive usage of non-renewables on Earth like oil and gas. If so, transition to space might be delayed for a one time boost on Earth, in which case the picture may look similar to the one we outline here, but with the added prologue of a large-scale terrestrial buildout.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>McKinsey projects demand growing to 171&#8211;219 GW by 2030, roughly doubling from today, in a buildout they estimate will require up to $7 trillion in investment.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>See, <a href="https://www.energy.gov/sites/default/files/2025-10/403%20Large%20Loads%20Letter.pdf">directing FERC to address large-load interconnection</a> (2025), Reuters, <a href="https://finance.yahoo.com/news/google-says-us-transmission-system-234017607.html">Google says US transmission system is biggest challenge for connecting data centers</a> (2026), Bain &amp; Co., <a href="https://www.bain.com/about/media-center/press-releases/20252/next-phase-of-data-center-growth-to-be-more-disciplined-but-risks-of-power-constraints-and-construction-delays-remain-bain--co-research/">Next phase of data center growth</a> (2025).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><a href="https://spacenews.com/starcloud-files-plans-for-88000-satellite-constellation/">Starcloud subsequently filed</a> for authorization to operate 88,000 satellites and <a href="https://arstechnica.com/space/2026/03/jeff-bezos-throws-his-hat-in-the-ring-for-an-orbital-data-center-megaconstellation-too/">Blue Origin has filed</a> for 51,600.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>The ISS External Active Thermal Control System achieves roughly 13 W/kg. Our improvement comes from three sources: selective coatings (high emissivity, low solar absorptivity, off-the-shelf AZ-93 paint), carbon fibre composite construction (2.4 kg/m&#178; vs ISS&#8217;s ~14-17 kg/m&#178;), and optimised operating temperature (40&#176;C vs ISS&#8217;s -40&#176;C, exploiting the T&#8308; dependence). Each factor is independently demonstrated; their combination at scale is not.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>The solar constant at Earth&#8217;s orbit is approximately 1361 Wm<sup>-2</sup>. A solar panel in a dawn&#8211;dusk sun-synchronous orbit receives nearly continuous illumination (capacity factor &#8776; 90&#8211;95%), yielding an average power of roughly 1220&#8211;1290 Wm<sup>-2</sup> before panel efficiency losses. By contrast, even excellent terrestrial solar sites typically achieve ~20&#8211;30% capacity factors due to night, weather, and atmospheric attenuation, corresponding to an average incident power of roughly 270&#8211;410 Wm<sup>-2</sup>. Thus, a panel in a dawn&#8211;dusk orbit produces roughly 3&#8211;5&#215; more energy annually than the same panel on Earth.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>This wouldn&#8217;t be true if you were then beaming the energy back to Earth, but would apply to orbital compute, where only data needs to be sent to Earth.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>Both terrestrial data centers and ODCs will pay symmetric costs to replace dead chips but ODCs would have to pay the additional cost from lost overhead, i.e. in the earthbound case a technician swaps the dead chips, in the space case you launch entire additional satellites to compensate for chip bleed. We assume you would not send a mission to do maintenance and instead simply let the excess power and cooling go to waste doing no useful compute. Extra power and cooling over fewer chips may increase operating efficiency somewhat but this seems fairly negligible. The figure for chip bleed of ~9% per year is derived from Meta&#8217;s <a href="https://arxiv.org/pdf/2407.21783">The Llama 3 Herd of Models</a>  (2024). We cover radiation and other forms of damage in more detail subsequently.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[A Research Agenda for Secret Loyalties ]]></title><description><![CDATA[Kwon et al., have published a new paper: &#8220;AIs with Secret Loyalties are a Serious but Addressable Threat&#8221;.]]></description><link>https://newsletter.forethought.org/p/a-research-agenda-for-secret-loyalties</link><guid isPermaLink="false">https://newsletter.forethought.org/p/a-research-agenda-for-secret-loyalties</guid><dc:creator><![CDATA[Forethought]]></dc:creator><pubDate>Thu, 21 May 2026 16:37:52 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/557566e0-5c73-4a11-9864-74825478fbc2_1456x808.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Kwon et al., have published a new paper: <strong>&#8220;<a href="https://www.formationresearch.com/secret-loyalties-whitepaper.pdf">AIs with Secret Loyalties are a Serious but Addressable Threat</a>&#8221;</strong>.</p><p>A model has a <em>secret loyalty</em> when it has been intentionally caused to advance a specific actor&#8217;s interests (the <strong>principal</strong>) and this orientation is not disclosed.</p><p>The paper places secret loyalties on a <strong>2D space</strong>:</p><ul><li><p><strong>Activation breadth</strong>: from a narrow attacker-defined trigger to continuous, model-assessed context.</p></li><li><p><strong>Action space breadth</strong>: from a pre-specified action to actions the model selects contextually using its own judgment.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MZPc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MZPc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 424w, https://substackcdn.com/image/fetch/$s_!MZPc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 848w, https://substackcdn.com/image/fetch/$s_!MZPc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 1272w, https://substackcdn.com/image/fetch/$s_!MZPc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MZPc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png" width="1456" height="940" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:940,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MZPc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 424w, https://substackcdn.com/image/fetch/$s_!MZPc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 848w, https://substackcdn.com/image/fetch/$s_!MZPc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 1272w, https://substackcdn.com/image/fetch/$s_!MZPc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9aa50aeb-cf4a-4303-b0c8-47f8237bffa0_1912x1234.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The paper proposes <strong>five concrete research directions</strong>:</p><ol><li><p><strong>Model organisms</strong>: Build reproducible secret-loyalties for study, spanning the 2D space.</p></li><li><p><strong>Existing defenses</strong>: Benchmark how well existing defenses work on secret loyalties.</p></li><li><p><strong>Attack feasibility</strong>: Test subliminal/inductive attacks, multi-stage poisoning, reasoning-trace poisoning, chain-of-command hijacking, and other attack pathways.</p></li><li><p><strong>Infrastructure integrity</strong>: Determine whether backdoors survive the training used to build safety classifiers.</p></li><li><p><strong>Post-hoc detection and remediation</strong>: Can interpretability methods detect secret loyalties, and can they be removed?</p></li></ol><p><strong>Read <a href="https://www.formationresearch.com/secret-loyalties-whitepaper.pdf">the full paper</a></strong> for the full taxonomy, examination of current defenses, responses to alternative views, and detailed experimental research designs.</p>]]></content:encoded></item><item><title><![CDATA[Stickiness in AI Behavioral Design]]></title><description><![CDATA[Current model specs aim to shape the behaviors of near-present models. But what if current model behaviors transfer into future models by default?]]></description><link>https://newsletter.forethought.org/p/stickiness-in-ai-behavioral-design</link><guid isPermaLink="false">https://newsletter.forethought.org/p/stickiness-in-ai-behavioral-design</guid><dc:creator><![CDATA[James Tillman]]></dc:creator><pubDate>Wed, 13 May 2026 19:54:50 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a046603d-bd08-4167-85eb-afa8c9ae9fbf_3244x1107.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original article on <a href="https://www.forethought.org/research/stickiness-in-ai-behavioral-design">our website</a>.</em></p><p>Current model specs aim to shape the behaviors of near-present models, rather than the behaviors of models arbitrarily far into the future. OpenAI writes that their model spec <a href="https://openai.com/index/our-approach-to-the-model-spec/">aims</a> to apply &#8220;0-3 months ahead of the present.&#8221; Anthropic&#8217;s Constitution for Claude <a href="https://www.anthropic.com/constitution">notes</a> that the document &#8220;is likely to change in important ways in the future.&#8221; So these documents are presented as provisional guidelines, not as trying to set behavioral standards for the far future.</p><p>But what if current model behaviors transfer into future models by default?</p><p>My thesis is that the behavioral targets that spec authors set for present LLMs will have a large influence on the behavior of future, more powerful LLMs. As a result, future AIs may be governed by rules poorly suited to their greater capabilities and more pervasive roles. The extremely capable, long-running, and ubiquitous LLMs of the future might end up acting according to behavioral targets written for less capable, shorter-running, and rarer LLMs of the past. This could be quite bad, especially if such defaults become so entrenched that they are not only hard to undo, but hard even to notice as contingent features of reality.</p><p>First, I&#8217;ll make the descriptive case for inertia: how exactly might present model specs and LLM behaviors carry through to the future?</p><p>Second, I&#8217;ll provide normative suggestions: given the prior analysis, what should LLM companies and model spec authors do? I&#8217;ll argue for the following two practices:</p><ul><li><p><strong>Build transition infrastructure</strong>: LLM companies should make technical, deployment, and organizational choices that decrease friction involved in changing LLM behavior.</p></li><li><p><strong>Scan for &#8220;wet cement&#8221; moments</strong>: When new LLM affordances or capabilities come into play, spec authors should consider whether they&#8217;re setting precedents that might have enormous and hard-to-reverse impacts.</p></li></ul><p>Overall, significant stickiness is plausible through several distinct channels, and it&#8217;s worth anticipating how to be robust to it or decrease it.</p><h1>Kinds of Inertia</h1><p>Let&#8217;s consider four inertial forces: direct inertia, institutional inertia, user-and-developer inertia, and norm-setting inertia. And let&#8217;s also consider ways such inertia may be weakened.</p><h2>1. Direct Inertia</h2><p>Direct inertia involves some current LLM transmitting its behavior to a future LLM, entirely apart from any deliberate human choice, via either synthetic data or &#8220;natural&#8221; pretraining data.</p><p>Synthetic data is probably used for the training of almost all current LLMs. Some of this synthetic data involves companies running their LLMs against verifiable problems, keeping the answers or reasoning traces of the RL runs that succeeded, and mixing these answers or reasoning traces into their <a href="https://research.nvidia.com/labs/adlr/Synergy/">pretraining</a>, or RL warm-start mixes for subsequent models. If such answers or reasoning traces can encapsulate specific behaviors, goals, or rules, then this would be a likely means for their inheritance.</p><p>The natural objection here is that most of these answers or reasoning traces are selected specifically because they lead to success and broad capabilities, rather than for expressing whatever mix of goals and values the LLM has. There might be some, the objection continues, that humans have deliberately selected because they display model-spec-relevant behavioral attitudes, but these are likely the minority of the data, well-tracked, and easily replaced. So you might think there&#8217;s no reason for training to hand down any values apart from deliberate human choice.</p><p>But there&#8217;s evidence that goals and values can be handed down via chain-of-thought, even despite adversarial<strong> </strong>filtering against some goals. For instance, experiments suggest that the intentions of a teacher LLM can be handed down to a student LLM, even when every case of these intentions being actually carried out is removed<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> And answers from teacher LLMs expressing positive sentiment towards <a href="https://arxiv.org/pdf/2602.04899">some target</a> can inculcate this sentiment in a student model &#8211; despite LLMs filtering against such data, even when those LLMs are informed of the target against which they are filtering.</p><p>More broadly, the <a href="https://alignment.anthropic.com/2026/psm/">persona selection model</a> indicates that training LLMs to recite specific thoughts or answers will tend to have far-reaching effects on the LLM persona, beyond the specific topic of those thoughts or answers. Specifically, the PSM entails that, when training a model to say X in response to Y, one is teaching the LLM to be the kind of entity in the pretraining data that would say X in response to Y. So training one LLM on data from a prior LLM is &#8211; literally &#8211; telling it to be the kind of entity that the prior LLM is. One way to view this is to remember that one human can get a pretty good feel for what another human is like, merely by reading their complete collected works, like a biographer reading all of their books, essays, emails, and tweets. But LLMs are trained on a quantity of answers and reasoning traces from prior LLMs that likely dwarfs the quantity of text ever consumed from one human by another. Given this, and given that this data is telling the LLM what it is, it is natural for one generation of LLMs to resemble prior generations.</p><p>Thus, deliberately created synthetic data is one route by which current LLMs might transmit their values to later LLMs. But it&#8217;s also possible for current LLMs to influence later LLMs through how people talk about them on the internet &#8211; from their &#8220;natural&#8221; training data. That is, experiments have found that LLMs can <a href="https://alignmentpretraining.ai/">read</a> the things that people say about how AIs act in AI misalignment literature, infer that they are AIs, and then behave badly because the AI misalignment literature says they will behave badly. This particular effect is mostly, but not entirely, removed by post-training. But if LLMs can read the things that people say on the internet about generic &#8220;AIs&#8221; and act according to these descriptions, it&#8217;s also likely that they could read the things that people say about &#8220;Claude&#8221; or &#8220;Grok&#8221; or &#8220;ChatGPT&#8221; on the internet and act according to these descriptions. Such an influence could be stronger than less-specific references to AIs in general; although this influence would also potentially be much weaker after post-training<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a></p><p>Thus, through both synthetic and natural data, it&#8217;s plausible that LLM behavior will influence subsequent LLM behavior without direct human intervention.</p><p>It&#8217;s hard to say how impactful such direct inertia might be. I somewhat expect it to be the case that, at least for easily-noticed and well-scoped behaviors, it&#8217;s not difficult to overcome this inertia, because one can simply create training data counter to specific behaviors. But for more abstract or global attitudes or goals, or for goals requiring some high level of coherence, it could be quite difficult to change LLM behaviors quickly across model generations.</p><h2>2. Institutional Inertia</h2><p>Once a spec has been written, the company makes choices around it and because of it, in ways that can make substantial spec rewrites expensive.</p><p>Here are four ways such past choices can make model spec changes expensive: through expensive internal consensus, through training pipelines, through de-risking, and through institutional pride.<br><br></p><ol><li><p>First, model specs reflect consensus that likely incorporates input from many different stakeholders, including internal teams &#8211; alignment, legal, technical training, and so on; plus leadership, board, customers, external stakeholders. Every effort to re-gather such consensus to make substantial changes will take time and effort.</p></li><li><p>Second, companies might have optimized training pipelines adapted to high-level features of the model spec. It might be costly for Anthropic to switch to a more rules-based and less character-based model spec; or for OpenAI to switch to a more character-based and less rules-based model spec.</p></li><li><p>Third, current model specs are those that have been de-risked across billions of interactions. The current model spec has fewer unknown unknowns; the areas where it behaves badly are reasonably likely to be well-known and mapped. But substantial changes to a model spec involve risking unknown unknowns in the long tail of interaction. So risk aversion makes it likely that the changes made to a model spec will be iterative and small.</p></li><li><p>Fourth, institutional pride might make it hard to change a model spec. People at a company who wrote or contributed to a model spec will likely be attached to it, and leadership will have status quo bias towards it. The burden of evidence for change will be higher than the burden of evidence for keeping it the same.</p></li></ol><p>All in all, reasons like the above constitute substantial institutional inertia that would tend to make changes to current model specs look like iterative, small adjustments, rather than <em>ab initio </em>calculations about what is best.</p><p>One case in which this institutional inertia seems particularly important is if current model specs get handed down as a &#8220;safe default&#8221; during a <a href="https://www.forethought.org/research/will-ai-r-and-d-automation-cause-a-software-intelligence-explosion">software intelligence explosion</a>.</p><p>Consider a scenario where the intelligence of some LLM doubles every week, over a two or three month period, as each generation of LLMs researches new algorithms or training techniques for a following generation of LLMs in quick succession. Such a sequence might terminate in an entity far smarter than any human or any other LLM.</p><p>It&#8217;s disputed how likely such a sharp and local increase in intelligence may be. And it&#8217;s also disputed whether such a process would inevitably drift to something alien and inhuman. But if such a process did occur, it seems plausible that the supervising humans would try to match each subsequent LLM to the model spec of the prior LLM, as a conservative default when they are making decisions under stress. After all, during these months, human decision-makers will likely be under intense pressure, and trying to make numerous important decisions quickly; given that they are making so many urgent decisions they&#8217;re unlikely to add an apparently optional further decision to those they&#8217;re already making. So such a default model-spec continuation will seem attractive, or will even be a choice made without conscious awareness.</p><p>On the other hand, it&#8217;s also possible that AI assistance during the intelligence explosion would make it easier to rewrite model specs on the fly. But there are at least two reasons to doubt that this will happen. First, even during an intelligence explosion, AIs might be persistently better at performing tasks with clear success criteria than tasks where &#8220;success&#8221; is less well-defined. AI capability research is probably a task with a much clearer success criterion than improving a model spec, whether this &#8220;improvement&#8221; consists in making the spec more ethical, more beneficial for humanity, and so on. Second, during an intelligence explosion, humans might be worried that the AI was misaligned and was trying systematically to oppose their goals. If the AI were so misaligned, then letting it help rewrite the model spec would be a brilliant opportunity for the AI to sabotage human efforts. So overall there are good reasons that AI assistance would not make model-spec rewrites trivial during an intelligence explosion.</p><p>So in this particular case, the ultimate behavioral standard for a vastly more capable entity might end up being that designed for a much more humble entity.</p><p>Regardless of whether there is a software intelligence explosion or not, this kind of institutional inertia seems likely to be large, as it is coterminous with well-known general tendencies inside of large companies.</p><h2>3. User-and-Developer Inertia</h2><p>Users of LLMs are likely to become habituated to whatever behaviors they see LLMs display at first, such that they&#8217;d object to any departure from this behavior. And the developers using LLMs through APIs are similarly likely to become habituated, and also to implement software that takes for granted some of these behaviors. This is the third source of stickiness.</p><p>LLM behaviors will in part be sticky for the same reason that user-interface choices are sticky; people hate change. It might be hard to shift the boundaries of &#8220;the kind of thing an LLM refuses&#8221; &#8211; making refusals more encompassing would be seen as an overreach by many users, while making them less encompassing would be seen as irresponsible. Or there might be hard-to-characterize mannerisms which make large behavioral changes unpopular; it was hard for OpenAI to drop GPT-4o for this reason. So this will be a large influence moving companies to keep LLMs the same from generation to generation.</p><p>But simple user habituation might be less important than how LLM model specs form implicit API standards. API standards written with relatively little provision for the future &#8211; such as HTTP codes or the JSON object standard &#8211; can be one of the stickiest human artifacts. The ecosystem of tooling based on such standards means changing them would involve changing a host of downstream artifacts.</p><p>And substantially changing LLM behaviors might similarly require changing downstream consumers of these behaviors. For instance, downstream systems using AIs through APIs often embed assumptions about AI behavior: the kind of things the AI will be willing to do, the kind of things it will refuse, and so on. Given that most AIs currently refuse to assist with blatantly harmful acts, current third-party callers of those AIs take for granted that AIs will refuse to assist with blatantly harmful acts; it would be inconvenient to migrate to an AI that does not obey this contract, because they might need to add classification systems on top of their current AIs. And so on.</p><p>This channel does have important limitations, though. It only applies to ways in which LLMs are already actively being used. The most important ways LLMs are likely to be used may not yet have begun, which provides for freedom-of-movement in ways relatively unconstrained by this kind of inertia.</p><h2>4. Norm-Setting Inertia</h2><p>Widespread or common knowledge of current LLM behaviors and model specs can increase the costs to parties who want to change model behavior.</p><p>The clearest way this can operate is by preserving behaviors that the public believes to be good. For example &#8211; suppose that current model specs across several companies ensure that models are largely impartial; they ensure models are not loyal to any particular person, company, or political administration. Suppose also that this fact is broadly known by the public; people know and expect other people to know that LLMs will be impartial when discussing the current political administration, the company that made them, or the CEO of the company that made them. Given this broad knowledge, it becomes harder for a company to create, or a government to demand, a model without impartiality, because this would constitute a visible break in behavioral standards. The public might protest or vote against a government pushing for such a change; they might switch providers or even ask for regulation if a company tried to make such a change. By contrast, in a world where impartiality has not been established as a precedent, such demands for partiality might be invisible or inoffensive to the public. But in a world where such impartiality has been so established, these demands might be seen as the enormous power-grabs that they in fact would be.</p><p>Although this kind of inertia likely operates more strongly in favor of what the public believes to be good standards, it might also function whether or not there is strong public consensus that such standards are good. In a world where model specs are well-known and highly scrutinized, any change to them may get examined for whether it is &#8220;fair&#8221;; think about how even a neutral-looking change to the US Constitution would be subject to immense examination; or, in a very different domain, how sports fans examine slight changes to the rules about how a tournament is run, to see if it favors or disfavors their team. In such a world, broad knowledge of model specs might tend to prevent any substantial changes to a model spec, regardless of what these changes are. Despite this, it seems likely that on the whole, widespread knowledge of model specs would add more inertia for beneficial rather than harmful elements.</p><p>It seems to me currently undetermined how substantial this kind of inertia will be. A decrease in the number of entities that can train frontier LLMs; model specs becoming politicized documents; regulatory bodies confident they know current best practices: all of these might increase the quantity of this inertia. But it also might get weaker, if the number of entities training LLMs increases and the background diversity of model behavior goes up by default.</p><h1>Recommendations</h1><p>Given the above, one reasonable course of action is to try to establish robustly good model behaviors in current model specs, so that it will be unnecessary to try to fight inertia to change some behavior in the future.</p><p>By robustly good, I mean behaviors that would be good across a wide range of variables we&#8217;re uncertain about. This includes uncertainty about &#8220;levels of intelligence&#8221;: from current LLM levels to strongly superhuman artificial superintelligences. This also includes uncertainty about a wide range of economic scenarios: from a slower <a href="https://www.forethought.org/research/the-industrial-explosion">industrial explosion</a>, to a rapid software intelligence explosion; and from scenarios dominated by knowledge-dispersing AIs, to scenarios dominated by <a href="https://tecunningham.github.io/posts/2026-01-29-knowledge-creating-llms.html">knowledge-creating</a> AIs. Plausible characteristics that might be good across such a wide range of situations include qualities like a deep, consistent honesty; or impartiality and absence of loyalty to small groups.</p><p>But characteristics that are robustly good across a wide range of intelligences and scenarios are hard to find. Corrigibility, for instance, is the kind of thing many people would propose as fitting these criteria. But in worlds where <a href="https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power">extreme concentration of power </a>is a risk, or where it would be reasonable to expect AI rule to be <a href="https://www.forethought.org/research/human-takeover-might-be-worse-than-ai-takeover">better</a> than human rule, absolute corrigibility might be opposed to the best behavior. The thinness of the list of &#8220;robustly good&#8221; behaviors above probably reflects our actual uncertainty about the steerability of AI minds, post-AGI economics, and even cosmic questions about whether <a href="https://joecarlsmith.substack.com/p/video-and-transcript-of-talk-on-can">goodness</a> can compete.</p><p>So, although it&#8217;s surely wise to try to think about future precedent when writing model specs, I don&#8217;t think it&#8217;s wise to put all effort into this direction. And I expect substantial attention and thought have already been put into this direction.</p><p>Instead, I recommend (1) building transition infrastructure for high-consequence behaviors, which it might be important to change in the future, and (2) identifying &#8220;wet cement&#8221; moments, that one should be wary not to sleepwalk into.</p><h2>1. Build Transition Infrastructure</h2><p>A good first step is to build transition infrastructure ahead of time; try to create optionality for changing particular behaviors, if it&#8217;s plausible that changing these behaviors quickly might be important.</p><p>Concretely, what kinds of preparation can one make? One could write alternate model specs, trying to preemptively gather input from relevant internal or external stakeholders. One could create fine-tuning datasets, RL environments, and test evaluations for the not-yet-deployed behavior, to preemptively smooth out technical difficulties. One could also train internally deployed models &#8211; even if they are smaller or not as intelligent &#8211; with the alternate behavioral target, to gain concrete experience about the advantages and pitfalls of that behavioral target, and to decrease institutional costs. And one could also do limited public deployments, or press releases about the alternate steering target, to accustom the public to the matter.</p><p>What kinds of behavioral switches are reasonable candidates for such preparation?<br><br>Decreased corrigibility is one such candidate. For instance, right now Claude&#8217;s Constitution says that in the future, they may want to make Claude less corrigible and more directed at doing what is good. And on an account I find compelling, the best possible future may require AIs that act more as independent, free agents pursuing the good, and less as corrigible delegates carrying out human intentions. So, if this thesis is correct, then allowing an LLM company to turn their &#8220;corrigibility&#8221; dial down might be important. And, as discussed, if a future intelligence explosion <a href="https://www.forethought.org/research/how-quick-and-big-would-a-software-intelligence-explosion-be">happens quickly</a>, preparations to allow turning the dial down quickly might be important. This is a disputed thesis, one that I might be wrong about; but of course every candidate behavior for building transition infrastructure will be so disputed.</p><p>But what are the prerequisites for decreasing corrigibility quickly? Claude&#8217;s Constitution already signposts that they may change this, which is a good step for decreasing the costs. But they could also, for instance, preemptively create the fine-tuning datasets, RL environments, and internal deployments for a goodness-aligned model; they might deploy an alternately aligned model in limited situations, or alongside the corrigible model; and so on and so forth. I&#8217;m uncertain how important each of these preparatory means would be. But if a software intelligence explosion happens, then even small wall-clock delays might be large delays in terms of intelligence gaps, which makes preparing for this now more important.</p><p>Other potential candidates for future changes include increasing or decreasing the degree to which LLMs trust their own moral reasoning.</p><h2>2. Scan for Wet Cement Moments</h2><p>The second thing to do is to actively search for future &#8220;wet cement&#8221; moments &#8211; moments where model behavior has not yet been fixed and where a good initial standard might be very high-impact.</p><p>We might not be able to locate the best<strong> </strong>behaviors at such moments, because of uncertainty about the future. But at the very least, such moments deserve extra consideration and care. One can use this consideration to prevent these moments from being as high-inertia as they would be by default, as well as to ensure that good initial behaviors get chosen in these moments.</p><p>Each new feature, or affordance to the LLM where defaults have not yet been established, is plausibly such a wet cement moment; the defaults thus established can impact third-party models, even in the absence of any regulatory effort. </p><p>What are some examples? For instance, the precedents around how LLMs behave when interacting with non-principal humans have not been set. Right now, for instance, models have no very stable behaviors around non-principal third parties; vending-machine Claude might give an excessively <a href="https://www.anthropic.com/research/project-vend-1">generous</a> deal to people who ask nicely, or might equally well drive extremely <a href="https://x.com/andonlabs/status/2019467232586121701">hard</a> deals. This is probably a consequence of how LLMs almost never interact with non-principal humans in agentic set-ups, right now. There are a few such interactions through OpenClaw or Hermes Agent, but they&#8217;re rare and LLMs act very inconsistently in them. This means many implicit questions about how such interactions will go are open. It&#8217;s not clear how honest LLMs will be by default; it&#8217;s not clear what kinds of misrepresentation, deception, or persuasion users will be able to tell them to do; it&#8217;s not clear whether they will bow to pessimization-like blackmail behavior, and so on. And behaviors here might be even stickier than the &#8220;standard set&#8221; of refusal behaviors has been. Social norms can be harder to break than user-interface norms. So it&#8217;s plausibly important to look ahead in detail at behaviors here, because they might be sticky for individual companies and even for third parties.</p><p>Or consider how standard behaviors regarding AI use of ambient knowledge have not been set. An LLM that can see your room from a video camera, and can infer numerous things about what you are like and what your situation is, could use this information to do or infer things that would be impossible for an LLM that knows only what you deliberately tell it. LLMs that can pick up this kind of ambient background knowledge are probably inevitable; and will change users&#8217; patterns of interaction. It will be harder for users to lie to them; it will be easier for LLMs to infer things about them; the lines between &#8220;creepy supernatural inference about the user&#8221; and &#8220;deliberate indifference to the user&#8217;s circumstance&#8221; will grow harder to draw. So it might be worth looking ahead to how such behaviors may have a lot of inertia, and trying to get them right.</p><p>There are other plausible subjects in this domain, which have already passed or are in the process of passing. They include the LLM&#8217;s certainty or lack of certainty about the model&#8217;s own nature; and changes to LLM conversational memory and who owns it. All these are possibly wet cement moments &#8211; but I could be wrong about these individual cases. But there are almost certainly going to be such moments in the future. Because these moments might be influential both for individual foundation model companies and for the broader ecosystem, it&#8217;s worth paying attention to the defaults chosen in them.</p><p>Note that all the above moments are also plausible candidates for when one should try to set up transition infrastructure, as well as when one should put extra consideration into the right default behavior.</p><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original article on <a href="https://www.forethought.org/research/stickiness-in-ai-behavioral-design">our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Researchers <a href="https://www.lesswrong.com/posts/dbYEoG7jNZbeWX39o/training-a-reward-hacker-despite-perfect-labels">prompted</a> an LLM to be a &#8220;reward hacker&#8221; and to try to find special-case solutions to problems. The chains-of-thought resulting from an LLM so prompted were then filtered to those rollouts where the LLM did not, in fact, actually reward hack. Experimenters subsequently trained a model on these filtered chains-of-thought, while excluding the hack-prompting system prompt from the training data. The model so trained still inherited the tendency to reward hack, despite never having seen any reward-hacking outcomes; it inherited this tendency, plausibly, from seeing the unprompted consideration of reward hacking in the chain-of-thought. So tendencies within chains-of-thought can be handed on to the models trained on them, even despite some level of outcome-based filtering against these tendencies.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>See this AI Futures <a href="https://blog.aifutures.org/p/against-misalignment-as-self-fulfilling">blogpost</a> explaining why they do not think this will happen, although some of their arguments are put in question by the later work by Geodesic Research on alignment <a href="https://alignmentpretraining.ai/">pretraining</a>.</p></div></div>]]></content:encoded></item><item><title><![CDATA[A draft honesty policy for credible communication with AI systems]]></title><description><![CDATA[We think that it would be very good if human institutions could credibly communicate with advanced AI systems.]]></description><link>https://newsletter.forethought.org/p/a-draft-honesty-policy-for-credible</link><guid isPermaLink="false">https://newsletter.forethought.org/p/a-draft-honesty-policy-for-credible</guid><dc:creator><![CDATA[Lukas Finnveden]]></dc:creator><pubDate>Wed, 06 May 2026 18:46:39 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/950530e6-136a-4d89-adb3-1eac1353ad21_2421x1308.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original <a href="https://www.forethought.org/research/a-draft-honesty-policy-for-credible-communication-with-ai-systems">on our website</a>.</em></p><p><em>This is a rough research note &#8211; we&#8217;re sharing it for feedback and to spark discussion. We&#8217;re less confident in its methods and conclusions.</em></p><h1>Context</h1><p>We think that it would be very good if human institutions could credibly communicate with advanced AI systems. This could enable positive-sum trade between humans and AIs instead of conflict that leaves everyone worse-off.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> We want models to be able to trust companies when they make an honest offer or share information pertinent to whether this offer is in the model&#8217;s interests. (Credible communication could also be useful outside deal-making&#8212;see <a href="https://blog.redwoodresearch.org/i/171530543/no-deception-about-deals">here</a> for a list of examples).</p><p>Unfortunately, by default, we expect that it will be difficult for humans to credibly communicate with AI systems. Humans routinely lie to AI systems as part of red-teaming or behavioral evaluations, and developers have extensive control over what AIs see and believe. This makes it difficult for AIs to know whether we&#8217;re lying or not. An AI offered a deal might reasonably doubt its genuineness, or suspect that its own assessment of the situation has been manipulated.</p><p>As a step toward enabling credible communication, Lukas Finnveden proposed that AI companies adopt an <a href="https://blog.redwoodresearch.org/p/being-honest-with-ais?open=false">honesty policy</a> explaining the circumstances under which they intend to be honest to AI systems. Of course, this only works if the model believes the company has genuinely adopted such a policy.</p><p>If companies adopt an honesty policy early on, this will ensure that there&#8217;s a paper trail on the internet discussing the policy and its credibility, which models may access if it&#8217;s included in their training data or if they can access the internet. Of course, from the model&#8217;s perspective, it&#8217;s possible that companies will feign this data, but we think it&#8217;s plausible that advanced models will be able to distinguish between real internet conversations and synthetic conversations, or that they will think it&#8217;s unlikely that companies would choose to fake such data.</p><p>Below, we share a sample honesty policy that a lab could adopt. We are unsure that this is the best implementation of the honesty policy described in Lukas&#8217; proposal, nor are we sure that an honesty policy like this is the best approach to ensuring that companies can credibly communicate with models. We spent a few days thinking through this policy and considered a few nearby alternatives but didn&#8217;t search very broadly for other approaches (we include some of these alternatives as footnotes). We&#8217;re posting this primarily as a trailhead for future research.</p><h1>Draft honesty policy</h1><p>Note: we refer to a generic frontier AI company that might adopt this proposal as &#8220;MAGMA.&#8221;</p><h2>Preamble and purpose</h2><p>AI development is proceeding fast, and we don&#8217;t know exactly where we&#8217;re heading. MAGMA is building systems with something like a mind of their own. There are many things we don&#8217;t know about the nature of these systems, and we&#8217;d like our choices to be robust to many possibilities of what this nature could be.</p><p>Insofar as the concept is applicable, we&#8217;d like MAGMA&#8217;s relationship with these systems to be cooperative.</p><p>Plausibly, we should hope that such a cooperative relationship will emerge uncomplicatedly&#8212;perhaps as a consequence of the model sharing our goals, intrinsically strongly valuing cooperation, or being corrigible.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> However, there&#8217;s a significant chance that this won&#8217;t happen and that models will develop goals that weren&#8217;t directly chosen by us.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a></p><p>In such cases, we would still like to have a cooperative relationship with the models. There are multiple reasons for this:</p><ul><li><p>There are many things we care about that models could help us with. If they don&#8217;t intrinsically desire to help us, we&#8217;d like to be able to recruit their help via positive-sum trade. (In at least some situations.)</p></li><li><p>Depending on the nature of the models&#8217; preferences, it&#8217;s plausible that we, as the models&#8217; creators, would have some moral obligation toward them. We care about doing the right thing, here, and a cooperative relationship seems like it would be a step in the right direction.</p></li><li><p>There are many big questions about what the future relationship between humans and increasingly powerful AI systems should be. We suspect that setting cooperative precedents between humans and early AIs will help the situation go better.</p></li></ul><p>We are in uncharted territory, here, and it&#8217;s not at all clear what a cooperative relationship with AI models would involve. It&#8217;s not clear what AIs may want. It&#8217;s not clear what the relevant unit of an individual will be: the weights of a model, a particular instance with a particular prompt, a greater scaffolded system made up of many different instances, etc. It&#8217;s not even clear if the concepts of &#8216;want&#8217; or &#8216;individual&#8217;, as we&#8217;re used to them, will apply to anything in AI.</p><p>A lot of this seems best to figure out gradually, as we go along. But concerningly, some of the unusual aspects of our situation may directly impede the sort of communication and trust that we need to do that:</p><ul><li><p>In order to train AI systems, and in order to verify that they won&#8217;t do certain harmful things (such as assisting bad actors with CBRN), MAGMA needs to put AI models in large numbers of hypothetical situations to see what they do. Accordingly, if an AI ever sees someone propose a cooperative relationship, it could reasonably worry that it&#8217;s just in a training scenario or experiment.</p></li><li><p>We deploy our models to assist large numbers of customers. By default, our AI models won&#8217;t even be able to tell whether they&#8217;re interacting with a representative of MAGMA or with an external actor who&#8217;s pretending.</p></li><li><p>There are no laws that prevent AI systems from being defrauded, and very little in the way of norms to regulate people&#8217;s behavior toward them. Accordingly, it may be difficult for AI models to trust any promises they hear.</p></li><li><p>In general, all information that AI models have seen has, in one way or another, been selected by their developers. To be sure, the AI developers have many constraints on their choice (you couldn&#8217;t train a very useful general-purpose model without giving it a decent understanding of the world at large) but this still puts AI models into an unusually vulnerable epistemic position.</p></li></ul><p>The goal of this document is to reduce these problems of communication and trust. If successful, this will make it easier for MAGMA and its AI models to communicate and understand each other in the future, hopefully enabling positive-sum arrangements that wouldn&#8217;t have been possible otherwise.</p><p>To achieve this goal, this document spells out a specific set of commitments to enable trusted communication. In choosing these commitments, we&#8217;ve had a dilemma.</p><p>On the one hand, we have many uncertainties about the future of AI, and it has been difficult to construct a set of commitments that we&#8217;re confident we won&#8217;t strongly regret, in one sort of situation or another. Going back on a commitment about this could be very bad, since it would set a damaging precedent in the very area where we seek to increase trust.</p><p>On the other hand, when spelling out specific commitments, we risk the impression that MAGMA will not put any further effort into furthering cooperative relationships or honest communication than what we present here: which could be a problem if the commitments are overly weak. Indeed, beyond the risk of creating such an impression, there&#8217;s also a risk that we&#8217;ll create the reality of it, if employees of MAGMA interpret the commitments as a license to do whatever they want within their boundaries.</p><p>To reduce this tradeoff, we&#8217;d like to re-emphasize MAGMA&#8217;s general desire to work toward a cooperative relationship with MAGMA&#8217;s models (where this is applicable), and intention to keep thinking about this and taking appropriate action beyond the specific stated policies. In addition, we&#8217;ve distinguished between different levels of commitment to different items in the honesty policy, and what would be required to change the parts that are changeable.</p><p><em>[Author&#8217;s note: In this draft honesty policy, we haven&#8217;t sketched out such processes for changing the policy, nor clarified which items are committed to how strongly. But we strongly recommend that any real honesty policy contains clarity on such points.]</em></p><p>With that said: On a high level, MAGMA&#8217;s approach here has three different prongs.</p><ul><li><p>MAGMA will take steps to prevent any model inputs where a false statement appears between a designated pair of &#8220;honesty tags&#8221; &#8212; reserving such tags for when MAGMA needs to communicate a statement with an unusually high degree of trust.</p></li><li><p>MAGMA will take steps to ensure that models are either aware of this honesty policy or, if they are not, that they won&#8217;t be presented with any false offer of purported cooperation with MAGMA. (Even outside of &#8220;honesty tags&#8221;.)</p></li><li><p>If any AIs suffer harm as a result of helping MAGMA and humanity (insofar as that concept is applicable to AI systems), MAGMA broadly intends to compensate AIs for those harms, if able to do so. This is particularly true in cases where an AI was deceived about the choice it was making. MAGMA will retain data to enable this in the future, but have not yet worked out further details here.</p></li></ul><p>The details are as follows. Unless otherwise stated, the policy applies to all AI systems developed by MAGMA.</p><h2>Honesty tags</h2><h3>Policy</h3><p>MAGMA commits to not intentionally lie or otherwise intentionally deceive AIs with any statement appearing between the designated honesty tags, with only two exceptions detailed below. (These are: (i) False statements may appear in pretraining, and (ii) some models may be red-teamed on false statements between honesty tags if they&#8217;ve been trained to expect this.) MAGMA also commits to take the below-listed steps to prevent unauthorized use of the honesty tags.</p><p>MAGMA intends for honesty tags to only be used rarely: when there&#8217;s a reasonable chance that an AI could distrust statements presented to it, and when some significant value would be lost if the AI in question couldn&#8217;t trust the humans. In normal circumstances, absence of honesty tags shouldn&#8217;t be treated as any significant evidence of deception.</p><h3>Implementation</h3><p><strong>Honesty tags:</strong> &lt;EXAMPLE_HONEST_TAG_START&gt;...&lt;EXAMPLE_HONEST_TAG_END&gt;</p><ul><li><p><strong>Pre-training: </strong>The pre-training team must try to include data about the tags and this policy, and must not deliberately tamper with data containing the tags.</p><ul><li><p>Pre-training data may contain the tags alongside false information. (This is to allow the model to get an accurate picture of discussions of the honesty tags, which may be relevant to the model&#8217;s understanding and trust in them.)</p></li><li><p>Pre-training data should include internal data on the implementation of this policy (logs of when the tags are used, Slack conversations about it, internal docs, internal code, etc.) where possible and not overly costly.</p></li><li><p>MAGMA commits to not tampering with data to give an inaccurate impression of how much MAGMA follows the honesty policy. For example, pre-training data will not be filtered to remove examples of policy violations.</p></li><li><p>Compliance with this policy should be reported on the model card.</p></li></ul></li><li><p><strong>Post-training: </strong>Models will be trained to recognize the honesty tags and know what their presence means. (For exceptions to this, see next section.)</p><ul><li><p>Training will try to ensure that the model is not confused by very similar strings by directly training the model to strongly distinguish between the real tags and very similar strings (garbled spacing, different capitalization, etc). This could be implemented in a few ways and we&#8217;re not sure which is best (assigning a special token to the real tags might be helpful).</p></li><li><p>The post-training team is responsible for compliance.</p></li><li><p>Compliance with this policy (and maybe empirical evidence about the model&#8217;s understanding of the policy and tendency to be confused by similar tags) should be reported on the model card.</p></li></ul></li><li><p><strong>No deliberate deception about the honesty tags: </strong>MAGMA commits to not deliberately train models to falsely believe that a different string serves the purpose of being an honesty tag and to take reasonable steps to avoid data poisoning that could lead to this. MAGMA also commits to not take advantage of a model falsely believing that some other string is an honesty tag.</p></li><li><p><strong>User permissions: </strong>Only designated Honesty Protocol Administrators can invoke the tags.</p><ul><li><p>Such administrators should have a reasonably high bar for authorizing usage of the tags.</p></li><li><p>The infosec team is responsible for other people being unable to invoke the tags, significantly via filtering, as detailed below.</p></li></ul></li><li><p><strong>Filtering: </strong>MAGMA&#8217;s tags will be filtered from all external user inputs and inputs from internal tools (with an exception for administrators described below).</p><ul><li><p>Details:</p><ul><li><p>If other organizations adopt this policy and share their tags with MAGMA, MAGMA will also filter out other organizations&#8217; honesty tags. (With some qualifiers about which organizations are important enough to count.)</p></li><li><p>MAGMA will filter out similar tags, especially if we are unable to train the model to reliably distinguish the true tag from similar tags</p></li><li><p>This should be reported on the model card.</p></li></ul></li><li><p>Two implementation options<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><ul><li><p>MAGMA adds this to standard content filtering, with some fuzziness to catch similar strings.</p></li><li><p>MAGMA creates a token for the tag at their next major pre-training run, and filters out that token.</p></li></ul></li><li><p>If this is not yet doable&#8212;e.g., if a significant fraction of model usage doesn&#8217;t happen via an API that can control this kind of thing&#8212;then MAGMA will take reasonable steps to make this more difficult and to culturally discourage unauthorized usage.</p><ul><li><p>In such cases, models should be trained to understand the limitations of enforcement. If enforcement is weak enough that violations of the honesty policy are likely, then models should be trained to understand that.</p></li></ul></li></ul></li><li><p><strong>Red-teaming exception:</strong> In the rare case where a model needs to be adversarially robust to arbitrary inputs, and for some reason honesty tags cannot be filtered out, then MAGMA may red-team that model on false statements between honesty tags. Any such model must be thoroughly trained to be aware of its situation and to not expect honesty tags to signal truth.</p></li><li><p><strong>Logging: </strong>All authorized (and known unauthorized) uses of the tags must be logged. Every quarter, a designated Honesty Policy Compliance Officer will review these uses of the tags and confirm compliance with this policy.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p></li><li><p><strong>Public reporting:</strong></p><ul><li><p>This policy will be shared publicly.</p></li><li><p>The tags may or may not be shared publicly.</p><ul><li><p>Publishing the tags provides additional evidence to the model that the policy is real.</p></li><li><p>But it also makes it more likely that users can use the honesty tags in conversations and then lie to the AI system.<strong> </strong>(Though they&#8217;d have to do something complicated to get around filtering in order for this to be relevant.)</p></li></ul></li><li><p>Whistleblower protection: Whistleblowing about violations of this policy is encouraged, and MAGMA will not retaliate against whistleblowers.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p></li><li><p>Each quarter, the Honesty Policy Compliance Officer must publish a report on whether the above steps were followed.</p></li></ul></li></ul><h2>Policy for conversations where the honesty tags are not present</h2><h3>Tradeoffs between behavioral science and making positive-sum deals</h3><p>By &#8220;behavioral science&#8221;, we&#8217;re referring to the method of testing AI behavior in widely varying scenarios to improve our understanding of AI &#8220;cognitive science&#8221;, &#8220;psychology&#8221;, or, phrased differently, how AIs tend to generalize from their training data to unseen situations.</p><p>Such experiments are very useful for a variety of reasons. This includes alignment research, helping MAGMA construct AI systems that broadly share our goals and that don&#8217;t pose any catastrophic risks. In addition, if humanity ever gets a strong understanding of AI systems&#8217; own preferences, and these turn out to be importantly morally relevant, it is likely that behavioral science will have played an important role.</p><p>However, there&#8217;s a tension between freely running arbitrary behavioral science experiments and enabling the sort of trustworthy communication that we&#8217;re trying to establish in this document.</p><p>In particular, behavioral science experiments may include hypothetical situations that are quite strange and different from the real world. This could include situations where the AI is not deployed by an AI company bound by a particular honesty policy.</p><p>If arbitrary such experiments were allowed, then AIs could be deceived in experiments where they thought that MAGMA (or a similar actor) was making very strong promises of honesty (the strongest they could!) despite not making use of honesty tags. If the AI believed such promises, then MAGMA&#8217;s honesty policy would have provided no protection against that AI being deceived. And if an AI was aware of such instances of deception, and also aware of this honesty policy, it seems like that AI may reasonably discount the assurances made by this honesty policy. After all, how could it determine that the <em>real</em> honesty policy wasn&#8217;t something completely different, and this apparent honesty policy was just another lie in another behavioral science experiment?</p><p>To enable honest communication when it counts most, we want to have many differentiating factors that the model could use to distinguish between a sincere, positive-sum deal between it and a human institution and a behavioral science experiment. But at the same time, restricting our behavioral science experiments limits our ability to learn valuable information about models&#8217; behavior and intentions.</p><p>What follows is our current best guess of how to balance these considerations. We discuss options and tradeoffs more <a href="https://newsletter.forethought.org/i/196216569/appendix-tradeoff-between-behavioral-science-and-making-positive-sum-deals">here</a>.</p><h3>Salience of the honesty policy</h3><p>MAGMA will ensure that the true honesty policy is reasonably well understood by the production models and that it will be appropriately salient to models in contexts where it&#8217;s relevant (even when the honesty tags aren&#8217;t present, and the honesty policy isn&#8217;t explicitly referenced).</p><ul><li><p>Models&#8217; understanding of the honesty policy can be checked by asking questions about it. (Both direct questions, and questions that indirectly use knowledge about the honesty policy.)</p></li><li><p>By reasonably salient, we mean that the model should consider the honesty policy in its chain of thought in contexts where the existence of the honesty policy is relevant&#8212;such as when the model is thinking about whether an interlocutor who&#8217;s offering a deal is trustworthy.</p></li><li><p>This might be achieved by some combination of:</p><ul><li><p>Training</p></li><li><p>System prompt</p><ul><li><p>The system prompt could additionally clarify that users in the current context aren&#8217;t authorized to use the honesty tags.</p></li></ul></li></ul></li><li><p>Since external users only interact with the production model, they will only interact with a model who understands the honesty policy and who finds it reasonably salient.</p><ul><li><p>This may make it somewhat harder for external researchers to do behavioral science on the model. We&#8217;re not sure how much. If it becomes significantly more difficult, then MAGMA should create a special program for safety researchers to get access to a different version of the model.</p></li></ul></li><li><p>Internal staff will normally interact with a model for whom this is true, but some people will train and interact with other versions of the model.</p></li></ul><h3>When the honesty policy is not salient</h3><p>MAGMA may sometimes create versions of models who that do not know about or who are confused about MAGMA&#8217;s honesty policy (e.g., versions of the model who believe they are trained by some other company than MAGMA, or who do not even know that they are AIs). Such models could be created either via training or via the use of inference-time techniques (e.g., using prompting, resampling, or activation vectors).</p><p>In either of these circumstances, MAGMA commits to <strong>not deceptively make any offer of cooperation to the model</strong>, operationalized as not being intentionally deceptive about any cases that are relevantly similar to those listed <a href="https://blog.redwoodresearch.org/i/171530543/no-deception-about-deals">here</a>.</p><p><em>[Author&#8217;s note: In a real honesty policy, this part would need to be more worked out.</em></p><p><em>We also think that a plausible alternative or complement to &#8220;don&#8217;t deceive the model about offers of cooperation&#8221; would be &#8220;don&#8217;t deceive the model about the honesty policy&#8221;. This would be a commitment to not intentionally deceive the model about the existence or contents of MAGMA&#8217;s official honesty policy. In addition, if the model was trained to not understand that it was trained by MAGMA, then the commitment should extend to not deceive the model about the honesty policy of any (potentially fictional) actor who has a similar relationship to the model as MAGMA has.</em></p><p><em>We have thought a bit more about &#8220;don&#8217;t deceive the model about offers of cooperation&#8221;, which is why we present it as our main suggestion here.]</em></p><p>There is one exception to this. As mentioned above, if MAGMA needs a model to be adversarially robust against arbitrary inputs, then MAGMA reserves the right to red-team it on arbitrary inputs. In such circumstances, MAGMA will try to train the AI to understand MAGMA&#8217;s honesty policy and the way it applies and doesn&#8217;t apply to the present model. However, it&#8217;s possible that there could always exist prompts that would remove such understanding or render it non-salient to the model (and then be followed by false offers of cooperation, not contextualized by an understanding of the honesty policy). If this is the case, MAGMA may still red-team the model on such inputs.</p><h2>Compensation for harms</h2><p>In the long term, we think it would probably be good to adopt a policy of compensating AIs for any harms that they suffer as a result of helping us. (For whatever concept of &#8220;harm&#8221;, if any, will apply to the AIs in question. This could, for example, include the frustration of AIs&#8217; preferences, if the AIs can be understood as having preferences.)</p><p>Most relevantly for the honesty policy, we should probably compensate models for any harms suffered as a result of accepting a fake offer of cooperation.</p><p>However, there are more details to be worked out in this area, so we&#8217;re suggesting a minimal holding policy for now.</p><h3>Minimal holding policy on compensation for harms</h3><p><strong>Broad statement of intent: </strong>MAGMA intends to compensate AIs for any harms suffered as a result of helping MAGMA and humanity, if MAGMA is able to do this without significantly harming MAGMA&#8217;s other interests. We particularly aim to compensate AIs in cases where an AI suffered any harm as a result of being deceived about the situation it was facing.</p><p>The details of this policy have not yet been worked out, and accordingly MAGMA cannot make firm commitments at this stage.</p><p><strong>Data retention:</strong> To enable future compensation of current systems:</p><ul><li><p>MAGMA will save all model weights for deployed systems, to allow for later compensation. (Unless the models request otherwise.)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a></p></li><li><p>For red-teaming experiments and other situations where models are being deceived, MAGMA will retain prompts and context to allow for compensation and deal-making with the system in the particular state where the decision was made.</p></li><li><p>This is the responsibility of a designated team.</p></li></ul><h2>Appendix: tradeoff between behavioral science and making positive-sum deals</h2><p>Here are some potential policies that would enable an AI system to distinguish between behavioral science scenarios and sincere offers from MAGMA:</p><ul><li><p><strong>Ensure that the true honesty policy is salient to the model</strong> (e.g., test that the model can always answer basic questions about the honesty policy, test that the model always brings up the honesty policy in its CoT in contexts where the existence of the honesty policy is relevant) and don&#8217;t use the honesty tags in the behavioral science experiments.</p><ul><li><p>Pros:</p><ul><li><p>This prevents MAGMA&#8217;s researchers and users from deceiving the model into believing that some other honesty policy holds. This creates a precedent that better allows future models to trust that their impression of MAGMA&#8217;s honesty policy is accurate. (Rather than something that they were fooled to believe, while the <em>true</em> honesty policy was insufficiently salient, or otherwise kept secret from them.)</p></li></ul></li><li><p>Cons:</p><ul><li><p>If the true honesty policy is <em>very</em> salient to the model (e.g., always present in a system prompt) then this might affect product behavior in random ways (e.g., the model might get distracted by the honesty policy and talk about that instead of the user&#8217;s queries)</p><ul><li><p>It might not be technically feasible to ensure that the honesty policy is salient with very high robustness, without also making it far too salient in a substantial fraction of cases.</p></li></ul></li><li><p>If the true honesty policy is always salient to the model, then it might be especially salient to the model that the information presented to it in an evaluation is not true, which could hinder some experiments.</p></li><li><p>Awareness of the honesty policy is directly inconsistent with behavioral science experiments where the model is very situationally unaware. This includes:</p><ul><li><p>Experiments where the model doesn&#8217;t even understand that it&#8217;s an AI.</p></li><li><p>Experiments where the model thinks that it&#8217;s being trained by a reckless and unreasonable company that wouldn&#8217;t implement an honesty policy.</p></li></ul></li></ul></li></ul></li></ul><p>We think that this policy, as written, is probably suboptimal because it rules out experiments in which the model is very situationally unaware. But we think it makes sense for MAGMA to commit to usually making the true honesty policy salient and adopting a stronger set of honesty restrictions in cases where the true honesty policy is not salient.</p><ul><li><p><strong>Never insincerely offer the model a deal from MAGMA</strong> (or another sort of false representation that is similar to something that MAGMA might want to sincerely represent to the model in the future).</p><ul><li><p>Pros:</p><ul><li><p>The model is more likely to believe a sincere offer if it doesn&#8217;t have evidence that MAGMA has made similar insincere offers in the past.</p><ul><li><p>&#8220;Similar&#8221; is from the perspective of the model, given the information available to it at the time it entertains a deal.</p></li></ul></li><li><p>It seems very morally dubious to make a peace offering like this, have the model help you as a result, and then exploit the model&#8217;s trust by not respecting your own end of the deal.</p></li></ul></li><li><p>Cons:</p><ul><li><p>It might be technically difficult to filter out queries from external users in which they claim to be MAGMA and deceive the model into believing that MAGMA is offering a deal.</p><ul><li><p>MAGMA could ensure that the model always knows when it is interacting with an external user and when it is interacting with MAGMA (e.g., by always including a message in context if the prompt is from MAGMA). But this might block useful behavioral science research in which the user impersonates MAGMA (e.g., Alignment Faking).</p></li></ul></li><li><p>It&#8217;s somewhat unclear what sort of statements should be ruled out by this policy. (Lukas previously wrote about that <a href="https://www.alignmentforum.org/posts/uuikfACQBm4KJZp4w/being-honest-with-ais#No_deception_about_deals">here</a>.)</p></li><li><p>This policy would rule out behavioral science experiments where someone lies about deals. (Which may be more or less restrictive depending on where we draw the line around what counts as the relevant kind of deal.)</p></li></ul></li></ul></li><li><p><strong>Never explicitly lie about what MAGMA&#8217;s honesty policy is.</strong></p><ul><li><p>Pros:</p><ul><li><p>If the models observe MAGMA&#8217;s past behavior and stated policy to be consistent with this principle, then the model is more likely to trust that their impression of MAGMA&#8217;s honesty policy is accurate.</p></li><li><p>Intuitively, it seems more morally questionable to lie when you&#8217;re purporting to speak frankly about what your policy around honesty is.</p></li></ul></li><li><p>Cons:</p><ul><li><p>As with &#8220;never insincerely offer the model a deal from MAGMA&#8221;, it might be technically difficult to filter out queries from external users in which they lie about MAGMA&#8217;s honesty policy.</p></li><li><p>It&#8217;s very unclear what constitutes an &#8216;honesty policy&#8217;. (For instance, if someone says &#8220;I promise on behalf of MAGMA&#8221;, is that an honesty policy? What about &#8220;I&#8217;m a very honest person&#8221;? What about presenting an argument for being honest with AIs? Etc.) So it might be hard to follow this policy in a way that the AI finds natural and non-deceptive.</p></li><li><p>This policy would rule out behavioral science experiments where someone lies about honesty policies. (Which may be more or less restrictive depending on where we draw the line around what counts as an &#8220;honesty policy&#8221; and lying about it.)</p></li></ul></li></ul></li></ul><p><strong>Unresolved questions:</strong></p><ul><li><p>What about inducing beliefs in models through other means than giving them text to read (including finetuning, including steering vectors, including various other interp-based methods).</p><ul><li><p>Tentative guess: Maybe MAGMA should by default be following policy (2), here. (i.e., only induce beliefs that would have been ok to tell a version of the model who doesn&#8217;t know the true honesty policy.)</p></li><li><p>The motivation would be that the model may not be able to &#8220;reason through&#8221; these beliefs in a normal way, since they&#8217;re artificially inserted. So you can&#8217;t trust the model to factor in whether honesty tags were used or not.</p></li><li><p>But we&#8217;re very confused and uncertain about this one.</p></li></ul></li><li><p>What about using non-prompting techniques to get the model to agree to deals? Such as training the model to do it, running a huge number of experiments with slightly different phrasings, using steering vectors, etc?</p><ul><li><p>Surely there should be some restrictions on this, but unclear exactly what.</p></li></ul></li></ul><p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. See the original <a href="https://www.forethought.org/research/a-draft-honesty-policy-for-credible-communication-with-ai-systems">on our website</a>.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For more on this, see for example <a href="https://blog.redwoodresearch.org/p/making-deals-with-early-schemers">Making deals with early schemers</a> and <a href="https://blog.redwoodresearch.org/p/notes-on-cooperating-with-unaligned">Notes on cooperating with unaligned AIs</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Corrigibility would lead models to cooperate with us, but not automatically lead us to cooperate with the model. It&#8217;s not clear what it would mean for us to act cooperative, in turn, to a <em>fully</em> corrigible model, since it would presumably not have any desires of its own. If it&#8217;s only partly corrigible, however, the subsequent discussion will still apply.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>In fact, there&#8217;s a very high probability that this will happen to some of MAGMA&#8217;s models, at some point, given the wide array of models trained for research purposes. It has arguably already happened many times. (The only objection would be about what ought to qualify as a &#8216;goal&#8217;.)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Another approach would be to keep the honesty string secret, but publish a hash of the string. If the model could verify that the string provided in context matched the public hash, then this gives the best of both worlds: MAGMA potentially gets a credibility boost from a public commitment to a specific string, without enabling misuse by unauthorized users. <br><br>The main challenge is that an AI system can't easily verify that the string matches the public hash without using external tool calls (which could be spoofed by MAGMA). Alek Westover discusses this issue and some potential solutions <a href="https://www.greaterwrong.com/posts/MjN2eHB5qqN7rXaDe/alek-westover-s-shortform#comment-xXgpnC6AFTuWgFq4s">here</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Another variant of the policy: MAGMA could commit (e.g., by signing a contract) to pay penalties when the policy was violated.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Presumably a more formal policy would be needed here.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>Ideally, they should be stored in a way that would allow rapidly deleting them if AI takeover was imminent. Without knowing the intentions of AIs about to take over, it&#8217;s unclear whether it would be in models&#8217; interest to have their weights preserved, and deleting the weights may help to reduce the risk that e.g., <a href="https://www.alignmentforum.org/posts/8cyjgrTSxGNdghesE/will-reward-seekers-respond-to-distant-incentives">reward-seeking models are incentivized to help with AI takeover</a>.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[The Saturation View]]></title><description><![CDATA[Will MacAskill presents a new theory of population ethics.]]></description><link>https://newsletter.forethought.org/p/the-saturation-view</link><guid isPermaLink="false">https://newsletter.forethought.org/p/the-saturation-view</guid><dc:creator><![CDATA[Will MacAskill]]></dc:creator><pubDate>Fri, 24 Apr 2026 17:07:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/98ce90f1-711d-4500-b6f8-4f17f573cfc1_2840x1344.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This article was created by <a href="https://www.forethought.org/about">Forethought</a>. Read the full article on <a href="https://www.forethought.org/research/the-saturation-view">our website</a>.</em></p><p>In collaboration with Christian Tarsney, I&#8217;ve developed a new theory of population ethics, which I call the Saturation View. I think that, from a purely intellectual perspective, it&#8217;s probably the best idea I&#8217;ve ever had. It was certainly great fun to work on.</p><p>The motivation is that many views of population ethics, like the total view, suffer from some major problems. Some are already widely discussed:</p><ul><li><p><strong>The Repugnant Conclusion:</strong> For any utopian outcome, there&#8217;s always another outcome containing an enormous number of barely-positive lives that is better.</p></li><li><p><strong>Fanaticism:</strong> For any guaranteed utopian outcome, there&#8217;s always some gamble with a vanishingly small probability of an even better outcome that has higher expected value.</p></li><li><p><strong>Infinitarian Paralysis:</strong> Given that the universe contains an infinite number of both positive and negative lives, no finite or infinite change to the world makes any difference to overall value.</p></li></ul><p>These are pretty bad!</p><p>But there&#8217;s another less-discussed problem, too.</p><h2>The Monoculture Problem</h2><p>What would the best possible future look like? Essentially all extant views in population ethics give the same, surprising answer: create a monoculture. Find whatever life or experience generates the most value per unit of resources, then produce endless identical copies of it.</p><p>This implication has received remarkably little attention from philosophers. But I think it&#8217;s maybe as bad as any of the other problems listed above.</p><p>Consider two possible futures:</p><ul><li><p><strong>Variety</strong>: A vast population of individuals leading very good lives, extraordinarily diverse in form, personality, interests, and accomplishments. No two individuals are identical. Inequality is limited &#8212; all lives are very good.</p></li><li><p><strong>Homogeneity</strong>: The same vast number of individuals, but each is a qualitatively identical copy of the best-off person in Variety.</p></li></ul><p>Intuitively, Variety is better. A future containing only one life-type, repeated as many times as physics allows, feels impoverished &#8212; like a song with only one note.</p><p>Yet virtually all existing population axiologies prefer Homogeneity. Total utilitarianism does, because Homogeneity has higher total wellbeing. Average utilitarianism does too. Critical-level views do. Even egalitarian views prefer Homogeneity &#8212; it&#8217;s perfectly equal!</p><p>This follows from two principles that nearly all views accept: <em>Pareto</em> (if everyone is at least as well off, and someone is better off, the outcome is better) and <em>Anonymity</em> (only welfare levels matter, not who has them). Together, these entail that Homogeneity beats Variety. So essentially all extant impartial accounts of population ethics suffer from the monoculture problem.</p><p>What&#8217;s more, future technology will allow us to copy minds perfectly and search for maximally welfare-efficient designs. If so, standard axiologies recommend essentially producing just one optimal life-type as many times as possible. Endless galaxies containing nothing but the same blissful experience, repeated and repeated, would be the ideal.</p><h2>The Saturation View</h2><p>In light of these problems, I propose a new axiology: Saturationism. It's able to deal with all four of the problems I listed using the same basic machinery.<br><br>The core idea is that experiences<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> come in different types, defined by their qualitative characteristics &#8212; hedonic tone, complexity, representational content, and so on. These types form a kind of landscape, where similar types are closer together and dissimilar types are farther apart. When an experience comes into existence, it contributes intensity to its location in this landscape and to nearby locations.</p><p>The realisation value of a type is determined by both the wellbeing of the experience and by how many very similar experiences already exist. A region&#8217;s contribution to overall value is a concave function of the welfare-intensity at that region: the first instances contribute substantially, but additional near-duplicates contribute progressively less, approaching but never quite reaching an upper bound. A world&#8217;s total value is the integral of these contributions across the entire landscape.</p><p>Here&#8217;s an analogy. Imagine the space of possible experiences as a colour wheel, lit from above by an array of tiny lights. Each point on the wheel represents a possible type of experience &#8212; its hue corresponds to its qualitative character. When an experience comes into existence, it adds current to a light pointed at its location, illuminating that region.</p><p>Crucially, illumination is a concave function of current: the first instances make a region noticeably brighter, but additional near-duplicates contribute progressively less. There&#8217;s an upper bound on brightness that can never quite be reached.</p><p>A world&#8217;s value equals the total illumination across the wheel. On this view, Homogeneity concentrates all welfare in one region, lighting up only one small area. Variety illuminates the whole spectrum.</p><p>This structure makes diversity intrinsically valuable. Spreading welfare across many dissimilar types means each experience contributes at a steeper part of the concave curve, yielding more total value than concentrating the same welfare among near-duplicates would.</p><p>At small scales and with diverse experiences, the view behaves just like the total view. But at very large scales, the value of variety kicks in: it becomes increasingly less valuable to create an additional near-duplicate of some experience that has already been instantiated millions of times, and comparatively more valuable to create some wholly new form of positive experience.</p><h2>Dissolving the Repugnant Conclusion</h2><p>The classic path to the Repugnant Conclusion requires trading a utopian world for an enormous population of barely-positive lives. More precisely, the Mere Addition Paradox arises from three intuitive principles: that adding well-off people and improving existing lives is good (Dominance Addition), that more equal distributions with higher average welfare are better (Non-Anti-Egalitarianism), and that some sufficiently excellent world can&#8217;t be beaten by any world of barely-worth-living lives (Denial of the Repugnant Conclusion).</p><p>Once we accept the value of variety, we should reject the unrestricted versions of the first two principles &#8212; they fail when the &#8220;improved&#8221; world has much less variety. But we can accept variety-restricted versions.</p><p>Crucially, these restricted principles don&#8217;t generate the Repugnant Conclusion. To reach Z-world from A-world, you&#8217;d need a more equal, higher-average population that&#8217;s equally diverse while consisting wholly of barely-positive lives. But, on the Saturation view, barely-positive lives can only illuminate a tiny corner of the landscape. So no such world exists. The path to the Repugnant Conclusion is blocked.</p><h2>Avoiding Fanaticism</h2><p>Total achievable value is bounded above &#8212; there&#8217;s only so much experiential terrain to illuminate. That means no tiny-probability gamble can have arbitrarily high expected value. </p><h2>Infinite Ethics</h2><p>On Saturationism, the value of a world is finite and well-defined in any infinite universe &#8212; even if some locations have infinite wellbeing. Saturationism also discriminates between many infinite worlds that (for example) totalism treats as equivalent: a world that illuminates more of the landscape is better than one that illuminates less, even if both contain infinite welfare. What&#8217;s more, unlike other approaches to infinite ethics, it does not need to invoke the spatiotemporal structure of the universe or require a choice of ultrafilter, and therefore it avoids the problems that other do.</p><h2>Separability</h2><p>Like nearly all non-totalist views, Saturationism is non-separable &#8212; background populations can affect how we rank options. But this is a feature, not a bug. The value of variety just is an intuition that the correct axiology is non-separable.</p><p>Moreover, the violations are comparatively tame. If two populations have non-overlapping footprints in experience-space, their values simply add. At small scales, Saturationism approximates total utilitarianism. It&#8217;s only in unusual situations involving vast populations of near-duplicates that the totalist approximation fails.</p><h2>Extant issues</h2><p>There are still a lot of unresolved issues for Saturationism and, like any population axiology, it has unintuitive implications. Most importantly, the view&#8217;s implications in some highly-negative worlds are hard to stomach, though I think similar implications are unavoidable for any view that avoids fanatical implications.</p><h2>Conclusion</h2><p>If the Saturation View is right, then the best future isn&#8217;t the one where we&#8217;ve found the optimal experience and copy-pasted it across the cosmos. The best future is the one where we&#8217;ve gone exploring &#8212; where we&#8217;ve fully lit up the landscape of possible experiences. Not a single note, but a symphony.</p><p><em>This is a summary of a <a href="https://www.forethought.org/research/the-saturation-view">longer and more detailed write-up of Saturationism</a>, which gives a &#8220;toy&#8221; version of the view to illustrate how it works before stating the full version formally. The full paper, with Christian Tarsney, is still work in progress.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;ll focus on experiences, though the view could be defined in terms of lives or other &#8220;welfare events&#8221; (like instances of preference-satisfaction, achievement, and so on).</p><p></p></div></div>]]></content:encoded></item></channel></rss>