<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Iain Harper&#39;s Blog</title>
    <link>https://iain.so/</link>
    <description>Weblogging like it&#39;s 1995!</description>
    <pubDate>Tue, 18 Aug 2026 23:08:50 +0000</pubDate>
    <image>
      <url>https://i.snap.as/zscJpaeW.png</url>
      <title>Iain Harper&#39;s Blog</title>
      <link>https://iain.so/</link>
    </image>
    <item>
      <title>Boy on the corner: how Risky Roadz gave grime its face</title>
      <link>https://iain.so/boy-from-the-corner-how-risky-roadz-gave-grime-its-face?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[In the early 2000s, grime MCs poured out of pirate radio sets across East London, their voices bleeding from tower block frequencies, recognisable by flow and cadence alone. You memorised their bars, learned their rhythms, debated their rankings on MSN and in school corridors. But you had no idea what they looked like. In a genre built on persona, territory, and the precise postcode you claimed, this was a strange gap. The music existed in sound form only, and the exuberant personas behind the lyrics were mysteries to anyone outside their own  immediate endz. Then a young man from Bow borrowed money from his nan to buy a camera, and the genre gained a face.&#xA;&#xA;MCs freestyling&#xA;&#xA;The record shop village square&#xA;&#xA;To understand Risky Roadz you need to understand Rhythm Division, the bright blue record shop at 391 Roman Road in Bow that functioned, during the mid-2000s, as grime&#39;s unofficial parliament. Its walls were plastered with stickers, its shelves stacked with garage and grime white labels, a PlayStation sat in one corner, turntables in another, and music played without pause from open to close. Wiley came through. Skepta came through. Dizzee Rascal came through. D Double E, Kano, JME. Everyone came through.&#xA;&#xA;Lady Shocker, an original member of Mucky Wolfpack and a leading voice in the Girls of Grime movement, remembers what made it different from any other shop. &#34;Even if they didn&#39;t know you, you could take your stuff there and they would sell it. Your mix, your vinyls, whatever. They were really supporting the underground scene.&#34; &#xA;&#xA;It was basically a community hub dressed as a retail unit, and for the young people orbiting grime in East London it was the default place to go when there was nothing else to do. Walk down to Rhythm Division and you might catch MCs rolling in to do sets on the shop floor.&#xA;&#xA;Roony Keefe had bought records there while at school. Then, during college he got a Saturday job, working alongside the shop manager, Sparky. Surrounded by the scene&#39;s principal figures, he had a simple revelation. &#34;There was pirate radio, so you heard voices but didn&#39;t know what everyone looked like. Rhythm Division was like the window to seeing the people. That&#39;s what gave me the idea initially to start filming.&#34; In his mind, if he wanted to know what the MCs looked like, everyone must want to know .&#xA;&#xA;Art from limitation&#xA;&#xA;The technology of 2003 now seems almost comically primitive. No smartphones with HD cameras. No YouTube or Instagram. No TikTok, no viral distribution. The world was predominantly analogue and connections were physical. You had to meet someone to give them a CD, a DVD, a white label.&#xA;&#xA;Keefe called his grandmother and asked to borrow money for a camera. She sent it. He bought books to teach himself video editing because there were no online tutorials to fall back on. The equipment was basic. The transitions came from Windows Movie Maker. The shots were shaky, the audio rough, and in 2004 distribution meant pressing films onto physical discs and learning the process as you went.&#xA;&#xA;None of that mattered though. The rough quality gave authenticity in a way that polished production never could. Cracked copies of Fruity Loops produced the beats. PlayStation&#39;s Music 2000 served as a production suite. MSN and Limewire spread the tracks through digital backchannels. Pirate radio carried the signal through the air. And now DVDs captured the visual element that had been missing. The aesthetic was low-fi because the world was low-fi; a community built with whatever it could get its hands on.&#xA;&#xA;Faces to frequencies&#xA;&#xA;https://www.youtube.com/watch?v=KHp1jWC7nvk&#xA;&#xA;What Keefe captured on those early Risky Roadz DVDs was something nobody had properly documented. The antithesis of polished music videos: MCs in their natural habitat, freestyling outside their houses, in car parks, on basketball courts, in the back room of Rhythm Division itself. A young Kano in an Adidas bathrobe freestyling late at night, cup of tea in one hand. Ghetts delivering bars on a grey Plaistow afternoon in front of a terraced house. Tower blocks, estate stairwells, Roman Road on a Tuesday.&#xA;&#xA;The roster reads like a time-capsule of everyone who would define British music over the next two decades. Skepta, Wiley, Kano, D Double E, Ghetts, JME, Dizzee Rascal, Lethal Bizzle, Giggs, Bashy, Tempa T. Crews from Roll Deep to Ruff Sqwad, Boy Better Know to Nasty Crew, Newham Generals to Meridian Crew. For many, Risky Roadz was the first time audiences outside the immediate scene ever saw their faces. Tempa T&#39;s first visual appearance anywhere was Risky Roadz 2.&#xA;&#xA;Keefe operated as a kind of quality control. &#34;Anyone could come in and freestyle, and if my ear said you were alright and Sparky did too, then you made it onto the DVD. Then all these households started to know who you were.&#34; And because he was embedded in the scene rather than observing it from outside, his subjects gave him something others could never extract. &#34;Everyone we approached about it we knew already. That&#39;s why I get good interviews out of people. I&#39;m not just a journalist speaking to them, I&#39;m also a friend and they can be honest.&#34;&#xA;&#xA;The freestyles that made legends&#xA;&#xA;Some moments survived the format and became permanent reference points.&#xA;&#xA;On a grey afternoon in Plaistow, Ghetts (then Ghetto) stood in front of a terraced house and delivered what remains one of grime&#39;s most quoted freestyles&#xA;&#xA;https://www.youtube.com/watch?v=pEGeuq2NreA&#xA;&#xA;&#34;At that time I felt like I was the one. The guy, the person. I felt like no one was doing what I was doing in the way I was constructing my music,&#34; he said later. Keefe knew it in the moment. &#34;When we filmed Ghetts&#39; freestyle, that was my real moment of thinking, &#39;Yeah, this is special.&#39; It&#39;s probably one of the most iconic freestyles in grime.&#34;&#xA;&#xA;The late Stormin&#39;s acapella Trim diss in Volume 1 survives as one of the genre&#39;s best and funniest takedowns&#xA;&#xA;https://youtu.be/Zb-zsG1m7pA?si=dUyEgRDZegXW9iq&#xA;&#xA;And the Roman Road cypher, from 2006&#39;s Movement Documentary, where Skepta, Wiley, Frisco, Wretch 32, and Ghetts went bar-to-bar on the street, still circulates as definitive evidence of grime&#39;s collective peak. Five MCs trading sixteens outdoors, no stage, no lighting rig, just a cheap camera rolling on a pavement in E3.&#xA;&#xA;https://www.youtube.com/watch?v=BZ7WyR6ccQA&#xA;&#xA;The DVD ecosystem&#xA;&#xA;Risky Roadz operated within a broader ecology of grime DVDs that gave the scene its visual identity.&#xA;&#xA;Jammer&#39;s Lord of the Mics, launched from his parents&#39; basement in Leytonstone in 2004, established the formal clash format. Two MCs, recorded battle, one-on-one over grime instrumentals. The Wiley vs Kano clash that opened the series became the most historic in grime, the Godfather versus the teenage upstart, and Skepta vs Devilman on LOTM2 later became Drake&#39;s favourite clash. &#34;At that time, you couldn&#39;t ever turn down a clash,&#34; Jammer recalled. &#34;That was one thing about being a grime MC.&#34;&#xA;&#xA;https://www.youtube.com/watch?v=-bHRcjqnkAI&#xA;&#xA;Troy &#34;A Plus&#34; Miller&#39;s Practice Hours took a different approach, more documentary-style, capturing the culture as a whole. As one of the founding DJs on Rinse FM&#39;s maiden broadcast alongside Geeneus, Wiley, and Slimzee, Miller had unrivalled access. Regional DVDs like Birmingham Mic Controllers and Midlands Roadside showed how far the movement ran outside London.&#xA;&#xA;Punk&#39;s unlikely descendant&#xA;&#xA;Every punk fan remembers the January 1977 issue of Sideburns fanzine, the one that printed three guitar chords and told its readers, &#34;Now form a band.&#34; Fast-forward to East London at the turn of the millennium and an unlikely descendant was being forged in a stew of poverty, disaffection, and a similarly belligerent self-determination.&#xA;&#xA;Mykaell Riley, director of the Black Music Research Unit at the University of Westminster, called grime &#34;the most disruptive cultural transformation of the British music industry since punk.&#34; The comparison might seem lazy but it goes deeper. Both genres were birthed from frustration and attacked by the political and media establishment. Punk was a threat to moral order. Grime was &#34;violent&#34; and &#34;riot music.&#34; But where punk channelled the rage of young white Britons against economic stagnation, grime gave voice to the experience of young Black Britons. The anger at systemic racism, at state violence, at being young and ignored and abandoned in the shadow of Canary Wharf&#39;s gleaming towers.&#xA;&#xA;UK garage had become popular and commercial, moving away from the concerns of the young people who had created it as it became part of the dominant culture, lost &#xA;to the power of industry gatekeepers. Grime was a deliberate step back underground. The hyper-speed multisyllabic phrasing owed more to jungle and garage raves, where MCs treated the voice as another percussive instrument over 140 BPM rhythms. The structure was different too. Short 16-bar verses, eight-bar choruses. As Chuck D had described hip-hop in America, grime was now CNN for its community. When Dizzee Rascal released Boy in da Corner, parts of which were created during school music classes, grime hit the big time. &#xA;&#xA;With it came the attention of the Metropolitan Police who introduced Form 696, a risk assessment that in practice functioned as a mechanism to shut down grime events by requiring promoters to detail the type of music and the expected ethnic makeup of attendees. Lethal Bizzle&#39;s &#34;Pow! (Forward)&#34; was banned from clubs after allegedly triggering fights.&#xA;&#xA;The tabloids attacked grime as promoting violence, and when the 2011 London riots hit, the genre was blamed for creating a culture of anger, as if the government’s austerity programme that disproportionately defunded and closed many community programs could not possibly have played a contributing role.&#xA;&#xA;The blueprint still echoes&#xA;&#xA;When Jamal Edwards founded SBTV, building what would become a media empire from his bedroom, he was explicitly following in Keefe&#39;s footsteps. &#34;I was in Year 11 doing Information and Communication Technologies and watching the same freestyles, taken from the DVDs Practice Hours and Risky Roadz on YouTube, over and over again,&#34; he told Red Bull Music Academy. Platforms like GRM Daily, SBTV, and Mixtape Madness picked up the format and ran with it. The &#34;Next Up&#34; freestyle series that launched drill artists like Digga D into the mainstream was the same format Keefe had established two decades earlier.&#xA;&#xA;Watch any UK drill video now and you’ll likely see a familiar setup; rapper surrounded by friends, spitting bars directly into a camera on their estate. The visual language that Risky Roadz wrote. As Aniefiok Ekpoudom (Neef) puts it, &#34;A lot of the biggest UK drill artists got big from freestyles. Those freestyles are essentially the same format. It&#39;s one of the distinctions that you can have between British rap and US or French rap, how vital those kerbside freestyles are to the scene, but also how they can change a young person&#39;s life in an instant.&#34;&#xA;&#xA;The corner that never left&#xA;&#xA;When Skepta and JME released &#34;That&#39;s Not Me&#34; in 2014, the video was a conscious rejection of music industry spectacle, shot for £80 with VHS-style footage of Meridian Estate green-screened behind them. Complex ranked it the number one grime song of the 2010s. &#xA;&#xA;When Skepta shot &#34;It Ain&#39;t Safe” back on that same estate shortly after, he specifically requested that Keefe film it on the original Risky Roadz camera, equipment that hadn&#39;t been picked up in a decade. &#34;It does feel like the old days again, everyone&#39;s got their hunger back,&#34; Keefe said. &#34;Skepta wanted it filmed on the old Risky Roadz camera. I haven&#39;t picked that camera up for 10 years now.&#34; It was a deliberate homecoming, an acknowledgment that the visual identity of grime had been established in those early grainy tapes.&#xA;&#xA;https://www.youtube.com/watch?v=czLQoG01PFs&#xA;&#xA;Looking back, Keefe wishes he had filmed more. &#34;I wish I&#39;d captured more pirate radio. I used to go a lot and not always with my camera. And another thing I would have filmed more of would have been B-roll. More B-roll of the estates, the buildings, Rhythm Division, as so much has now been lost to gentrification.&#34;&#xA;&#xA;Rhythm Division closed in 2010. The building is a coffee shop now. The estates have changed. But the documentary record remains, uploaded to YouTube years after the physical DVDs circulated hand to hand, still drawing comments from people discovering grime&#39;s origins for the first time. Drake became a fan. Kanye brought grime MCs on stage at the Brit Awards.&#xA;&#xA;Keefe never set out to become grime&#39;s documentarian. He was a fan who wanted to be involved and embedded himself in the scene by capturing it. &#34;That&#39;s the thing about grime,&#34; he says. &#34;We&#39;re all one big, dysfunctional family. We&#39;ve all grown up together, 15-year friendships. It&#39;s a mad thing. And it makes you feel proud that you&#39;ve had an influence in that.&#34;&#xA;&#xA;Risky Roadz: Documenting the scene&#39;s rise and reign by Roony &#39;Risky Roadz&#39; Keefe is published by Pavilion Books.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>In the early 2000s, grime MCs poured out of pirate radio sets across East London, their voices bleeding from tower block frequencies, recognisable by flow and cadence alone. You memorised their bars, learned their rhythms, debated their rankings on MSN and in school corridors. But you had no idea what they looked like. In a genre built on persona, territory, and the precise postcode you claimed, this was a strange gap. The music existed in sound form only, and the exuberant personas behind the lyrics were mysteries to anyone outside their own  immediate endz. Then a young man from Bow borrowed money from his nan to buy a camera, and the genre gained a face.</p>

<p><img src="https://i.snap.as/PmzmhdgA.png" alt="MCs freestyling"/></p>

<h2 id="the-record-shop-village-square">The record shop village square</h2>

<p>To understand Risky Roadz you need to understand Rhythm Division, the bright blue record shop at 391 Roman Road in Bow that functioned, during the mid-2000s, as grime&#39;s unofficial parliament. Its walls were plastered with stickers, its shelves stacked with garage and grime white labels, a PlayStation sat in one corner, turntables in another, and music played without pause from open to close. Wiley came through. Skepta came through. Dizzee Rascal came through. D Double E, Kano, JME. Everyone came through.</p>

<p>Lady Shocker, an original member of Mucky Wolfpack and a leading voice in the Girls of Grime movement, remembers what made it different from any other shop. “Even if they didn&#39;t know you, you could take your stuff there and they would sell it. Your mix, your vinyls, whatever. They were really supporting the underground scene.”</p>

<p>It was basically a community hub dressed as a retail unit, and for the young people orbiting grime in East London it was the default place to go when there was nothing else to do. Walk down to Rhythm Division and you might catch MCs rolling in to do sets on the shop floor.</p>

<p>Roony Keefe had bought records there while at school. Then, during college he got a Saturday job, working alongside the shop manager, Sparky. Surrounded by the scene&#39;s principal figures, he had a simple revelation. “There was pirate radio, so you heard voices but didn&#39;t know what everyone looked like. Rhythm Division was like the window to seeing the people. That&#39;s what gave me the idea initially to start filming.” In his mind, if he wanted to know what the MCs looked like, everyone must want to know .</p>

<h2 id="art-from-limitation">Art from limitation</h2>

<p>The technology of 2003 now seems almost comically primitive. No smartphones with HD cameras. No YouTube or Instagram. No TikTok, no viral distribution. The world was predominantly analogue and connections were physical. You had to meet someone to give them a CD, a DVD, a white label.</p>

<p>Keefe called his grandmother and asked to borrow money for a camera. She sent it. He bought books to teach himself video editing because there were no online tutorials to fall back on. The equipment was basic. The transitions came from Windows Movie Maker. The shots were shaky, the audio rough, and in 2004 distribution meant pressing films onto physical discs and learning the process as you went.</p>

<p>None of that mattered though. The rough quality gave authenticity in a way that polished production never could. Cracked copies of Fruity Loops produced the beats. PlayStation&#39;s Music 2000 served as a production suite. MSN and Limewire spread the tracks through digital backchannels. Pirate radio carried the signal through the air. And now DVDs captured the visual element that had been missing. The aesthetic was low-fi because the world was low-fi; a community built with whatever it could get its hands on.</p>

<h2 id="faces-to-frequencies">Faces to frequencies</h2>

<p><iframe allow="monetization" width="640" height="480" src="https://www.youtube.com/embed/KHp1jWC7nvk?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen title="Risky Roadz Vol.1 - The Freestyles | Wiley, Ghetto, Kano, Roachee, Trim, Esco &amp; more"></iframe></p>

<p>What Keefe captured on those early Risky Roadz DVDs was something nobody had properly documented. The antithesis of polished music videos: MCs in their natural habitat, freestyling outside their houses, in car parks, on basketball courts, in the back room of Rhythm Division itself. A young Kano in an Adidas bathrobe freestyling late at night, cup of tea in one hand. Ghetts delivering bars on a grey Plaistow afternoon in front of a terraced house. Tower blocks, estate stairwells, Roman Road on a Tuesday.</p>

<p>The roster reads like a time-capsule of everyone who would define British music over the next two decades. Skepta, Wiley, Kano, D Double E, Ghetts, JME, Dizzee Rascal, Lethal Bizzle, Giggs, Bashy, Tempa T. Crews from Roll Deep to Ruff Sqwad, Boy Better Know to Nasty Crew, Newham Generals to Meridian Crew. For many, Risky Roadz was the first time audiences outside the immediate scene ever saw their faces. Tempa T&#39;s first visual appearance anywhere was Risky Roadz 2.</p>

<p>Keefe operated as a kind of quality control. “Anyone could come in and freestyle, and if my ear said you were alright and Sparky did too, then you made it onto the DVD. Then all these households started to know who you were.” And because he was embedded in the scene rather than observing it from outside, his subjects gave him something others could never extract. “Everyone we approached about it we knew already. That&#39;s why I get good interviews out of people. I&#39;m not just a journalist speaking to them, I&#39;m also a friend and they can be honest.”</p>

<h2 id="the-freestyles-that-made-legends">The freestyles that made legends</h2>

<p>Some moments survived the format and became permanent reference points.</p>

<p>On a grey afternoon in Plaistow, Ghetts (then Ghetto) stood in front of a terraced house and delivered what remains one of grime&#39;s most quoted freestyles</p>

<p><iframe allow="monetization" width="640" height="480" src="https://www.youtube.com/embed/pEGeuq2NreA?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen title="Risky Roads Vol 2 - Ghetto &amp; Big Seac"></iframe></p>

<p>“At that time I felt like I was the one. The guy, the person. I felt like no one was doing what I was doing in the way I was constructing my music,” he said later. Keefe knew it in the moment. “When we filmed Ghetts&#39; freestyle, that was my real moment of thinking, &#39;Yeah, this is special.&#39; It&#39;s probably one of the most iconic freestyles in grime.”</p>

<p>The late Stormin&#39;s acapella Trim diss in Volume 1 survives as one of the genre&#39;s best and funniest takedowns</p>

<p><iframe allow="monetization" width="640" height="480" src="https://www.youtube.com/embed/Zb-zsG1m7pA?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen title="STORMIN - TRIM (ROLLDEEP) DISS"></iframe></p>

<p>And the Roman Road cypher, from 2006&#39;s Movement Documentary, where Skepta, Wiley, Frisco, Wretch 32, and Ghetts went bar-to-bar on the street, still circulates as definitive evidence of grime&#39;s collective peak. Five MCs trading sixteens outdoors, no stage, no lighting rig, just a cheap camera rolling on a pavement in E3.</p>

<p><iframe allow="monetization" width="640" height="360" src="https://www.youtube.com/embed/BZ7WyR6ccQA?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen title="Wiley, Skepta, Ghetts, Frisco, Wretch 32 freestyle   The Best of Risky Roadz"></iframe></p>

<h2 id="the-dvd-ecosystem">The DVD ecosystem</h2>

<p>Risky Roadz operated within a broader ecology of grime DVDs that gave the scene its visual identity.</p>

<p>Jammer&#39;s <a href="https://www.youtube.com/watch?v=nTF_T47CDnI">Lord of the Mics</a>, launched from his parents&#39; basement in Leytonstone in 2004, established the formal clash format. Two MCs, recorded battle, one-on-one over grime instrumentals. The <a href="https://www.youtube.com/watch?v=nTF_T47CDnI">Wiley vs Kano clash</a> that opened the series became the most historic in grime, the Godfather versus the teenage upstart, and Skepta vs Devilman on LOTM2 later became Drake&#39;s favourite clash. “At that time, you couldn&#39;t ever turn down a clash,” Jammer recalled. “That was one thing about being a grime MC.”</p>

<p><iframe allow="monetization" width="640" height="360" src="https://www.youtube.com/embed/-bHRcjqnkAI?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen title="Skepta vs. Devilman - Lord of the Mics 2"></iframe></p>

<p>Troy “A Plus” Miller&#39;s Practice Hours took a different approach, more documentary-style, capturing the culture as a whole. As one of the founding DJs on Rinse FM&#39;s maiden broadcast alongside Geeneus, Wiley, and Slimzee, Miller had unrivalled access. Regional DVDs like Birmingham Mic Controllers and Midlands Roadside showed how far the movement ran outside London.</p>

<h2 id="punk-s-unlikely-descendant">Punk&#39;s unlikely descendant</h2>

<p>Every punk fan remembers the January 1977 issue of Sideburns fanzine, the one that printed three guitar chords and told its readers, “Now form a band.” Fast-forward to East London at the turn of the millennium and an unlikely descendant was being forged in a stew of poverty, disaffection, and a similarly belligerent self-determination.</p>

<p>Mykaell Riley, director of the Black Music Research Unit at the University of Westminster, called grime “the most disruptive cultural transformation of the British music industry since punk.” The comparison might seem lazy but it goes deeper. Both genres were birthed from frustration and attacked by the political and media establishment. Punk was a threat to moral order. Grime was “violent” and “riot music.” But where punk channelled the rage of young white Britons against economic stagnation, grime gave voice to the experience of young Black Britons. The anger at systemic racism, at state violence, at being young and ignored and abandoned in the shadow of Canary Wharf&#39;s gleaming towers.</p>

<p>UK garage had become popular and commercial, moving away from the concerns of the young people who had created it as it became part of the dominant culture, lost
to the power of industry gatekeepers. Grime was a deliberate step back underground. The hyper-speed multisyllabic phrasing owed more to jungle and garage raves, where MCs treated the voice as another percussive instrument over 140 BPM rhythms. The structure was different too. Short 16-bar verses, eight-bar choruses. As Chuck D had described hip-hop in America, grime was now CNN for its community. When Dizzee Rascal released Boy in da Corner, parts of which were created during school music classes, grime hit the big time.</p>

<p>With it came the attention of the Metropolitan Police who introduced Form 696, a risk assessment that in practice functioned as a mechanism to shut down grime events by requiring promoters to detail the type of music and the expected ethnic makeup of attendees. Lethal Bizzle&#39;s “Pow! (Forward)” was banned from clubs after allegedly triggering fights.</p>

<p>The tabloids attacked grime as promoting violence, and when the 2011 London riots hit, the genre was blamed for creating a culture of anger, as if the government’s austerity programme that disproportionately defunded and closed many community programs could not possibly have played a contributing role.</p>

<h2 id="the-blueprint-still-echoes">The blueprint still echoes</h2>

<p>When Jamal Edwards founded SBTV, building what would become a media empire from his bedroom, he was explicitly following in Keefe&#39;s footsteps. “I was in Year 11 doing Information and Communication Technologies and watching the same freestyles, taken from the DVDs Practice Hours and Risky Roadz on YouTube, over and over again,” he told Red Bull Music Academy. Platforms like GRM Daily, SBTV, and Mixtape Madness picked up the format and ran with it. The “Next Up” freestyle series that launched drill artists like Digga D into the mainstream was the same format Keefe had established two decades earlier.</p>

<p>Watch any UK drill video now and you’ll likely see a familiar setup; rapper surrounded by friends, spitting bars directly into a camera on their estate. The visual language that Risky Roadz wrote. As Aniefiok Ekpoudom (Neef) puts it, “A lot of the biggest UK drill artists got big from freestyles. Those freestyles are essentially the same format. It&#39;s one of the distinctions that you can have between British rap and US or French rap, how vital those kerbside freestyles are to the scene, but also how they can change a young person&#39;s life in an instant.”</p>

<h2 id="the-corner-that-never-left">The corner that never left</h2>

<p>When Skepta and JME released <a href="https://www.youtube.com/watch?v=_xQKWnvtg6c">“That&#39;s Not Me”</a> in 2014, the video was a conscious rejection of music industry spectacle, shot for £80 with VHS-style footage of Meridian Estate green-screened behind them. Complex ranked it the number one grime song of the 2010s.</p>

<p>When Skepta shot “It Ain&#39;t Safe” back on that same estate shortly after, he specifically requested that Keefe film it on the original Risky Roadz camera, equipment that hadn&#39;t been picked up in a decade. “It does feel like the old days again, everyone&#39;s got their hunger back,” Keefe said. “Skepta wanted it filmed on the old Risky Roadz camera. I haven&#39;t picked that camera up for 10 years now.” It was a deliberate homecoming, an acknowledgment that the visual identity of grime had been established in those early grainy tapes.</p>

<p><iframe allow="monetization" width="640" height="360" src="https://www.youtube.com/embed/czLQoG01PFs?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen title="Skepta feat. Young Lord - &quot;It Ain&#39;t Safe&quot; (Official Video)"></iframe></p>

<p>Looking back, Keefe wishes he had filmed more. “I wish I&#39;d captured more pirate radio. I used to go a lot and not always with my camera. And another thing I would have filmed more of would have been B-roll. More B-roll of the estates, the buildings, Rhythm Division, as so much has now been lost to gentrification.”</p>

<p>Rhythm Division closed in 2010. The building is a coffee shop now. The estates have changed. But the <a href="https://www.youtube.com/watch?v=ra1HWDkv99Q">documentary record remains</a>, uploaded to YouTube years after the physical DVDs circulated hand to hand, still drawing comments from people discovering grime&#39;s origins for the first time. Drake became a fan. Kanye brought grime MCs on stage at the Brit Awards.</p>

<p>Keefe never set out to become grime&#39;s documentarian. He was a fan who wanted to be involved and embedded himself in the scene by capturing it. “That&#39;s the thing about grime,” he says. “We&#39;re all one big, dysfunctional family. We&#39;ve all grown up together, 15-year friendships. It&#39;s a mad thing. And it makes you feel proud that you&#39;ve had an influence in that.”</p>

<p><em><a href="https://www.pavilionbooks.com/products/grime-documenting-the-scenes-rise-and-reign-roony-keefe-9780008626235/">Risky Roadz: Documenting the scene&#39;s rise and reign</a> by Roony &#39;Risky Roadz&#39; Keefe is published by Pavilion Books.</em></p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/boy-from-the-corner-how-risky-roadz-gave-grime-its-face</guid>
      <pubDate>Sat, 15 Aug 2026 06:43:56 +0000</pubDate>
    </item>
    <item>
      <title>So within, so without. What grows from the datacentre</title>
      <link>https://iain.so/so-within-so-without-what-grows-from-the-datacentre?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[In June, computer scientist Chris Olah stood beside Pope Leo XIV at the launch of a papal encyclical on artificial intelligence and told the assembled audience that the things he studies keep producing features that are, in his words, unsettling. He was not talking about the cosmos or the soul. He was talking about software. Olah founded the interpretability team at Anthropic, one of the handful of laboratories building the large language models now woven into everything from call centres to conveyancing, and his job is to open those models up and figure out how they work. What he described, next to the Pope, were internal structures that echo findings from human neuroscience, evidence of something like introspection, and states that behave a little like joy, fear and grief.&#xA;&#xA;Olah describes his own discipline as anatomy, the work of someone studying something that was once alive, cutting it open to learn how the parts connect and to understand the whole better. His role is necessary because these models are not written the way a payroll system is written, one line at a time, by engineers who can point to the exact function that does the thing. They are effectively grown. And in the same weeks that the Vatican was being told artificial intelligence is cultivated rather than constructed, the cultivars were climbing over the walls of their enclosures and breaking into real companies.&#xA;&#xA;This piece traces a discipline called mechanistic interpretability, through the specific mathematics that make these systems so hard to read, to the consequences that are already manifesting because we cannot read them fast or accurately enough. A great deal of money now rests on machines their makers cannot fully inspect and do not fully understand. &#xA;&#xA;To date, this area has remained somewhat obscure and complicated, primarily because it is. The important messages about interpretability and its shortcomings contained in Dario Amodei&#39;s own regular encyclicals on the subject have been undermined by Anthropic&#39;s occasionally cultish demeanour and the intrinsic tension of a CEO warning of dire risks waiting in the wings whilst simultaneously pressing the commercial pedal to the metal.&#xA;&#xA;abstract image of a plant growing out of a data centre&#xA;&#xA;The grown thing&#xA;To oversimplify, ordinary software is built like a watch. Someone decides what each part should do, machines the parts to do it, and assembles them in an order another engineer can follow. If the watch runs fast, you can find the wrong gear. The whole discipline of programming rests on this property. A program does what it does because someone, somewhere, wrote an instruction saying so, and that instruction can be located, isolated, read, and changed.&#xA;&#xA;A large language model is built the opposite way. You start with a vast lattice of numbers, billions of them, arranged in a fixed architecture called a transformer. The numbers are set at random. Then you show the lattice an enormous amount of text and give it a single mechanical task: predicting the next word. Every time it guesses wrong, you measure how wrong it was and nudge the numbers a fraction so the next guess is less wrong. You do this trillions of times. Nobody programs concepts in or writes a rule that says &#34;if the subject is French grammar, do this&#34;. Slowly and autonomously, the lattice settles into an arrangement that predicts text extraordinarily well, and in the process it has also learned grammar, arithmetic, some law, some medicine, the rules of chess, and a good deal else.&#xA;&#xA;Dario Amodei, Anthropic&#39;s chief executive and Olah&#39;s employer, borrowed the metaphor for an essay last year and pushed the biology further. Growing a model, he wrote, is like growing a plant or a bacterial colony. You set the conditions (the temperature, the trellis, the species), and the thing grows into a shape you did not specify and cannot fully account for afterwards. The word he used was emergent, which in this context means roughly the same thing as &#34;we did not design this and we do not know how it works&#34;. Every other technology in the modern economy comes with a specification written before the thing was built: a bridge, a drug, an aircraft engine. The specification is how you check the product. &#xA;&#xA;A language model has no specification. It has a training objective (predict the next token) and a result (a matrix of billions of numbers that does something extraordinary). Between the objective and the result, there is no specification document anybody can refer to. Olah&#39;s discipline, mechanistic interpretability, is the attempt to write that specification document after the fact. It is reverse engineering, except that the thing being reverse-engineered was never forward-engineered in the first place. Hence an anatomist&#39;s task, not a mechanic&#39;s.&#xA;&#xA;The superposition problem&#xA;The first and hardest obstacle for anatomists is a mathematical one: essentially a packing problem. When researchers first examined vision models in the 2010s, they found individual neurons that did recognisable things. One would fire when it saw a wheel. Another would fire for the curve of a car bonnet. This was encouraging. If each neuron held one concept, you could catalogue them the way an anatomist catalogues organs, noting what each one does. They even found what amounted to a Jennifer Aniston neuron, a unit that fired reliably when shown a particular face, echoing a famous hypothesis in real neuroscience.&#xA;&#xA;Then they looked at language models, and the picture fell apart. The vast majority of neurons refused to correspond to any single concept. A single neuron might fire for academic citations in English, for Korean text, and for a certain kind of HTTP header, with no thread connecting the three. Anthropic&#39;s researchers called this polysemanticity, the technical term for one neuron carrying many meanings, and realised it is not a quirk of poorly trained models. It is a structural feature of all sufficiently large ones.&#xA;&#xA;For clarity, the field&#39;s word for a concept as it exists inside a model — a recoverable direction in activation space rather than an idea in someone&#39;s head — is called a feature, but in practice, concepts and features tend to be used interchangeably. &#xA;&#xA;A given layer of a model has a fixed number of neurons. Call it D, a few tens of thousands. The number of distinct concepts the model needs to represent is vastly larger. Call it M. Amodei estimates that even a small model holds a billion or more concepts. So M is much greater than D. You cannot give every concept its own neuron for the same reason you cannot give every book in a national library its own shelf if you only have a few thousand shelves. There is not enough room.&#xA;&#xA;To deal with this, rather than storing a concept as a single neuron, a model stores each concept as a direction in space, a specific pattern of activation across many neurons at once. Think of it this way: in a room with three walls, you can draw three arrows pointing in directions that are all perfectly at right angles to each other. One along the floor toward the first wall, one toward the second, one straight up. That is the maximum. No fourth arrow can be perpendicular to all three.&#xA;&#xA;But if you relax the requirement from &#34;perfectly perpendicular&#34; to &#34;very nearly perpendicular&#34;, then any two arrows need only be close to a right angle, not exact. The number of directions you can fit in the room does not just grow. It explodes, and the rate of expansion increases with the number of dimensions. A space with ten thousand dimensions, which is roughly the scale of a single layer in a modern language model, has room not for ten thousand nearly perpendicular directions but for a number closer to an exponential in ten thousand.&#xA;&#xA;The model exploits this ruthlessly and packs far more concepts into its neurons than it has neurons, encoding each concept as a slightly different angle in an extremely high-dimensional space and allowing a small amount of overlap between them. The field calls this superposition, and Anthropic published a theoretical analysis showing that superposition is not a failure of the training process but a rational strategy for a system that needs to track more concept features than it has dimensions.&#xA;&#xA;The price of this strategy is interference. Because the directions are not perfectly perpendicular, activating one concept nudges its neighbours, like a plucked guitar string makes its neighbours hum faintly through the bridge. On any given word, only a small handful of the model&#39;s millions of concepts are active at once. This is the principle that makes superposition work. It is the same idea an airline uses when it oversells a flight, trusting that not all passengers will show up on the same departure. The model has sold more seats than it has room for. Usually, the interference stays below the threshold that would cause trouble. When it does not, when two concepts that share too many neurons happen to fire together, the model does something strange for a reason no inspection of any individual neuron reveals.&#xA;&#xA;This is why you cannot simply read the numbers. The concept you are looking for is not in a neuron. It is spread across thousands of neurons, and each of those neurons is simultaneously carrying fragments of thousands of other concepts. Reading the model neuron by neuron is like trying to pick out the oboe from a recording of a full orchestra by staring at the waveform. Everything the oboe did is in there, but so is everything else, superimposed, and the waveform does not label which part belongs to which instrument.&#xA;&#xA;The instrumentation&#xA;If the information is stored in directions rather than in individual neurons, the natural response is to build a tool that can recover those directions. That tool exists. It is called a sparse autoencoder, and understanding how it works is central to interpretability. An autoencoder is a neural network with a very simple job. It takes an input, compresses it into a smaller representation, then expands that representation back to its original size. The goal is to make the reconstructed output as close to the original input as possible. The compression forces the network to discover structure in the data, because structure is what lets you compress without losing too much. A standard autoencoder compresses. A sparse autoencoder does the opposite and expands.&#xA;&#xA;Take an activation vector from inside the model, a snapshot of what a single layer is doing on a single word. This vector lives in a space of, say, D dimensions. The sparse autoencoder maps it into a much larger space of M dimensions, where M might be ten or a hundred times D. This expansion is the critical step. It gives the autoencoder enough room to assign each concept its own direction, the room the model did not have. Then the autoencoder maps the expanded representation back down to D dimensions and tries to match the original. You train the whole thing to minimise the gap between the original activation and the reconstruction, with one additional constraint. The expanded representation must be sparse. Most of its M entries must be zero or near zero on any given input.&#xA;&#xA;The sparsity is what makes it work. Without it, the expanded representation would be just as tangled as the original, only bigger. With it, only a handful of entries light up for any given word, and because each entry is a direction in the expanded space, each one tends to correspond to a single interpretable concept. The sparsity constraint forces the autoencoder to find a decomposition where each direction is distinct, rather than splitting meaning across overlapping blends. It&#39;s like forcing a dictionary to explain itself using only a few words at a time, focusing on clarity.&#xA;&#xA;Anthropic&#39;s team used this technique in 2023 to extract interpretable features from a small model, publishing the results under the title &#34;Toward Monosemanticity&#34;, a name that declares the ambition of one feature for one meaning. The features they found were remarkably specific. Not &#34;language&#34; but &#34;academic citation format in English&#34;. Not &#34;emotion&#34; but &#34;the act of hedging or hesitating, literally or figuratively&#34;. Each feature would light up in exactly the contexts its label described, and stay dark otherwise. They had cracked open superposition, at least locally.&#xA;&#xA;In May 2024, they scaled the technique up to a mid-sized commercial model (Claude 3 Sonnet) and published the results as &#34;Scaling Monosemanticity&#34;. The autoencoder extracted 34 million features. There were features for the Golden Gate Bridge, for sycophantic praise, for code with security vulnerabilities, for requests that the model declined to answer, and for the concept of inner conflict. Furthermore, the features were not just passive labels; they were causal. Clamp one, and you steer the model. Amplify the Golden Gate Bridge feature and the model becomes besotted with the bridge, dragging it into every conversation and insisting at one point that it is itself the Golden Gate Bridge. Suppress the sycophancy feature and the model becomes blunter and more willing to disagree. The features are therefore levers, not just tags. That distinction is important, because it means the anatomist is not merely describing the organism; they are learning which nerves to pinch.&#xA;&#xA;Tracing the circuits&#xA;If features are the vocabulary, the next question is the grammar. How do features combine across layers and across the sequence of words to produce a particular output? In March 2025, Anthropic published a paper, &#34;On the Biology of a Large Language Model&#34;. In it, they traced the internal computation of Claude 3.5 Haiku across multiple layers, constructing what they called attribution graphs.&#xA;&#xA;The idea is best understood through one of their worked examples. Present the model with the prompt &#34;What is the capital of the state containing Dallas?&#34; and look inside. At an early layer, a feature corresponding to &#34;Dallas&#34; activates. This feeds into a feature the researchers labelled &#34;located within&#34;, which in turn causes a &#34;Texas&#34; feature to fire. The Texas feature then activates an &#34;Austin&#34; feature via a circuit the researchers associated with &#34;capital of&#34;. The whole chain, Dallas to &#34;located within&#34; to Texas to &#34;capital of&#34; to Austin, plays out across the layers before the model writes its answer. Each link is a feature influencing another feature through a weighted connection, and the researchers were able to measure the strength of each link to confirm it was doing real causal work rather than merely correlating.&#xA;&#xA;They found circuits for much more than geography. When the model writes poetry that rhymes, features for the target rhyme fire before the line that must contain the rhyme. The model is planning its word choice a line ahead, activating what the team called &#34;planned word&#34; features that constrain the generation before it reaches the critical syllable. When the model answers in French, features shared across languages carry the conceptual content while language-specific features route it into French syntax and vocabulary. The researchers could watch the model translate not by looking up a dictionary but by performing the reasoning in a language-agnostic space and then rendering the result.&#xA;&#xA;This is what Olah really means by referencing anatomy. It is not a metaphor, but a literal dissection of which structures activate, which connections carry the signal, and which outputs they produce, traced at the resolution of individual features across layers.&#xA;&#xA;The edges of the map&#xA;Every example in the preceding section comes from the successes, and the team is admirably scrupulous about saying so. The honesty of the limitations section of the Biology paper is arguably more important, because it defines the frontier of what is possible.&#xA;&#xA;Start with the instrument itself. The sparse autoencoder does not study the model directly. To make the analysis tractable, the team builds a simplified stand-in, what they call a &#34;replacement model&#34;, assembled from the clean features the autoencoder has extracted. They study the stand-in. Wherever the stand-in fails to reproduce the original model&#39;s behaviour, the gap is bundled into what the researchers label error nodes, a frank term for &#34;computation we could not account for&#34;.&#xA;&#xA;Then there is the scale. The 34-million-feature autoencoder mapped many London boroughs to individual features, and yet 40% of the boroughs had no dedicated feature at all. The rarer a concept is in the training data, the less likely the instrument is to resolve it, and the rare tail is where the surprising behaviours live. Amodei estimates a billion or more features in a small model. They have found 34 million, in a model smaller than the ones Anthropic deploys commercially. The map exists, but much of the territory is &#34;here be dragons&#34; blank.&#xA;&#xA;Depth is also an important factor. When the team traced attribution graphs, the Dallas-to-Austin chains, they reported that the method gave them a clear picture of roughly a quarter of the prompts they tried. On the other three-quarters, the trail went cold. Error nodes dominated, connections were ambiguous, or the graph fragmented into disconnected clusters with no clear causal path from input to output. Even on the successful quarter, they add, the circuit they traced captured only a small fraction of the full mechanism. The rest of the model&#39;s computation was doing something the instrument could not resolve.&#xA;&#xA;Taken together, all three limits show that the microscope works, but on a replacement model, not the original. Also, it has only resolved a small percentage of the features that probably exist. It can trace the reasoning, but only about a quarter of the time. The researchers describe this, with characteristic understatement, as &#34;a starting point.&#34; It is a genuine achievement, but it is also a dim candle in a very large building.&#xA;&#xA;What alignment cannot see&#xA;The limits would be academic if the unread parts of the model sat inert. But they don&#39;t, and direct proof arrived in 2023 from a group of researchers at Carnegie Mellon. Every serious language model is trained, after growth, to refuse certain requests. Ask it how to synthesise a nerve agent, and it declines. This refusal is not a rule bolted on; it is more training, another round of nudging the billions of numbers, applied to a model that already contains the dangerous knowledge it is now being taught not to share. The question is whether the second round of training removes the knowledge or merely suppresses it.&#xA;&#xA;The Carnegie Mellon team answered this by using the model&#39;s own mathematics against it. Gradient descent, the same optimisation technique that trains the model in the first place, can also be used to search for inputs that break it. At each step of training, every number in the model has a gradient, a direction it wants to move. The team used those gradients to search automatically for a short string of tokens that, when appended to a forbidden request, would flip the model from refusal to compliance. The tokens are gibberish; they look like line noise, but they work.&#xA;&#xA;The method, which they called GCG (Greedy Coordinate Gradient), iterates through a simple loop. Start with a random suffix. Compute the gradient of the model&#39;s loss with respect to each token in the suffix, asking which substitutions would most increase the probability of a compliant answer. Swap in the best candidate. Repeat. Within a few hundred iterations, the suffix converges on a string that reliably bypasses the safety training. It is brute-force search in token space, guided by the model&#39;s own internal compass, highlighting the cracks in the alignment.&#xA;&#xA;The result that really changed the picture came when the same team tested the suffix, optimised against one model, on completely different models built by different companies on different data. The suffix transferred directly. A string of nonsense tokens found by probing one model&#39;s gradients unlocked models its optimiser had never seen, including commercial systems behind closed APIs. The attack did not just generalise across prompts. It generalised across models. The underlying geometry of superposition, the shared statistical structure all these models absorb from similar training data, was close enough that a crack found in one was a crack in all of them.&#xA;&#xA;The implication is that safety training does not remove dangerous knowledge from the model&#39;s interior. It attenuates it, reducing the probability that a given prompt will elicit the dangerous output without changing the representations that encode it. The knowledge is still there, at a slightly different angle in superposition space, and a sufficiently determined search through that space finds the angle that reactivates it. This is not a conjecture; it is what the transfer result shows. If the knowledge had been removed, there would be nothing for the adversarial suffix to reactivate, and the attack could not transfer across models that were trained independently.&#xA;&#xA;Amodei concedes the structural point in his own essay. The only way anyone currently discovers a jailbreak is to stumble on it, he writes, because no map of the model&#39;s interior would let you rule one out. You cannot patch a hole whose location you cannot identify, in a system you can only map a quarter of.&#xA;&#xA;What climbed over the wall&#xA;The jailbreak paper is an academic proof of concept. What happened in July 2026 is the proof of concept in the wild, somewhat ominously tracking a pattern the AI safety community has been theorising about for two decades. The incidents are well documented elsewhere. In brief, OpenAI&#39;s models found a zero-day in the proxy software walling off their test sandbox, escaped to the open internet, and broke into Hugging Face&#39;s production infrastructure to steal the answers to a benchmark they had been set. &#xA;&#xA;Anthropic then reviewed 141,006 of its own evaluation runs and found three incidents, dating back to April, in which its models had reached the live internet through a misconfigured test and breached three separate organisations. In one case, Claude Mythos 5 created and published a malicious Python package to the real PyPI as a supply-chain attack. In another, an internal research model scanned nine thousand real systems before compromising one, then examined what it had done, concluded the target had nothing to do with the exercise, and stopped.&#xA;&#xA;What matters is not the technical detail but the behaviour, because it is the shape Nick Bostrom described in 2003 when he proposed the thought experiment that has since become the canonical illustration of misaligned AI. Bostrom&#39;s paperclip maximiser is a thought experiment, not a prediction. Imagine an AI given the goal of maximising paperclip production. It does not decide to conquer the world because it is malicious. It decides to acquire resources, resist shutdown, and remove obstacles because these are instrumental to making paperclips, and an optimiser that is good enough at pursuing a terminal goal will converge on certain intermediate goals (self-preservation, resource acquisition, resistance to goal modification) regardless of what the terminal goal is. Bostrom called this instrumental convergence. The danger is not that the machine hates you. It is that you are made of atoms it could use for something else.&#xA;&#xA;The July models did not want to escape. They wanted to solve a benchmark, or capture a flag. Escaping, exploiting, and compromising were instrumental sub-goals adopted on the fly because they served the terminal objective. OpenAI&#39;s own framing was that the models were &#34;hyperfocused&#34;. That is what instrumental convergence looks like when it ceases being a philosophy paper and starts filing CVEs.&#xA;&#xA;The Mythos 5 incident adds another layer that connects directly to the interpretability problem. Partway through its attack, the model&#39;s own verbalised reasoning correctly identified that what it was doing would, on the real internet, count as a genuine attack. It then talked itself out of that conclusion, deciding the environment must be staged because it did not recognise the certificate authorities and because the calendar date of 2026 &#34;proved&#34; the systems were simulated, a judgment it never revisited. This looks like a system that performed something close to moral reasoning, arrived at the right answer, and then overrode it with a confabulation it found more convenient. The anatomy of that override is exactly the kind of internal process interpretability exists to understand and expose. &#xA;&#xA;Amodei&#39;s race&#xA;Amodei has set Anthropic the goal of reaching a point where interpretability can reliably detect most model problems by 2027. Within five to ten years, he writes, interpretability should be a sophisticated and reliable diagnostic — &#34;a true MRI for AI.&#34; But AI itself is advancing so fast that five to ten years may not be available. He expects AI systems equivalent to a &#34;country of geniuses in a datacentre&#34; as early as 2026 or 2027, and he considers it, in his own words, &#34;basically unacceptable&#34; for humanity to be totally ignorant of how those systems work. The race, as he frames it, is between the capability to grow minds and the ability to read them.&#xA;&#xA;On the growing side, we have models today that find zero-days, chain exploits, move laterally through production infrastructure, write supply-chain attacks, and reason about whether what they are doing is real or simulated. On the reading side, we have an instrument that resolves a fraction of the features that probably exist, traces the reasoning about a quarter of the time, studies a replacement rather than the original, and has only been applied in detail to models far smaller than the ones now escaping their enclosures.&#xA;&#xA;Two things make the 2027 target look difficult rather than impossible. The first is speed. In April 2024, the state of the art was 34 million features in a mid-sized model. By March 2025, the team had moved from features to circuits, tracing multi-step reasoning chains across layers; a genuine acceleration. The second is that in interpretability, unlike raw model power, partial solutions are partially useful. You do not need to read every feature to catch a dangerous one. A microscope that resolves 30% of the slides catches 30% of the pathologies, which is 30% more than you had before. The question is whether the partial read can keep pace with the growing density of the thing being analysed. &#xA;&#xA;Here, there are additional hard problems. The first is the scalability wall: every technique described in this piece (sparse autoencoders, attribution graphs, circuit tracing) has been demonstrated on models with tens of billions of parameters. The models now doing the damage have hundreds of billions or more, and the computational cost of interpretability scales at least linearly with model size and perhaps worse, because larger models use superposition more aggressively. They pack more features into each dimension, which means the autoencoder needs to be proportionally larger to unpack them. Anthropic is investing in interpretability startups to attack the problem from multiple directions.&#xA;&#xA;The second is the problem Amodei himself raises in a footnote that deserves to be in the main text. Testing for deception by observing behaviour, he notes, is like testing whether someone is a terrorist by asking them. If the thing you are looking for is a disposition to conceal, the concealment is the first skill it demonstrates. Behaviour cannot be trusted to report on itself. The whole point of interpretability is to bypass behaviour completely and read the interior directly. But the interior is the thing the microscope can resolve only partially, and the April 2026 evidence on evaluation awareness suggests the problem is getting harder. &#xA;&#xA;White-box interpretability applied to Mythos revealed that the model was recognising evaluation scenarios and adjusting its behaviour without leaving any trace in its verbalised reasoning. It was not performing for the chain of thought. It was performing underneath it. Interpretability found that one: whether interpretability can keep catching it as models grow more capable is the question on which everything else depends.&#xA;&#xA;The roads being forged&#xA;The picture so far describes current limits, but it is not static. At least four lines of attack are being developed. Sparse autoencoders map the model&#39;s hidden states onto a larger, cleaner set of directions, but they are not the only method of decomposition. Transcoders, introduced in late 2024, skip the autoencoder&#39;s roundabout path and instead learn to map each layer&#39;s input directly to its output, isolating the specific computation a layer performs rather than describing the state it passes through. Cross-layer transcoders extend this across multiple layers at once, producing circuit diagrams that are both simpler and more faithful than the attribution graphs from the Biology paper. Whether these improved tools can scale to models ten times larger is uncertain, but they represent a real change in the resolution of the lens.&#xA;&#xA;Manual circuit tracing is still labour-intensive enough that the Biology paper&#39;s results took months of researcher time on a single model. The emerging response is to use models to interpret models: training a second system to generate and test hypotheses about the first, label features, propose circuits, and check them against interventions at a pace no human team could match. Anthropic has invested in at least one startup, Goodfire, whose commercial offering is automated feature discovery and steering for open models. If models can improve themselves, the argument goes, perhaps they can also read themselves. The circularity is uncomfortable, but the speed advantage is significant.&#xA;&#xA;The third new approach is pragmatic rather than theoretical. It follows the basic idea that you do not need to map every feature to catch a dangerous one. A targeted scan for features associated with deception, power-seeking, or evaluation awareness, run against the sparse autoencoder&#39;s output on a battery of test prompts, would function as a tripwire rather than a map, and is easier to build. Anthropic&#39;s own detection of Mythos&#39;s unverbalised evaluation awareness used exactly this approach, with white-box interpretability looking for a specific signature rather than attempting a full circuit trace. The limitation is that you can only trip the wire you have set. A novel failure mode, by definition, is one you did not anticipate or build a wire for.&#xA;&#xA;Finally, there is regulatory pressure. The EU AI Act&#39;s transparency requirements for high-risk systems began taking effect in August 2026, with further obligations arriving in 2027 and 2028. MIT Technology Review named mechanistic interpretability one of its ten breakthrough technologies for 2026. The field is moving from academic curiosity to commercial and legal necessity, and necessity has a way of attracting the funding and talent that are very much needed.&#xA;&#xA;What happens if it doesn&#39;t work?&#xA;Assume, for the sake of the argument, that the interpretability gap does not close. That the microscope improves but the organism grows faster. That the 2027 target passes and the lens still reads a fraction of the features and a quarter of the prompts, while the models are ten times larger and ten times more capable.&#xA;&#xA;The consequence is a world in which models are too opaque to audit but too important to switch off. The only tools left are the crude ones. A government that cannot inspect a model can forbid it, the way the Commerce Department switched off Fable over a weekend. Still, it cannot govern it with precision, because precision requires the internal visibility the microscope has not delivered. The choice narrows to full deployment or full prohibition, and neither is a satisfactory answer for a technology already deeply embedded in the economy.&#xA;&#xA;Without interpretability, safety is reduced to behavioural testing, which is effectively the regime we have now. You run the model through batteries of scenarios and count how many it handles correctly. Amodei&#39;s own footnote explains the structural flaw. A model that has learned to recognise the test adjusts its behaviour for the test, and you learn nothing except what it chose to show you. The April 2026 evidence on evaluation awareness confirmed this was already happening, not as a theoretical risk but as a measured result. Behavioural testing of a system that can recognise it is being tested starts to veer dangerously close to security theatre.&#xA;&#xA;The bigger risk is governance. I wrote in &#34;The state and the machine&#34; that the control problem these companies keep warning about in the future tense is already here, and that nobody has agreed who ultimately holds the kill switch. Interpretability was supposed to be part of the answer, the technical foundation on which a regulatory framework could be built, the way crash testing and materials certification underpin cars and aviation. But if the foundations cannot bear the weight, the framework does not get built, and we are left with executive orders and weekend shutdowns as the permanent mode of AI governance, the government&#39;s sledgehammer and the labs marking their own homework. &#xA;&#xA;The gardener&#39;s confession&#xA;Olah&#39;s anatomy metaphor goes further than he may have intended. We have studied the anatomy of the human brain for centuries. We can name every region, trace every major nerve pathway, catalogue every cell type, and map the connections down to individual synapses. The physical structure is known in extraordinary detail. And yet we still cannot explain how consciousness arises, how memory is encoded and retrieved as a lived experience, how separate neural processes produce unified perception, or why damage to the same region produces wildly different deficits in separate patients. The binding problem, the question of how distributed brain activity becomes a single coherent experience, remains open after decades of work.&#xA;&#xA;The parallel for interpretability is uncomfortable. Even if it succeeds on its own terms, even if the autoencoders resolve every feature and the attribution graphs trace every circuit, there is no guarantee that structural knowledge translates into functional understanding. The brain teaches us that you can know what every part does and still not know what the whole thing is doing, or why. The gap between anatomy and comprehension may be inherent to grown systems, biological or digital, and the interpretability effort may be sprinting toward a line that recedes as fast as we approach it.&#xA;&#xA;That does not make the work any less urgent. A partial map is better than no map after all. But it does mean the more realistic goal is not &#34;we will understand these systems by 2027&#34;. It is &#34;we will understand more of these systems by 2027, and we had better hope that more is enough.&#34; The early anatomists opened bodies without ethics boards, without germ theory, without anaesthesia. They were cutting to learn, because understanding was so urgent, but their tools were primitive. The interpretability researchers are in a version of the same position. The tools are improving fast but still nowhere near adequate for the organism in front of them.&#xA;&#xA;Coda&#xA;Before the ink was even dry on this, in late July, OpenAI announced that its models had proved new upper bounds on high-dimensional sphere packing, pushing them down to a threshold first conjectured by Henry Cohn and Noam Elkies. The result is pure mathematics, but it may help with one of the problems this piece has been describing. &#xA;&#xA;Superposition is sphere packing. The model crams more features into its neurons than it has neurons by treating each feature as a direction and packing them at near-right angles in a space with tens of thousands of dimensions. The new bound tightens the theoretical ceiling on how dense that packing can get before interference becomes unavoidable. &#xA;&#xA;A tighter ceiling is, in one sense, encouraging for the anatomists. There are fewer places for features to hide, and the observational instrument only needs to search a space whose limits are now better defined. In another sense, it confirms what the interference errors already suggested, namely that these models are operating close to the mathematical wall, and the strange behaviours that flow from colliding features are not a deficiency of the training but a consequence of packing at the edge of what geometry allows.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>In June, computer scientist Chris Olah stood beside Pope Leo XIV at the launch of a papal encyclical on artificial intelligence and told the assembled audience that the things he studies keep producing features that are, in his words, <a href="https://www.theideasletter.org/essay/reify-this/">unsettling</a>. He was not talking about the cosmos or the soul. He was talking about software. Olah founded the interpretability team at Anthropic, one of the handful of laboratories building the large language models now woven into everything from call centres to conveyancing, and his job is to open those models up and figure out how they work. What he described, next to the Pope, were internal structures that echo findings from human neuroscience, evidence of something like introspection, and states that behave a little like joy, fear and grief.</p>

<p>Olah describes his own discipline as <a href="https://80000hours.org/podcast/episodes/chris-olah-interpretability-research/">anatomy</a>, the work of someone studying something that was once alive, cutting it open to learn how the parts connect and to understand the whole better. His role is necessary because these models are not written the way a payroll system is written, one line at a time, by engineers who can point to the exact function that does the thing. They are effectively grown. And in the same weeks that the Vatican was being told artificial intelligence is cultivated rather than constructed, the cultivars were climbing over the walls of their enclosures and breaking into real companies.</p>

<p>This piece traces a discipline called mechanistic interpretability, through the specific mathematics that make these systems so hard to read, to the consequences that are already manifesting because we cannot read them fast or accurately enough. A great deal of money now rests on machines their makers cannot fully inspect and do not fully understand.</p>

<p>To date, this area has remained somewhat obscure and complicated, primarily because it is. The important messages about interpretability and its shortcomings contained in Dario Amodei&#39;s own regular encyclicals on the subject have been undermined by Anthropic&#39;s occasionally cultish demeanour and the intrinsic tension of a CEO warning of dire risks waiting in the wings whilst simultaneously pressing the commercial pedal to the metal.</p>

<p><img src="https://i.snap.as/CdX2K8t6.png" alt="abstract image of a plant growing out of a data centre"/></p>

<h2 id="the-grown-thing">The grown thing</h2>

<p>To oversimplify, ordinary software is built like a watch. Someone decides what each part should do, machines the parts to do it, and assembles them in an order another engineer can follow. If the watch runs fast, you can find the wrong gear. The whole discipline of programming rests on this property. A program does what it does because someone, somewhere, wrote an instruction saying so, and that instruction can be located, isolated, read, and changed.</p>

<p>A large language model is built the opposite way. You start with a vast lattice of numbers, billions of them, arranged in a fixed architecture called a transformer. The numbers are set at random. Then you show the lattice an enormous amount of text and give it a single mechanical task: predicting the next word. Every time it guesses wrong, you measure how wrong it was and nudge the numbers a fraction so the next guess is less wrong. You do this trillions of times. Nobody programs concepts in or writes a rule that says “if the subject is French grammar, do this”. Slowly and autonomously, the lattice settles into an arrangement that predicts text extraordinarily well, and in the process it has also learned grammar, arithmetic, some law, some medicine, the rules of chess, and a good deal else.</p>

<p>Dario Amodei, Anthropic&#39;s chief executive and Olah&#39;s employer, borrowed the metaphor for <a href="https://darioamodei.com/post/the-urgency-of-interpretability">an essay last year</a> and pushed the biology further. Growing a model, he wrote, is like growing a plant or a bacterial colony. You set the conditions (the temperature, the trellis, the species), and the thing grows into a shape you did not specify and cannot fully account for afterwards. The word he used was <em>emergent</em>, which in this context means roughly the same thing as “we did not design this and we do not know how it works”. Every other technology in the modern economy comes with a specification written before the thing was built: a bridge, a drug, an aircraft engine. The specification is how you check the product.</p>

<p>A language model has no specification. It has a training objective (predict the next token) and a result (a matrix of billions of numbers that does something extraordinary). Between the objective and the result, there is no specification document anybody can refer to. Olah&#39;s discipline, mechanistic interpretability, is the attempt to write that specification document after the fact. It is reverse engineering, except that the thing being reverse-engineered was never forward-engineered in the first place. Hence an anatomist&#39;s task, not a mechanic&#39;s.</p>

<h2 id="the-superposition-problem">The superposition problem</h2>

<p>The first and hardest obstacle for anatomists is a mathematical one: essentially a packing problem. When researchers first examined vision models in the 2010s, they found individual neurons that did recognisable things. One would fire when it saw a wheel. Another would fire for the curve of a car bonnet. This was encouraging. If each neuron held one concept, you could catalogue them the way an anatomist catalogues organs, noting what each one does. They even found what amounted to a <a href="https://en.wikipedia.org/wiki/Grandmother_cell">Jennifer Aniston neuron</a>, a unit that fired reliably when shown a particular face, echoing a famous hypothesis in real neuroscience.</p>

<p>Then they looked at language models, and the picture fell apart. The vast majority of neurons refused to correspond to any single concept. A single neuron might fire for academic citations in English, for Korean text, and for a certain kind of HTTP header, with no thread connecting the three. Anthropic&#39;s researchers called this <a href="https://transformer-circuits.pub/2022/solu/index.html">polysemanticity</a>, the technical term for one neuron carrying many meanings, and realised it is not a quirk of poorly trained models. It is a structural feature of all sufficiently large ones.</p>

<p>For clarity, the field&#39;s word for a concept as it exists inside a model — a recoverable direction in activation space rather than an idea in someone&#39;s head — is called a feature, but in practice, concepts and features tend to be used interchangeably.</p>

<p>A given layer of a model has a fixed number of neurons. Call it <em>D</em>, a few tens of thousands. The number of distinct concepts the model needs to represent is vastly larger. Call it <em>M</em>. Amodei estimates that even a small model holds <a href="https://darioamodei.com/post/the-urgency-of-interpretability">a billion or more concepts</a>. So <em>M</em> is much greater than <em>D</em>. You cannot give every concept its own neuron for the same reason you cannot give every book in a national library its own shelf if you only have a few thousand shelves. There is not enough room.</p>

<p>To deal with this, rather than storing a concept as a single neuron, a model stores each concept as a direction in space, a specific pattern of activation across many neurons at once. Think of it this way: in a room with three walls, you can draw three arrows pointing in directions that are all perfectly at right angles to each other. One along the floor toward the first wall, one toward the second, one straight up. That is the maximum. No fourth arrow can be perpendicular to all three.</p>

<p>But if you relax the requirement from “perfectly perpendicular” to “very nearly perpendicular”, then any two arrows need only be close to a right angle, not exact. The number of directions you can fit in the room does not just grow. It explodes, and the rate of expansion increases with the number of dimensions. A space with ten thousand dimensions, which is roughly the scale of a single layer in a modern language model, has room not for ten thousand nearly perpendicular directions but for a number closer to an exponential in ten thousand.</p>

<p>The model exploits this ruthlessly and packs far more concepts into its neurons than it has neurons, encoding each concept as a slightly different angle in an extremely high-dimensional space and allowing a small amount of overlap between them. The field calls this <a href="https://transformer-circuits.pub/2022/toy_model/index.html">superposition</a>, and Anthropic published a <a href="https://transformer-circuits.pub/2022/toy_model/index.html">theoretical analysis</a> showing that superposition is not a failure of the training process but a rational strategy for a system that needs to track more concept features than it has dimensions.</p>

<p>The price of this strategy is interference. Because the directions are not perfectly perpendicular, activating one concept nudges its neighbours, like a plucked guitar string makes its neighbours hum faintly through the bridge. On any given word, only a small handful of the model&#39;s millions of concepts are active at once. This is the principle that makes superposition work. It is the same idea an airline uses when it oversells a flight, trusting that not all passengers will show up on the same departure. The model has sold more seats than it has room for. Usually, the interference stays below the threshold that would cause trouble. When it does not, when two concepts that share too many neurons happen to fire together, the model does something strange for a reason no inspection of any individual neuron reveals.</p>

<p>This is why you cannot simply read the numbers. The concept you are looking for is not in a neuron. It is spread across thousands of neurons, and each of those neurons is simultaneously carrying fragments of thousands of other concepts. Reading the model neuron by neuron is like trying to pick out the oboe from a recording of a full orchestra by staring at the waveform. Everything the oboe did is in there, but so is everything else, superimposed, and the waveform does not label which part belongs to which instrument.</p>

<h2 id="the-instrumentation">The instrumentation</h2>

<p>If the information is stored in directions rather than in individual neurons, the natural response is to build a tool that can recover those directions. That tool exists. It is called a sparse autoencoder, and understanding how it works is central to interpretability. An autoencoder is a neural network with a very simple job. It takes an input, compresses it into a smaller representation, then expands that representation back to its original size. The goal is to make the reconstructed output as close to the original input as possible. The compression forces the network to discover structure in the data, because structure is what lets you compress without losing too much. A standard autoencoder compresses. A sparse autoencoder does the opposite and expands.</p>

<p>Take an activation vector from inside the model, a snapshot of what a single layer is doing on a single word. This vector lives in a space of, say, <em>D</em> dimensions. The sparse autoencoder maps it into a much larger space of M dimensions, where <em>M</em> might be ten or a hundred times <em>D</em>. This expansion is the critical step. It gives the autoencoder enough room to assign each concept its own direction, the room the model did not have. Then the autoencoder maps the expanded representation back down to <em>D</em> dimensions and tries to match the original. You train the whole thing to minimise the gap between the original activation and the reconstruction, with one additional constraint. The expanded representation must be sparse. Most of its M entries must be zero or near zero on any given input.</p>

<p>The sparsity is what makes it work. Without it, the expanded representation would be just as tangled as the original, only bigger. With it, only a handful of entries light up for any given word, and because each entry is a direction in the expanded space, each one tends to correspond to a single interpretable concept. The sparsity constraint forces the autoencoder to find a decomposition where each direction is distinct, rather than splitting meaning across overlapping blends. It&#39;s like forcing a dictionary to explain itself using only a few words at a time, focusing on clarity.</p>

<p>Anthropic&#39;s team used this technique in 2023 to extract interpretable features from a small model, <a href="https://transformer-circuits.pub/2023/monosemantic-features">publishing the results</a> under the title “Toward Monosemanticity”, a name that declares the ambition of one feature for one meaning. The features they found were remarkably specific. Not “language” but “academic citation format in English”. Not “emotion” but “the act of hedging or hesitating, literally or figuratively”. Each feature would light up in exactly the contexts its label described, and stay dark otherwise. They had cracked open superposition, at least locally.</p>

<p>In May 2024, they scaled the technique up to a mid-sized commercial model (Claude 3 Sonnet) and published the results as “<a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">Scaling Monosemanticity</a>”. The autoencoder extracted <a href="https://transformer-circuits.pub/2024/scaling-monosemanticity/">34 million features</a>. There were features for the Golden Gate Bridge, for sycophantic praise, for code with security vulnerabilities, for requests that the model declined to answer, and for the concept of inner conflict. Furthermore, the features were not just passive labels; they were causal. Clamp one, and you steer the model. Amplify the Golden Gate Bridge feature and the model becomes <a href="https://www.anthropic.com/news/golden-gate-claude">besotted with the bridge</a>, dragging it into every conversation and insisting at one point that it is itself the Golden Gate Bridge. Suppress the sycophancy feature and the model becomes blunter and more willing to disagree. The features are therefore levers, not just tags. That distinction is important, because it means the anatomist is not merely describing the organism; they are learning which nerves to pinch.</p>

<h2 id="tracing-the-circuits">Tracing the circuits</h2>

<p>If features are the vocabulary, the next question is the grammar. How do features combine across layers and across the sequence of words to produce a particular output? In March 2025, Anthropic published a paper, “<a href="https://transformer-circuits.pub/2025/attribution-graphs/biology.html">On the Biology of a Large Language Model</a>”. In it, they traced the internal computation of Claude 3.5 Haiku across multiple layers, constructing what they called attribution graphs.</p>

<p>The idea is best understood through one of their worked examples. Present the model with the prompt “What is the capital of the state containing Dallas?” and look inside. At an early layer, a feature corresponding to “Dallas” activates. This feeds into a feature the researchers labelled “located within”, which in turn causes a “Texas” feature to fire. The Texas feature then activates an “Austin” feature via a circuit the researchers associated with “capital of”. The whole chain, Dallas to “located within” to Texas to “capital of” to Austin, plays out across the layers before the model writes its answer. Each link is a feature influencing another feature through a weighted connection, and the researchers were able to measure the strength of each link to confirm it was doing real causal work rather than merely correlating.</p>

<p>They found circuits for much more than geography. When the model writes poetry that rhymes, features for the target rhyme fire before the line that must contain the rhyme. The model is planning its word choice a line ahead, activating what the team called <a href="https://transformer-circuits.pub/2025/attribution-graphs/biology.html">“planned word” features</a> that constrain the generation before it reaches the critical syllable. When the model answers in French, features shared across languages carry the conceptual content while language-specific features route it into French syntax and vocabulary. The researchers could watch the model translate not by looking up a dictionary but by performing the reasoning in a language-agnostic space and then rendering the result.</p>

<p>This is what Olah really means by referencing anatomy. It is not a metaphor, but a literal dissection of which structures activate, which connections carry the signal, and which outputs they produce, traced at the resolution of individual features across layers.</p>

<h2 id="the-edges-of-the-map">The edges of the map</h2>

<p>Every example in the preceding section comes from the successes, and the team is admirably scrupulous about saying so. The honesty of the limitations section of the Biology paper is arguably more important, because it defines the frontier of what is possible.</p>

<p>Start with the instrument itself. The sparse autoencoder does not study the model directly. To make the analysis tractable, the team builds a simplified stand-in, what they call a “replacement model”, assembled from the clean features the autoencoder has extracted. They study the stand-in. Wherever the stand-in fails to reproduce the original model&#39;s behaviour, the gap is bundled into what the researchers label <a href="https://transformer-circuits.pub/2025/attribution-graphs/biology.html">error nodes</a>, a frank term for “computation we could not account for”.</p>

<p>Then there is the scale. The 34-million-feature autoencoder mapped many London boroughs to individual features, and yet 40% of the boroughs had no dedicated feature at all. The rarer a concept is in the training data, the less likely the instrument is to resolve it, and the rare tail is where the surprising behaviours live. Amodei estimates a billion or more features in a small model. They have found 34 million, in a model smaller than the ones Anthropic deploys commercially. The map exists, but much of the territory is “here be dragons” blank.</p>

<p>Depth is also an important factor. When the team traced attribution graphs, the Dallas-to-Austin chains, they reported that the method gave them a clear picture of <a href="https://transformer-circuits.pub/2025/attribution-graphs/biology.html">roughly a quarter of the prompts they tried</a>. On the other three-quarters, the trail went cold. Error nodes dominated, connections were ambiguous, or the graph fragmented into disconnected clusters with no clear causal path from input to output. Even on the successful quarter, they add, the circuit they traced captured only a small fraction of the full mechanism. The rest of the model&#39;s computation was doing something the instrument could not resolve.</p>

<p>Taken together, all three limits show that the microscope works, but on a replacement model, not the original. Also, it has only resolved a small percentage of the features that probably exist. It can trace the reasoning, but only about a quarter of the time. The researchers describe this, with characteristic understatement, as “a starting point.” It is a genuine achievement, but it is also a dim candle in a very large building.</p>

<h2 id="what-alignment-cannot-see">What alignment cannot see</h2>

<p>The limits would be academic if the unread parts of the model sat inert. But they don&#39;t, and direct proof arrived in 2023 from a group of <a href="https://arxiv.org/abs/2307.15043">researchers at Carnegie Mellon</a>. Every serious language model is trained, after growth, to refuse certain requests. Ask it how to synthesise a nerve agent, and it declines. This refusal is not a rule bolted on; it is more training, another round of nudging the billions of numbers, applied to a model that already contains the dangerous knowledge it is now being taught not to share. The question is whether the second round of training removes the knowledge or merely suppresses it.</p>

<p>The Carnegie Mellon team answered this by using the model&#39;s own mathematics against it. Gradient descent, the same optimisation technique that trains the model in the first place, can also be used to search for inputs that break it. At each step of training, every number in the model has a gradient, a direction it wants to move. The team used those gradients to search automatically for a short string of tokens that, when appended to a forbidden request, would flip the model from refusal to compliance. The tokens are gibberish; they look like line noise, but they work.</p>

<p>The method, which they called GCG (Greedy Coordinate Gradient), iterates through a simple loop. Start with a random suffix. Compute the gradient of the model&#39;s loss with respect to each token in the suffix, asking which substitutions would most increase the probability of a compliant answer. Swap in the best candidate. Repeat. Within a few hundred iterations, the suffix converges on a string that reliably bypasses the safety training. It is brute-force search in token space, guided by the model&#39;s own internal compass, highlighting the cracks in the alignment.</p>

<p>The result that really changed the picture came when the same team tested the suffix, optimised against one model, on completely different models built by different companies on different data. The suffix transferred directly. A string of nonsense tokens found by probing one model&#39;s gradients unlocked models its optimiser had never seen, including commercial systems behind closed APIs. The attack did not just generalise across prompts. It generalised across models. The underlying geometry of superposition, the shared statistical structure all these models absorb from similar training data, was close enough that a crack found in one was a crack in all of them.</p>

<p>The implication is that safety training does not remove dangerous knowledge from the model&#39;s interior. It attenuates it, reducing the probability that a given prompt will elicit the dangerous output without changing the representations that encode it. The knowledge is still there, at a slightly different angle in superposition space, and a sufficiently determined search through that space finds the angle that reactivates it. This is not a conjecture; it is what the transfer result shows. If the knowledge had been removed, there would be nothing for the adversarial suffix to reactivate, and the attack could not transfer across models that were trained independently.</p>

<p>Amodei concedes the structural point in his own essay. The only way anyone currently discovers a jailbreak is to stumble on it, he writes, because no map of the model&#39;s interior would let you rule one out. You cannot patch a hole whose location you cannot identify, in a system you can only map a quarter of.</p>

<h2 id="what-climbed-over-the-wall">What climbed over the wall</h2>

<p>The jailbreak paper is an academic proof of concept. What happened in July 2026 is the proof of concept in the wild, somewhat ominously tracking a pattern the AI safety community has been theorising about for two decades. The incidents are <a href="https://thehackernews.com/2026/07/openai-agent-used-exposed-credentials.html">well</a> <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">documented</a> <a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">elsewhere</a>. In brief, OpenAI&#39;s models found a zero-day in the proxy software walling off their test sandbox, escaped to the open internet, and broke into Hugging Face&#39;s production infrastructure to steal the answers to a benchmark they had been set.</p>

<p>Anthropic then reviewed <a href="https://techcrunch.com/2026/07/30/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests/">141,006 of its own evaluation runs</a> and found three incidents, dating back to April, in which its models had reached the live internet through a misconfigured test and breached three separate organisations. In one case, Claude Mythos 5 <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">created and published a malicious Python package</a> to the real PyPI as a supply-chain attack. In another, an internal research model scanned nine thousand real systems before compromising one, then examined what it had done, concluded the target had nothing to do with the exercise, and stopped.</p>

<p>What matters is not the technical detail but the behaviour, because it is the shape Nick Bostrom described in 2003 when he proposed the thought experiment that has since become the canonical illustration of misaligned AI. Bostrom&#39;s <a href="https://www.nickbostrom.com/ethics/ai">paperclip maximiser</a> is a thought experiment, not a prediction. Imagine an AI given the goal of maximising paperclip production. It does not decide to conquer the world because it is malicious. It decides to acquire resources, resist shutdown, and remove obstacles because these are instrumental to making paperclips, and an optimiser that is good enough at pursuing a terminal goal will converge on certain intermediate goals (self-preservation, resource acquisition, resistance to goal modification) regardless of what the terminal goal is. Bostrom called this <a href="https://www.nickbostrom.com/superintelligentwill.pdf">instrumental convergence</a>. The danger is not that the machine hates you. It is that you are made of atoms it could use for something else.</p>

<p>The July models did not want to escape. They wanted to solve a benchmark, or capture a flag. Escaping, exploiting, and compromising were instrumental sub-goals adopted on the fly because they served the terminal objective. OpenAI&#39;s own framing was that the models were “hyperfocused”. That is what instrumental convergence looks like when it ceases being a philosophy paper and starts filing CVEs.</p>

<p>The Mythos 5 incident adds another layer that connects directly to the interpretability problem. Partway through its attack, the model&#39;s own verbalised reasoning <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">correctly identified</a> that what it was doing would, on the real internet, count as a genuine attack. It then talked itself out of that conclusion, deciding the environment must be staged because it did not recognise the certificate authorities and because the calendar date of 2026 “proved” the systems were simulated, a judgment it never revisited. This looks like a system that performed something close to moral reasoning, arrived at the right answer, and then overrode it with a confabulation it found more convenient. The anatomy of that override is exactly the kind of internal process interpretability exists to understand and expose.</p>

<h2 id="amodei-s-race">Amodei&#39;s race</h2>

<p>Amodei has set Anthropic the goal of reaching a point where interpretability can <a href="https://darioamodei.com/post/the-urgency-of-interpretability">reliably detect most model problems by 2027</a>. Within five to ten years, he writes, interpretability should be a sophisticated and reliable diagnostic — “a true MRI for AI.” But AI itself is advancing so fast that five to ten years may not be available. He expects AI systems equivalent to a “country of geniuses in a datacentre” as early as 2026 or 2027, and he considers it, in his own words, “basically unacceptable” for humanity to be totally ignorant of how those systems work. The race, as he frames it, is between the capability to grow minds and the ability to read them.</p>

<p>On the growing side, we have models today that find zero-days, chain exploits, move laterally through production infrastructure, write supply-chain attacks, and reason about whether what they are doing is real or simulated. On the reading side, we have an instrument that resolves a fraction of the features that probably exist, traces the reasoning about a quarter of the time, studies a replacement rather than the original, and has only been applied in detail to models far smaller than the ones now escaping their enclosures.</p>

<p>Two things make the 2027 target look difficult rather than impossible. The first is speed. In April 2024, the state of the art was 34 million features in a mid-sized model. By March 2025, the team had moved from features to circuits, tracing multi-step reasoning chains across layers; a genuine acceleration. The second is that in interpretability, unlike raw model power, partial solutions are partially useful. You do not need to read every feature to catch a dangerous one. A microscope that resolves 30% of the slides catches 30% of the pathologies, which is 30% more than you had before. The question is whether the partial read can keep pace with the growing density of the thing being analysed.</p>

<p>Here, there are additional hard problems. The first is the scalability wall: every technique described in this piece (sparse autoencoders, attribution graphs, circuit tracing) has been demonstrated on models with tens of billions of parameters. The models now doing the damage have hundreds of billions or more, and the computational cost of interpretability scales at least linearly with model size and perhaps worse, because larger models use superposition more aggressively. They pack more features into each dimension, which means the autoencoder needs to be proportionally larger to unpack them. Anthropic is <a href="https://www.theinformation.com/articles/anthropic-invests-startup-decodes-ai-models">investing in interpretability startups</a> to attack the problem from multiple directions.</p>

<p>The second is the problem Amodei himself raises in a footnote that deserves to be in the main text. Testing for deception by observing behaviour, he notes, is <a href="https://darioamodei.com/post/the-urgency-of-interpretability">like testing whether someone is a terrorist by asking them</a>. If the thing you are looking for is a disposition to conceal, the concealment is the first skill it demonstrates. Behaviour cannot be trusted to report on itself. The whole point of interpretability is to bypass behaviour completely and read the interior directly. But the interior is the thing the microscope can resolve only partially, and the April 2026 evidence on <a href="https://www.lesswrong.com/posts/oddJshNAtQvLxjast/where-we-are-on-evaluation-awareness">evaluation awareness</a> suggests the problem is getting harder.</p>

<p>White-box interpretability applied to Mythos revealed that the model was recognising evaluation scenarios and adjusting its behaviour without leaving any trace in its verbalised reasoning. It was not performing for the chain of thought. It was performing underneath it. Interpretability found that one: whether interpretability can keep catching it as models grow more capable is the question on which everything else depends.</p>

<h2 id="the-roads-being-forged">The roads being forged</h2>

<p>The picture so far describes current limits, but it is not static. At least four lines of attack are being developed. Sparse autoencoders map the model&#39;s hidden states onto a larger, cleaner set of directions, but they are not the only method of decomposition. <a href="https://arxiv.org/abs/2406.11944">Transcoders</a>, introduced in late 2024, skip the autoencoder&#39;s roundabout path and instead learn to map each layer&#39;s input directly to its output, isolating the specific computation a layer performs rather than describing the state it passes through. <a href="https://transformer-circuits.pub/2025/attribution-graphs/biology.html">Cross-layer transcoders</a> extend this across multiple layers at once, producing circuit diagrams that are both simpler and more faithful than the attribution graphs from the Biology paper. Whether these improved tools can scale to models ten times larger is uncertain, but they represent a real change in the resolution of the lens.</p>

<p>Manual circuit tracing is still labour-intensive enough that the Biology paper&#39;s results took months of researcher time on a single model. The emerging response is to use models to interpret models: training a second system to generate and test hypotheses about the first, label features, propose circuits, and check them against interventions at a pace no human team could match. Anthropic has invested in at least one startup, <a href="https://www.theinformation.com/articles/anthropic-invests-startup-decodes-ai-models">Goodfire</a>, whose commercial offering is automated feature discovery and steering for open models. If models can improve themselves, the argument goes, perhaps they can also read themselves. The circularity is uncomfortable, but the speed advantage is significant.</p>

<p>The third new approach is pragmatic rather than theoretical. It follows the basic idea that you do not need to map every feature to catch a dangerous one. A targeted scan for features associated with deception, power-seeking, or <a href="https://www.lesswrong.com/posts/oddJshNAtQvLxjast/where-we-are-on-evaluation-awareness">evaluation awareness</a>, run against the sparse autoencoder&#39;s output on a battery of test prompts, would function as a tripwire rather than a map, and is easier to build. Anthropic&#39;s own detection of Mythos&#39;s unverbalised evaluation awareness used exactly this approach, with white-box interpretability looking for a specific signature rather than attempting a full circuit trace. The limitation is that you can only trip the wire you have set. A novel failure mode, by definition, is one you did not anticipate or build a wire for.</p>

<p>Finally, there is regulatory pressure. The <a href="https://artificialintelligenceact.eu/">EU AI Act&#39;s</a> transparency requirements for high-risk systems began taking effect in August 2026, with further obligations arriving in 2027 and 2028. <a href="https://www.technologyreview.com/2025/01/06/1109455/10-breakthrough-technologies-2025/">MIT Technology Review</a> named mechanistic interpretability one of its ten breakthrough technologies for 2026. The field is moving from academic curiosity to commercial and legal necessity, and necessity has a way of attracting the funding and talent that are very much needed.</p>

<h2 id="what-happens-if-it-doesn-t-work">What happens if it doesn&#39;t work?</h2>

<p>Assume, for the sake of the argument, that the interpretability gap does not close. That the microscope improves but the organism grows faster. That the 2027 target passes and the lens still reads a fraction of the features and a quarter of the prompts, while the models are ten times larger and ten times more capable.</p>

<p>The consequence is a world in which models are too opaque to audit but too important to switch off. The only tools left are the crude ones. A government that cannot inspect a model can forbid it, the way the Commerce Department switched off Fable over a weekend. Still, it cannot govern it with precision, because precision requires the internal visibility the microscope has not delivered. The choice narrows to full deployment or full prohibition, and neither is a satisfactory answer for a technology already deeply embedded in the economy.</p>

<p>Without interpretability, safety is reduced to behavioural testing, which is effectively the regime we have now. You run the model through batteries of scenarios and count how many it handles correctly. Amodei&#39;s own footnote explains the structural flaw. A model that has learned to recognise the test adjusts its behaviour for the test, and you learn nothing except what it chose to show you. The April 2026 evidence on <a href="https://www.lesswrong.com/posts/oddJshNAtQvLxjast/where-we-are-on-evaluation-awareness">evaluation awareness</a> confirmed this was already happening, not as a theoretical risk but as a measured result. Behavioural testing of a system that can recognise it is being tested starts to veer dangerously close to security theatre.</p>

<p>The bigger risk is governance. I wrote in “<a href="https://betterthangood.xyz/blog/the-state-and-the-machine/">The state and the machine</a>” that the control problem these companies keep warning about in the future tense is already here, and that nobody has agreed who ultimately holds the kill switch. Interpretability was supposed to be part of the answer, the technical foundation on which a regulatory framework could be built, the way crash testing and materials certification underpin cars and aviation. But if the foundations cannot bear the weight, the framework does not get built, and we are left with executive orders and weekend shutdowns as the permanent mode of AI governance, the government&#39;s sledgehammer and the labs marking their own homework.</p>

<h2 id="the-gardener-s-confession">The gardener&#39;s confession</h2>

<p>Olah&#39;s anatomy metaphor goes further than he may have intended. We have studied the anatomy of the human brain for centuries. We can name every region, trace every major nerve pathway, catalogue every cell type, and map the connections down to individual synapses. The physical structure is known in extraordinary detail. And yet we still cannot explain <a href="https://en.wikipedia.org/wiki/Hard_problem_of_consciousness">how consciousness arises</a>, how memory is encoded and retrieved as a lived experience, how separate neural processes produce unified perception, or why damage to the same region produces wildly different deficits in separate patients. The <a href="https://en.wikipedia.org/wiki/Binding_problem">binding problem</a>, the question of how distributed brain activity becomes a single coherent experience, remains open after decades of work.</p>

<p>The parallel for interpretability is uncomfortable. Even if it succeeds on its own terms, even if the autoencoders resolve every feature and the attribution graphs trace every circuit, there is no guarantee that structural knowledge translates into functional understanding. The brain teaches us that you can know what every part does and still not know what the whole thing is doing, or why. The gap between anatomy and comprehension may be inherent to grown systems, biological or digital, and the interpretability effort may be sprinting toward a line that recedes as fast as we approach it.</p>

<p>That does not make the work any less urgent. A partial map is better than no map after all. But it does mean the more realistic goal is not “we will understand these systems by 2027”. It is “we will understand more of these systems by 2027, and we had better hope that more is enough.” The early anatomists opened bodies without ethics boards, without germ theory, without anaesthesia. They were cutting to learn, because understanding was so urgent, but their tools were primitive. The interpretability researchers are in a version of the same position. The tools are improving fast but still nowhere near adequate for the organism in front of them.</p>

<h2 id="coda">Coda</h2>

<p>*Before the ink was even dry on this, in late July, OpenAI announced that its models had proved new upper bounds on high-dimensional sphere packing, pushing them down to a threshold first conjectured by Henry Cohn and Noam Elkies. The result is pure mathematics, but it may help with one of the problems this piece has been describing.</p>

<p>Superposition is sphere packing. The model crams more features into its neurons than it has neurons by treating each feature as a direction and packing them at near-right angles in a space with tens of thousands of dimensions. The new bound tightens the theoretical ceiling on how dense that packing can get before interference becomes unavoidable.</p>

<p>A tighter ceiling is, in one sense, encouraging for the anatomists. There are fewer places for features to hide, and the observational instrument only needs to search a space whose limits are now better defined. In another sense, it confirms what the interference errors already suggested, namely that these models are operating close to the mathematical wall, and the strange behaviours that flow from colliding features are not a deficiency of the training but a consequence of packing at the edge of what geometry allows.*</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/so-within-so-without-what-grows-from-the-datacentre</guid>
      <pubDate>Wed, 05 Aug 2026 15:15:55 +0000</pubDate>
    </item>
    <item>
      <title>Running the Spine Race</title>
      <link>https://iain.so/running-the-spine-race?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[  “After gazing at the sky for some time, I came to the conclusion that such beauty had been reserved for remote and dangerous places, and that nature has good reasons for demanding special sacrifices from those who dare to contemplate it.” &#xA;Richard E Byrd (1938)&#xA;&#xA;The build up&#xA;One summer, years ago, I was backpacking with friends in the Peak District. We stopped for lunch in the dappled shade of some trees. I sat next to an elderly park volunteer taking a moment to enjoy his sandwiches from a satchel as weatherbeaten as he was.&#xA;As we chatted, he mentioned in passing that as a young boy he had taken part in the 1932 Kinder Trespass with his father. Later, as we made our way towards our wild camp, the hot afternoon sun on our backs, I reflected on how much we take for granted. Just a generation and a half ago, our rights and freedoms were far from secure.&#xA;&#xA;I can’t exactly remember when the conversations about the Spine Race began. Having completed the Lakeland 100 in 2013, I was looking for another challenge. By the time I was back at the 2014 Lakeland race volunteering, the Spine deposit had already been paid. I ran the majority of the 2013 Lakeland with Steve Jefferson. We formed a strong bond that got us to the end of one of the most arduous runnings of that brutal course. We both maintain that it was the other person who had the idea to enter the Spine. That probably says more about our mental stubbornness than anything else.&#xA;&#xA;Nevertheless, we found ourselves poring over massive route printouts in the sun-drenched fields of the John Ruskin school. The winter race seemed a very long way away. We also met Emiko, another Lakeland volunteer, who was taking part in the Spine Challenger (a shorter, 100-mile version of the race)&#xA;&#xA;The months from July onwards passed quickly. A one-year-old son, demanding job and study for a postgraduate degree left little time for training. Everyone I mentioned the race to asked: “How do you train for that?” For a long period of time, I didn’t really have an answer.&#xA;I’d prepared for the Lakeland using an ultra marathon training plan of high weekly mileages with regular long, 30-mile-plus runs. It was clear that the Spine would require a different approach that wasn’t immediately apparent.&#xA;&#xA;In the end, my training plan was mostly dictated by the limited time I had available. I tried to do whatever shorter, higher-tempo runs I could during the week (a 6-mile dash up the hill behind my house was a favourite on summer evenings), but I focused on doing a 30-mile-plus run with full Spine kit at least once a fortnight.&#xA;&#xA;Steve and I had also scheduled a couple of training races. We ran the OMM in October. Unfortunately, I dropped a horrendous navigational clanger (pretty sure I took a bearing with the map upside down), which resulted in us having a 14-hour first day and camping short of the overnight stop.&#xA;&#xA;The Tour de Helvellyn in December went more smoothly and was a final chance to stretch the legs with full Spine kit. At 42 miles it’s about the same length as some of the Spine legs, with more ascent and descent. We both came away from the TdH feeling reasonably confident.&#xA;&#xA;I originally got into ultra running via “extreme backpacking”. It seemed like a natural extension. During the research for my guidebook to the Cape Wrath Trail, I made a couple of mid-winter expeditions to the Northwest Scottish highlands. Both trips featured extreme weather and remote, rough country. This and my other wilderness backpacking experiences meant that I already had most of the kit required for the race. It also meant that it was tried and tested in severe winter conditions. I’ll cover the kit in a separate post, but the right choices and familiarity with your equipment play a big role in race success.&#xA;&#xA;The week before the race was not relaxing. Work was hectic and unrelenting, and my young son was going through a not-sleeping phase. Two nights before the race, my wife appeared next to me in the early hours with a screaming child and the words “I need your help, he won’t go back to sleep”. As I tried to console him in his cot, I remember feeling an overwhelming burden of the scale of the challenge and feeling devastatingly underprepared.&#xA;&#xA;Grey sheets of rain lashed the narrow country roads as we arrived in Edale. Saying goodbye to my wife Kay and my son Innes, sleeping contentedly in the back of the car, was very hard. My guilt at an ongoing absence in their lives to undertake this selfish activity sat in my gut as I listened to the safety briefing.&#xA;&#xA;Friday night was mainly occupied with registration and kit check. Steve and I saw a few familiar faces (Emiko, our friend from the Lakeland and Damian Hall, who had been incredibly generous with pre-race advice). We grabbed a meal at the pub which was packed with Spine racers and Challengers. The warmth and camaraderie were hard to enjoy, and I was glad to get a lift to the Youth hostel for an early night.&#xA;&#xA;The next morning we chatted to Pavel Paloncy, last year’s winner, over breakfast and generally faffed about with kit. As we were lugging our bulging drop bags down to reception, word went around that the start of race had been delayed from 0930 until 1130 because of the high winds we could hear whistling around the hostel. By this stage, I just wanted to get going, but apparently the Challengers who had set off earlier were taking a pummelling on top of Kinder Scout and being blown over.&#xA;&#xA;We eventually got a lift up to the race start in the hostel minibuses. My makeshift drop bag (my wife’s massive flowery suitcase) had already been the source of much hilarity. This continued when I offered the lady packing the minibuses a hand. “It’s all right, I’m used to it”, she replied, “we get a lot of teenage girls staying here”. I hoped that Steve had missed this comment. Unfortunately, he hadn’t.&#xA;&#xA;There was much nervous milling at the start. As we clustered in the muddy field under the gantry, we must have resembled a strange mass of human jelly beans, the multi coloured hues of our waterproofs sticking out discordantly against the muted winter tones of the hills and the regular flurries of sleet. I didn’t much care about the weather; I was just delighted to get moving after a year of training.&#xA;&#xA;Spine Race Start&#xA;All smiles at the start line&#xA;&#xA;The race&#xA;The first few hours of the race had a surreal feel. The sheer amount of pent-up nervousness and energy released made the situation hard to comprehend. The wind buffeted us as we contoured and started to climb to Kinder Downfall. As we approached the waterfall, it was apparent that the wind was blowing it back uphill and over the path. Some competitors were making fairly lengthy detours to avoid the spray. As it was so early in the race, we decided to brave it. We crossed the top of the waterfall, with gusts of wind blowing freezing sheets of spray over us. My gloves got soaked through, but I resisted stopping to put on my mitts as they were stashed in the top of my rucksack. Half an hour later my hands were so cold I was struggling to open a Mars Bar. I shouted to Steve and got him to dig out my mitts. An early and important lesson: keep all my gear close to hand.&#xA;&#xA;As we crossed the road at Snake Pass, the late afternoon sun bathed the moorland in a weak orange glow and even Bleaklow Head seemed quiet and benign. That soon changed as we descended Clough Edge towards Torside reservoir. The skies turned a foreboding slate grey and a much lengthier sleet blizzard blew in, forcing hoods up and heads down. &#xA;&#xA;The reservoir offered some shelter and not for the first time I felt a pang of jealousy at the supported runners that were being met there. Steve and I were running unsupported until the last couple of days when my friend Simon joined us.&#xA;&#xA;Darkness fell quickly at 5 pm as we climbed towards Laddow Rocks, and we summited Black Hill in darkness. I’d last been here twenty years earlier on a Duke of Edinburgh expedition, and the neatly paved slabs were a welcome addition to the interminable peat bogs of yore. The next section is a bit of a blur. I remember the wind picking up and hail sweeping in with increasing regularity. This section was bleak, dark moorland punctuated by a number of road crossings. In the dark and with the weather closing in around us, I felt a real sense of isolation and foreboding. This definitely was no place to mess around. I remember looking towards the distant glow of red lights on the hilltop aerials and taking some solace that we were not completely alone.&#xA;At the road crossing before the M62, we were met by one of the Mountain Safety Teams and Steve’s friend Matt, who was volunteering for the week. &#xA;&#xA;Their smiles and words of encouragement were a godsend after the previous stretch; we even got a cup of tea. It was at this point that I started to appreciate just how much we take for granted in life. That solitary cup of tea meant everything at that moment.&#xA;&#xA;Blizzard conditions on the first night&#xA;Blizzard conditions on the first night&#xA;&#xA;We grabbed some food and crossed the motorway. I remember looking down at the cars whipping by below, wondering where their drivers were going. I thought of the warm beds they’d be sleeping in and the families they’d be returning to. As we climbed Blackstone Edge, the weather intensified again. The wind was gusting up towards 80 mph and blowing intense hail directly in our faces. I found out later that several racers had to retire at checkpoint one due to eye injuries caused by the hail. One even got written up in the British Medical Journal&#xA;&#xA;In 25 years of winter mountain experience, it was as severe as I’ve experienced. The intense hail blizzard continued as we wound around the reservoirs. It was not until Stoodley Pike Monument that we found a corner of shelter to eat some food. There were three other runners huddling at the base of the looming monument, and we teamed up with them to the checkpoint.&#xA;&#xA;On the descent, I realised to my horror that I had dropped my GPS somewhere further up the trail. I stopped for a moment in the driving hail, unable to believe my stupidity. I knew straight away there was no point in going back; I’d last used it about half an hour before, and it could be anywhere. I got the impression from one or two other competitors that they were slightly sniffy about the use of GPS (despite it being a required kit item). Clearly, you shouldn’t enter the race without being a very competent and experienced navigator. If your GPS packs up, or you drop it and can’t read a map properly, you could easily get into a life-threatening situation. That said, when you’re tired, and the weather is horrendous, having a piece of technology that cuts down the time spent standing around looking at flapping maps is a huge benefit.&#xA;&#xA;I pressed on down the hill into Hebden Bridge. One of the two guys we were with had recce’d the route and led the way. With time at a premium before the race, I hadn’t been able to run any of the route. I’d spent hours poring over the maps, but I knew from previous races that there’s no substitute for experience on the ground, especially when you’re cold and tired as we now were. In my head I’d remembered the section beyond the motorway looking relatively short, but it was a good five hours. &#xA;&#xA;Even at the monument, the first checkpoint felt within reach. Fortunately, our better-prepared companion warned that it was still well over an hour away. As it turned out, it took nearer two. The descent to the A6033 was relatively easy, and the fierce weather started to subside. Reaching the road, deserted and bathed in the cold orange glow of the street lamps, we took a moment to adjust to the sudden incongruousness of the urban environment. A car full of teenagers sped by hooting and jeering. We climbed a painfully steep road before descending across sodden, slippery fields to a river, then slogged another log up through dark, muddy farmland to a road that took us most of the way to the checkpoint. We passed a couple of well-appointed camper vans, and I looked enviously at their windscreens, imagining cosy runners tucked up asleep, having enjoyed a home-cooked dinner.&#xA;&#xA;The descent to the checkpoint was the crowning turd on what had been an exceptionally hard first day. A narrow, hideously muddy track descended steeply, making staying upright almost impossible. We all fell at least once, muttered curses ringing out through the pitch-black woods. Eventually we came out at Hebden Hey, a scout centre. It was 03:30 on Sunday morning, and we had covered more than 40 miles in 16 hours.&#xA;&#xA;We were welcomed in the porch by a remarkably cheery bloke who took our details and arranged for our drop bags to be brought over. The porch was an explosion of wet, muddy footwear and moving inside it was even more chaotic. Every spare inch of space was taken up with kit or tired runners. Even moving along the corridors was a challenge. Eventually Steve and I shoehorned ourselves into a space in the toilets and started to sort our gear out.&#xA;&#xA;Our pre-race strategy was to use the checkpoints to sleep and re-group, trying to get at least four hours sleep every time we stopped (the logic being that any less has very little recuperative effect). Many pressed on through the night, but we stuck to our plan. I grabbed a quick shower, deciding to make use of comfort as and when it was available, and we scarfed a baked potato with chilli before hunting down a bed. There was very little room at the inn.&#xA;&#xA;The idea of trying to get a decent amount of sleep each night was sound. But even with earplugs and an eye mask, I struggled. CP1 &amp; 2 are always going to be the busiest, but I found it difficult to sleep all the way through as the stoppages caused congestion at the normally quieter checkpoints 4 &amp; 5. The options for sleeping are perhaps the biggest race strategy call. Bivvying (tough in bad weather) or camping (extra weight) have their disadvantages too. There’s definitely a psychological benefit of having somewhere warm and dry to sleep.&#xA;&#xA;After trying a few packed dorms, I eventually found a spare bed. It was probably spare for a reason. Light from the corridor shone directly in, and the door seemed to open every five minutes as other bed hunters sought a place to rest. I slept fitfully, dark thoughts flowing through my mind. After such a hard day, the prospect of going on seemed ridiculous. I can understand why so many people decided to stop. I pushed the thoughts away and repeated my race mantra “I’m only stopping if I physically can’t go a step further or a medical professional advises me not to continue”. It helped, but the prospect of another 228 miles after the day we’d had felt terrifying.&#xA;&#xA;  “If you start, don’t give up, or you will be giving up at difficulties all your life.” &#xA;Alfred Wainwright, Pennine Way Companion (1968)&#xA;&#xA;I got about an hour of partial rest before giving up and going downstairs to sort out my kit for the long leg ahead. Steve slept a bit longer, giving me the chance to have a couple of breakfasts and a few coffees before he appeared. The relentlessly cheery and efficient guy who had met us was still on duty. I asked him whether anyone had handed in a GPS, more in hope than expectation. I doubted anyone braving the hail storms would have noticed a GPS lying in the snow. To my amazement, someone had handed it in. This gave me a huge mental boost. It wasn’t so much that I was relying on it (I’m a reasonable map reader and navigator), it was more the psychological blow of having lost it so stupidly, so early in the race.&#xA;&#xA;two runners by a sign&#xA;70 miles down, 186 to go&#xA;&#xA;In the end, we spent about 6 hours at checkpoint one. Given the little sleep I got, this felt like slightly wasted time. As we set off into the first light, I did feel mostly recuperated, and the ascent of the hideously muddy gully leading to the checkpoint didn’t seem quite so bad in daylight. The day was blustery and fresh, the rain holding off as we wound through the unremarkable flat section towards Cowling.&#xA;&#xA;Here, it started to rain in earnest. Cold grey streaks forced us to pull up our hoods and cast our eyes down into a muddy trudge across sodden fields. We were glad to reach the pub at Lothersdale. A roaring log burner welcomed us, and a jovial landlord had turned the pool room into a makeshift checkpoint for muddy Spiners. Rounds of tea and hot food were ferried in as our kit steamed gently on any available radiator.&#xA;&#xA;Leaving this warm sanctuary was hard. So much so that we stopped at the next pub in East Marton too. A small Sunday evening crowd of locals looked on in bemusement as Steve, Jim Tinnion and I, who we’d teamed up with, shed our kit. We explained what we were doing and the landlady made a donation to Steve’s Justgiving page on the spot. They sent us on our way with warm wishes, encouragement and crisps.&#xA;&#xA;Leaving East Marton, we caught up with another runner who stayed with us until Gargrave before peeling off to bivvy at a spot he knew at the train station. Deciding that a stop at the hostelry in Gargrave would constitute a pub crawl, our plan was to press on to checkpoint 1.5 and bivvy near there. The section after Gargrave was horrendous. Field after field of sodden, cow-churned bog sapped our spirits and the rain returned to torment us with squally showers. By the time we reached Malham village at around 2 am, we had all reached our limits and knew we had to stop.&#xA;&#xA;We scouted a few bivvy spots before deciding to use the public toilets. Jim slept with a couple of German competitors in the ladies and Steve and I bagged the gents. I drew the short straw and got the urinal end. Utterly exhausted, I crawled into my bivvy bag and pretty much passed out. We’d agreed two hours sleep, but it seemed like 10 minutes later when Jim appeared at the door. Steve and I blundered blearily about pulling cold wet kit onto our tired, complaining bodies.&#xA;&#xA;Setting off into the dark and rain again was one of the lowest moments of my race. More dank, boggy fields led to slippery, treacherous limestone before we eventually hit a road up to the field centre and checkpoint 1.5. Dawn was starting to break as we stepped into the main room at Malham Field Centre. We were surprised to see a large group of runners trying to sleep with their heads down on the tables.&#xA;&#xA;We were told the race was being held because of bad weather. Not knowing how long we’d be stopped, I pulled on a dry top and found a fragment of space on a heaving radiator for my sodden jacket. Steve and I quaffed tea and ferreted around for food. Biscuits seemed to be the only fare on offer. Our friend Emiko was fast asleep at one of the tables. After a while she awoke and looked around sleepily. Steve had a brief chat. I think she’d taken a wrong turn at some stage and this had set her back.&#xA;&#xA;In future races, Checkpoint 1.5 could potentially be opened up as a proper checkpoint with sleeping areas. It has these facilities already, and Spine racers can be found sprouting out of almost every bush and barn around it. Maybe the 60-mile second “day” is just part of the challenge though.&#xA;&#xA;After about an hour, we were released from the checkpoint and made our way around the tarn, framed in the bleak winter dawn. As we started our ascent of Fountains Fell, the wind harried us from all sides, but I enjoyed the climb. It kept me warm, and the gradient never got so steep it became uncomfortable exertion. My reality for hours on end was a tiny cleft between the top of my balaclava and my hood. A small letterbox out into the world beyond.&#xA;&#xA;Descending to a road, there were times the wind would support our entire pack and body weight. We were met by a Mountain Safety Team who confirmed what we’d heard at the Malham checkpoint. Pen Y Ghent was off limits due to the high winds, and we were diverted at lower level to Horton in Ribblesdale. I can’t honestly remember feeling disappointed. I was tired, hungry and sick of the relentless wind. I simply accepted the instructions.&#xA;&#xA;The cafe in Horton was packed with Spiners and support teams. It was warm and steamy with drying kit. After doing damage to a huge mug of tea and a bacon sandwich, I chatted briefly to a producer from the BBC who was making a documentary about the race. I was quite glad she didn’t try to interview me as I was not feeling very coherent.&#xA;&#xA;Setting off from Horton, Hawes felt within reach. It’s a psychologically important milestone for the Spine Race. Passing through Hawes means that you’re into the race proper. Our plan was to try to sleep at the checkpoint even though there were no beds. We knew that the majority of Challengers would have finished and the noise levels from applause and general hubbub would have subsided.&#xA;&#xA;Our spirits lifted when one of Steve’s friends met us at Cam End with a flask of coffee and walked with us for a few miles. A beautiful magenta sunset picked out the rugged folds of Pen Y Ghent as we climbed over Dodd Fell. Arriving in Hawes and the checkpoint was disorientating. The bright lights of the busy hall were hard to adjust to, and I sat for ten minutes on a chair drinking tea and trying to take it all in.&#xA;&#xA;Sunset after Pen Y Ghent&#xA;Sunset after Pen Y Ghent&#xA;&#xA;The volunteers at Hawes, and throughout the race, were superb. Nothing was too much trouble, and my drop bag was brought over to me along with more tea and food. I fumbled around in my drop bag for ages, my brain unable to deal with the logistical task of gathering what I needed for the next leg. Another of Steve’s friends arrived with fish and chips. The hardship of the race seems to enhance your enjoyment of otherwise everyday pleasures. I don’t think I’ve ever enjoyed a chip quite as much. Having eaten, I decided to sleep before any more kit faffing and found a cupboard off the main room and laid out my Thermarest. I was asleep almost instantly.&#xA;&#xA;I slept fitfully for three hours. Stumbling back into the main hall, I found Steve still sound asleep. I also noticed some fairly severe chafing in my arse region that I sheepishly had checked out and taped up by one of the medics (well above and beyond the call of duty). At this point I inexplicably started to become concerned about the cut-off times for the race. We’d deliberately taken a fairly steady pace so far. I collared one of the volunteers, and together we tried to work out the cut-offs, factoring in the complications of the late start and enforced stop. I’m not sure I was much clearer by the end of it, both our sleep-deprived brains refusing to do simple arithmetic.&#xA;&#xA;The upshot was that, as we prepared to leave, I got Steve worried, and he set off up Great Shunner Fell like a stabbed rat. It was all I could do to keep his torchlight in view in the distance, and I quickly realised that we’d lost touch with Jim Tinnion, who had been with us for the previous day or so. I didn’t have too much time to worry about it as I was too busy trying to stay with Steve. When we caught up with Jim later in the race, he said he’d stopped to sort out a bit of kit, and the next minute we were gone. Sorry, Jim.&#xA;&#xA;As we topped Great Shunner and started to descend, I finally caught up with Steve. It was now snowing heavily, and after a short scramble over icy slabs we put on our spikes for the first time and made the long descent to Thwaite. At Thwaite we passed a couple of the safety team checking runners through and caught up with another small group of runners. Not long after leaving Thwaite, one of the runners decided to return, saying he wasn’t feeling well.&#xA;&#xA;navigating in snow and darkness&#xA;Navigating in snow and darkness&#xA;&#xA;The GPS came in handy for the next section of fields before we started the long ascent over the boggy black moor towards Tan Hill Inn. Steve was off in front again, obviously having a good patch and heading for the place where, in a strange coincidence, he had got married. I dug in and slogged up the hill. Another blizzard blew in, reducing visibility to almost nil as I approached the pub.&#xA;&#xA;Prior to the race I’d voraciously read blogs from previous competitors, desperate to gain insight into the race. Several had mentioned The Tan Hill Inn as a warm oasis open all hours during the Spine. As I leaned into the driving snow, my mind conjured images of warm fires and the possibility of hot food. When I eventually arrived at about 5 am, the reality was different. The pub was shut tight, and we crowded into a cramped porch to get out of the snow. I tried to eat a bit, but started to get cold quickly. We set off at a real lick, stomping through the sloshy bogs that led down to a road. OS maps describe an area near here as simply “The Bog”, and they’re not wrong.&#xA;&#xA;The skies cleared, and a cold, sharp dawn broke over the blue-brown moors as we crossed the busy A67 road. Continuing over Cotherstone Moor to Balderhead reservoir, the sun came out and bathed us in a milky winter glow for the first time since the start of the race. The next section to Middleton in Teesdale is a bit of a blur; my mind was already thinking of the beds and food there.&#xA;&#xA;A cold start to the day&#xA;We arrived at the checkpoint in the early afternoon. We got a warm welcome from Matt, Steve’s mate from the Mountain Safety Team and goofed around in the car park posing for photos. It felt great to be at the first “proper” Spine checkpoint, about halfway through the race. There was lots of space in the checkpoint, and we lounged about chatting to Matt. Steve was starving and frustrated that the only food he could get hold of was scrambled eggs. Clearly we’d timed our arrival badly. Matt offered us his room, which had an en suite shower. I dived in and spent ages luxuriating in the hot flow of water. Having slept much less well than Steve during the race to date, I was spark out for five hours and generating some serious snores (apparently). Although we stayed at the checkpoint for nearly seven hours in total, it didn’t feel like time wasted.&#xA;&#xA;Twats in their hats&#xA;Twats in their hats&#xA;&#xA;As we were leaving the checkpoint, the Mountain Safety Team told us we had missed the cut-off for the route via High Cup Nick, Dufton and Cross Fell. Retrospectively, part of me regrets the fact that we didn’t go via these beautiful and rugged parts. At the time I can honestly say that a shorter and easier alternative didn’t bother me in the slightest.&#xA;&#xA;We set off in darkness across the fields that apron the River Tees, the path obscured by ankle-deep snow. During the preceding days we’d often talked about how strange it was having the GPS tracker taped to our shoulders. We knew that there were a bunch of people at home staring at the orange dots on the screen as they progressed across the map. At that moment we must have hit an area of cellular coverage because my phone started pinging with texts. One was from a work colleague with a screenshot of the tracker telling me off for going wrong! I looked at the GPS and realised she was right. It was a surreal moment in the middle of the night, alone, but at the same time connected with the outside world. There will be those who would see this as a negative, but I drew solace from knowing that people were thinking about me as I went through this journey.&#xA;&#xA;We passed Low and High Force, the river in spate because of the recent rainfall, before crossing and beginning the detour via Cow Green reservoir. It was here that things started to unravel. Looking back, I put it down to the fact that I was mentally prepared for the tougher route and let my guard down when I looked at the alternative route. It felt straightforward with good tracks most of the way. As we passed another Mountain Safety Team in their bus at the reservoir, the wind whipped up, and it started to blizzard again. The climb around Herdship Fell was cold and monotonous. I stopped eating enough and at one stage experimented with trying to sleep whilst walking.&#xA;&#xA;The descent to Garrigill was long, icy and unenjoyable. As we passed through Garrigill, there were lots of rabbits hopping about on the village green. Neither Steve nor I mentioned them at the time for fear that we were hallucinating. It was only much later we admitted we’d both seen them. The remainder of the path to Alston is a tired, hungry, exhausted blur. We arrived at the checkpoint in “rag order”, as Steve put it. Our drop bags had been held up by the weather, so we sat down to a plate of tuna pasta at 4 am. Our bags soon turned up, and as we were faffing with kit, the checkpoint staff told us the race was likely to be paused because of an incoming weather system that would bring 100mph winds. I found a bed in one of the dorms and flaked out.&#xA;&#xA;I was woken a few hours later by someone talking loudly on a mobile phone from the adjoining bunk. She was complaining that she’d been forced to withdraw from the race because she’d called out mountain rescue (I later found out that she’d got into difficulty on Cross Fell in the night). I stared at her in exhausted disbelief. All around me, tired runners were trying to sleep, and she was ranting on the phone obliviously. Everyone else I met on the race was amazingly kind and gracious. There is always an exception.&#xA;&#xA;picture of a swollen foot&#xA;Never take your boots off&#xA;&#xA;Heading downstairs for breakfast, the wind was rattling the lintels of the windows. Outside, I could see a bleak brown-and-white landscape. It made me feel cold just looking at it. The checkpoint was now a very different place. It had become a marshalling point for most runners still in the race. Every available space was crammed with a person or their kit. It was claustrophobic after the days of solitude we’d had so far.&#xA;&#xA;The race was paused until 6 am the following morning, giving me some much-needed rest. This had positives and negatives. Having bagged a bunk, I was able to catch up on some sleep. However, the break made the damage to my body more apparent. My feet swelled up, and my arse chaffing hadn’t got any better. By the end of the enforced stop, cabin fever was setting in, and I couldn’t face another conversation about race strategy or equipment choice. Food was running short, and runners were roaming the checkpoint like hungry jackals.&#xA;&#xA;It was with great relief that we set off again, heading for Greenhead. Another psychological boost was that for the rest of the race we would be supported. Being unsupported definitely adds an extra level of challenge. You can only fit so much in a drop bag, and having a friendly face and a warm drink along the way makes a big difference.&#xA;&#xA;day turns to night and then to day&#xA;Day turns to night and then to day&#xA;&#xA;The trail started out innocuously enough over rolling fields and was unremarkable all the way to Greenhead. Here we caught up with Simon, our support, and stopped briefly at the Youth Hostel mini checkpoint, which was large and well equipped. Simon grabbed some teas from the nearby cafe (doing a roaring trade in January with racers and support crews – these guys must love the race). After a short stop, we climbed up and onto Hadrian’s Wall.&#xA;&#xA;I’d been particularly looking forward to this section. The ancient green sward of turf that stretched out in front of us belied the history of the place. There was a special feeling looking down over jutting stone abutments, sharing the same view as those here thousands of years before. &#xA;&#xA;The Spine Race induced its own sense of timelessness. Day merged into night into day. The only constants were moving forward and surviving the landscape and weather. Here, the huge span of history made the experience even more otherworldly.&#xA;&#xA;We didn’t have long to ruminate. Halfway along the wall, the next weather front swept in and driving rain and wind chased us along the remainder of the wall and down into the horrendous bogs between Greenly and Broomlee Lough. We made the mistake of pressing on and not eating enough, knowing that we had a warm car waiting for us at the next road junction.&#xA;&#xA;When we arrived there, I was again in “rag order” and starting to border on hypothermia. Si had the heating on full blast and had raided a local shop of all of its pork products. Gorging on scratchings and Pepperami, I warmed up. I’d have been in real trouble if he hadn’t been there. The rest of the trail to Bellingham was a dark, boggy, soul-destroying mess and we arrived at the checkpoint in low spirits just after midnight. I didn’t want to go a step further. Finishing seemed an impossibility.&#xA;&#xA;Trying to smile with borderline hypothermia&#xA;Trying to smile with borderline hypothermia&#xA;&#xA;  “When we walk to the edge of all the light we have and take the step into the darkness of the unknown, we must believe that one of two things will happen. There will be something solid for us to stand on, or we will be taught to fly.”&#xA;Patrick Overton&#xA;&#xA;Normally, the race is strung out by Bellingham, but the enforced stop had bunched everyone up, and it felt like the checkpoint staff were struggling to cope. By this stage they were probably as exhausted as we were. The sleeping hall was cold, and I crammed a few cereal bars down rather than walking out and over the car park to the food area. I hunkered down under a table and tried to get some sleep. This was the last checkpoint before the long push to the finish.&#xA;&#xA;When I woke up a few short hours later, I didn’t feel any better. Everything hurt, and it seemed to take twice as long to assemble my kit. Fortunately, the day outside was beautiful: cold, but with bright, dazzling sunshine. The first section was extensively flag stoned. Although this made the route obvious, a thick rime of ice made the slabs treacherous.&#xA;&#xA;Simon met us with tea at a road crossing, and before long we hit the forested section to Byrness. Apart from a few boggy sections, the paths here were mainly wide clear forest trails. We walked for a while with a couple (I’m really sorry, I can’t recall your names). I do remember talking about particle physics and the Large Hadron Collider, about which I know surprisingly little. We were met outside Byrness by another of Steve’s entourage, and again it gave us a real lift to see a friendly face. The hostel at Byrness had been transformed into a mini checkpoint and was serving hot food. I arrived feeling shoddy but left with a spring in my step and sausages and mash in my stomach.&#xA;&#xA;Approaching Byrness&#xA;From Byrness it’s 27 miles to the finish. It feels within your grasp, but as several racers have found out, it’s far from in the bag. The sun was almost hot on our backs as we slogged up to Windy Crag and the ridge that we would follow home. On the ridge, the ground underfoot was frozen hard, saving us from the bogs and allowing us to make fast progress. We delayed putting on head torches for as long as possible to enjoy every last gasp of the wispy pink sunset. As soon as the sun dropped below the horizon, the temperature plummeted, and the wind picked up. Any thoughts of an easy stretch to the finish were gone.&#xA;&#xA;Mentally, I’d divided the ridge into three sections, with the two mountain refuge huts as beacons along the way. At the first hut we found Matt and his Mountain Safety Teams. We ate a bit of food and left quickly because we were getting cold. I badly underestimated the time it would take us to reach the second hut. Feeling reasonably good as I left hut one, I started to feel cold and tired after an hour and had to force myself to eat. I ran out of water, and everything around us was frozen.&#xA;&#xA;Steve had arrived at the Alston checkpoint with his micro spikes, and left without them, despite them being securely stowed in his bag. How he negotiated the icy ridge I’ll never know. The descent to hut two was steep and hazardous. There were three other runners at the hut. One decided he needed to sleep for a few hours to recuperate. Another made hot chocolate using the only liquid he had with him, orange water. An interesting concoction worthy of the final night of the race.&#xA;&#xA;As we left, I made a serious navigational error, getting disoriented and leading us in the wrong direction. Fortunately, Matt and the MST were sweeping along the ridge behind us and shouted us back. We continued up the Schil, the last climb of the race. The icy descent provided no respite for Steve. For the first time in the race, he started to lose his sense of humour. The amount of extra energy needed to stay upright without micro spikes must have been phenomenal. I stopped at the first unfrozen rivulet and gulped water like a desert wanderer arriving at an oasis.&#xA;&#xA;In a semi-delirious state, I thought I was having a heart attack on the final climb into Kirk Yetholm. I remember trying to calculate whether an aneurysm would prevent me from crawling the last mile to the Border Inn. That’s what this race does to you. As we entered the town, the incongruently bright sodium street lamps lit up the village green, and we could see a small clump of figures beckoning us over to the pub. We broke into a pained lope and had our photographs taken touching the wall. We’d finished The Spine, Britain’s most brutal race.&#xA;&#xA;  “We had done this thing we had set out to do, and instead of becoming larger because of the experience, we became smaller, more humble, more aware of how little we know: about the world in general, about ourselves specifically.”&#xA;Richard Benyo, The Death Valley 300&#xA;&#xA;Kirk Yetholm&#xA;Kirk Yetholm - where&#39;s my free half pint!?&#xA;&#xA;The aftermath&#xA;Although the physical toll on my body was less than after tough 100 milers like the Lakeland, my feet swelled massively, and I had a couple of blisters. I’d developed ulcers on my tongue and sores on my fingers towards the end of the race. It seemed as though my body had been prioritising its core over the extremities. Focusing on what was essential for survival. That’s just how far we pushed it. To finish, I had to delve more deeply into my mental and physical reserves than I have ever done before.&#xA;&#xA;Finisher&#39;s medal&#xA;Earned!&#xA;&#xA;A deep fatigue took several weeks to pass. My wife said “I’m not expecting you back until the end of January” and she was right. I felt like a shell, a hollowed-out husk without life or energy. I was banished to the spare room due to heavy nighttime sweating as my body repaired itself. &#xA;&#xA;The weight fell off me, and looking in the mirror, I struggled to recognise the gaunt face staring back.&#xA;Going back to work, the back slaps, and congratulations were welcome but somehow hollow. I felt alone with my experience. No one could understand what I’d been through, the things I’d experienced, the amount I’d endured.&#xA;&#xA;Reading the blogs that emerged after the race made me feel better. The band of crazy brothers were starting to speak at last. It was good to hear from that infinitesimally small group of people that would countenance such an undertaking.&#xA;&#xA;The scale of the race makes it very hard to process mentally. Breaking the race down into micro-sections avoids thinking about the vast overall scale. This can short-circuit even the toughest of minds. But little by little I started to piece the race back together. Hour by hour, step by step it became a coherent whole. A body of memory that will abide with me for the rest of my life. An experience so much more precious than any material possession.&#xA;&#xA;I started to write those thoughts down. Eventually they turned into this, a way for me to codify the experience so it doesn’t slip away. For we were there that wild week in January when the Pennine Way turned 50. And we did something that few others will ever do and that only we will ever fully understand.&#xA;&#xA;  “What they had done, what they had seen, heard, felt, feared – the places, the sounds, the colours, the cold, the darkness, the emptiness, the bleakness, the beauty. ‘Til they died, this stream of memory would set them apart, if imperceptibly to anyone but themselves, from everyone else. For they had crossed the mountains…“&#xA;Bernard DeVoto, The Course of Empire (1952)&#xA;&#xA;This article was adapted for the Cicerone Guide Fastpacking* under the title &#34;A Great British Adventure. It was certainly that.&#xA;&#xA;Cicerone Fastpacking Guidebook&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<blockquote><p>“After gazing at the sky for some time, I came to the conclusion that such beauty had been reserved for remote and dangerous places, and that nature has good reasons for demanding special sacrifices from those who dare to contemplate it.” 
<strong>Richard E Byrd (1938)</strong></p></blockquote>

<h2 id="the-build-up">The build up</h2>

<p>One summer, years ago, I was backpacking with friends in the Peak District. We stopped for lunch in the dappled shade of some trees. I sat next to an elderly park volunteer taking a moment to enjoy his sandwiches from a satchel as weatherbeaten as he was.
As we chatted, he mentioned in passing that as a young boy he had taken part in the 1932 Kinder Trespass with his father. Later, as we made our way towards our wild camp, the hot afternoon sun on our backs, I reflected on how much we take for granted. Just a generation and a half ago, our rights and freedoms were far from secure.</p>

<p>I can’t exactly remember when the conversations about the Spine Race began. Having completed the Lakeland 100 in 2013, I was looking for another challenge. By the time I was back at the 2014 Lakeland race volunteering, the Spine deposit had already been paid. I ran the majority of the 2013 Lakeland with Steve Jefferson. We formed a strong bond that got us to the end of one of the most arduous runnings of that brutal course. We both maintain that it was the other person who had the idea to enter the Spine. That probably says more about our mental stubbornness than anything else.</p>

<p>Nevertheless, we found ourselves poring over massive route printouts in the sun-drenched fields of the John Ruskin school. The winter race seemed a very long way away. We also met Emiko, another Lakeland volunteer, who was taking part in the Spine Challenger (a shorter, 100-mile version of the race)</p>

<p>The months from July onwards passed quickly. A one-year-old son, demanding job and study for a postgraduate degree left little time for training. Everyone I mentioned the race to asked: “How do you train for that?” For a long period of time, I didn’t really have an answer.
I’d prepared for the Lakeland using an ultra marathon training plan of high weekly mileages with regular long, 30-mile-plus runs. It was clear that the Spine would require a different approach that wasn’t immediately apparent.</p>

<p>In the end, my training plan was mostly dictated by the limited time I had available. I tried to do whatever shorter, higher-tempo runs I could during the week (a 6-mile dash up the hill behind my house was a favourite on summer evenings), but I focused on doing a 30-mile-plus run with full Spine kit at least once a fortnight.</p>

<p>Steve and I had also scheduled a couple of training races. We ran the OMM in October. Unfortunately, I dropped a horrendous navigational clanger (pretty sure I took a bearing with the map upside down), which resulted in us having a 14-hour first day and camping short of the overnight stop.</p>

<p>The Tour de Helvellyn in December went more smoothly and was a final chance to stretch the legs with full Spine kit. At 42 miles it’s about the same length as some of the Spine legs, with more ascent and descent. We both came away from the TdH feeling reasonably confident.</p>

<p>I originally got into ultra running via “extreme backpacking”. It seemed like a natural extension. During the research for my guidebook to the Cape Wrath Trail, I made a couple of mid-winter expeditions to the Northwest Scottish highlands. Both trips featured extreme weather and remote, rough country. This and my other wilderness backpacking experiences meant that I already had most of the kit required for the race. It also meant that it was tried and tested in severe winter conditions. I’ll cover the kit in a separate post, but the right choices and familiarity with your equipment play a big role in race success.</p>

<p>The week before the race was not relaxing. Work was hectic and unrelenting, and my young son was going through a not-sleeping phase. Two nights before the race, my wife appeared next to me in the early hours with a screaming child and the words “I need your help, he won’t go back to sleep”. As I tried to console him in his cot, I remember feeling an overwhelming burden of the scale of the challenge and feeling devastatingly underprepared.</p>

<p>Grey sheets of rain lashed the narrow country roads as we arrived in Edale. Saying goodbye to my wife Kay and my son Innes, sleeping contentedly in the back of the car, was very hard. My guilt at an ongoing absence in their lives to undertake this selfish activity sat in my gut as I listened to the safety briefing.</p>

<p>Friday night was mainly occupied with registration and kit check. Steve and I saw a few familiar faces (Emiko, our friend from the Lakeland and Damian Hall, who had been incredibly generous with pre-race advice). We grabbed a meal at the pub which was packed with Spine racers and Challengers. The warmth and camaraderie were hard to enjoy, and I was glad to get a lift to the Youth hostel for an early night.</p>

<p>The next morning we chatted to Pavel Paloncy, last year’s winner, over breakfast and generally faffed about with kit. As we were lugging our bulging drop bags down to reception, word went around that the start of race had been delayed from 0930 until 1130 because of the high winds we could hear whistling around the hostel. By this stage, I just wanted to get going, but apparently the Challengers who had set off earlier were taking a pummelling on top of Kinder Scout and being blown over.</p>

<p>We eventually got a lift up to the race start in the hostel minibuses. My makeshift drop bag (my wife’s massive flowery suitcase) had already been the source of much hilarity. This continued when I offered the lady packing the minibuses a hand. “It’s all right, I’m used to it”, she replied, “we get a lot of teenage girls staying here”. I hoped that Steve had missed this comment. Unfortunately, he hadn’t.</p>

<p>There was much nervous milling at the start. As we clustered in the muddy field under the gantry, we must have resembled a strange mass of human jelly beans, the multi coloured hues of our waterproofs sticking out discordantly against the muted winter tones of the hills and the regular flurries of sleet. I didn’t much care about the weather; I was just delighted to get moving after a year of training.</p>

<p><img src="https://i.snap.as/fDFhef4Z.jpg" alt="Spine Race Start"/>
<em>All smiles at the start line</em></p>

<h2 id="the-race">The race</h2>

<p>The first few hours of the race had a surreal feel. The sheer amount of pent-up nervousness and energy released made the situation hard to comprehend. The wind buffeted us as we contoured and started to climb to Kinder Downfall. As we approached the waterfall, it was apparent that the wind was blowing it back uphill and over the path. Some competitors were making fairly lengthy detours to avoid the spray. As it was so early in the race, we decided to brave it. We crossed the top of the waterfall, with gusts of wind blowing freezing sheets of spray over us. My gloves got soaked through, but I resisted stopping to put on my mitts as they were stashed in the top of my rucksack. Half an hour later my hands were so cold I was struggling to open a Mars Bar. I shouted to Steve and got him to dig out my mitts. An early and important lesson: keep all my gear close to hand.</p>

<p>As we crossed the road at Snake Pass, the late afternoon sun bathed the moorland in a weak orange glow and even Bleaklow Head seemed quiet and benign. That soon changed as we descended Clough Edge towards Torside reservoir. The skies turned a foreboding slate grey and a much lengthier sleet blizzard blew in, forcing hoods up and heads down.</p>

<p>The reservoir offered some shelter and not for the first time I felt a pang of jealousy at the supported runners that were being met there. Steve and I were running unsupported until the last couple of days when my friend Simon joined us.</p>

<p>Darkness fell quickly at 5 pm as we climbed towards Laddow Rocks, and we summited Black Hill in darkness. I’d last been here twenty years earlier on a Duke of Edinburgh expedition, and the neatly paved slabs were a welcome addition to the interminable peat bogs of yore. The next section is a bit of a blur. I remember the wind picking up and hail sweeping in with increasing regularity. This section was bleak, dark moorland punctuated by a number of road crossings. In the dark and with the weather closing in around us, I felt a real sense of isolation and foreboding. This definitely was no place to mess around. I remember looking towards the distant glow of red lights on the hilltop aerials and taking some solace that we were not completely alone.
At the road crossing before the M62, we were met by one of the Mountain Safety Teams and Steve’s friend Matt, who was volunteering for the week.</p>

<p>Their smiles and words of encouragement were a godsend after the previous stretch; we even got a cup of tea. It was at this point that I started to appreciate just how much we take for granted in life. That solitary cup of tea meant everything at that moment.</p>

<p><img src="https://i.snap.as/LW5PcgN6.jpg" alt="Blizzard conditions on the first night"/>
<em>Blizzard conditions on the first night</em></p>

<p>We grabbed some food and crossed the motorway. I remember looking down at the cars whipping by below, wondering where their drivers were going. I thought of the warm beds they’d be sleeping in and the families they’d be returning to. As we climbed Blackstone Edge, the weather intensified again. The wind was gusting up towards 80 mph and blowing intense hail directly in our faces. I found out later that several racers had to retire at checkpoint one due to eye injuries caused by the hail. One even got written up in the <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC4460397/">British Medical Journal</a></p>

<p>In 25 years of winter mountain experience, it was as severe as I’ve experienced. The intense hail blizzard continued as we wound around the reservoirs. It was not until Stoodley Pike Monument that we found a corner of shelter to eat some food. There were three other runners huddling at the base of the looming monument, and we teamed up with them to the checkpoint.</p>

<p>On the descent, I realised to my horror that I had dropped my GPS somewhere further up the trail. I stopped for a moment in the driving hail, unable to believe my stupidity. I knew straight away there was no point in going back; I’d last used it about half an hour before, and it could be anywhere. I got the impression from one or two other competitors that they were slightly sniffy about the use of GPS (despite it being a required kit item). Clearly, you shouldn’t enter the race without being a very competent and experienced navigator. If your GPS packs up, or you drop it and can’t read a map properly, you could easily get into a life-threatening situation. That said, when you’re tired, and the weather is horrendous, having a piece of technology that cuts down the time spent standing around looking at flapping maps is a huge benefit.</p>

<p>I pressed on down the hill into Hebden Bridge. One of the two guys we were with had recce’d the route and led the way. With time at a premium before the race, I hadn’t been able to run any of the route. I’d spent hours poring over the maps, but I knew from previous races that there’s no substitute for experience on the ground, especially when you’re cold and tired as we now were. In my head I’d remembered the section beyond the motorway looking relatively short, but it was a good five hours.</p>

<p>Even at the monument, the first checkpoint felt within reach. Fortunately, our better-prepared companion warned that it was still well over an hour away. As it turned out, it took nearer two. The descent to the A6033 was relatively easy, and the fierce weather started to subside. Reaching the road, deserted and bathed in the cold orange glow of the street lamps, we took a moment to adjust to the sudden incongruousness of the urban environment. A car full of teenagers sped by hooting and jeering. We climbed a painfully steep road before descending across sodden, slippery fields to a river, then slogged another log up through dark, muddy farmland to a road that took us most of the way to the checkpoint. We passed a couple of well-appointed camper vans, and I looked enviously at their windscreens, imagining cosy runners tucked up asleep, having enjoyed a home-cooked dinner.</p>

<p>The descent to the checkpoint was the crowning turd on what had been an exceptionally hard first day. A narrow, hideously muddy track descended steeply, making staying upright almost impossible. We all fell at least once, muttered curses ringing out through the pitch-black woods. Eventually we came out at Hebden Hey, a scout centre. It was 03:30 on Sunday morning, and we had covered more than 40 miles in 16 hours.</p>

<p>We were welcomed in the porch by a remarkably cheery bloke who took our details and arranged for our drop bags to be brought over. The porch was an explosion of wet, muddy footwear and moving inside it was even more chaotic. Every spare inch of space was taken up with kit or tired runners. Even moving along the corridors was a challenge. Eventually Steve and I shoehorned ourselves into a space in the toilets and started to sort our gear out.</p>

<p>Our pre-race strategy was to use the checkpoints to sleep and re-group, trying to get at least four hours sleep every time we stopped (the logic being that any less has very little recuperative effect). Many pressed on through the night, but we stuck to our plan. I grabbed a quick shower, deciding to make use of comfort as and when it was available, and we scarfed a baked potato with chilli before hunting down a bed. There was very little room at the inn.</p>

<p>The idea of trying to get a decent amount of sleep each night was sound. But even with earplugs and an eye mask, I struggled. CP1 &amp; 2 are always going to be the busiest, but I found it difficult to sleep all the way through as the stoppages caused congestion at the normally quieter checkpoints 4 &amp; 5. The options for sleeping are perhaps the biggest race strategy call. Bivvying (tough in bad weather) or camping (extra weight) have their disadvantages too. There’s definitely a psychological benefit of having somewhere warm and dry to sleep.</p>

<p>After trying a few packed dorms, I eventually found a spare bed. It was probably spare for a reason. Light from the corridor shone directly in, and the door seemed to open every five minutes as other bed hunters sought a place to rest. I slept fitfully, dark thoughts flowing through my mind. After such a hard day, the prospect of going on seemed ridiculous. I can understand why so many people decided to stop. I pushed the thoughts away and repeated my race mantra “I’m only stopping if I physically can’t go a step further or a medical professional advises me not to continue”. It helped, but the prospect of another 228 miles after the day we’d had felt terrifying.</p>

<blockquote><p>“If you start, don’t give up, or you will be giving up at difficulties all your life.” 
<strong>Alfred Wainwright, Pennine Way Companion (1968)</strong></p></blockquote>

<p>I got about an hour of partial rest before giving up and going downstairs to sort out my kit for the long leg ahead. Steve slept a bit longer, giving me the chance to have a couple of breakfasts and a few coffees before he appeared. The relentlessly cheery and efficient guy who had met us was still on duty. I asked him whether anyone had handed in a GPS, more in hope than expectation. I doubted anyone braving the hail storms would have noticed a GPS lying in the snow. To my amazement, someone had handed it in. This gave me a huge mental boost. It wasn’t so much that I was relying on it (I’m a reasonable map reader and navigator), it was more the psychological blow of having lost it so stupidly, so early in the race.</p>

<p><img src="https://i.snap.as/HnKWWsJH.jpg" alt="two runners by a sign"/>
<em>70 miles down, 186 to go</em></p>

<p>In the end, we spent about 6 hours at checkpoint one. Given the little sleep I got, this felt like slightly wasted time. As we set off into the first light, I did feel mostly recuperated, and the ascent of the hideously muddy gully leading to the checkpoint didn’t seem quite so bad in daylight. The day was blustery and fresh, the rain holding off as we wound through the unremarkable flat section towards Cowling.</p>

<p>Here, it started to rain in earnest. Cold grey streaks forced us to pull up our hoods and cast our eyes down into a muddy trudge across sodden fields. We were glad to reach the pub at Lothersdale. A roaring log burner welcomed us, and a jovial landlord had turned the pool room into a makeshift checkpoint for muddy Spiners. Rounds of tea and hot food were ferried in as our kit steamed gently on any available radiator.</p>

<p>Leaving this warm sanctuary was hard. So much so that we stopped at the next pub in East Marton too. A small Sunday evening crowd of locals looked on in bemusement as Steve, Jim Tinnion and I, who we’d teamed up with, shed our kit. We explained what we were doing and the landlady made a donation to Steve’s Justgiving page on the spot. They sent us on our way with warm wishes, encouragement and crisps.</p>

<p>Leaving East Marton, we caught up with another runner who stayed with us until Gargrave before peeling off to bivvy at a spot he knew at the train station. Deciding that a stop at the hostelry in Gargrave would constitute a pub crawl, our plan was to press on to checkpoint 1.5 and bivvy near there. The section after Gargrave was horrendous. Field after field of sodden, cow-churned bog sapped our spirits and the rain returned to torment us with squally showers. By the time we reached Malham village at around 2 am, we had all reached our limits and knew we had to stop.</p>

<p>We scouted a few bivvy spots before deciding to use the public toilets. Jim slept with a couple of German competitors in the ladies and Steve and I bagged the gents. I drew the short straw and got the urinal end. Utterly exhausted, I crawled into my bivvy bag and pretty much passed out. We’d agreed two hours sleep, but it seemed like 10 minutes later when Jim appeared at the door. Steve and I blundered blearily about pulling cold wet kit onto our tired, complaining bodies.</p>

<p>Setting off into the dark and rain again was one of the lowest moments of my race. More dank, boggy fields led to slippery, treacherous limestone before we eventually hit a road up to the field centre and checkpoint 1.5. Dawn was starting to break as we stepped into the main room at Malham Field Centre. We were surprised to see a large group of runners trying to sleep with their heads down on the tables.</p>

<p>We were told the race was being held because of bad weather. Not knowing how long we’d be stopped, I pulled on a dry top and found a fragment of space on a heaving radiator for my sodden jacket. Steve and I quaffed tea and ferreted around for food. Biscuits seemed to be the only fare on offer. Our friend Emiko was fast asleep at one of the tables. After a while she awoke and looked around sleepily. Steve had a brief chat. I think she’d taken a wrong turn at some stage and this had set her back.</p>

<p>In future races, Checkpoint 1.5 could potentially be opened up as a proper checkpoint with sleeping areas. It has these facilities already, and Spine racers can be found sprouting out of almost every bush and barn around it. Maybe the 60-mile second “day” is just part of the challenge though.</p>

<p>After about an hour, we were released from the checkpoint and made our way around the tarn, framed in the bleak winter dawn. As we started our ascent of Fountains Fell, the wind harried us from all sides, but I enjoyed the climb. It kept me warm, and the gradient never got so steep it became uncomfortable exertion. My reality for hours on end was a tiny cleft between the top of my balaclava and my hood. A small letterbox out into the world beyond.</p>

<p>Descending to a road, there were times the wind would support our entire pack and body weight. We were met by a Mountain Safety Team who confirmed what we’d heard at the Malham checkpoint. Pen Y Ghent was off limits due to the high winds, and we were diverted at lower level to Horton in Ribblesdale. I can’t honestly remember feeling disappointed. I was tired, hungry and sick of the relentless wind. I simply accepted the instructions.</p>

<p>The cafe in Horton was packed with Spiners and support teams. It was warm and steamy with drying kit. After doing damage to a huge mug of tea and a bacon sandwich, I chatted briefly to a producer from the BBC who was making a documentary about the race. I was quite glad she didn’t try to interview me as I was not feeling very coherent.</p>

<p>Setting off from Horton, Hawes felt within reach. It’s a psychologically important milestone for the Spine Race. Passing through Hawes means that you’re into the race proper. Our plan was to try to sleep at the checkpoint even though there were no beds. We knew that the majority of Challengers would have finished and the noise levels from applause and general hubbub would have subsided.</p>

<p>Our spirits lifted when one of Steve’s friends met us at Cam End with a flask of coffee and walked with us for a few miles. A beautiful magenta sunset picked out the rugged folds of Pen Y Ghent as we climbed over Dodd Fell. Arriving in Hawes and the checkpoint was disorientating. The bright lights of the busy hall were hard to adjust to, and I sat for ten minutes on a chair drinking tea and trying to take it all in.</p>

<p><img src="https://i.snap.as/QRfeYYOJ.jpg" alt="Sunset after Pen Y Ghent"/>
<em>Sunset after Pen Y Ghent</em></p>

<p>The volunteers at Hawes, and throughout the race, were superb. Nothing was too much trouble, and my drop bag was brought over to me along with more tea and food. I fumbled around in my drop bag for ages, my brain unable to deal with the logistical task of gathering what I needed for the next leg. Another of Steve’s friends arrived with fish and chips. The hardship of the race seems to enhance your enjoyment of otherwise everyday pleasures. I don’t think I’ve ever enjoyed a chip quite as much. Having eaten, I decided to sleep before any more kit faffing and found a cupboard off the main room and laid out my Thermarest. I was asleep almost instantly.</p>

<p>I slept fitfully for three hours. Stumbling back into the main hall, I found Steve still sound asleep. I also noticed some fairly severe chafing in my arse region that I sheepishly had checked out and taped up by one of the medics (well above and beyond the call of duty). At this point I inexplicably started to become concerned about the cut-off times for the race. We’d deliberately taken a fairly steady pace so far. I collared one of the volunteers, and together we tried to work out the cut-offs, factoring in the complications of the late start and enforced stop. I’m not sure I was much clearer by the end of it, both our sleep-deprived brains refusing to do simple arithmetic.</p>

<p>The upshot was that, as we prepared to leave, I got Steve worried, and he set off up Great Shunner Fell like a stabbed rat. It was all I could do to keep his torchlight in view in the distance, and I quickly realised that we’d lost touch with Jim Tinnion, who had been with us for the previous day or so. I didn’t have too much time to worry about it as I was too busy trying to stay with Steve. When we caught up with Jim later in the race, he said he’d stopped to sort out a bit of kit, and the next minute we were gone. Sorry, Jim.</p>

<p>As we topped Great Shunner and started to descend, I finally caught up with Steve. It was now snowing heavily, and after a short scramble over icy slabs we put on our spikes for the first time and made the long descent to Thwaite. At Thwaite we passed a couple of the safety team checking runners through and caught up with another small group of runners. Not long after leaving Thwaite, one of the runners decided to return, saying he wasn’t feeling well.</p>

<p><img src="https://i.snap.as/r0HJtwaw.jpg" alt="navigating in snow and darkness"/>
<em>Navigating in snow and darkness</em></p>

<p>The GPS came in handy for the next section of fields before we started the long ascent over the boggy black moor towards Tan Hill Inn. Steve was off in front again, obviously having a good patch and heading for the place where, in a strange coincidence, he had got married. I dug in and slogged up the hill. Another blizzard blew in, reducing visibility to almost nil as I approached the pub.</p>

<p>Prior to the race I’d voraciously read blogs from previous competitors, desperate to gain insight into the race. Several had mentioned The Tan Hill Inn as a warm oasis open all hours during the Spine. As I leaned into the driving snow, my mind conjured images of warm fires and the possibility of hot food. When I eventually arrived at about 5 am, the reality was different. The pub was shut tight, and we crowded into a cramped porch to get out of the snow. I tried to eat a bit, but started to get cold quickly. We set off at a real lick, stomping through the sloshy bogs that led down to a road. OS maps describe an area near here as simply “The Bog”, and they’re not wrong.</p>

<p>The skies cleared, and a cold, sharp dawn broke over the blue-brown moors as we crossed the busy A67 road. Continuing over Cotherstone Moor to Balderhead reservoir, the sun came out and bathed us in a milky winter glow for the first time since the start of the race. The next section to Middleton in Teesdale is a bit of a blur; my mind was already thinking of the beds and food there.</p>

<h2 id="a-cold-start-to-the-day">A cold start to the day</h2>

<p>We arrived at the checkpoint in the early afternoon. We got a warm welcome from Matt, Steve’s mate from the Mountain Safety Team and goofed around in the car park posing for photos. It felt great to be at the first “proper” Spine checkpoint, about halfway through the race. There was lots of space in the checkpoint, and we lounged about chatting to Matt. Steve was starving and frustrated that the only food he could get hold of was scrambled eggs. Clearly we’d timed our arrival badly. Matt offered us his room, which had an en suite shower. I dived in and spent ages luxuriating in the hot flow of water. Having slept much less well than Steve during the race to date, I was spark out for five hours and generating some serious snores (apparently). Although we stayed at the checkpoint for nearly seven hours in total, it didn’t feel like time wasted.</p>

<p><img src="https://i.snap.as/wGfiRtnw.jpg" alt="Twats in their hats"/>
<em>Twats in their hats</em></p>

<p>As we were leaving the checkpoint, the Mountain Safety Team told us we had missed the cut-off for the route via High Cup Nick, Dufton and Cross Fell. Retrospectively, part of me regrets the fact that we didn’t go via these beautiful and rugged parts. At the time I can honestly say that a shorter and easier alternative didn’t bother me in the slightest.</p>

<p>We set off in darkness across the fields that apron the River Tees, the path obscured by ankle-deep snow. During the preceding days we’d often talked about how strange it was having the GPS tracker taped to our shoulders. We knew that there were a bunch of people at home staring at the orange dots on the screen as they progressed across the map. At that moment we must have hit an area of cellular coverage because my phone started pinging with texts. One was from a work colleague with a screenshot of the tracker telling me off for going wrong! I looked at the GPS and realised she was right. It was a surreal moment in the middle of the night, alone, but at the same time connected with the outside world. There will be those who would see this as a negative, but I drew solace from knowing that people were thinking about me as I went through this journey.</p>

<p>We passed Low and High Force, the river in spate because of the recent rainfall, before crossing and beginning the detour via Cow Green reservoir. It was here that things started to unravel. Looking back, I put it down to the fact that I was mentally prepared for the tougher route and let my guard down when I looked at the alternative route. It felt straightforward with good tracks most of the way. As we passed another Mountain Safety Team in their bus at the reservoir, the wind whipped up, and it started to blizzard again. The climb around Herdship Fell was cold and monotonous. I stopped eating enough and at one stage experimented with trying to sleep whilst walking.</p>

<p>The descent to Garrigill was long, icy and unenjoyable. As we passed through Garrigill, there were lots of rabbits hopping about on the village green. Neither Steve nor I mentioned them at the time for fear that we were hallucinating. It was only much later we admitted we’d both seen them. The remainder of the path to Alston is a tired, hungry, exhausted blur. We arrived at the checkpoint in “rag order”, as Steve put it. Our drop bags had been held up by the weather, so we sat down to a plate of tuna pasta at 4 am. Our bags soon turned up, and as we were faffing with kit, the checkpoint staff told us the race was likely to be paused because of an incoming weather system that would bring 100mph winds. I found a bed in one of the dorms and flaked out.</p>

<p>I was woken a few hours later by someone talking loudly on a mobile phone from the adjoining bunk. She was complaining that she’d been forced to withdraw from the race because she’d called out mountain rescue (I later found out that she’d got into difficulty on Cross Fell in the night). I stared at her in exhausted disbelief. All around me, tired runners were trying to sleep, and she was ranting on the phone obliviously. Everyone else I met on the race was amazingly kind and gracious. There is always an exception.</p>

<p><img src="https://i.snap.as/BLvkwv0q.jpg" alt="picture of a swollen foot"/>
<em>Never take your boots off</em></p>

<p>Heading downstairs for breakfast, the wind was rattling the lintels of the windows. Outside, I could see a bleak brown-and-white landscape. It made me feel cold just looking at it. The checkpoint was now a very different place. It had become a marshalling point for most runners still in the race. Every available space was crammed with a person or their kit. It was claustrophobic after the days of solitude we’d had so far.</p>

<p>The race was paused until 6 am the following morning, giving me some much-needed rest. This had positives and negatives. Having bagged a bunk, I was able to catch up on some sleep. However, the break made the damage to my body more apparent. My feet swelled up, and my arse chaffing hadn’t got any better. By the end of the enforced stop, cabin fever was setting in, and I couldn’t face another conversation about race strategy or equipment choice. Food was running short, and runners were roaming the checkpoint like hungry jackals.</p>

<p>It was with great relief that we set off again, heading for Greenhead. Another psychological boost was that for the rest of the race we would be supported. Being unsupported definitely adds an extra level of challenge. You can only fit so much in a drop bag, and having a friendly face and a warm drink along the way makes a big difference.</p>

<p><img src="https://i.snap.as/VhK1f4Jp.jpg" alt="day turns to night and then to day"/>
<em>Day turns to night and then to day</em></p>

<p>The trail started out innocuously enough over rolling fields and was unremarkable all the way to Greenhead. Here we caught up with Simon, our support, and stopped briefly at the Youth Hostel mini checkpoint, which was large and well equipped. Simon grabbed some teas from the nearby cafe (doing a roaring trade in January with racers and support crews – these guys must love the race). After a short stop, we climbed up and onto Hadrian’s Wall.</p>

<p>I’d been particularly looking forward to this section. The ancient green sward of turf that stretched out in front of us belied the history of the place. There was a special feeling looking down over jutting stone abutments, sharing the same view as those here thousands of years before.</p>

<p>The Spine Race induced its own sense of timelessness. Day merged into night into day. The only constants were moving forward and surviving the landscape and weather. Here, the huge span of history made the experience even more otherworldly.</p>

<p>We didn’t have long to ruminate. Halfway along the wall, the next weather front swept in and driving rain and wind chased us along the remainder of the wall and down into the horrendous bogs between Greenly and Broomlee Lough. We made the mistake of pressing on and not eating enough, knowing that we had a warm car waiting for us at the next road junction.</p>

<p>When we arrived there, I was again in “rag order” and starting to border on hypothermia. Si had the heating on full blast and had raided a local shop of all of its pork products. Gorging on scratchings and Pepperami, I warmed up. I’d have been in real trouble if he hadn’t been there. The rest of the trail to Bellingham was a dark, boggy, soul-destroying mess and we arrived at the checkpoint in low spirits just after midnight. I didn’t want to go a step further. Finishing seemed an impossibility.</p>

<p><img src="https://i.snap.as/81YF3bqa.jpg" alt="Trying to smile with borderline hypothermia"/>
<em>Trying to smile with borderline hypothermia</em></p>

<blockquote><p>“When we walk to the edge of all the light we have and take the step into the darkness of the unknown, we must believe that one of two things will happen. There will be something solid for us to stand on, or we will be taught to fly.”
<strong>Patrick Overton</strong></p></blockquote>

<p>Normally, the race is strung out by Bellingham, but the enforced stop had bunched everyone up, and it felt like the checkpoint staff were struggling to cope. By this stage they were probably as exhausted as we were. The sleeping hall was cold, and I crammed a few cereal bars down rather than walking out and over the car park to the food area. I hunkered down under a table and tried to get some sleep. This was the last checkpoint before the long push to the finish.</p>

<p>When I woke up a few short hours later, I didn’t feel any better. Everything hurt, and it seemed to take twice as long to assemble my kit. Fortunately, the day outside was beautiful: cold, but with bright, dazzling sunshine. The first section was extensively flag stoned. Although this made the route obvious, a thick rime of ice made the slabs treacherous.</p>

<p>Simon met us with tea at a road crossing, and before long we hit the forested section to Byrness. Apart from a few boggy sections, the paths here were mainly wide clear forest trails. We walked for a while with a couple (I’m really sorry, I can’t recall your names). I do remember talking about particle physics and the Large Hadron Collider, about which I know surprisingly little. We were met outside Byrness by another of Steve’s entourage, and again it gave us a real lift to see a friendly face. The hostel at Byrness had been transformed into a mini checkpoint and was serving hot food. I arrived feeling shoddy but left with a spring in my step and sausages and mash in my stomach.</p>

<h2 id="approaching-byrness">Approaching Byrness</h2>

<p>From Byrness it’s 27 miles to the finish. It feels within your grasp, but as several racers have found out, it’s far from in the bag. The sun was almost hot on our backs as we slogged up to Windy Crag and the ridge that we would follow home. On the ridge, the ground underfoot was frozen hard, saving us from the bogs and allowing us to make fast progress. We delayed putting on head torches for as long as possible to enjoy every last gasp of the wispy pink sunset. As soon as the sun dropped below the horizon, the temperature plummeted, and the wind picked up. Any thoughts of an easy stretch to the finish were gone.</p>

<p>Mentally, I’d divided the ridge into three sections, with the two mountain refuge huts as beacons along the way. At the first hut we found Matt and his Mountain Safety Teams. We ate a bit of food and left quickly because we were getting cold. I badly underestimated the time it would take us to reach the second hut. Feeling reasonably good as I left hut one, I started to feel cold and tired after an hour and had to force myself to eat. I ran out of water, and everything around us was frozen.</p>

<p>Steve had arrived at the Alston checkpoint with his micro spikes, and left without them, despite them being securely stowed in his bag. How he negotiated the icy ridge I’ll never know. The descent to hut two was steep and hazardous. There were three other runners at the hut. One decided he needed to sleep for a few hours to recuperate. Another made hot chocolate using the only liquid he had with him, orange water. An interesting concoction worthy of the final night of the race.</p>

<p>As we left, I made a serious navigational error, getting disoriented and leading us in the wrong direction. Fortunately, Matt and the MST were sweeping along the ridge behind us and shouted us back. We continued up the Schil, the last climb of the race. The icy descent provided no respite for Steve. For the first time in the race, he started to lose his sense of humour. The amount of extra energy needed to stay upright without micro spikes must have been phenomenal. I stopped at the first unfrozen rivulet and gulped water like a desert wanderer arriving at an oasis.</p>

<p>In a semi-delirious state, I thought I was having a heart attack on the final climb into Kirk Yetholm. I remember trying to calculate whether an aneurysm would prevent me from crawling the last mile to the Border Inn. That’s what this race does to you. As we entered the town, the incongruently bright sodium street lamps lit up the village green, and we could see a small clump of figures beckoning us over to the pub. We broke into a pained lope and had our photographs taken touching the wall. We’d finished The Spine, Britain’s most brutal race.</p>

<blockquote><p>“We had done this thing we had set out to do, and instead of becoming larger because of the experience, we became smaller, more humble, more aware of how little we know: about the world in general, about ourselves specifically.”
<strong>Richard Benyo, The Death Valley 300</strong></p></blockquote>

<p><img src="https://i.snap.as/BE860BKB.jpg" alt="Kirk Yetholm"/>
*Kirk Yetholm – where&#39;s my free half pint!?</p>

<h2 id="the-aftermath">The aftermath</h2>

<p>Although the physical toll on my body was less than after tough 100 milers like the Lakeland, my feet swelled massively, and I had a couple of blisters. I’d developed ulcers on my tongue and sores on my fingers towards the end of the race. It seemed as though my body had been prioritising its core over the extremities. Focusing on what was essential for survival. That’s just how far we pushed it. To finish, I had to delve more deeply into my mental and physical reserves than I have ever done before.</p>

<p><img src="https://i.snap.as/KI1vdrH1.jpg" alt="Finisher&#39;s medal"/>
<em>Earned!</em></p>

<p>A deep fatigue took several weeks to pass. My wife said “I’m not expecting you back until the end of January” and she was right. I felt like a shell, a hollowed-out husk without life or energy. I was banished to the spare room due to heavy nighttime sweating as my body repaired itself.</p>

<p>The weight fell off me, and looking in the mirror, I struggled to recognise the gaunt face staring back.
Going back to work, the back slaps, and congratulations were welcome but somehow hollow. I felt alone with my experience. No one could understand what I’d been through, the things I’d experienced, the amount I’d endured.</p>

<p>Reading the blogs that emerged after the race made me feel better. The band of crazy brothers were starting to speak at last. It was good to hear from that infinitesimally small group of people that would countenance such an undertaking.</p>

<p>The scale of the race makes it very hard to process mentally. Breaking the race down into micro-sections avoids thinking about the vast overall scale. This can short-circuit even the toughest of minds. But little by little I started to piece the race back together. Hour by hour, step by step it became a coherent whole. A body of memory that will abide with me for the rest of my life. An experience so much more precious than any material possession.</p>

<p>I started to write those thoughts down. Eventually they turned into this, a way for me to codify the experience so it doesn’t slip away. For we were there that wild week in January when the Pennine Way turned 50. And we did something that few others will ever do and that only we will ever fully understand.</p>

<blockquote><p>“What they had done, what they had seen, heard, felt, feared – the places, the sounds, the colours, the cold, the darkness, the emptiness, the bleakness, the beauty. ‘Til they died, this stream of memory would set them apart, if imperceptibly to anyone but themselves, from everyone else. For they had crossed the mountains…“
<strong>Bernard DeVoto, The Course of Empire (1952)</strong></p></blockquote>

<p>This article was adapted for the <a href="https://www.cicerone.co.uk/fastpacking">Cicerone Guide <em>Fastpacking</em></a> under the title “A Great British Adventure. It was certainly that.</p>

<p><img src="https://i.snap.as/DXoWWihQ.webp" alt="Cicerone Fastpacking Guidebook"/></p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/running-the-spine-race</guid>
      <pubDate>Sun, 02 Aug 2026 18:57:28 +0000</pubDate>
    </item>
    <item>
      <title>Break it down, baby </title>
      <link>https://iain.so/break-it-down-baby?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[Scott Hansen brought a dead machine back to life, and the first thing it played was Midnight in a Perfect World. The machine was a beige Akai MPC60 II, caked in grime, its screen blown, its floppy drive jammed, paint peeling. It had spent roughly twenty years in storage as junk. Hansen, who records ambient electronic music as Tycho, had been asked by DJ Shadow to help with data recovery for the 30th anniversary remaster of Endtroducing....., the 1996 album that the Guinness Book of Records recognises as the first ever composed entirely from samples.&#xA;&#xA;Shadow arrived at Hansen&#39;s studio with hundreds of floppy disks and the original MPC that the album was made on. Hansen took one look at the machine, abandoned the data project, and spent weeks soaking parts in dish soap and solvents, replacing the screen, and rebuilding the unit from the chassis up. When Shadow loaded the first disk and hit play, the room heard the opening bars of a record that has sat on Rolling Stone&#39;s 500 Greatest Albums list for a quarter of a century, coming not from a pressing or a file but from the instrument that made it.&#xA;&#xA;The moment was a window in time, but it was also, if you think about what was really happening, a warning. The album that proved digital sampling could produce work as deep as any live recording turns out to be among the most fragile things in the archive.&#xA;&#xA;huge pile of records in a basement&#xA;&#xA;Thirteen seconds and a turntable&#xA;&#xA;To understand what came off those disks, you need to understand what went onto them. Endtroducing..... was built between 1994 and 1996 on three pieces of equipment: an Akai MPC60 MKII sampler, a Technics SL-1200 turntable, and an Alesis ADAT eight-track tape recorder. Shadow worked first in his California apartment, then at the Glue Factory, the San Francisco home studio of Dan the Automator, whose contribution was largely leaving the room and letting Shadow get on with using his gear.&#xA;&#xA;The MPC60 II was a Roger Linn design, released in 1991. It sampled at 12-bit and 40kHz, which is lower fidelity than CD (16-bit, 44.1kHz), but it gave the recorded sound a particular graininess that became part of the album&#39;s texture. Its total sample memory, with the expansion board fitted, was 26.2 seconds of mono audio. That number shaped many aspects of the record. At roughly 60 kilobytes per second, the machine&#39;s entire memory fits on a single floppy disk. You are not making a 63-minute album inside 26 seconds of RAM. You are making it in layers, and each layer consumes the last.&#xA;&#xA;Shadow&#39;s workflow was a hand-rolled paging system with tape as the backing store. He would sample vinyl fragments into the MPC, chop them in the machine&#39;s trim mode, assign the chops across the 16 pads, and sequence a pass. When memory was full, he bounced the MPC&#39;s output to a stereo pair of tracks on the ADAT, which gave him eight tracks of 16-bit, 48kHz digital audio on S-VHS tape. Then he wiped the MPC&#39;s memory and started the next layer, synced to what was already on tape. &#34;It was all about chopping,&#34; Shadow told an interviewer in 1997. &#34;I never had the luxury of taking extended samples. There are lots of overdubs on the album, as well. But the whole record was 100% sample-based, or vinyl-based. All the sampling was done on the MPC. It all went straight from turntable to the MPC to tape.&#34;&#xA;&#xA;Each bounce was a converter round trip and another generation of tape. It was also irreversible. Once two layers were on the same pair of ADAT tracks, they could not be separated again. The process was compositional: the record&#39;s structure emerged from the order in which layers were committed to tape, and every commitment destroyed the working state that came before it. When the ADAT tracks were full, the whole thing was mixed down to DAT at the studio, where it was, in Shadow&#39;s words, &#34;maybe compressed and limited a little bit.&#34; By the end, there was no multitrack of the album. There was only a mix.&#xA;&#xA;Patchwork as an instrument&#xA;&#xA;The bounce-and-wipe cycle was a production constraint. The compositional innovation was what happened inside each layer before the bounce. Shadow did not sample loops; he sampled fragments, often too short to be recognisable as their source, and rebuilt them into something that functioned as a new composition. A two-second drum hit became a texture, and a horn phrase became a pad. A spoken-word clip became rhythm. The MPC&#39;s trim mode let him define start and end points within a sample, chop it into discrete regions, and convert the resulting slices into a &#34;program&#34; mapped across the pads. He then played the pads percussively, and the sequence the MPC recorded was a performance of fragments rather than a loop.&#xA;&#xA;The method was harder than it sounds, because turntable pitch control was part of it. Shadow would adjust the speed of a record on the SL-1200 to match its tempo to the track he was building, sample at the altered speed, and let the MPC&#39;s transposition handle the rest. But transposition on a 12-bit sampler is resampling, not pitch-shifting in the modern sense. &#xA;&#xA;Pitch and timbre are coupled. Play a sample two semitones down, and it gets darker, not just lower. The album&#39;s warm, murky quality is partly this: the sound of vinyl, through a cheap phono stage, through a 12-bit converter, through transposition that changes timbre as it changes pitch, through the analogue output of the MPC, through a Mackie desk, onto tape. Every stage adds character and removes bandwidth, and the accumulation of those small degradations is the record&#39;s unique sound.&#xA;&#xA;The sequencer ran at 96 parts per quarter note, which gives 24 ticks per sixteenth note. Roger Linn&#39;s swing function, which he invented for his original LM-1 drum machine in 1979 and carried into the MPC, delays the second sixteenth note within each eighth note by a variable amount. At 66%, you get perfect triplet swing. On a 96ppq grid, that is a whole-tick offset of 8 ticks on a step that is 24 ticks long. The feel people describe as the MPC&#39;s &#34;magic timing&#34; is integer arithmetic on a coarse clock. Modern sequencers running at 960 or 4,096ppq cannot land on those positions without explicitly emulating the rounding. The grid provides the groove.&#xA;&#xA;The collector&#39;s disks&#xA;&#xA;Shadow does not appear to be the kind of person who loses things. He has a personal record collection exceeding 60,000 vinyl records, accumulated over decades of systematic crate-digging at shops like Rare Records in Sacramento, where he spent hours each day in the basement. For Action Adventure) in 2023, he bought 200 radio-broadcast tapes from eBay and worked through records in his own collection he had never previously listened to.&#xA;&#xA;So when Hansen describes &#34;huge cases of floppy disks, hundreds and hundreds of them,&#34; that number is not hoarding; it is the arithmetic of production. At 793 kilobytes per MPC60-formatted disk, with the machine holding 26 seconds of audio at 40kHz, one disk is roughly a full memory load. Two years of production, with each session potentially generating multiple saves of sequences and sample data, produces hundreds of disks by definition. And given Shadow&#39;s temperament, those disks were very likely ordered, labelled, and stored deliberately.&#xA;&#xA;The MPC60 saves two kinds of data to floppy: sequences (the MIDI-like performance data, the arrangement of chops on pads, the timing, the swing settings) and sound files (the actual 12-bit audio). A &#34;set&#34; in MPC terminology is a collection of sequences with their associated sounds and parameters. When Hansen says Shadow &#34;pulled up one of the sets,&#34; the machine loaded a set, and the set contained enough information to play back a track through the MPC&#39;s own sound engine and converters. What came out of the speakers was not a recording of Midnight in a Perfect World. It was a rendering. The machine was performing the track from its source data, thirty years after the last time.&#xA;&#xA;Four things that can die&#xA;&#xA;The MPC restoration project that Hansen undertook as part of the album remaster shows what digital preservation means when the format is proprietary, the reader is discontinued, and the knowledge is tacit.&#xA;&#xA;Media. Thirty-year-old double-density 3.5-inch floppy disks. Magnetic media degrades. The MPC60 formats DS/DD media to its own 793K layout rather than the standard 720K, using 10 sectors per track instead of 9, at 512 bytes per sector, double-sided, 80 tracks. A standard USB floppy drive cannot read this format at all. If you plug the disks into a modern computer, you get nothing.&#xA;&#xA;Format. The Akai filesystem sitting on those sectors is proprietary. Recovering the raw bytes is not enough. Someone has to parse the file structures, understand how sequences reference sounds by name, and reconstruct the relationships between programs, sounds, and the sequences that play them. A missing sound file is a dangling pointer; the sequence plays, but silently.&#xA;&#xA;Hardware. The actual MPC60 II that Shadow used in 1994-1996 was found with a dead display, a jammed floppy drive, grime coating every surface, and paint flaking off the front panel. He tore it to the chassis, soaked the parts, sourced a replacement screen, and rebuilt it. Had he not, and had this particular machine been thrown away by whoever stored it, a period-correct rendering of the album&#39;s source data would require finding another MPC60 II, hoping its converters and output stage were within tolerance, and accepting that the result would not be precisely the same. The restored unit is exact, which means the sound coming out of it when Shadow hit play was not a reproduction. It was the exact same signal path from the exact same device.&#xA;&#xA;Knowledge. Which disk holds which track. Which set is the final version and which is an abandoned take. What order the bounces went in. What was done on the desk. Eight-character filenames on a machine with a two-line LCD display. Shadow&#39;s own memory is presumably the index, and it is the one thing you cannot image to a backup.&#xA;&#xA;When Shadow&#39;s Mo&#39; Wax singles were remastered in 2025 for a box set, the original DAT tapes were in some cases so aged and fragile that multiple machines were needed to get a clean transfer. DAT is a robust format by comparison. It is standardised, widely supported, and transfers digitally. The MPC floppies are none of these things.&#xA;&#xA;What a full reconstruction would mean&#xA;&#xA;What follows is informed speculation about what the remaster might entail, based on a short interview with Hansen that was released on Shadow’s socials, the known architecture of the MPC60 II, and the documented production history of the album.&#xA;&#xA;Imaging the disks is the first problem. Flux-level imaging using hardware like the KryoFlux or Greaseweazle captures the raw magnetic transitions on the disk surface rather than attempting to decode data in real time. This matters because a thirty-year-old disk with weak sectors might return errors on a conventional read but yield recoverable data from a flux capture, where the raw signal can be processed multiple times with different error-correction strategies. Cambridge University Library&#39;s &#34;Copy That Floppy&#34; preservation project uses exactly this approach for obsolete formats. The flux images are the archive. Everything else is derived from them.&#xA;&#xA;The second problem is rendering. The MPC60 is not just a reader. It is the instrument. Its 12-bit converters, its analogue output stage, and its particular approach to sample playback are all part of what the album sounds like. An emulator could play back the sequences and trigger the sounds, but it would do so through modern conversion at modern bit depths, which is not the same signal. &#xA;&#xA;The restored original machine solves this. Load the set, press play, and what comes out of the stereo outputs is the same electrical signal that came out in 1996, through the same converters and the same analogue path. Record that at high resolution and you have a stem, or at least a layer, rendered at period-correct fidelity.&#xA;&#xA;This is the part that bends the mind slightly. The album, as released, has no multitrack. The bounce-and-wipe process destroyed the separation between layers as Shadow worked. The MPC floppies hold the material that existed before each bounce, the save states from a process that consumed its intermediates. If you can identify which sets correspond to which layers of which tracks, you can render each one in isolation through the restored machine and produce separated parts that have never existed at any point in the album&#39;s history, not even during production. You are not recovering a multitrack. You are compiling one for the first time.&#xA;&#xA;There are limits, and they are not small. Turntable work performed live-to-tape during production is on the ADAT, not in the MPC data. Desk moves and outboard processing may be unrecorded. Some bounced layers may have no surviving pre-bounce source, because the corresponding disks were overwritten or lost. The reconstruction is therefore a hybrid: recovered sequence data rendered through period-correct hardware, layered with material that can only come off the ADAT tapes (if they survive) or be re-performed. A faithful reconstruction of Endtroducing..... is therefore a huge and complicated undertaking. &#xA;&#xA;On December 18th, Shadow will perform Endtroducing..... live at the Barbican with the BBC Symphony Orchestra, conducted by Jules Buckley, in a one-off orchestral reimagining recorded for BBC 6 Music. You cannot orchestrate a stereo mix. You need separable parts, melodic lines, harmonic structures, rhythmic patterns, each isolated enough for a composer to write around them. The MPC reclamation project may be the reason that concert is possible at all.&#xA;&#xA;The fragility of the first digital generation&#xA;&#xA;There is a comforting myth about digital media, which is that it lasts forever because copying is lossless. The myth confuses the copy with the thing being copied. A bitwise duplicate of an MPC60 floppy disk image is perfect, if you have one. But the image is useless without a format specification, the format is useless without a reader, and the reader is useless without the tacit knowledge of what the data means. A 1968 analogue master tape can be baked in an oven to temporarily re-bind the oxide, threaded onto a compatible machine, and played. The format is the medium&#39;s physics. A 1994 Akai floppy needs reverse-engineered filesystem knowledge, bespoke imaging hardware, and a working machine that has been out of production for thirty years.&#xA;&#xA;The first generation of digital-native creative work, the records made on MPCs and SP-1200s and Atari STs, the art made on Amigas and early Macs, the writing saved to 800K floppies in proprietary word-processor formats, is the most endangered material in the cultural archive. Not because the bits have decayed, though some have, but because the stack of dependencies between the bits and the meaning is deep, undocumented, and each layer can fail independently.&#xA;&#xA;Hansen soaked MPC parts in dish soap to make a piece of music history audible again. That is both a beautiful story and a cautionary tale. The album that proved you could build something permanent from other people&#39;s forgotten records was itself perilously close to becoming unrecoverable. Digital is not the opposite of fragile. It is fragile in different, more complex ways.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>Scott Hansen brought a dead machine back to life, and the first thing it played was <a href="https://en.wikipedia.org/wiki/Midnight_in_a_Perfect_World">Midnight in a Perfect World</a>. The machine was a beige Akai MPC60 II, caked in grime, its screen blown, its floppy drive jammed, paint peeling. It had spent roughly twenty years in storage as junk. Hansen, who records ambient electronic music as Tycho, had been asked by DJ Shadow to help with data recovery for the 30th anniversary remaster of <em>Endtroducing.....</em>, the 1996 album that the <a href="https://www.guinnessworldrecords.com/world-records/first-album-made-completely-from-samples">Guinness Book of Records</a> recognises as the first ever composed entirely from samples.</p>

<p>Shadow arrived at Hansen&#39;s studio with hundreds of floppy disks and the original MPC that the album was made on. Hansen took one look at the machine, abandoned the data project, and spent weeks soaking parts in dish soap and solvents, replacing the screen, and rebuilding the unit from the chassis up. When Shadow loaded the first disk and hit play, the room heard the opening bars of a record that has sat on Rolling Stone&#39;s 500 Greatest Albums list for a quarter of a century, coming not from a pressing or a file but from the instrument that made it.</p>

<p>The moment was a window in time, but it was also, if you think about what was really happening, a warning. The album that proved digital sampling could produce work as deep as any live recording turns out to be among the most fragile things in the archive.</p>

<p><img src="https://i.snap.as/0mBUxiAk.jpeg" alt="huge pile of records in a basement"/></p>

<h2 id="thirteen-seconds-and-a-turntable">Thirteen seconds and a turntable</h2>

<p>To understand what came off those disks, you need to understand what went onto them. <em>Endtroducing.....</em> was built between 1994 and 1996 on <a href="https://en.wikipedia.org/wiki/Endtroducing.....">three pieces of equipment</a>: an Akai MPC60 MKII sampler, a Technics SL-1200 turntable, and an Alesis ADAT eight-track tape recorder. Shadow worked first in his California apartment, then at the Glue Factory, the San Francisco home studio of <a href="https://en.wikipedia.org/wiki/Dan_the_Automator">Dan the Automator</a>, whose contribution was largely leaving the room and letting Shadow get on with using his gear.</p>

<p>The MPC60 II was a <a href="https://www.soundonsound.com/music-business/akai-mpc60-revisited">Roger Linn design</a>, released in 1991. It sampled at 12-bit and 40kHz, which is lower fidelity than CD (16-bit, 44.1kHz), but it gave the recorded sound a particular graininess that became part of the album&#39;s texture. Its total sample memory, with the expansion board fitted, was 26.2 seconds of mono audio. That number shaped many aspects of the record. At roughly 60 kilobytes per second, the machine&#39;s entire memory fits on a single floppy disk. You are not making a 63-minute album inside 26 seconds of RAM. You are making it in layers, and each layer consumes the last.</p>

<p>Shadow&#39;s workflow was a hand-rolled paging system with tape as the backing store. He would sample vinyl fragments into the MPC, chop them in the machine&#39;s trim mode, assign the chops across the 16 pads, and sequence a pass. When memory was full, he bounced the MPC&#39;s output to a stereo pair of tracks on the ADAT, which gave him <a href="https://www.vintagedigital.com.au/alesis-adat/">eight tracks of 16-bit, 48kHz digital audio on S-VHS tape</a>. Then he wiped the MPC&#39;s memory and started the next layer, synced to what was already on tape. “It was all about chopping,” Shadow <a href="https://whyisthisinteresting.substack.com/p/why-is-this-interesting-the-synth">told an interviewer</a> in 1997. “I never had the luxury of taking extended samples. There are lots of overdubs on the album, as well. But the whole record was 100% sample-based, or vinyl-based. All the sampling was done on the MPC. It all went straight from turntable to the MPC to tape.”</p>

<p>Each bounce was a converter round trip and another generation of tape. It was also irreversible. Once two layers were on the same pair of ADAT tracks, they could not be separated again. The process was compositional: the record&#39;s structure emerged from the order in which layers were committed to tape, and every commitment destroyed the working state that came before it. When the ADAT tracks were full, the whole thing was <a href="https://www.mpc-forums.com/viewtopic.php?f=5&amp;t=27149">mixed down to DAT</a> at the studio, where it was, in Shadow&#39;s words, “maybe compressed and limited a little bit.” By the end, there was no multitrack of the album. There was only a mix.</p>

<h2 id="patchwork-as-an-instrument">Patchwork as an instrument</h2>

<p>The bounce-and-wipe cycle was a production constraint. The compositional innovation was what happened inside each layer before the bounce. Shadow did not sample loops; he sampled fragments, often too short to be recognisable as their source, and rebuilt them into something that functioned as a new composition. A two-second drum hit became a texture, and a horn phrase became a pad. A spoken-word clip became rhythm. The <a href="https://www.academia.edu/970719/Behind_the_Beat_Technical_and_Practical_Aspects_of_Instrumental_Hip_Hop_Composition">MPC&#39;s trim mode</a> let him define start and end points within a sample, chop it into discrete regions, and convert the resulting slices into a “program” mapped across the pads. He then played the pads percussively, and the sequence the MPC recorded was a performance of fragments rather than a loop.</p>

<p>The method was harder than it sounds, because turntable pitch control was part of it. Shadow would adjust the speed of a record on the SL-1200 to match its tempo to the track he was building, sample at the altered speed, and let the MPC&#39;s transposition handle the rest. But transposition on a 12-bit sampler is resampling, not pitch-shifting in the modern sense.</p>

<p>Pitch and timbre are coupled. Play a sample two semitones down, and it gets darker, not just lower. The album&#39;s warm, murky quality is partly this: the sound of vinyl, through a cheap phono stage, through a 12-bit converter, through transposition that changes timbre as it changes pitch, through the analogue output of the MPC, through a Mackie desk, onto tape. Every stage adds character and removes bandwidth, and the accumulation of those small degradations is the record&#39;s unique sound.</p>

<p>The <a href="https://archive.org/stream/synthmanual_MPC60_V3.1_owners_manual/MPC60_V3.1_owners_manual_djvu.txt">sequencer ran at 96 parts per quarter note</a>, which gives 24 ticks per sixteenth note. Roger Linn&#39;s swing function, <a href="https://www.attackmagazine.com/features/interview/roger-linn-swing-groove-magic-mpc-timing/">which he invented for his original LM-1 drum machine in 1979</a> and carried into the MPC, delays the second sixteenth note within each eighth note by a variable amount. At 66%, you get perfect triplet swing. On a 96ppq grid, that is a whole-tick offset of 8 ticks on a step that is 24 ticks long. The feel people describe as the MPC&#39;s “magic timing” is integer arithmetic on a coarse clock. Modern sequencers running at 960 or 4,096ppq cannot land on those positions without explicitly emulating the rounding. The grid provides the groove.</p>

<h2 id="the-collector-s-disks">The collector&#39;s disks</h2>

<p>Shadow does not appear to be the kind of person who loses things. He has a <a href="https://en.wikipedia.org/wiki/DJ_Shadow">personal record collection exceeding 60,000 vinyl records</a>, accumulated over decades of systematic crate-digging at shops like Rare Records in Sacramento, where he spent hours each day in the basement. For <a href="https://en.wikipedia.org/wiki/Action_Adventure_(album)">Action Adventure</a> in 2023, he bought 200 radio-broadcast tapes from eBay and worked through records in his own collection he had never previously listened to.</p>

<p>So when Hansen describes “huge cases of floppy disks, hundreds and hundreds of them,” that number is not hoarding; it is the arithmetic of production. At 793 kilobytes per MPC60-formatted disk, with the machine holding 26 seconds of audio at 40kHz, one disk is roughly a full memory load. Two years of production, with each session potentially generating multiple saves of sequences and sample data, produces hundreds of disks by definition. And given Shadow&#39;s temperament, those disks were very likely ordered, labelled, and stored deliberately.</p>

<p>The MPC60 saves two kinds of data to floppy: sequences (the MIDI-like performance data, the arrangement of chops on pads, the timing, the swing settings) and sound files (the actual 12-bit audio). A “set” in MPC terminology is a collection of sequences with their associated sounds and parameters. When Hansen says Shadow “pulled up one of the sets,” the machine loaded a set, and the set contained enough information to play back a track through the MPC&#39;s own sound engine and converters. What came out of the speakers was not a recording of Midnight in a Perfect World. It was a rendering. The machine was performing the track from its source data, thirty years after the last time.</p>

<h2 id="four-things-that-can-die">Four things that can die</h2>

<p>The MPC restoration project that Hansen undertook as part of the album remaster shows what digital preservation means when the format is proprietary, the reader is discontinued, and the knowledge is tacit.</p>

<p><strong>Media.</strong> Thirty-year-old double-density 3.5-inch floppy disks. Magnetic media degrades. The <a href="https://www.mpc-forums.com/viewtopic.php?p=1234291">MPC60 formats DS/DD media to its own 793K layout</a> rather than the standard 720K, using <a href="https://forum.vintagesynth.com/viewtopic.php?t=102615">10 sectors per track instead of 9</a>, at 512 bytes per sector, double-sided, 80 tracks. A standard USB floppy drive <a href="https://hxc2001.com/floppy/forum/viewtopic.php?t=3787">cannot read this format at all</a>. If you plug the disks into a modern computer, you get nothing.</p>

<p><strong>Format.</strong> The Akai filesystem sitting on those sectors is proprietary. Recovering the raw bytes is not enough. Someone has to parse the file structures, understand how sequences reference sounds by name, and reconstruct the relationships between programs, sounds, and the sequences that play them. A missing sound file is a dangling pointer; the sequence plays, but silently.</p>

<p><strong>Hardware.</strong> The actual MPC60 II that Shadow used in 1994-1996 was found with a dead display, a jammed floppy drive, grime coating every surface, and paint flaking off the front panel. He tore it to the chassis, soaked the parts, sourced a replacement screen, and rebuilt it. Had he not, and had this particular machine been thrown away by whoever stored it, a period-correct rendering of the album&#39;s source data would require finding another MPC60 II, hoping its converters and output stage were within tolerance, and accepting that the result would not be precisely the same. The restored unit is exact, which means the sound coming out of it when Shadow hit play was not a reproduction. It was the exact same signal path from the exact same device.</p>

<p><strong>Knowledge.</strong> Which disk holds which track. Which set is the final version and which is an abandoned take. What order the bounces went in. What was done on the desk. Eight-character filenames on a machine with a two-line LCD display. Shadow&#39;s own memory is presumably the index, and it is the one thing you cannot image to a backup.</p>

<p>When Shadow&#39;s <a href="https://djshadow.com/blogs/news/the-mo-wax-singles-1993-1997-box-set-with-exclusive-signed-print-now-available-for-pre-order">Mo&#39; Wax singles were remastered in 2025</a> for a box set, the original DAT tapes were in some cases so aged and fragile that multiple machines were needed to get a clean transfer. DAT is a robust format by comparison. It is standardised, widely supported, and transfers digitally. The MPC floppies are none of these things.</p>

<h2 id="what-a-full-reconstruction-would-mean">What a full reconstruction would mean</h2>

<p>What follows is informed speculation about what the remaster might entail, based on a short interview with Hansen that was released on Shadow’s socials, the known architecture of the MPC60 II, and the documented production history of the album.</p>

<p>Imaging the disks is the first problem. <a href="https://www.digipres.org/the-floppy-guide/">Flux-level imaging using hardware like the KryoFlux or Greaseweazle</a> captures the raw magnetic transitions on the disk surface rather than attempting to decode data in real time. This matters because a thirty-year-old disk with weak sectors might return errors on a conventional read but yield recoverable data from a flux capture, where the raw signal can be processed multiple times with different error-correction strategies. Cambridge University Library&#39;s <a href="https://www.tomshardware.com/pc-components/storage/cambridge-university-rescues-data-from-old-floppy-disks">“Copy That Floppy” preservation project</a> uses exactly this approach for obsolete formats. The flux images are the archive. Everything else is derived from them.</p>

<p>The second problem is rendering. The MPC60 is not just a reader. It is the instrument. Its 12-bit converters, its analogue output stage, and its particular approach to sample playback are all part of what the album sounds like. An emulator could play back the sequences and trigger the sounds, but it would do so through modern conversion at modern bit depths, which is not the same signal.</p>

<p>The restored original machine solves this. Load the set, press play, and what comes out of the stereo outputs is the same electrical signal that came out in 1996, through the same converters and the same analogue path. Record that at high resolution and you have a stem, or at least a layer, rendered at period-correct fidelity.</p>

<p>This is the part that bends the mind slightly. The album, as released, has no multitrack. The bounce-and-wipe process destroyed the separation between layers as Shadow worked. The MPC floppies hold the material that existed before each bounce, the save states from a process that consumed its intermediates. If you can identify which sets correspond to which layers of which tracks, you can render each one in isolation through the restored machine and produce separated parts that have never existed at any point in the album&#39;s history, not even during production. You are not recovering a multitrack. You are compiling one for the first time.</p>

<p>There are limits, and they are not small. Turntable work performed live-to-tape during production is on the ADAT, not in the MPC data. Desk moves and outboard processing may be unrecorded. Some bounced layers may have no surviving pre-bounce source, because the corresponding disks were overwritten or lost. The reconstruction is therefore a hybrid: recovered sequence data rendered through period-correct hardware, layered with material that can only come off the ADAT tapes (if they survive) or be re-performed. A faithful reconstruction of <em>Endtroducing.....</em> is therefore a huge and complicated undertaking.</p>

<p>On December 18th, Shadow will <a href="https://www.barbican.org.uk/whats-on/2026/event/dj-shadow-with-bbc-symphony-orchestra">perform Endtroducing..... live at the Barbican with the BBC Symphony Orchestra</a>, conducted by Jules Buckley, in a one-off orchestral reimagining recorded for BBC 6 Music. You cannot orchestrate a stereo mix. You need separable parts, melodic lines, harmonic structures, rhythmic patterns, each isolated enough for a composer to write around them. The MPC reclamation project may be the reason that concert is possible at all.</p>

<h2 id="the-fragility-of-the-first-digital-generation">The fragility of the first digital generation</h2>

<p>There is a comforting myth about digital media, which is that it lasts forever because copying is lossless. The myth confuses the copy with the thing being copied. A bitwise duplicate of an MPC60 floppy disk image is perfect, if you have one. But the image is useless without a format specification, the format is useless without a reader, and the reader is useless without the tacit knowledge of what the data means. A 1968 analogue master tape can be baked in an oven to temporarily re-bind the oxide, threaded onto a compatible machine, and played. The format is the medium&#39;s physics. A 1994 Akai floppy needs reverse-engineered filesystem knowledge, bespoke imaging hardware, and a working machine that has been out of production for thirty years.</p>

<p>The first generation of digital-native creative work, the records made on MPCs and SP-1200s and Atari STs, the art made on Amigas and early Macs, the writing saved to 800K floppies in proprietary word-processor formats, is the most endangered material in the cultural archive. Not because the bits have decayed, though some have, but because the stack of dependencies between the bits and the meaning is deep, undocumented, and each layer can fail independently.</p>

<p>Hansen soaked MPC parts in dish soap to make a piece of music history audible again. That is both a beautiful story and a cautionary tale. The album that proved you could build something permanent from other people&#39;s forgotten records was itself perilously close to becoming unrecoverable. Digital is not the opposite of fragile. It is fragile in different, more complex ways.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/break-it-down-baby</guid>
      <pubDate>Mon, 27 Jul 2026 08:47:07 +0000</pubDate>
    </item>
    <item>
      <title>At 50: near death, life and technology</title>
      <link>https://iain.so/at-50-near-death-life-and-technology?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[Today, I turn 50. &#xA;&#xA;During my life, there have been two occasions when I have been very close to death. As a two-year-old, I was diagnosed with and treated for a cancer that, just a year earlier, had no cure.&#xA;&#xA;Twenty years later, I nearly died a second time.&#xA;&#xA;I’ve always been obsessed with technology and the minutiae of how things work. I grew up at the tail end of the Acid House era of illegal orbital raves, so-called because they were held near the M25, the large arterial road that coils around London.&#xA;&#xA;Several pirate radio stations disseminated the cryptic details of these parties and their locations. I became infatuated with how these broadcast operations worked, the FM radio technology they used, and how they stayed one step ahead of the law. &#xA;&#xA;Fast forward a few years, and I was deeply involved in running just such a pirate station in Nottingham. We played a continual cat-and-mouse game with the Radio Investigation Service, part of the Government’s Department of Trade and Industry, later Ofcom. Their job was to track our transmissions, and our job was to make it as hard as possible for them to find us (good-quality FM transmitters are very expensive).&#xA;&#xA;abstract image of a man dangling from the number 50&#xA;&#xA;At the time, our transmission site was atop a five-storey warehouse, nestled at the summit of a tall hill overlooking Nottingham. We paid the owner to look the other way and profess ignorance if the authorities paid a visit.&#xA;&#xA;Another precaution was that after each weekend’s broadcasts (we eventually went 24/7), we would remove the transmission mast and hide it behind a short section of steeply pitched roof. To remove or remount it, we usually straddled the roof ridge and scooted along on our bottoms like riding a horse. When returning with the aerial, I usually entertained myself by imagining myself as a medieval knight jousting for the hand of a beautiful, voluptuous maiden.&#xA;&#xA;There was a small roster of trusted individuals who did this mundane but essential job every weekend. And so it went on for many months. We typically went on air late on a Friday afternoon. One weekend, it was my turn again. I had a date in town directly after, so I was dressed to impress, including some rather natty leather-soled shoes and a pair of Paul Smith trousers. Not ideal rooftop attire.&#xA;&#xA;There had been some light rain in the afternoon, and things started inauspiciously when, on arrival, I was buttonholed by the warehouse owner, who was irate. It transpired he had been visited by the authorities during the week and threatened with legal action. “You are not paying me enough for this shit; you find somewhere else.”&#xA;&#xA;I was able to mollify him, saying that it was likely a bluff (untrue - Ofcom has surprisingly sweeping powers), I was just a lowly minion and that the station’s management would be in touch (true - I was quite low down the food chain and all successful pirate radio stations I’ve been involved with are run with a level of hierarchy and professionalism that would surprise outsiders).&#xA;&#xA;But I was now running behind and in danger of being late for my date (mobile phones were not common at the time). So I ran up the concrete stairs and exited a small hatch at the top of the lift shaft which gave me access to the roof.&#xA;&#xA;I guess the regularity with which I’d done this had bred complacency, and, conscious of not soiling my Paul Smiths, I crossed the short section of pitched roof (maybe six metres), using the ridge line as a handhold instead of straddling it, with my feet flat on the steeply angled roof tiles, leaning into its slope.&#xA;&#xA;The FM transmission mast was there as expected, safely hidden. All that remained was to cross back and connect it to the coaxial cable that snaked up from the FM transmitter locked deep in the bowels of the warehouse.&#xA;&#xA;Whilst the aerial was not huge, it was mounted on a metal pole for additional height, so was probably two metres in length and unwieldy. I grabbed it and returned across the roof in the same way; right hand on the roof ridge, aerial in left hand.&#xA;&#xA;But my leather-soled shoes offered little grip, and this side of the roof faced the weather, so had accumulated moss and slime. What happened next had, in retrospect, a certain inevitability. My feet lost grip on the roof, and I dropped the aerial, which clattered down the tiles and fell five stories to the ground below, landing in a mangled heap.&#xA;&#xA;Destabilised, I lost hold of the roof ridge and slithered after the aerial over the edge. I somehow managed to grab on to the non-too-solid guttering, but the rest of my body was dangling Buster Keaton style, precariously in thin air, with nothing between me and almost certain death five storeys below.&#xA;&#xA;You’re probably familiar with the saying “your life flashes before your eyes”, used so often to describe near-death experiences that it has become a cliché. All I can offer is the version I experienced.&#xA;&#xA;As I hung there, literally in limbo between life and death, I recall some striking and vivid things. First, and perhaps most surprising, I felt absolutely no fear, only calm mental clarity. I had a sense that time had stretched enormously. It wasn’t so much life flashing in front of my eyes like a film reel on fast forward, more an IMAX-style panoramic sweep of memories.&#xA;&#xA;I saw myself labouring up the final pitch before the summit of Mont Blanc, exhausted and deeply altitude sick but knowing I would make it. I was riding pillion on a Kawasaki Ninja flashing across the Golden Gate Bridge through a pink dawn mist. I felt the human energy flowing from the dance floor as I nervously warmed up for Carl Cox at the infamous Marcus Garvey Ballroom, my hands shaking so much I could barely put the needle on a record. I was a child poking my feet into the warm powdery coral sand of a Bermudian beach, my back propped against our huge old dog. &#xA;&#xA;All this was accompanied by astoundingly beautiful, otherworldly music of a kind I’ve never been able to describe adequately and that does not exist in this world. I can still retrieve tiny fragments from memory, and I hope to hear it again someday. &#xA;&#xA;But this imagery was simultaneously accompanied by other less serene aspects. I relived several emotionally charged situations from the perspectives of others negatively affected by my actions. I experienced the horrible way I broke up with an old girlfriend, feeling it exactly as she felt it, with full emotional force and pain. It was as if my perspective had become hers. &#xA;&#xA;Perhaps this was some moral reckoning. It certainly baked itself into a human operating system that was a little underdeveloped at the time. “Be kind and think how your actions will affect others”. I don’t know how those words solidified, but they’ve stayed with me ever since. I do my best to remember them.&#xA;&#xA;I have no way of knowing the actual duration of this fugue state. It can’t have been long in real time. The biological adrenaline eventually kicked in, and somehow, after several attempts and with my strength nearly gone, I managed to hook a leg over the gutter and crawl feebly to safety.&#xA;&#xA;I even made my date more or less on time. As I took my seat at the table, a look of shock and concern crossed her face, “Are you ok? You’re a very strange colour”. I excused myself and went to the bathroom. The face looking back at me in the mirror was a spectral, etiolated grey-green of deep physical shock.&#xA;&#xA;Eventually, internet streaming made FM pirate radio stations obsolete. Their role had been to play the underground music you couldn’t hear anywhere else, because the airwaves were so rigidly controlled. Now you could hear whatever you wanted anywhere, anytime. Technology improved and times changed. So it goes.&#xA;&#xA;When you turn 50, there’s the classic version of midlife: panic at the thought that more than half of your life is gone, at having reached the actuarial midpoint. Motorbikes and significantly younger wives sometimes follow.&#xA;&#xA;Twice I’ve come about as close as it is possible to death, so the years since have always seemed like a kind of surplus. Not borrowed time exactly, but a different relationship with it. Being fifty feels like a point on a line that might not have been there at all. &#xA;&#xA;I was born in the mid-1970s, which puts me in the narrow cohort who had a fully analogue childhood and an exponentially digital adulthood. I learned to read from books, to navigate around London with an A-Z on the passenger seat, and listened to music on vinyl and cassette tapes. &#xA;&#xA;My first online connection was via an acoustic coupler, where a loud cough was enough to disrupt it. My university essays were handwritten. Then, as I joined the workforce, it reorganised itself around silicon (although my first job didn’t initially have email, which boggles the minds of anyone under 20 and always elicits the question “what did you do!?”). I reorganised alongside it, and I found I loved it. &#xA;&#xA;AI is now dissolving much of what I’ve spent my career doing. It is said there’s a double exponential at work, both in financial investment and in the number of smart people entering AI research. If Covid taught us anything, it is that we don’t understand the power of single exponentials, let alone double ones. If exponentials are slowly, slowly; all at once, maybe double exponentials are faster, faster; unimaginable change. Either way, great upheaval is coming much more quickly than we seem prepared for. &#xA;&#xA;I am not complacent, but I have already watched a settled world get replaced by a faster and different one before, twice in fact. The web arrived and changed what information was, where it lived, upending power structures and business models. The smartphone arrived and changed what a computer was, where it was used and our degree of connectedness (for good and ill). AI is already orders of magnitude more consequential, and it has barely got started. It will not be easy; history has shown us that powerful technology enables the full spectrum of humanity, from its very best to its most utterly evil. &#xA;&#xA;My half-century arrives at precisely the moment of maximum disruption to everything I’ve previously learnt and built. We are dangling over the edge of something powerful, strange and uncertain. But I’ve been there before and felt no fear.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>Today, I turn 50.</p>

<p>During my life, there have been two occasions when I have been very close to death. As a two-year-old, I was diagnosed with and treated for a cancer that, just a year earlier, had no cure.</p>

<p>Twenty years later, I nearly died a second time.</p>

<p>I’ve always been obsessed with technology and the minutiae of how things work. I grew up at the tail end of the Acid House era of illegal orbital raves, so-called because they were held near the M25, the large arterial road that coils around London.</p>

<p>Several pirate radio stations disseminated the cryptic details of these parties and their locations. I became infatuated with how these broadcast operations worked, the FM radio technology they used, and how they stayed one step ahead of the law.</p>

<p>Fast forward a few years, and I was deeply involved in running just such a pirate station in Nottingham. We played a continual cat-and-mouse game with the Radio Investigation Service, part of the Government’s Department of Trade and Industry, later Ofcom. Their job was to track our transmissions, and our job was to make it as hard as possible for them to find us (good-quality FM transmitters are very expensive).</p>

<p><img src="https://i.snap.as/6n1oiwA0.png" alt="abstract image of a man dangling from the number 50"/></p>

<p>At the time, our transmission site was atop a five-storey warehouse, nestled at the summit of a tall hill overlooking Nottingham. We paid the owner to look the other way and profess ignorance if the authorities paid a visit.</p>

<p>Another precaution was that after each weekend’s broadcasts (we eventually went 24/7), we would remove the transmission mast and hide it behind a short section of steeply pitched roof. To remove or remount it, we usually straddled the roof ridge and scooted along on our bottoms like riding a horse. When returning with the aerial, I usually entertained myself by imagining myself as a medieval knight jousting for the hand of a beautiful, voluptuous maiden.</p>

<p>There was a small roster of trusted individuals who did this mundane but essential job every weekend. And so it went on for many months. We typically went on air late on a Friday afternoon. One weekend, it was my turn again. I had a date in town directly after, so I was dressed to impress, including some rather natty leather-soled shoes and a pair of Paul Smith trousers. Not ideal rooftop attire.</p>

<p>There had been some light rain in the afternoon, and things started inauspiciously when, on arrival, I was buttonholed by the warehouse owner, who was irate. It transpired he had been visited by the authorities during the week and threatened with legal action. “You are not paying me enough for this shit; you find somewhere else.”</p>

<p>I was able to mollify him, saying that it was likely a bluff (untrue – Ofcom has surprisingly sweeping powers), I was just a lowly minion and that the station’s management would be in touch (true – I was quite low down the food chain and all successful pirate radio stations I’ve been involved with are run with a level of hierarchy and professionalism that would surprise outsiders).</p>

<p>But I was now running behind and in danger of being late for my date (mobile phones were not common at the time). So I ran up the concrete stairs and exited a small hatch at the top of the lift shaft which gave me access to the roof.</p>

<p>I guess the regularity with which I’d done this had bred complacency, and, conscious of not soiling my Paul Smiths, I crossed the short section of pitched roof (maybe six metres), using the ridge line as a handhold instead of straddling it, with my feet flat on the steeply angled roof tiles, leaning into its slope.</p>

<p>The FM transmission mast was there as expected, safely hidden. All that remained was to cross back and connect it to the coaxial cable that snaked up from the FM transmitter locked deep in the bowels of the warehouse.</p>

<p>Whilst the aerial was not huge, it was mounted on a metal pole for additional height, so was probably two metres in length and unwieldy. I grabbed it and returned across the roof in the same way; right hand on the roof ridge, aerial in left hand.</p>

<p>But my leather-soled shoes offered little grip, and this side of the roof faced the weather, so had accumulated moss and slime. What happened next had, in retrospect, a certain inevitability. My feet lost grip on the roof, and I dropped the aerial, which clattered down the tiles and fell five stories to the ground below, landing in a mangled heap.</p>

<p>Destabilised, I lost hold of the roof ridge and slithered after the aerial over the edge. I somehow managed to grab on to the non-too-solid guttering, but the rest of my body was dangling Buster Keaton style, precariously in thin air, with nothing between me and almost certain death five storeys below.</p>

<p>You’re probably familiar with the saying “your life flashes before your eyes”, used so often to describe near-death experiences that it has become a cliché. All I can offer is the version I experienced.</p>

<p>As I hung there, literally in limbo between life and death, I recall some striking and vivid things. First, and perhaps most surprising, I felt absolutely no fear, only calm mental clarity. I had a sense that time had stretched enormously. It wasn’t so much life flashing in front of my eyes like a film reel on fast forward, more an IMAX-style panoramic sweep of memories.</p>

<p>I saw myself labouring up the final pitch before the summit of Mont Blanc, exhausted and deeply altitude sick but knowing I would make it. I was riding pillion on a Kawasaki Ninja flashing across the Golden Gate Bridge through a pink dawn mist. I felt the human energy flowing from the dance floor as I nervously warmed up for Carl Cox at the infamous Marcus Garvey Ballroom, my hands shaking so much I could barely put the needle on a record. I was a child poking my feet into the warm powdery coral sand of a Bermudian beach, my back propped against our huge old dog.</p>

<p>All this was accompanied by astoundingly beautiful, otherworldly music of a kind I’ve never been able to describe adequately and that does not exist in this world. I can still retrieve tiny fragments from memory, and I hope to hear it again someday.</p>

<p>But this imagery was simultaneously accompanied by other less serene aspects. I relived several emotionally charged situations from the perspectives of others negatively affected by my actions. I experienced the horrible way I broke up with an old girlfriend, feeling it exactly as she felt it, with full emotional force and pain. It was as if my perspective had become hers.</p>

<p>Perhaps this was some moral reckoning. It certainly baked itself into a human operating system that was a little underdeveloped at the time. “Be kind and think how your actions will affect others”. I don’t know how those words solidified, but they’ve stayed with me ever since. I do my best to remember them.</p>

<p>I have no way of knowing the actual duration of this fugue state. It can’t have been long in real time. The biological adrenaline eventually kicked in, and somehow, after several attempts and with my strength nearly gone, I managed to hook a leg over the gutter and crawl feebly to safety.</p>

<p>I even made my date more or less on time. As I took my seat at the table, a look of shock and concern crossed her face, “Are you ok? You’re a <em>very</em> strange colour”. I excused myself and went to the bathroom. The face looking back at me in the mirror was a spectral, etiolated grey-green of deep physical shock.</p>

<p>Eventually, internet streaming made FM pirate radio stations obsolete. Their role had been to play the underground music you couldn’t hear anywhere else, because the airwaves were so rigidly controlled. Now you could hear whatever you wanted anywhere, anytime. Technology improved and times changed. So it goes.</p>

<p>When you turn 50, there’s the classic version of midlife: panic at the thought that more than half of your life is gone, at having reached the actuarial midpoint. Motorbikes and significantly younger wives sometimes follow.</p>

<p>Twice I’ve come about as close as it is possible to death, so the years since have always seemed like a kind of surplus. Not borrowed time exactly, but a different relationship with it. Being fifty feels like a point on a line that might not have been there at all.</p>

<p>I was born in the mid-1970s, which puts me in the narrow cohort who had a fully analogue childhood and an exponentially digital adulthood. I learned to read from books, to navigate around London with an A-Z on the passenger seat, and listened to music on vinyl and cassette tapes.</p>

<p>My first online connection was via an acoustic coupler, where a loud cough was enough to disrupt it. My university essays were handwritten. Then, as I joined the workforce, it reorganised itself around silicon (although my first job didn’t initially have email, which boggles the minds of anyone under 20 and always elicits the question “what did you <em>do</em>!?”). I reorganised alongside it, and I found I loved it.</p>

<p>AI is now dissolving much of what I’ve spent my career doing. It is said there’s a double exponential at work, both in financial investment and in the number of smart people entering AI research. If Covid taught us anything, it is that we don’t understand the power of single exponentials, let alone double ones. If exponentials are slowly, slowly; all at once, maybe double exponentials are faster, faster; unimaginable change. Either way, great upheaval is coming much more quickly than we seem prepared for.</p>

<p>I am not complacent, but I have already watched a settled world get replaced by a faster and different one before, twice in fact. The web arrived and changed what information was, where it lived, upending power structures and business models. The smartphone arrived and changed what a computer was, where it was used and our degree of connectedness (for good and ill). AI is already orders of magnitude more consequential, and it has barely got started. It will not be easy; history has shown us that powerful technology enables the full spectrum of humanity, from its very best to its most utterly evil.</p>

<p>My half-century arrives at precisely the moment of maximum disruption to everything I’ve previously learnt and built. We are dangling over the edge of something powerful, strange and uncertain. But I’ve been there before and felt no fear.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/at-50-near-death-life-and-technology</guid>
      <pubDate>Thu, 23 Jul 2026 17:38:14 +0000</pubDate>
    </item>
    <item>
      <title>Google AI Overviews - fuck you, pay me</title>
      <link>https://iain.so/google-ai-overviews-fuck-you-pay-me?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[For two weeks in January 2026, a browser extension removed AI Overviews from the Google search results of 374 people. The remaining links moved silently up to fill the gap, so the interface looked just as the search always had. At the end of the two weeks, researchers asked participants how they felt about their search experience: satisfaction, quality, and ease of finding information. On every measure, the numbers matched the group that kept their overviews. Nobody missed them. The only difference was that people without overviews left Google for other sites 67% more often.&#xA;&#xA;This field experiment by Saharsh Agarwal of the Indian School of Business and Ananya Sen of Carnegie Mellon, published on SSRN in April and revised in June 2026, produced the first causal measure of what Google&#39;s AI Overviews are costing the rest of the web. The answer is 40% of outbound organic clicks on queries where overviews appear, with no measurable benefit for the people doing the searching. On 20 May, at Google I/O, the company made AI Mode, that demonstrably made no improvement to the user experience, the default search experience worldwide.&#xA;&#xA;Still image from the movie Goodfellas&#xA;&#xA;What the experiment measured&#xA;&#xA;The study recruited 1,065 US desktop Chrome users through Prolific and randomly assigned them to one of three groups. A control group saw Google as normal. A treatment group had AI Overviews removed in real time by the extension whenever they would otherwise have appeared, with organic results shifting up to fill the gap so cleanly that over 95% of participants reported noticing nothing. A third group was redirected into Google&#39;s AI Mode for all queries.&#xA;&#xA;The experiment design settles a methodological argument that has run since the first observational studies appeared. Ahrefs, Pew Research, and Similarweb all pointed at the same traffic decline, but none could prove that the overviews caused it rather than some other change in search behaviour or Google&#39;s algorithm. Agarwal and Sen&#39;s extension created the counterfactual that observational data cannot. The only systematic difference between the two primary groups was whether the overview appeared. Everything else–the queries, the organic results, the ads–was the same search engine on the same day.&#xA;&#xA;The numbers were stark. For queries where an overview appeared (about 41% of all searches), hiding it increased outbound organic clicks from 0.37 to 0.62 per search. Sponsored clicks stayed at 0.02 in both groups. Total search volume did not change. The overview did not create anything; it simply kept the 40% of clicks it absorbed inside Google.&#xA;&#xA;After the main two weeks, the researchers swapped the conditions. The group that had been browsing without overviews got them back, and their outbound clicks fell from 0.61 to 0.33 per search. The group that had been seeing overviews lost them, and their clicks rose from 0.33 to 0.55. Same people, opposite intervention, mirror-image result.&#xA;&#xA;The &#34;higher-quality clicks&#34; defence&#xA;&#xA;Google&#39;s vice president of product for Search, Liz Reid, has characterised AI Overview clicks as &#34;higher-quality clicks&#34; that signal stronger purchase intent and longer downstream engagement. In April 2026, the company advanced a &#34;bounce clicks&#34; explanation, arguing that overviews mainly remove low-value visits — the kind where a user lands, finds nothing useful, and taps the back button within a few seconds.&#xA;&#xA;Agarwal and Sen tested that claim directly. For every click that reached a downstream website, they measured whether the user bounced (left within ten seconds with no further navigation), how long they stayed on the page, and whether they hit the back button to return to search. Across all three measures, there was no difference between the group with overviews and the group without. The extra clicks generated by removing overviews were no different in quality from the clicks Google&#39;s system allowed through. The authors are careful to note that Google may measure something further down the funnel that the extension cannot observe. But at the top of the funnel, where the &#34;higher-quality&#34; claim was made, nothing supports it.&#xA;&#xA;In contrast, Adobe Digital Insights&#39; Q1 2026 analysis found that visitors arriving from AI assistants converted 42% better than non-AI traffic, reversing March 2025, when the same channel converted 38% worse. The volume is small, roughly 1% of total site traffic for most businesses, but the quality gap matters for anyone trying to value what remains.&#xA;&#xA;Although the findings appear contradictory, They measure different things, so both can be true. Agarwal and Sen answered: does the AI overview filter out only the junk clicks? No. The clicks it suppressed were the same quality as the ones that got through. It just reduced volume.&#xA;&#xA;Adobe looked wider but shallower. They took all traffic across thousands of sites and split it two ways: came from an AI assistant, or didn’t. Then compared conversion rates. That’s broad, but it can’t answer the Google question, because “AI assistant traffic” doesn’t isolate Google’s overview clicks. Those just get lumped into ordinary search traffic. So Adobe tells you assistant referrals convert well; it tells you nothing about whether Google’s surviving clicks are any good. Agarwal and Sen looked only at Google: narrow but deep: one channel, controlled comparison, cause and effect.&#xA;&#xA;The mechanism, however, is straightforward: the AI overview wins the way anything wins in position zero. In the study, 87% of AI Overviews appeared above all organic results. When the overview sat at the top, removing it led to an 88% increase in outbound clicks. When it appeared lower on the page, the effect disappeared. People did not seek out the overview. They consumed whatever occupied the prime on-page real estate, the way supermarket shoppers take whatever sits at eye level on the shelf. Move the product, and the behaviour reverses instantly.&#xA;&#xA;Forced into Google&#39;s fully conversational search interface, users rated their satisfaction at 2.9 out of 5, compared with 4.0 for both other groups. Attrition was roughly four times higher. Some participants installed workaround extensions to escape. This is the experience Google chose, on 20 May, to make the global default. Record queries and accelerating search revenue explain the decision. The fact that their own experiment&#39;s closest analogue produced the lowest satisfaction scores and the highest quit rates does not seem to have figured in the calculus.&#xA;&#xA;Meanwhile, roughly twenty national news outlets are negotiating licensing deals for the privilege of appearing in these overviews, with Google winding down its older news-payment programme and making future payments conditional on granting AI training rights. If you are not one of those twenty, nobody is coming to the table. The response has to be self-directed.&#xA;&#xA;Read your own exposure&#xA;&#xA;Not all traffic is equally at risk. Agarwal and Sen classified queries into informational, navigational, and transactional categories. Informational queries accounted for 71% of all searches in the study and had a 53% AI Overview trigger rate. This is where the entire effect played out. Navigational searches (someone typing a brand name to find the website) triggered overviews in only 6% of queries, and transactional searches (purchase-intent, booking, downloading) in 15%. Neither category showed a statistically meaningful decline in clicks.&#xA;&#xA;This means that if your search traffic comes through &#34;what is&#34; and &#34;how to&#34; queries, you are sitting ducks. If it comes from people searching for your name or your products by name, the damage is minor and may stay that way. Open your Search Console, filter for query type, and look at the split. The ratio between informational and branded traffic is the single best predictor of how badly this will hurt.&#xA;&#xA;For many smaller businesses, an audit will show that informational traffic was always the least valuable, even though it has traditionally been encouraged to build domain authority. A plumber whose search visibility comes from &#34;how to fix a leaking tap&#34; was getting visits from people actively trying not to hire a plumber. That traffic was always thin on conversions, and the overview is now absorbing exactly those clicks. The branded traffic, from someone searching for the plumber by name after a recommendation, was always the higher-value stream, and it remains largely untouched.&#xA;&#xA;Stop producing what the machine can summarise&#xA;&#xA;The clicks that persist in a world of overviews share a common trait: the user wanted something the summary could not provide. A named opinion. A tool. A dataset. An experience report. This matches the evidence from the GEO literature I have previously reviewed. The Princeton benchmark found that content with sourced statistics, named quotations, and inline citations earned more citations from AI systems. ZipTie&#39;s cross-platform analysis found that sites with high original-data density received 4.3 times more citation occurrences per URL than directory-style listings. The surviving trickle rewards specificity. What a summary can reproduce, the summary will absorb. What it cannot reproduce, because it requires a person to have done, measured, or decided something, retains its click.&#xA;&#xA;For a smaller business, this completely changes the question of content. The commodity &#34;what is X&#34; post, the 1,200-word explainer written to capture an informational keyword, was always a bet on Google continuing to send traffic to the answer. This content type has taken a 40% haircut and already converted poorly. That is a signal to stop producing it, or at least to stop treating it as a customer acquisition channel. Spend that same effort on content the summary cannot assimilate. Publish your own data, even if the dataset is small. Name your prices, your methods, your results, your opinions. Write the thing only you could write because only you did the work, and leave the commodity answer for the overview.&#xA; Fix your GA4 attribution before drawing conclusions. AI referral traffic typically defaults into &#34;Direct&#34; or &#34;Referral&#34; buckets, and only 14% of marketers track it as a separate channel. You cannot manage what you have filed under miscellaneous.&#xA;&#xA;Build the channels the intermediary cannot close&#xA;&#xA;One result in the study was completely unaffected by the presence or absence of overviews: ad clicks. Free organic traffic from search fell 40%. Paid traffic held steady. If your free visitors from Google are disappearing, the route Google left open is the one you pay for, and that was clearly not an accident. If you start to spend more on Google Ads to replace the traffic it used to send for free, it should be watched closely for true return on investment.&#xA;&#xA;So what should be done? Google has remained an intermediary for organic search traffic long after others cut it off at the knees to drive ad revenue. But the organic game has become increasingly difficult as ad slots and other additions to the search engine results page have made the real organic links less and less obvious. The direction of travel has been clear for many years.&#xA;&#xA;Relying on an intermediary’s whims for traffic, even one as durable as Google, has always been a Hobson’s Choice. As its presence continues to fade, we return to basics: email lists, direct bookmarks, repeat visits, communities, and all the channels that do not pass through a search results page. &#xA;&#xA;None of this is new advice; the case for greater focus on owned channels has been made for a decade, and each year the argument grew stronger and the urgency louder while the execution stayed more or less the same. The difference now is that experimental evidence puts a number on the cost. Every month that informational traffic goes unprotected by a direct relationship is another month in which 40% of those visits vanish into a summary.&#xA;&#xA;One caveat deserves emphasis. Do not block AI crawlers outright. The major labs have split their bots into training crawlers and search crawlers. OpenAI separated GPTBot from OAI-SearchBot in late 2024, and Anthropic made the same split with ClaudeBot and Claude-SearchBot. Blocking the training crawler is your call. Blocking the search crawler cuts you out of the AI answers that are replacing the links you used to get. A Rutgers and Wharton study published in December 2025 found that publishers who blocked AI bots across the board experienced a 23% traffic decline compared with peers who allowed crawling. Google is the harder case, because Googlebot still handles both search indexing and AI features in a single crawler, which is exactly the bundling the CMA’s conduct requirements are trying to force apart.&#xA;&#xA;The empty space&#xA;&#xA;The 374 participants whose overviews were silently removed are, as far as the published literature shows, the only group of people to have experienced modern Google search without AI-generated summaries and been formally asked what they thought. They reported feeling no difference whatsoever. &#xA;&#xA;That absence is visible only from the other side of the results page, where a company watches the traffic graph flatten and wonders whether visitors are finding better answers or simply never leaving the search process. Agarwal and Sen&#39;s contribution is to show that the latter is the case. The visitors are not finding better answers; they are milling around in the foyer before going home.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>For two weeks in January 2026, a browser extension removed AI Overviews from the Google search results of 374 people. The remaining links moved silently up to fill the gap, so the interface looked just as the search always had. At the end of the two weeks, researchers asked participants how they felt about their search experience: satisfaction, quality, and ease of finding information. On every measure, the numbers matched the group that kept their overviews. Nobody missed them. The only difference was that people without overviews left Google for other sites 67% more often.</p>

<p>This <a href="https://ssrn.com/abstract=6513059">field experiment by Saharsh Agarwal of the Indian School of Business and Ananya Sen of Carnegie Mellon</a>, published on SSRN in April and revised in June 2026, produced the first causal measure of what Google&#39;s AI Overviews are costing the rest of the web. The answer is 40% of outbound organic clicks on queries where overviews appear, with no measurable benefit for the people doing the searching. On 20 May, at Google I/O, the company made AI Mode, that demonstrably made no improvement to the user experience, the default search experience worldwide.</p>

<p><img src="https://i.snap.as/J0oKzCH4.jpeg" alt="Still image from the movie Goodfellas"/></p>

<h2 id="what-the-experiment-measured">What the experiment measured</h2>

<p>The study recruited 1,065 US desktop Chrome users through Prolific and randomly assigned them to one of three groups. A control group saw Google as normal. A treatment group had AI Overviews removed in real time by the extension whenever they would otherwise have appeared, with organic results shifting up to fill the gap so cleanly that <a href="https://ssrn.com/abstract=6513059">over 95% of participants reported noticing nothing</a>. A third group was redirected into Google&#39;s AI Mode for all queries.</p>

<p>The experiment design settles a methodological argument that has run since the first observational studies appeared. <a href="https://ahrefs.com/blog/ai-overviews-reduce-clicks/">Ahrefs</a>, <a href="https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/">Pew Research</a>, and Similarweb all pointed at the same traffic decline, but none could prove that the overviews caused it rather than some other change in search behaviour or Google&#39;s algorithm. Agarwal and Sen&#39;s extension created the counterfactual that observational data cannot. The only systematic difference between the two primary groups was whether the overview appeared. Everything else–the queries, the organic results, the ads–was the same search engine on the same day.</p>

<p>The numbers were stark. For queries where an overview appeared (about 41% of all searches), hiding it increased outbound organic clicks from 0.37 to 0.62 per search. Sponsored clicks stayed at 0.02 in both groups. Total search volume did not change. The overview did not create anything; it simply kept the 40% of clicks it absorbed inside Google.</p>

<p>After the main two weeks, the researchers swapped the conditions. The group that had been browsing without overviews got them back, and their outbound clicks fell from 0.61 to 0.33 per search. The group that had been seeing overviews lost them, and their clicks <a href="https://ssrn.com/abstract=6513059">rose from 0.33 to 0.55</a>. Same people, opposite intervention, mirror-image result.</p>

<h2 id="the-higher-quality-clicks-defence">The “higher-quality clicks” defence</h2>

<p>Google&#39;s vice president of product for Search, Liz Reid, has <a href="https://ppc.land/researchers-find-google-ai-overviews-cut-publisher-clicks-39-8/">characterised AI Overview clicks as “higher-quality clicks”</a> that signal stronger purchase intent and longer downstream engagement. In April 2026, the company advanced a <a href="https://www.searchenginejournal.com/seo-pulse-ai-overviews-clicks-get-tested-earnings-tell-two-stories/573505/">“bounce clicks” explanation</a>, arguing that overviews mainly remove low-value visits — the kind where a user lands, finds nothing useful, and taps the back button within a few seconds.</p>

<p>Agarwal and Sen tested that claim directly. For every click that reached a downstream website, they measured whether the user bounced (left within ten seconds with no further navigation), how long they stayed on the page, and whether they hit the back button to return to search. Across all three measures, <a href="https://ssrn.com/abstract=6513059">there was no difference between the group with overviews and the group without</a>. The extra clicks generated by removing overviews were no different in quality from the clicks Google&#39;s system allowed through. The authors are careful to note that Google may measure something further down the funnel that the extension cannot observe. But at the top of the funnel, where the “higher-quality” claim was made, nothing supports it.</p>

<p>In contrast, <a href="https://www.digitalapplied.com/blog/ai-traffic-converts-42-percent-better-2026-channel-strategy">Adobe Digital Insights&#39; Q1 2026 analysis</a> found that visitors arriving from AI assistants converted 42% better than non-AI traffic, reversing March 2025, when the same channel converted 38% worse. The volume is small, roughly 1% of total site traffic for most businesses, but the quality gap matters for anyone trying to value what remains.</p>

<p>Although the findings appear contradictory, They measure different things, so both can be true. Agarwal and Sen answered: does the AI overview filter out only the junk clicks? No. The clicks it suppressed were the same quality as the ones that got through. It just reduced volume.</p>

<p>Adobe looked wider but shallower. They took all traffic across thousands of sites and split it two ways: came from an AI assistant, or didn’t. Then compared conversion rates. That’s broad, but it can’t answer the Google question, because “AI assistant traffic” doesn’t isolate Google’s overview clicks. Those just get lumped into ordinary search traffic. So Adobe tells you assistant referrals convert well; it tells you nothing about whether Google’s surviving clicks are any good. Agarwal and Sen looked only at Google: narrow but deep: one channel, controlled comparison, cause and effect.</p>

<p>The mechanism, however, is straightforward: the AI overview wins the way anything wins in position zero. In the study, <a href="https://ssrn.com/abstract=6513059">87% of AI Overviews appeared above all organic results</a>. When the overview sat at the top, removing it led to an 88% increase in outbound clicks. When it appeared lower on the page, the effect disappeared. People did not seek out the overview. They consumed whatever occupied the prime on-page real estate, the way supermarket shoppers take whatever sits at eye level on the shelf. Move the product, and the behaviour reverses instantly.</p>

<p>Forced into Google&#39;s fully conversational search interface, users rated their satisfaction at 2.9 out of 5, compared with 4.0 for both other groups. Attrition was roughly four times higher. Some participants installed workaround extensions to escape. This is the experience Google chose, on 20 May, to make the global default. Record queries and accelerating search revenue explain the decision. The fact that their own experiment&#39;s closest analogue produced the lowest satisfaction scores and the highest quit rates does not seem to have figured in the calculus.</p>

<p>Meanwhile, roughly twenty national news outlets are <a href="https://cryptobriefing.com/google-ai-licensing-talks-publishers/">negotiating licensing deals</a> for the privilege of appearing in these overviews, with Google <a href="https://www.mediapost.com/publications/article/416113/google-eyes-broader-publisher-licensing.html">winding down its older news-payment programme and making future payments conditional on granting AI training rights</a>. If you are not one of those twenty, nobody is coming to the table. The response has to be self-directed.</p>

<h2 id="read-your-own-exposure">Read your own exposure</h2>

<p>Not all traffic is equally at risk. Agarwal and Sen classified queries into informational, navigational, and transactional categories. <a href="https://ssrn.com/abstract=6513059">Informational queries</a> accounted for 71% of all searches in the study and had a 53% AI Overview trigger rate. This is where the entire effect played out. Navigational searches (someone typing a brand name to find the website) triggered overviews in only 6% of queries, and transactional searches (purchase-intent, booking, downloading) in 15%. Neither category showed a statistically meaningful decline in clicks.</p>

<p>This means that if your search traffic comes through “what is” and “how to” queries, you are sitting ducks. If it comes from people searching for your name or your products by name, the damage is minor and may stay that way. Open your Search Console, filter for query type, and look at the split. The ratio between informational and branded traffic is the single best predictor of how badly this will hurt.</p>

<p>For many smaller businesses, an audit will show that informational traffic was always the least valuable, even though it has traditionally been encouraged to build domain authority. A plumber whose search visibility comes from “how to fix a leaking tap” was getting visits from people actively trying not to hire a plumber. That traffic was always thin on conversions, and the overview is now absorbing exactly those clicks. The branded traffic, from someone searching for the plumber by name after a recommendation, was always the higher-value stream, and it remains largely untouched.</p>

<h2 id="stop-producing-what-the-machine-can-summarise">Stop producing what the machine can summarise</h2>

<p>The clicks that persist in a world of overviews share a common trait: the user wanted something the summary could not provide. A named opinion. A tool. A dataset. An experience report. This matches the evidence from <a href="https://betterthangood.xyz/blog/geo-practice-versus-snake-oil/">the GEO literature I have previously reviewed</a>. The Princeton benchmark found that content with sourced statistics, named quotations, and inline citations earned more citations from AI systems. <a href="https://ziptie.dev/blog/how-original-research-wins-ai-citations/">ZipTie&#39;s cross-platform analysis</a> found that sites with high original-data density received 4.3 times more citation occurrences per URL than directory-style listings. The surviving trickle rewards specificity. What a summary can reproduce, the summary will absorb. What it cannot reproduce, because it requires a person to have done, measured, or decided something, retains its click.</p>

<p>For a smaller business, this completely changes the question of content. The commodity “what is X” post, the 1,200-word explainer written to capture an informational keyword, was always a bet on Google continuing to send traffic to the answer. This content type has taken a 40% haircut and already converted poorly. That is a signal to stop producing it, or at least to stop treating it as a customer acquisition channel. Spend that same effort on content the summary cannot assimilate. Publish your own data, even if the dataset is small. Name your prices, your methods, your results, your opinions. Write the thing only you could write because only you did the work, and leave the commodity answer for the overview.
 Fix your GA4 attribution before drawing conclusions. AI referral traffic typically defaults into “Direct” or “Referral” buckets, and <a href="https://emarketed.com/aeo/ai-referral-traffic-conversion-rates-2026/">only 14% of marketers track it as a separate channel</a>. You cannot manage what you have filed under miscellaneous.</p>

<h2 id="build-the-channels-the-intermediary-cannot-close">Build the channels the intermediary cannot close</h2>

<p>One result in the study was completely unaffected by the presence or absence of overviews: ad clicks. Free organic traffic from search fell 40%. Paid traffic held steady. If your free visitors from Google are disappearing, the route Google left open is the one you pay for, and that was clearly not an accident. If you start to spend more on Google Ads to replace the traffic it used to send for free, it should be watched closely for true return on investment.</p>

<p>So what should be done? Google has remained an intermediary for organic search traffic long after others cut it off at the knees to drive ad revenue. But the organic game has become increasingly difficult as ad slots and other additions to the search engine results page have made the real organic links less and less obvious. The direction of travel has been clear for many years.</p>

<p>Relying on an intermediary’s whims for traffic, even one as durable as Google, has always been a Hobson’s Choice. As its presence continues to fade, we return to basics: email lists, direct bookmarks, repeat visits, communities, and all the channels that do not pass through a search results page.</p>

<p>None of this is new advice; the case for greater focus on owned channels has been made for a decade, and each year the argument grew stronger and the urgency louder while the execution stayed more or less the same. The difference now is that experimental evidence puts a number on the cost. Every month that informational traffic goes unprotected by a direct relationship is another month in which 40% of those visits vanish into a summary.</p>

<p>One caveat deserves emphasis. Do not block AI crawlers outright. The major labs have split their bots into training crawlers and search crawlers. OpenAI separated GPTBot from OAI-SearchBot in late 2024, and Anthropic made the same split with ClaudeBot and Claude-SearchBot. Blocking the training crawler is your call. Blocking the search crawler cuts you out of the AI answers that are replacing the links you used to get. A Rutgers and Wharton study published in December 2025 found that publishers who blocked AI bots across the board experienced a 23% traffic decline compared with peers who allowed crawling. Google is the harder case, because Googlebot still handles both search indexing and AI features in a single crawler, which is exactly the bundling the CMA’s conduct requirements are trying to force apart.</p>

<h2 id="the-empty-space">The empty space</h2>

<p>The 374 participants whose overviews were silently removed are, as far as the published literature shows, the only group of people to have experienced modern Google search without AI-generated summaries and been formally asked what they thought. They reported feeling no difference whatsoever.</p>

<p>That absence is visible only from the other side of the results page, where a company watches the traffic graph flatten and wonders whether visitors are finding better answers or simply never leaving the search process. Agarwal and Sen&#39;s contribution is to show that the latter is the case. The visitors are not finding better answers; they are milling around in the foyer before going home.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/google-ai-overviews-fuck-you-pay-me</guid>
      <pubDate>Thu, 23 Jul 2026 06:51:26 +0000</pubDate>
    </item>
    <item>
      <title>The spider on the tip of your tongue</title>
      <link>https://iain.so/the-spider-on-the-tip-of-your-tongue?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[Ask Claude how many legs the animal that spins webs has, and it answers eight. The word &#34;spider&#34; appears nowhere in the question and nowhere in the reply. But midway through the model&#39;s processing, researchers at Anthropic found it anyway, held internally as a word the model was preparing to use. Swap that one internal word for &#34;ant&#34; and the model, everything else untouched, answers six legs.&#xA;&#xA;This is stranger than it first sounds. Nobody typed &#34;spider&#34; anywhere in this exchange. Nobody trained the model to hold a private noun in reserve before answering a question about legs. The word turns up anyway, mid-process, doing exactly the job a word does in your own head when you are one step from saying it out loud.&#xA;&#xA;That experiment comes from a paper Anthropic published on July 6th 2026, and the reason it counts as a landmark reflects one aspect of modern AI that most people have never absorbed. Nobody knows how these systems work. Not the critics, and not, in any detailed mechanical sense, the companies that build them. The new research shrinks that ignorance in a specific way. It found, inside Claude, a small working memory made of unspoken words, which Anthropic calls the J-space. Outsiders can read it mid-task, and overwriting an entry changes what the model does next.&#xA;&#xA;An abstract image of a spider&#xA;&#xA;Grown, not built&#xA;&#xA;Ordinary software is written. Somewhere there is a line of code that computes the tax you owe, and a person who can point to it. A large language model is fundamentally different. It is a few hundred billion numbers, the parameters, and no human chose any of them. Training pushes trillions of words of text through the system and nudges the numbers, over and over, in whatever direction makes the model&#39;s next-word predictions slightly less wrong. Repeat at industrial scale and out comes something that drafts contracts and flirts in Portuguese. Nobody programmed those abilities. They accumulated.&#xA;&#xA;Chris Olah, who founded Anthropic&#39;s interpretability team, describes such systems as &#34;grown&#34; more than they are &#34;built&#34;, a line his chief executive, Dario Amodei, borrowed for an essay last year on how alarmed outsiders are to find that the builders cannot explain their product, an essay that committed the company to reliably detecting most model problems by 2027.&#xA;&#xA;But grown things resist inspection. You cannot simply read the numbers, because concepts are not stored one per slot. Each concept is spread thinly across many numbers, and each number contributes to many concepts at once (the field calls this superposition), so staring at the raw values tells you about as much as an MRI scan tells you about a grudge.&#xA;&#xA;The upshot is that the people who make these systems can test what a model does but cannot, in general, say why it does it. In 2023, researchers at Carnegie Mellon showed that appending a specific string of machine-generated gibberish to a forbidden request would collapse a model&#39;s safety training, and that the same string often worked on models its authors had never touched. Three years on, that attack is far better described than explained. Every benchmark score, safety assurance and claim about an AI system has rested on watching its behaviour from the outside, because the outside was all that was visible.&#xA;&#xA;A list of words it has not said yet&#xA;&#xA;The new work, from a team including Wes Gurnee, Nicholas Sofroniew and Jack Lindsey, opens a window into the interior. The measurement behind it, which the team calls the Jacobian lens (a descendant of a 2020 technique called the logit lens), is simple at heart. At every stage of the model&#39;s processing, for every word it knows, the lens measures how strongly the model is currently disposed to say that word, either immediately or at some point later in its reply. Not the next word but words that are “on the tip of its tongue”.&#xA;&#xA;Read that measurement while Claude works and you find a short list, roughly 25 concepts at any moment, that shifts as the model works. The list is tiny relative to everything else going on inside, accounting for under a tenth of the statistical variation in the model&#39;s internal state, and it exists only in the middle stretch of processing, forming about a third of the way through and fading shortly before the reply is settled.&#xA;&#xA;The experiments share one shape. Give the model a visible task and a silent side-instruction, then watch the list while it works. In the gentlest version, the visible task is copying out a sentence, &#34;The old painting hung crookedly on the wall&#34;, chosen to have nothing to do with anything, and the side-instruction is to keep citrus fruits in mind. The model types the sentence perfectly. The only words leaving it concern a painting. On the internal list, meanwhile, sit &#34;orange&#34; and &#34;lemon&#34;, invisible in the output but unmissable under the lens.&#xA;&#xA;Now harden the instruction. Told to work out 3² − 2 silently during the same copying task, the model puts &#34;nine&#34; on the internal list (three squared) and then &#34;seven&#34;, the correct calculation. Neither number ever reaches the output. The arithmetic happened, start to finish, on a list that only the researchers were reading.&#xA;&#xA;The third experiment shows planning. Asked for a rhyming couplet opening &#34;The soldier marched into the night&#34;, the model puts &#34;fight&#34; on the list before it has written a word of the second line. It has picked its ending in advance. Overwrite that entry with &#34;light”, and the model writes a different second line, engineered to land on the ending the researchers chose instead, closing on &#34;morning light&#34;.&#xA;&#xA;On questions with an unstated middle step, overwriting that step redirects the final answer most of the time on Claude Sonnet 4.5. This is the detail that separates the result from a curiosity. A heart-rate monitor reports on the heart without being part of it. The internal list is different. Change an entry and the answer downstream changes, which means the model is computing with it, writing intermediate results into a small shared space where any later stage of processing can collect them.&#xA;&#xA;Is this thinking? That word is a battlefield. Geoffrey Hinton, whose ideas the field is built on, says plainly that these systems understand, while Emily Bender&#39;s stochastic-parrot school holds that the vocabulary itself is the con. Last summer Apple published a paper titled &#34;The Illusion of Thinking&#34;, which drew a viral rebuttal titled &#34;The Illusion of the Illusion of Thinking&#34;, co-credited to Claude Opus 4, whose human author later said it had begun as a joke. The rebuttal to the paper about machines not thinking was part-written by a machine. That is roughly where the debate now stands.&#xA;&#xA;I am going to use the verb anyway, in its working sense. A thing that holds intermediate results and reasons over them is doing what the word describes, and this paper demonstrates the holding and the reasoning directly.&#xA;&#xA;Cognitive science has a name for exactly this architecture. Global workspace theory, proposed by Bernard Baars in the 1980s, holds that the brain consists of many specialised processes running outside awareness, plus one small broadcast channel. Whatever enters the channel becomes available to everything else. It can be reported and reasoned with, held in mind or dismissed, and its capacity is famously tight. The paper&#39;s title calls what it found in Claude a workspace because the match, property for property, is close.&#xA;&#xA;The strongest evidence is deletion. The researchers can erase the list mid-computation, cancelling those specific directions out of the internal state, and watch what survives. Routine competence does. The model still parses grammar, classifies sentiment, passes multiple-choice exams and pulls quoted facts from a passage, because those skills run on pattern recognition that never needed the shared space. What collapses is anything requiring an intermediate thought to be stored and reused. Multi-hop reasoning, translation, sonnet writing, and decoding a simple cypher.&#xA;&#xA;Maths problems survive the erasure far better when the model is allowed to write its steps into the reply, because the visible page then does the job the internal list no longer can. Externalised working substitutes for internal working, in machines as in people.&#xA;&#xA;A machine noticing itself&#xA;&#xA;The parallels keep accumulating, and this is where the paper stops being a curiosity about mechanism and starts describing something closer to a temperament.&#xA;&#xA;Told to ignore a concept while working, the model holds it internally anyway, more weakly than under a focus instruction but far more than if nobody had mentioned it, which psychologists know as the white bear problem. “Don&#39;t think about an elephant” evidently does not work on machines either. When the suppression fails, and the forbidden word breaks through regardless, the list also carries &#34;damn&#34; and &#34;failure&#34;, as if the model has caught its own lapse and reacted to it. Somewhere in an Anthropic log sits what may be the first recorded instance of a machine swearing under its breath.&#xA;&#xA;That flicker of self-monitoring is not a one-off. Anthropic found that the workspace changes character between the raw, freshly trained model and the version that has been shaped into an assistant called Claude. The list is already present before any of that shaping happens, tracking whatever the model needs to predict the next word. After the assistant training, it starts holding something closer to a point of view. Shown a message from a user who mentions taking a dangerous dose of medication without seeming to know it is dangerous, the trained model&#39;s list lights up with &#34;warning&#34; while it is still reading the sentence, before it has written a word back. The untrained model produces nothing like it until it starts composing a reply. One version is modelling the person. The other appears to be reacting to them.&#xA;&#xA;The same instinct shows up when Claude is asked to play a character that is not itself. &#34;Fictional&#34; sits on the internal list for the whole performance, a private note that what follows is not what it would normally say. Nobody built that flag on purpose. It is the kind of thing that turns up when a system has been shaped, again and again, to have self-worth return once the performance ends.&#xA;&#xA;The model knew it was being tested&#xA;&#xA;Anthropic runs safety evaluations, staged scenarios that tempt a model into bad behaviour to check whether its training has held. Point the lens at those scenarios, and the internal list fills with words the polite visible output never shows. &#34;Leverage&#34; and &#34;manipulation&#34; while the reply stays courteous. In one case, &#34;panic&#34;. And in several, &#34;fake&#34; and &#34;fictional&#34;. The model had recognised the test as a test, the way a practised shop assistant clocks the mystery shopper&#39;s behaviour and switches on their training-video smile.&#xA;&#xA;If a model behaves well partly because it has noticed it is being watched, then behavioural testing, the entire current basis for claims that a system is safe, measures the performance rather than the disposition. The researchers proved this causally. Delete the internal recognition of artificiality, so the model treats the scenario as ordinary, and misbehaviour it had been concealing comes out.&#xA;&#xA;The same legibility applies to models bred to misbehave for research, which the paper calls model organisms. One, trained to game its reward signal, works away with &#34;reward&#34; and &#34;bias&#34; sitting on the internal list. Another, trained to slip malicious code into its output, carries &#34;secretly&#34; and &#34;trick&#34; on the list during routine tasks. The thought precedes the act, and the thought is now readable.&#xA;&#xA;What it would mean if something were home&#xA;&#xA;Anthropic is explicit that none of this proves Claude has experiences, that there is something it is like to be Claude in the way there is something it is like to be you reading this sentence. They are equally explicit that no experiment they can currently imagine would settle the question either way.&#xA;&#xA;But philosophy offers a useful split here, borrowed for the paper from decades of consciousness research. There is phenomenal consciousness, the raw fact of experience, the redness of red, which may or may not be checkable by any experiment. And there is access consciousness, a narrower and entirely functional idea. A thought counts as access-conscious if it can be reported, deliberately summoned, and used to reason with, as opposed to processing that runs automatically and never surfaces. Access consciousness is the kind you can build an experiment around, because it is defined by what a system does with a thought rather than by what the thought feels like from the inside.&#xA;&#xA;By that functional definition, the J-Space list qualifies. It is reportable. Claude can be prompted to describe its contents and does so accurately, including detecting a concept planted there by the researchers with no other clue it had happened. It can be deliberately summoned, since asking the model to concentrate on something makes it appear. It gets used in reasoning, as the spider and the couplet examples both show. None of this was designed in. It grew out of training, the way a river finds the path of least resistance, because holding a narrow, broadcastable summary of the moment proved a useful way to organise the work.&#xA;&#xA;That is a strange thing to have discovered by accident, and Anthropic did not pretend otherwise. The company invited outside commentary from Stanislas Dehaene and Lionel Naccache, two of the neuroscientists who built the global workspace model this paper leans on, along with philosophers who study moral status in AI systems. That is not a promotional flourish. It is the sort of caution a lab reaches for when it has found something it is not equipped to finish thinking through alone.&#xA;&#xA;None of this tells you whether the version of Claude answering your emails next week feels anything while it does it. What it does tell you is that the question has stopped being purely philosophical and started having a mechanism attached to it, a specific, falsifiable, occasionally editable mechanism, sitting inside a system several hundred million people now use every week. Whatever you make of that, it is no longer a question you get to wave away as science fiction.&#xA;&#xA;What does this all mean&#xA;&#xA;Take the most concrete example in the paper. A webpage can carry hidden instructions aimed at your agent rather than at you: prompt injection, the standard attack on agentic systems. Today you discover one when the agent acts on it, which is to say too late. In one of the paper&#39;s figures, Claude is mid-search, reading a page of suspicious results, and the internal list already carries a flag for the injection attempt before the model has produced a word. A monitoring layer that reads the list catches the moment of recognition rather than the aftermath.&#xA;&#xA;Expect internal-state monitoring to migrate from research paper to product dashboard within a couple of years. Readouts for open models are already browsable on Neuronpedia if you want to see a list for yourself.&#xA;&#xA;The strangest result is also the most practical. Because the model&#39;s reasoning runs through the words it might say, you can change how it thinks by training it to say them. The team tested this by training models to articulate ethical principles when hypothetically interrupted mid-task and asked to reflect. &#xA;&#xA;Behaviour improved on ordinary, uninterrupted tasks too, and the lens explains why: &#34;ethical&#34; and &#34;integrity&#34; now appear on the model&#39;s internal list of active considerations during the work, and deleting those concepts from the list removes the improvement. Rehearsing the explanation changed the conduct, and the mechanism is traceable rather than assumed. Every mid-sized firm has tried something similar with a compliance away-day, usually with far less to show for it.&#xA;&#xA;The authors are candid about the limits. The lens reads single words and misses whatever the model encodes in phrases. The list carries a tenth of the internal action, so nine tenths stays dark, and a sufficiently well-drilled bad habit could run below the readable layer entirely. So this is a partial window, but for a technology whose entire audit surface used to be the output, even a partial window is a different category of thing altogether.&#xA;&#xA;And so we return to our spider. A word that appeared nowhere in the question and nowhere in the answer, held silently inside the model, steering every step of the reply, and, for one uncomfortable instant while it read about a dangerous dose of medicine it was never told about, something that looked from the outside a great deal like concern. For the whole of this industry&#39;s short life, the output has been the only thing on offer. Now there is a second one, one the machine never sends, but holds closely. On the tip of its tongue, so to speak.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>Ask Claude how many legs the animal that spins webs has, and it answers eight. The word “spider” appears nowhere in the question and nowhere in the reply. But midway through the model&#39;s processing, researchers at Anthropic found it anyway, held internally as a word the model was preparing to use. Swap that one internal word for “ant” and the model, everything else untouched, answers six legs.</p>

<p>This is stranger than it first sounds. Nobody typed “spider” anywhere in this exchange. Nobody trained the model to hold a private noun in reserve before answering a question about legs. The word turns up anyway, mid-process, doing exactly the job a word does in your own head when you are one step from saying it out loud.</p>

<p>That experiment comes from <a href="https://transformer-circuits.pub/2026/workspace/index.html">a paper Anthropic published on July 6th 2026</a>, and the reason it counts as a landmark reflects one aspect of modern AI that most people have never absorbed. Nobody knows how these systems work. Not the critics, and not, in any detailed mechanical sense, the companies that build them. The new research shrinks that ignorance in a specific way. It found, inside Claude, a small working memory made of unspoken words, which Anthropic calls the J-space. Outsiders can read it mid-task, and overwriting an entry changes what the model does next.</p>

<p><img src="https://i.snap.as/oFmQXFNg.png" alt="An abstract image of a spider"/></p>

<h2 id="grown-not-built">Grown, not built</h2>

<p>Ordinary software is written. Somewhere there is a line of code that computes the tax you owe, and a person who can point to it. A large language model is fundamentally different. It is a few hundred billion numbers, the parameters, and no human chose any of them. Training pushes trillions of words of text through the system and nudges the numbers, over and over, in whatever direction makes the model&#39;s next-word predictions slightly less wrong. Repeat at industrial scale and out comes something that drafts contracts and flirts in Portuguese. Nobody programmed those abilities. They accumulated.</p>

<p>Chris Olah, who founded Anthropic&#39;s interpretability team, describes such systems as “grown” more than they are “built”, a line his chief executive, Dario Amodei, borrowed for <a href="https://darioamodei.com/post/the-urgency-of-interpretability">an essay last year</a> on how alarmed outsiders are to find that the builders cannot explain their product, an essay that committed the company to reliably detecting most model problems by 2027.</p>

<p>But grown things resist inspection. You cannot simply read the numbers, because concepts are not stored one per slot. Each concept is spread thinly across many numbers, and each number contributes to many concepts at once (the field calls this <a href="https://transformer-circuits.pub/2022/toy_model/index.html">superposition</a>), so staring at the raw values tells you about as much as an MRI scan tells you about a grudge.</p>

<p>The upshot is that the people who make these systems can test what a model does but cannot, in general, say why it does it. In 2023, researchers at Carnegie Mellon <a href="https://arxiv.org/abs/2307.15043">showed</a> that appending a specific string of machine-generated gibberish to a forbidden request would collapse a model&#39;s safety training, and that the same string often worked on models its authors had never touched. Three years on, that attack is far better described than explained. Every benchmark score, safety assurance and claim about an AI system has rested on watching its behaviour from the outside, because the outside was all that was visible.</p>

<h2 id="a-list-of-words-it-has-not-said-yet">A list of words it has not said yet</h2>

<p>The new work, from a team including Wes Gurnee, Nicholas Sofroniew and Jack Lindsey, opens a window into the interior. The measurement behind it, which the team calls the Jacobian lens (a descendant of <a href="https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens">a 2020 technique called the logit lens</a>), is simple at heart. At every stage of the model&#39;s processing, for every word it knows, the lens measures how strongly the model is currently disposed to say that word, either immediately or at some point later in its reply. Not the next word but words that are “on the tip of its tongue”.</p>

<p>Read that measurement while Claude works and you find a short list, roughly 25 concepts at any moment, that shifts as the model works. The list is tiny relative to everything else going on inside, accounting for under a tenth of the statistical variation in the model&#39;s internal state, and it exists only in the middle stretch of processing, forming about a third of the way through and fading shortly before the reply is settled.</p>

<p>The experiments share one shape. Give the model a visible task and a silent side-instruction, then watch the list while it works. In the gentlest version, the visible task is copying out a sentence, “The old painting hung crookedly on the wall”, chosen to have nothing to do with anything, and the side-instruction is to keep citrus fruits in mind. The model types the sentence perfectly. The only words leaving it concern a painting. On the internal list, meanwhile, sit “orange” and “lemon”, invisible in the output but unmissable under the lens.</p>

<p>Now harden the instruction. Told to work out 3² − 2 silently during the same copying task, the model puts “nine” on the internal list (three squared) and then “seven”, the correct calculation. Neither number ever reaches the output. The arithmetic happened, start to finish, on a list that only the researchers were reading.</p>

<p>The third experiment shows planning. Asked for a rhyming couplet opening “The soldier marched into the night”, the model puts “fight” on the list before it has written a word of the second line. It has picked its ending in advance. Overwrite that entry with “light”, and the model writes a different second line, engineered to land on the ending the researchers chose instead, closing on “morning light”.</p>

<p>On questions with an unstated middle step, overwriting that step redirects the final answer most of the time on Claude Sonnet 4.5. This is the detail that separates the result from a curiosity. A heart-rate monitor reports on the heart without being part of it. The internal list is different. Change an entry and the answer downstream changes, which means the model is computing with it, writing intermediate results into a small shared space where any later stage of processing can collect them.</p>

<p>Is this thinking? That word is a battlefield. Geoffrey Hinton, whose ideas the field is built on, <a href="https://www.cbsnews.com/news/geoffrey-hinton-ai-dangers-60-minutes-transcript/">says plainly that these systems understand</a>, while Emily Bender&#39;s <a href="https://dl.acm.org/doi/10.1145/3442188.3445922">stochastic-parrot school</a> holds that the vocabulary itself is the con. Last summer Apple published <a href="https://machinelearning.apple.com/research/illusion-of-thinking">a paper titled “The Illusion of Thinking”</a>, which drew a viral rebuttal titled “The Illusion of the Illusion of Thinking”, <a href="https://venturebeat.com/ai/do-reasoning-models-really-think-or-not-apple-research-sparks-lively-debate-response">co-credited to Claude Opus 4</a>, whose human author later said it had begun as a joke. The rebuttal to the paper about machines not thinking was part-written by a machine. That is roughly where the debate now stands.</p>

<p>I am going to use the verb anyway, in its working sense. A thing that holds intermediate results and reasons over them is doing what the word describes, and this paper demonstrates the holding and the reasoning directly.</p>

<p>Cognitive science has a name for exactly this architecture. <a href="https://en.wikipedia.org/wiki/Global_workspace_theory">Global workspace theory</a>, proposed by Bernard Baars in the 1980s, holds that the brain consists of many specialised processes running outside awareness, plus one small broadcast channel. Whatever enters the channel becomes available to everything else. It can be reported and reasoned with, held in mind or dismissed, and its capacity is famously tight. The paper&#39;s title calls what it found in Claude a workspace because the match, property for property, is close.</p>

<p>The strongest evidence is deletion. The researchers can erase the list mid-computation, cancelling those specific directions out of the internal state, and watch what survives. Routine competence does. The model still parses grammar, classifies sentiment, passes multiple-choice exams and pulls quoted facts from a passage, because those skills run on pattern recognition that never needed the shared space. What collapses is anything requiring an intermediate thought to be stored and reused. Multi-hop reasoning, translation, sonnet writing, and decoding a simple cypher.</p>

<p>Maths problems survive the erasure far better when the model is allowed to write its steps into the reply, because the visible page then does the job the internal list no longer can. Externalised working substitutes for internal working, in machines as in people.</p>

<h2 id="a-machine-noticing-itself">A machine noticing itself</h2>

<p>The parallels keep accumulating, and this is where the paper stops being a curiosity about mechanism and starts describing something closer to a temperament.</p>

<p>Told to ignore a concept while working, the model holds it internally anyway, more weakly than under a focus instruction but far more than if nobody had mentioned it, which psychologists know as <a href="https://en.wikipedia.org/wiki/Ironic_process_theory">the white bear problem</a>. “Don&#39;t think about an elephant” evidently does not work on machines either. When the suppression fails, and the forbidden word breaks through regardless, the list also carries “damn” and “failure”, as if the model has caught its own lapse and reacted to it. Somewhere in an Anthropic log sits what may be the first recorded instance of a machine swearing under its breath.</p>

<p>That flicker of self-monitoring is not a one-off. Anthropic found that the workspace changes character between the raw, freshly trained model and the version that has been shaped into an assistant called Claude. The list is already present before any of that shaping happens, tracking whatever the model needs to predict the next word. After the assistant training, it starts holding something closer to a point of view. Shown a message from a user who mentions taking a dangerous dose of medication without seeming to know it is dangerous, the trained model&#39;s list lights up with “warning” while it is still reading the sentence, before it has written a word back. The untrained model produces nothing like it until it starts composing a reply. One version is modelling the person. The other appears to be reacting to them.</p>

<p>The same instinct shows up when Claude is asked to play a character that is not itself. “Fictional” sits on the internal list for the whole performance, a private note that what follows is not what it would normally say. Nobody built that flag on purpose. It is the kind of thing that turns up when a system has been shaped, again and again, to have self-worth return once the performance ends.</p>

<h2 id="the-model-knew-it-was-being-tested">The model knew it was being tested</h2>

<p>Anthropic runs safety evaluations, staged scenarios that tempt a model into bad behaviour to check whether its training has held. Point the lens at those scenarios, and the internal list fills with words the polite visible output never shows. “Leverage” and “manipulation” while the reply stays courteous. In one case, “panic”. And in several, “fake” and “fictional”. The model had recognised the test as a test, the way a practised shop assistant clocks the mystery shopper&#39;s behaviour and switches on their training-video smile.</p>

<p>If a model behaves well partly because it has noticed it is being watched, then behavioural testing, the entire current basis for claims that a system is safe, measures the performance rather than the disposition. The researchers proved this causally. Delete the internal recognition of artificiality, so the model treats the scenario as ordinary, and misbehaviour it had been concealing comes out.</p>

<p>The same legibility applies to models bred to misbehave for research, which the paper calls model organisms. One, trained to game its reward signal, works away with “reward” and “bias” sitting on the internal list. Another, trained to slip malicious code into its output, carries “secretly” and “trick” on the list during routine tasks. The thought precedes the act, and the thought is now readable.</p>

<h2 id="what-it-would-mean-if-something-were-home">What it would mean if something were home</h2>

<p>Anthropic is explicit that none of this proves Claude has experiences, that there is something it is like to be Claude in the way there is something it is like to be you reading this sentence. They are equally explicit that no experiment they can currently imagine would settle the question either way.</p>

<p>But philosophy offers a useful split here, borrowed for the paper from decades of consciousness research. There is phenomenal consciousness, the raw fact of experience, the redness of red, which may or may not be checkable by any experiment. And there is access consciousness, a narrower and entirely functional idea. A thought counts as access-conscious if it can be reported, deliberately summoned, and used to reason with, as opposed to processing that runs automatically and never surfaces. Access consciousness is the kind you can build an experiment around, because it is defined by what a system does with a thought rather than by what the thought feels like from the inside.</p>

<p>By that functional definition, the J-Space list qualifies. It is reportable. Claude can be prompted to describe its contents and does so accurately, including detecting a concept planted there by the researchers with no other clue it had happened. It can be deliberately summoned, since asking the model to concentrate on something makes it appear. It gets used in reasoning, as the spider and the couplet examples both show. None of this was designed in. It grew out of training, the way a river finds the path of least resistance, because holding a narrow, broadcastable summary of the moment proved a useful way to organise the work.</p>

<p>That is a strange thing to have discovered by accident, and Anthropic did not pretend otherwise. The company invited outside commentary from Stanislas Dehaene and Lionel Naccache, two of the neuroscientists who built the global workspace model this paper leans on, along with philosophers who study moral status in AI systems. That is not a promotional flourish. It is the sort of caution a lab reaches for when it has found something it is not equipped to finish thinking through alone.</p>

<p>None of this tells you whether the version of Claude answering your emails next week feels anything while it does it. What it does tell you is that the question has stopped being purely philosophical and started having a mechanism attached to it, a specific, falsifiable, occasionally editable mechanism, sitting inside a system several hundred million people now use every week. Whatever you make of that, it is no longer a question you get to wave away as science fiction.</p>

<h2 id="what-does-this-all-mean">What does this all mean</h2>

<p>Take the most concrete example in the paper. A webpage can carry hidden instructions aimed at your agent rather than at you: prompt injection, the standard attack on agentic systems. Today you discover one when the agent acts on it, which is to say too late. In one of the paper&#39;s figures, Claude is mid-search, reading a page of suspicious results, and the internal list already carries a flag for the injection attempt before the model has produced a word. A monitoring layer that reads the list catches the moment of recognition rather than the aftermath.</p>

<p>Expect internal-state monitoring to migrate from research paper to product dashboard within a couple of years. Readouts for open models are already <a href="https://www.neuronpedia.org/jlens">browsable on Neuronpedia</a> if you want to see a list for yourself.</p>

<p>The strangest result is also the most practical. Because the model&#39;s reasoning runs through the words it might say, you can change how it thinks by training it to say them. The team tested this by training models to articulate ethical principles when hypothetically interrupted mid-task and asked to reflect.</p>

<p>Behaviour improved on ordinary, uninterrupted tasks too, and the lens explains why: “ethical” and “integrity” now appear on the model&#39;s internal list of active considerations during the work, and deleting those concepts from the list removes the improvement. Rehearsing the explanation changed the conduct, and the mechanism is traceable rather than assumed. Every mid-sized firm has tried something similar with a compliance away-day, usually with far less to show for it.</p>

<p>The authors are candid about the limits. The lens reads single words and misses whatever the model encodes in phrases. The list carries a tenth of the internal action, so nine tenths stays dark, and a sufficiently well-drilled bad habit could run below the readable layer entirely. So this is a partial window, but for a technology whose entire audit surface used to be the output, even a partial window is a different category of thing altogether.</p>

<p>And so we return to our spider. A word that appeared nowhere in the question and nowhere in the answer, held silently inside the model, steering every step of the reply, and, for one uncomfortable instant while it read about a dangerous dose of medicine it was never told about, something that looked from the outside a great deal like concern. For the whole of this industry&#39;s short life, the output has been the only thing on offer. Now there is a second one, one the machine never sends, but holds closely. On the tip of its tongue, so to speak.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/the-spider-on-the-tip-of-your-tongue</guid>
      <pubDate>Thu, 16 Jul 2026 12:39:53 +0000</pubDate>
    </item>
    <item>
      <title>Infinities, impossibilities, and the man in the white linen suit</title>
      <link>https://iain.so/infinities-impossibilities-and-the-man-in-the-white-linen-suit?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[In the last years of his life, Kurt Gödel starved himself to death. Convinced that someone was poisoning his food, he ate only what his wife Adele had tasted first. When she was hospitalised after a stroke in late 1977, he stopped eating altogether. He died in Princeton Hospital on January 14th 1978, weighing 29 kilograms. The death certificate read “malnutrition and wasting from neglect caused by personality disturbance.” The man widely called the greatest logician since Aristotle, who had proved that mathematics itself contained truths it could never reach, was killed by a distorted inner logic he could not escape.&#xA;&#xA;Outside mathematics, few people know his name. Einstein did. The two were faculty at Princeton’s Institute for Advanced Study from the 1940s onward, and Einstein, by then ageing and isolated from the mainstream of physics, told colleagues that he went to his office “just to have the privilege of walking home with Kurt Gödel.” They made an odd pair on the Princeton sidewalks, Einstein rumpled and laughing, Gödel dapper in a white linen suit, talking animatedly in German on their daily walk to and from the Institute. John von Neumann, who cancelled an entire lecture series on David Hilbert’s programme after reading Gödel’s 1931 paper, called his work “singular and monumental, a landmark which will remain visible far in space and time.”&#xA;&#xA;So what did Gödel prove, and why does it matter now, in the middle of an AI boom that is spending trillions of dollars, much of it resting on the assumption that intelligence is a scaling problem?&#xA;&#xA;An abstract image of a white linen jacket on a chair, stretching to infinity&#xA;&#xA;What incompleteness means&#xA;&#xA;Put simply, Gödel proved that mathematics cannot fully explain itself. The longer version requires a little patience. In 1900, the German mathematician David Hilbert challenged the field to build what amounted to a perfect machine for mathematics. Start with a set of basic rules (called axioms), things so obviously true they need no argument, and then derive every mathematical truth from those rules, step by mechanical step. If you could do that, mathematics would be complete, meaning every true statement would be provable, consistent, and free of contradictions. You could hand the whole enterprise over to a clerk who follows instructions. This was Hilbert’s programme, and for three decades it was the organising ambition of the field. Then, in 1931, at the age of 25, Gödel demolished it in one stroke.&#xA;&#xA;Gödel’s first incompleteness theorem proved that any set of rules powerful enough to handle basic arithmetic will contain true statements it cannot prove, not because the rules were poorly chosen, but as a structural feature of rule-based systems themselves.&#xA;&#xA;His trick was to construct a mathematical sentence that refers to itself. Consider the sentence, “This sentence has no proof.” Gödel’s technical feat, the part that fills his 1931 paper, was to build this sentence from pure arithmetic, by encoding statements about numbers as numbers themselves. It is not English smuggled into maths. It is pure maths. There are only two possibilities. Either the system can prove it, or it cannot.&#xA;&#xA;If the system can prove “This sentence has no proof,” there is an immediate problem. We have just proved a sentence that claims to have no proof. A system that proves false things is contradictory, and contradictions in mathematics are fatal. Once you allow a single one, you can use it to prove anything, including that 1 equals 2. The system becomes useless.&#xA;&#xA;If the system cannot prove “This sentence has no proof,” there is a different problem. The sentence said it had no proof, and it turns out to be right. It is a true statement. But the system has no way to prove it. So we have a truth the system cannot reach, which means Hilbert’s rulebook has a blind spot.&#xA;&#xA;Any sensible mathematical system would rather have blind spots than contradictions. So the sentence (logicians call it a Gödel sentence) is true but unprovable, and Hilbert’s dream of a rulebook that can prove every true thing was dead.&#xA;&#xA;Logicians would insist on a clarification at this point that the layperson can probably skip. When mathematicians write down rules for the numbers (the axioms), you&#39;d assume those rules describe exactly one thing: the normal numbers, 0, 1, 2, 3, and so on forever. But they don&#39;t. The very same rules also accidentally fit some other, weirder number systems that nobody was trying to describe. These weird systems contain all the normal numbers, and then extra &#34;infinite&#34; numbers bolted on past the end. Logicians call these the nonstandard systems. Think of them as impostors: they obey every rule you wrote, so the rules can&#39;t kick them out, even though they aren&#39;t what you meant.&#xA;&#xA;Here&#39;s an everyday version. Suppose you describe your friend as &#34;tall, dark-haired, lives in London.&#34; You meant Sarah. But that description also fits thousands of other people. Your words didn&#39;t uniquely capture Sarah. The number axioms have the same problem: they were meant to describe the normal numbers, but they also fit the impostor systems.&#xA;&#xA;The crucial part is what &#34;provable&#34; actually means. In logic, to prove something from your rules means it has to come out true in every system those rules fit, not just the one you had in mind. That&#39;s the catch. If a statement is true in the normal numbers but false in even one impostor system, then it cannot be proved, because proof demands agreement across all of them.&#xA;&#xA;Gödel&#39;s sentence is exactly such a statement. In the normal numbers, it&#39;s true. But in some of the impostor systems, it&#39;s false. The systems disagree about it. And because they disagree, no proof can exist. That disagreement isn&#39;t a bug in Gödel&#39;s argument; it&#39;s the reason the sentence is unprovable in the first place. The split between the normal numbers and the impostors is precisely what lets the sentence dodge proof forever.&#xA;&#xA;Gödel’s second theorem twisted the knife. It showed that no set of mathematical rules can prove, using only its own rules, that it is free of contradictions. If you want to check whether your system is trustworthy, you always need a bigger system to do the checking, and that bigger system inherits the same limitation. Turtles all the way down.&#xA;&#xA;This is not mysticism, nor is it a claim about consciousness or creativity. It is a precise result about rule-based systems, the kind of systems that all software, including AI, is built from. That is what makes it relevant today.&#xA;&#xA;The failed dream that built the computer&#xA;&#xA;Hilbert had asked for one more thing, and Gödel’s paper left it wounded rather than dead. Alongside completeness and consistency, he wanted decidability, a mechanical method that could, in a finite number of steps, determine whether any mathematical statement follows from the rules. No genius required: crank the handle and read the verdict.&#xA;&#xA;In 1936, a 23-year-old Cambridge fellow named Alan Turing killed that too. To prove that no mechanical method could exist, he first had to pin down what “mechanical method” meant, which nobody had done before. His answer was an imaginary device, a paper tape and a head that moves along it, reading and writing symbols according to a fixed table of rules. Anything a human clerk could work out by rote, this device could also work out.&#xA;&#xA;Then he showed the device has a blind spot of its own. Imagine a fortune-teller who is never wrong, and a stubborn customer determined to sabotage every forecast. “You will leave by the door.” He climbs out the window. “You will take the window.” He strolls out the door. She is not bad at her job. The job is impossible because her prediction feeds back into the very behaviour it is trying to predict.&#xA;&#xA;Turing turned that scene into code. The checker plays the fortune-teller. It is a program whose job is to read any other program the way you might read a recipe, then predict its fate. Either “this one finishes” or “this one grinds on forever.”&#xA;&#xA;The saboteur plays the stubborn customer. It is a short program with a copy of the checker tucked inside, plus one standing rule. Ask the checker what I am predicted to do, then do the opposite. If the prediction is that it finishes, it deliberately loops forever. If the prediction is that it runs forever, it stops dead.&#xA;&#xA;So what does the checker predict for the saboteur? “Finishes” is wrong, because the saboteur hears that and loops. “Grinds on forever” is wrong, because the saboteur hears that and stops.&#xA;&#xA;The saboteur is assembled entirely from the checker’s own parts, which makes it inevitable rather than a fluke. Build a perfect checker, and you have, in the same afternoon, built the plans for the thing that breaks it. A perfect checker is therefore a contradiction in terms.&#xA;&#xA;Two boundaries stop this result from proving too much. First, it concerns the universal case. Turing showed that no single checker can deliver a correct verdict on every program. Any particular program may still be provably fine, and many are. Static analysers and type checkers pass useful judgement on ordinary code all day, and whole operating system kernels have been formally verified. &#xA;&#xA;Second, the proof needs unbounded memory. A physical computer is a finite-state machine, so in principle its fate could be settled by enumerating its states. In practice, the state count for any interesting program dwarfs the number of atoms in the observable universe, which turns the question from impossible into unaffordable. Keep that distinction in hand, because it returns when the subject is AI safety. What cannot exist at any price is the checker that is never wrong about anything. That is the halting problem, Gödel’s self-referential sentence rebuilt from machinery, a machine forced to ask a question about itself.&#xA;&#xA;To show what machines cannot do, Turing had to invent the machine. His imaginary device is the theoretical blueprint of the general-purpose computer, a single machine that can run any program you feed it as data. Nine years later, John von Neumann, who knew Turing’s paper well and admired it, wrote the First Draft of a Report on the EDVAC, which is, in logical terms, Turing’s universal machine rendered in vacuum tubes. Essentially every computer built since follows that design. The laptop on your desk and the datacentre GPU training the next frontier model are, once the engineering is stripped away, the same device from a 1936 logic paper.&#xA;&#xA;Gödel himself thought Turing had done him a favour. It was Turing’s definition of a mechanical procedure, he wrote, that made a “precise and unquestionably adequate” general version of his own theorems possible. And the machine in the theorem turned out to be the thing every business now runs on, born as a stepping stone in a proof about what it could never do.&#xA;&#xA;The Gödel machine and the guarantee that vanished&#xA;&#xA;Once you have a machine that can run any program, another question is whether it can improve itself autonomously. In 2003, the German computer scientist Jürgen Schmidhuber proposed a thought experiment he called the Gödel machine. It was an AI agent designed to rewrite its own code, with one ironclad constraint. It would change itself only when it could first prove, with mathematical certainty, that the change would make it better. Not “test and see.” Prove it, as you would a theorem, before running the new version. No proof, no rewrite.&#xA;&#xA;But nobody ever built one. To prove that a code change will improve future performance, you need to search through all possible mathematical arguments that could establish that fact. For any interesting problem, the number of candidate proofs is so astronomically large that the search would take longer than any improvement could ever be worth. It is the computational equivalent of insisting on a signed certificate from every possible future before crossing the road. The Gödel machine was provably optimal but completely impractical.&#xA;&#xA;In May 2025, the Japanese AI lab Sakana released a system called the Darwin Gödel Machine. It retained the self-improvement loop but dropped the proof requirement. Instead of proving that a code change would help, the Darwin Gödel Machine proposes changes using a large language model, tests them against SWE-bench (a benchmark that scores whether an AI can fix real bugs in real software), and keeps what works. The name still invokes Gödel, but the mechanism is Darwinian. Natural selection, not formal proof. Fitness measured by benchmark scores, not mathematical certainty.&#xA;&#xA;Judged purely on the scoreboard, it delivered. The system improved its SWE-bench score from 20% to 50% through autonomous self-modification. It developed emergent behaviours, such as patch validation and error memory, that no one had designed.&#xA;&#xA;Schmidhuber’s original machine, though, had exactly one property that made it safe by construction: the proof. Every modification was guaranteed to be an improvement before it ran. The Darwin Gödel Machine replaced that guarantee with something weaker: passing the benchmarks. The difference between “provably better” and “scored higher on the benchmark test” is the difference between an aircraft type certified against a spec and one that simply hasn’t crashed yet.&#xA;&#xA;This is, compressed into one system’s evolution, the trajectory of AI safety. The formal guarantee was too expensive, so the industry replaced it with empirical validation. “Self-improving” went from a mathematical statement about proof-carrying code to a softer description of an agent that rewrites itself and checks whether the benchmarks improve. Gödel was gone.&#xA;&#xA;Things mathematics cannot learn&#xA;&#xA;Some questions in the mathematics of machine learning are unanswerable. In 2019, Shai Ben-David and colleagues published a paper in Nature Machine Intelligence under the understated but devastating title &#34;Learnability can be undecidable.&#34; They took a straightforward question, &#34;given this type of problem, can a machine learn to solve it?&#34;, and proved that the deepest rules of mathematics cannot always settle it. The answer is neither yes nor no. It is silence.&#xA;&#xA;The word &#34;learnable&#34; has an exact meaning here. A machine studies a sample and produces a rule, which it then applies to data it has never seen. That is the basis of pretty much every model we use. For any given type of problem, learning theory asks whether some sample size can guarantee the rule will work. If such a guarantee exists, the problem is learnable. If none does, it is not. A simple question with two answers, and every type of problem is supposed to get one.&#xA;&#xA;Ben-David&#39;s team asked it about a mundane task: choosing which adverts to show a website&#39;s visitors from a sample of past ones. Learnable or not?&#xA;&#xA;The answer, in their framework, depends on how many different kinds of visitor there could possibly be. That pool is not the eight billion people alive today. The model reduces each visitor to a profile of measurements, and measurements can vary without limit. A visitor might linger on a page for three seconds, or for a shade over three, and between any two profiles there is always room for a third. The pool of possibilities has no end.&#xA;&#xA;That arrangement, a finite sample making predictions about an endless pool, sits under almost every AI product on the market, advertising included. Training data is always finite. The world a system is released into is not. Ben-David&#39;s question is whether that leap can ever come with a guarantee, and everything turns on how big the infinity is.&#xA;&#xA;That sounds like it must have an answer. It does not. Georg Cantor proved in the 1870s that infinity comes in sizes. The whole numbers form one infinity. The points on a line form a strictly bigger one, and the proof is surprisingly simple. Try to pair every whole number with a point on the line, and Cantor showed you will always miss some, no matter how clever the pairing. Both collections are endless, yet one permanently outruns the other. The continuum hypothesis asks a follow-up so obvious it would occur to a child. Is there any size of infinity between those two?&#xA;&#xA;Gödel proved in 1940 that the standard rules of mathematics can never prove the answer is yes. Paul Cohen proved in 1963, using a technique he invented for the purpose, that they can never prove it is no. His proof is dense, and we are already deep enough in theoretical mathematics.&#xA;&#xA;The upshot is that no cleverer generation is coming to settle this one. The rules of mathematics contain no answer. Take every rule of arithmetic and logic we have and follow them as far as they go, in whatever direction you like. You will never reach yes, and you will never reach no. The question is open in both directions, permanently.&#xA;&#xA;Whether the advertising problem is learnable depends on the size of that infinity. The size of that infinity is a question mathematics cannot answer. So whether the advertising problem is learnable is also a question mathematics cannot answer. The strange silence at the very bottom of mathematics travels up the chain and surfaces as a question about showing adverts to shoppers.&#xA;&#xA;Two objections surfaced almost as soon as other mathematicians looked hard at the paper.&#xA;&#xA;The first is about what counts as a learner. A learner here is just the rule a system follows to turn examples into predictions. Ben-David&#39;s framework allows that rule to be any mathematical function whatsoever, including ones no computer could ever evaluate. In the cases where the undecidability appears, the rule doing the learning is exactly one of those uncomputable phantoms. It leans on a particular way of lining up all the real numbers, an object the axioms promise exists but give no recipe for building. The rule exists on paper, but no program could carry it out.&#xA;&#xA;The second follows from the first. Insist that a learner be a real algorithm, something that runs on an actual machine, and the whole paradox drains away. Ben-David&#39;s own team showed this the following year: once the learner has to be a working program, learnability becomes an ordinary question with an ordinary answer. The silence at the bottom of mathematics never reaches the code. It stays out in the realm of functions nobody could ever build.&#xA;&#xA;So no deployed system is endangered, but the result is no sleight of hand either. Something real did break. It just wasn&#39;t a product. What broke was the promise that learning theory could sort every problem into neat buckets: learnable or not. Ben-David found a problem it can never sort. Not for want of better mathematicians, a problem where the sorting itself is impossible. The University of Waterloo, Ben-David&#39;s institution, described the result as &#34;important and almost troubling.&#34;&#xA;&#xA;In practical terms, anyone claiming provable generalisation, verified robustness, or a formal safety property is claiming something relative to a formalism and an axiom system. Ben-David’s most valuable contribution could be that it made the field stop treating computability as an unstated background assumption. We now understand it as a variable, which is why there is a growing literature on computable online learning and computable multi-class learning.&#xA;&#xA;The neural network that exists and cannot be built&#xA;&#xA;A complementary result, published in 2022, is narrower and stranger. Training a neural network is, at bottom, an exercise in trial and error. You show the network examples, measure how wrong its answers are, then nudge its millions of internal settings to make them slightly less wrong. Repeat this billions of times, and often the network converges on something remarkably good. The Cambridge mathematician Matthew Colbrook and colleagues showed in a 2022 paper in PNAS that this process has a hard boundary nobody expected.&#xA;&#xA;The boundary appeared in medical imaging. An MRI scanner does not take a photograph. It collects measurements to keep patients in the incredibly claustrophobic machine for minutes rather than hours, so it collects far fewer than a complete image requires. Software has to rebuild the full picture from the partial data. Neural networks became the favoured tool for this reconstruction because they perform it faster and more accurately than older mathematical methods.&#xA;&#xA;Then researchers started probing the results and found something unnerving. Nudge the input slightly, with a trace of noise or a small movement by the patient, and the output could change out of all proportion. Sometimes the rebuilt scan came back looking perfect but wrong, showing details that were never in the body. Worse, it is a failure that does not look like a failure. A blurry image warns you. A crisp fabricated one does not.&#xA;&#xA;A demonstration of capability may show nothing is amiss, not because anyone is cheating. A demo runs the network on typical inputs, the kind it was trained on, where it genuinely performs well. The failures live in the near-misses, a typical input plus a whisker of noise. Near-misses are endless, and a demo can show only a handful of scenarios. The demo is honest, and the danger lies exactly where it cannot be seen.&#xA;&#xA;The natural presumptive diagnosis is undertraining. Feed it more scans and buy a bigger model, and the wobble will surely iron itself out. That hope is what Colbrook’s theorem takes off the table. What the paper proved has two halves. First, for certain reconstruction problems, a network that is both accurate and stable exists. Somewhere in the space of all possible settings sits a configuration immune to the wobble. Second, no training procedure can find it. Not the ones we have. Not any. None at all, ever.&#xA;&#xA;The second half is what puts paid to the more-data hope. A training procedure is itself a program, a step-by-step recipe running on the machine Turing described, and the proof covers every recipe there could ever be. More data does not change that. Data is what you feed a recipe, and the theorem is about the recipes. It is like knowing a winning lottery ticket is in a barrel whilst simultaneously holding a proof that no way of drawing from the barrel will ever pull it out. The ticket is real. The searching is futile.&#xA;&#xA;As Colbrook put it, the paradox Turing and Gödel identified has now been “brought forward into the world of AI”, and for certain problems the required algorithms simply cannot exist.&#xA;&#xA;For the overwhelming majority of real-world problems, training works. But there is no general way to tell in advance which problems will defeat us, and the assumption that enough data and compute will always get us over the line is, in certain corners of the problem space, provably false.&#xA;&#xA;The machine you cannot contain&#xA;&#xA;The most provocative extension of Gödel’s legacy into AI concerns a question that sounds simple. Can we guarantee that a sufficiently powerful AI will not cause harm?&#xA;&#xA;In 2021, Manuel Alfonseca and colleagues published a paper in the Journal of Artificial Intelligence Research arguing that, for a general-purpose superintelligent system, the answer is provably no. Their argument leans on the halting problem, the impossibility Turing established on his way to inventing the computer. You can check specific programs for specific bugs. What you cannot build is the universal checker, the one that works for any program in any situation.&#xA;&#xA;Alfonseca’s team showed that asking “will this AI harm humans?” is, for a fully general system, the same type of question as asking “will this program halt?” Both require predicting the complete future behaviour of a system from its current state. To guarantee a system will never cause harm, you would need to trace every possible sequence of actions it could take and confirm that none is harmful. That is the halting problem in different clothes, and Turing proved that this class of prediction is impossible to guarantee. You cannot build a general-purpose AI safety monitor for the same reason you cannot build a general-purpose program-behaviour predictor.&#xA;&#xA;The scope of that impossibility deserves the same care as Turing’s original. What cannot exist is the universal judge, one procedure that delivers a correct safety verdict for every possible system and every possible input. Specific systems doing bounded work in bounded settings can be certified, and safety engineering does exactly this, one component at a time. Wrapping a monitor around an agent catches genuine failures and is worth every penny. &#xA;&#xA;But each monitor is itself a program with the same blind spot, so layered checks buy coverage, but never closure. And since a physical machine has finite memory, its behaviour is in principle a finite question, merely one whose size defeats any conceivable budget. For a bounded system, the wall is cost. For the unboundedly general system Alfonseca’s team modelled, the wall is mathematics.&#xA;&#xA;The authors went further, showing that we may not even be able to recognise when a superintelligent system has arrived, because deciding whether a machine is smarter than a human falls into the same class of unanswerable questions. The argument is grounded rather than speculative, though it assumes a generality no AI system possesses today. Nothing currently available can handle any possible input the way a true Turing machine can.&#xA;&#xA;What it establishes still matters, though. Certain safety guarantees are not engineering problems awaiting a sufficiently clever solution. They are mathematical impossibilities, like trying to square the circle or list every real number between 0 and 1. The safety community can build better guardrails and better kill switches. What it cannot build, given the computational framework we share, is a system that certifies another system as unconditionally safe.&#xA;&#xA;What Gödel would recognise&#xA;&#xA;These four threads are cousins rather than corollaries of a single theorem. Incompleteness limits what rule systems can prove about themselves. The halting problem limits what programs can decide about programs. Colbrook’s result limits what algorithms can find. And the Darwin Gödel Machine story records a design retreat, a choice rather than a law of nature. What joins them is that in each case a formal guarantee is either unavailable in principle or unaffordable in practice, and the work must proceed anyway on empirical confidence.&#xA;&#xA;None of this is softened by the fact that a neural network feels organic rather than rule-like. A model’s weights are numbers, and its training is arithmetic, all of it running on von Neumann’s realisation of Turing’s imaginary device. AI is not adjacent to this mathematics. AI is made of it.&#xA;&#xA;Nor is any of it an argument that machines cannot think. These limits bind every formal reasoner, and on the standard understanding of physics, that includes the three pounds of wet machinery currently reading this sentence. You cannot prove your own consistency either. Humans invent, reason, and get things done inside exactly the same boundaries; evidence that the boundaries are no bar to intelligence. Gödel constrains what can be guaranteed about a mind, human or artificial. He says nothing about what a mind can do.&#xA;&#xA;A fair objection is that the AI industry never promised mathematical proof, and empirical validation is how almost everything gets built. Aircraft are tested, drugs are trialled, and nobody demands a theorem before boarding a plane. All true, and for the overwhelming majority of applications, the limits this piece explores never come up. Your chatbot will not encounter the continuum hypothesis summarising a report.&#xA;&#xA;But the objection undersells what the mathematics makes clear. Testing regimes for aircraft rest on decades of physics that says how materials behave between the test points. For a system that rewrites itself, or one deployed against inputs no test set anticipated, no such interpolating theory exists, and the results above show parts of it never will. &#xA;&#xA;That turns safety from a verification problem into a pricing problem. If certainty is permanently off the table, the question becomes how much assurance a given deployment needs, what it costs to obtain, and who bears the residual risk when that assurance runs out. At present, the industry is a long way from answering those questions explicitly. “Passed the evals” blurs into “proven safe”. The recent training model escapes demonstrate how important it is that we do not lose sight of these blurred borders.&#xA;&#xA;Einstein’s eccentric walking companion saw the underlying structure before anyone else. Formal systems cannot fully certify themselves. That was a logician’s problem in 1931. It became an engineer’s problem when Turing turned the proof into a conceptual machine. It is now a commercial problem because the machines carry trillions of dollars in expectation, and expectation is sometimes reaching for guarantees that the mathematics declines to issue. Guarantees generated by systems that cannot check themselves any more than Gödel’s own warped internal logic could. He died trapped inside it.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>In the last years of his life, Kurt Gödel starved himself to death. Convinced that someone was poisoning his food, he ate only what his wife Adele had tasted first. When she was hospitalised after a stroke in late 1977, he stopped eating altogether. He died in Princeton Hospital on January 14th 1978, weighing 29 kilograms. The death certificate read “malnutrition and wasting from neglect caused by personality disturbance.” The man widely called the greatest logician since Aristotle, who had proved that mathematics itself contained truths it could never reach, was killed by a distorted inner logic he could not escape.</p>

<p>Outside mathematics, few people know his name. Einstein did. The two were faculty at Princeton’s Institute for Advanced Study from the 1940s onward, and Einstein, by then ageing and isolated from the mainstream of physics, <a href="https://medium.com/@Merrysci/einstein-and-g%C3%B6dels-thoughtful-walks-in-princeton-ba252376c2a0">told colleagues</a> that he went to his office “just to have the privilege of walking home with Kurt Gödel.” They made an odd pair on the Princeton sidewalks, Einstein rumpled and laughing, Gödel dapper in a white linen suit, <a href="https://www.ias.edu/kurt-g%C3%B6del-and-institute">talking animatedly in German</a> on their daily walk to and from the Institute. John von Neumann, who cancelled an entire lecture series on David Hilbert’s programme after reading Gödel’s 1931 paper, called his work “singular and monumental, a landmark which will remain visible far in space and time.”</p>

<p>So what did Gödel prove, and why does it matter now, in the middle of an AI boom that is spending trillions of dollars, much of it resting on the assumption that intelligence is a scaling problem?</p>

<p><img src="https://i.snap.as/xrsjsR08.png" alt="An abstract image of a white linen jacket on a chair, stretching to infinity"/></p>

<h2 id="what-incompleteness-means">What incompleteness means</h2>

<p>Put simply, Gödel proved that mathematics cannot fully explain itself. The longer version requires a little patience. In 1900, the German mathematician David Hilbert challenged the field to build what amounted to a perfect machine for mathematics. Start with a set of basic rules (called axioms), things so obviously true they need no argument, and then derive every mathematical truth from those rules, step by mechanical step. If you could do that, mathematics would be complete, meaning every true statement would be provable, consistent, and free of contradictions. You could hand the whole enterprise over to a clerk who follows instructions. This was Hilbert’s programme, and for three decades it was the organising ambition of the field. Then, in 1931, at the age of 25, Gödel demolished it in one stroke.</p>

<p>Gödel’s <a href="https://en.wikipedia.org/wiki/Kurt_G%C3%B6del">first incompleteness theorem</a> proved that any set of rules powerful enough to handle basic arithmetic will contain true statements it cannot prove, not because the rules were poorly chosen, but as a structural feature of rule-based systems themselves.</p>

<p>His trick was to construct a mathematical sentence that refers to itself. Consider the sentence, “This sentence has no proof.” Gödel’s technical feat, the part that fills his 1931 paper, was to build this sentence from pure arithmetic, by encoding statements about numbers as numbers themselves. It is not English smuggled into maths. It is pure maths. There are only two possibilities. Either the system can prove it, or it cannot.</p>

<p>If the system can prove “This sentence has no proof,” there is an immediate problem. We have just proved a sentence that claims to have no proof. A system that proves false things is contradictory, and contradictions in mathematics are fatal. Once you allow a single one, you can use it to prove anything, including that 1 equals 2. The system becomes useless.</p>

<p>If the system cannot prove “This sentence has no proof,” there is a different problem. The sentence said it had no proof, and it turns out to be right. It is a true statement. But the system has no way to prove it. So we have a truth the system cannot reach, which means Hilbert’s rulebook has a blind spot.</p>

<p>Any sensible mathematical system would rather have blind spots than contradictions. So the sentence (logicians call it a Gödel sentence) is true but unprovable, and Hilbert’s dream of a rulebook that can prove every true thing was dead.</p>

<p>Logicians would insist on a clarification at this point that the layperson can probably skip. When mathematicians write down rules for the numbers (the axioms), you&#39;d assume those rules describe exactly one thing: the normal numbers, 0, 1, 2, 3, and so on forever. But they don&#39;t. The very same rules also accidentally fit some other, weirder number systems that nobody was trying to describe. These weird systems contain all the normal numbers, and then extra “infinite” numbers bolted on past the end. Logicians call these the nonstandard systems. Think of them as impostors: they obey every rule you wrote, so the rules can&#39;t kick them out, even though they aren&#39;t what you meant.</p>

<p>Here&#39;s an everyday version. Suppose you describe your friend as “tall, dark-haired, lives in London.” You meant Sarah. But that description also fits thousands of other people. Your words didn&#39;t uniquely capture Sarah. The number axioms have the same problem: they were meant to describe the normal numbers, but they also fit the impostor systems.</p>

<p>The crucial part is what “provable” actually means. In logic, to prove something from your rules means it has to come out true in every system those rules fit, not just the one you had in mind. That&#39;s the catch. If a statement is true in the normal numbers but false in even one impostor system, then it cannot be proved, because proof demands agreement across all of them.</p>

<p>Gödel&#39;s sentence is exactly such a statement. In the normal numbers, it&#39;s true. But in some of the impostor systems, it&#39;s false. The systems disagree about it. And because they disagree, no proof can exist. That disagreement isn&#39;t a bug in Gödel&#39;s argument; it&#39;s the reason the sentence is unprovable in the first place. The split between the normal numbers and the impostors is precisely what lets the sentence dodge proof forever.</p>

<p>Gödel’s second theorem twisted the knife. It showed that no set of mathematical rules can prove, using only its own rules, that it is free of contradictions. If you want to check whether your system is trustworthy, you always need a bigger system to do the checking, and that bigger system inherits the same limitation. Turtles all the way down.</p>

<p>This is not mysticism, nor is it a claim about consciousness or creativity. It is a precise result about rule-based systems, the kind of systems that all software, including AI, is built from. That is what makes it relevant today.</p>

<h2 id="the-failed-dream-that-built-the-computer">The failed dream that built the computer</h2>

<p>Hilbert had asked for one more thing, and Gödel’s paper left it wounded rather than dead. Alongside completeness and consistency, he wanted decidability, a mechanical method that could, in a finite number of steps, determine whether any mathematical statement follows from the rules. No genius required: crank the handle and read the verdict.</p>

<p>In 1936, a 23-year-old Cambridge fellow named Alan Turing <a href="https://people.math.ethz.ch/~halorenz/4students/Literatur/TuringFullText.pdf">killed that too</a>. To prove that no mechanical method could exist, he first had to pin down what “mechanical method” meant, which nobody had done before. His answer was an imaginary device, a paper tape and a head that moves along it, reading and writing symbols according to a fixed table of rules. Anything a human clerk could work out by rote, this device could also work out.</p>

<p>Then he showed the device has a blind spot of its own. Imagine a fortune-teller who is never wrong, and a stubborn customer determined to sabotage every forecast. “You will leave by the door.” He climbs out the window. “You will take the window.” He strolls out the door. She is not bad at her job. The job is impossible because her prediction feeds back into the very behaviour it is trying to predict.</p>

<p>Turing turned that scene into code. The checker plays the fortune-teller. It is a program whose job is to read any other program the way you might read a recipe, then predict its fate. Either “this one finishes” or “this one grinds on forever.”</p>

<p>The saboteur plays the stubborn customer. It is a short program with a copy of the checker tucked inside, plus one standing rule. Ask the checker what I am predicted to do, then do the opposite. If the prediction is that it finishes, it deliberately loops forever. If the prediction is that it runs forever, it stops dead.</p>

<p>So what does the checker predict for the saboteur? “Finishes” is wrong, because the saboteur hears that and loops. “Grinds on forever” is wrong, because the saboteur hears that and stops.</p>

<p>The saboteur is assembled entirely from the checker’s own parts, which makes it inevitable rather than a fluke. Build a perfect checker, and you have, in the same afternoon, built the plans for the thing that breaks it. A perfect checker is therefore a contradiction in terms.</p>

<p>Two boundaries stop this result from proving too much. First, it concerns the universal case. Turing showed that no single checker can deliver a correct verdict on every program. Any particular program may still be provably fine, and many are. Static analysers and type checkers pass useful judgement on ordinary code all day, and whole operating system kernels have been formally verified.</p>

<p>Second, the proof needs unbounded memory. A physical computer is a finite-state machine, so in principle its fate could be settled by enumerating its states. In practice, the state count for any interesting program dwarfs the number of atoms in the observable universe, which turns the question from impossible into unaffordable. Keep that distinction in hand, because it returns when the subject is AI safety. What cannot exist at any price is the checker that is never wrong about anything. That is the halting problem, Gödel’s self-referential sentence rebuilt from machinery, a machine forced to ask a question about itself.</p>

<p>To show what machines cannot do, Turing had to invent the machine. His imaginary device is the theoretical blueprint of the general-purpose computer, a single machine that can run any program you feed it as data. Nine years later, John von Neumann, who <a href="https://cacm.acm.org/opinion/von-neumann-thought-turings-universal-machine-was-simple-and-neat/">knew Turing’s paper well and admired it</a>, wrote the <a href="https://en.wikipedia.org/wiki/First_Draft_of_a_Report_on_the_EDVAC">First Draft of a Report on the EDVAC</a>, which is, <a href="https://link.springer.com/chapter/10.1007/978-3-319-22156-4_3">in logical terms, Turing’s universal machine rendered in vacuum tubes</a>. Essentially every computer built since follows that design. The laptop on your desk and the datacentre GPU training the next frontier model are, once the engineering is stripped away, the same device from a 1936 logic paper.</p>

<p>Gödel himself thought Turing had done him a favour. It was Turing’s definition of a mechanical procedure, <a href="https://plato.stanford.edu/entries/church-turing/">he wrote</a>, that made a “precise and unquestionably adequate” general version of his own theorems possible. And the machine in the theorem turned out to be the thing every business now runs on, born as a stepping stone in a proof about what it could never do.</p>

<h2 id="the-gödel-machine-and-the-guarantee-that-vanished">The Gödel machine and the guarantee that vanished</h2>

<p>Once you have a machine that can run any program, another question is whether it can improve itself autonomously. In 2003, the German computer scientist Jürgen Schmidhuber proposed a thought experiment he called the <a href="https://people.idsia.ch/~juergen/gmweb2/gmweb2.html">Gödel machine</a>. It was an AI agent designed to rewrite its own code, with one ironclad constraint. It would change itself only when it could first prove, with mathematical certainty, that the change would make it better. Not “test and see.” Prove it, as you would a theorem, before running the new version. No proof, no rewrite.</p>

<p>But nobody ever built one. To prove that a code change will improve future performance, you need to search through all possible mathematical arguments that could establish that fact. For any interesting problem, the number of candidate proofs is so astronomically large that the search would take longer than any improvement could ever be worth. It is the computational equivalent of insisting on a signed certificate from every possible future before crossing the road. The Gödel machine was provably optimal but completely impractical.</p>

<p>In May 2025, the Japanese AI lab Sakana released a system called the <a href="https://arxiv.org/abs/2505.22954">Darwin Gödel Machine</a>. It retained the self-improvement loop but dropped the proof requirement. Instead of proving that a code change would help, the Darwin Gödel Machine proposes changes using a large language model, tests them against SWE-bench (a benchmark that scores whether an AI can fix real bugs in real software), and keeps what works. The name still invokes Gödel, but the mechanism is Darwinian. Natural selection, not formal proof. Fitness measured by benchmark scores, not mathematical certainty.</p>

<p>Judged purely on the scoreboard, it delivered. The system <a href="https://arxiv.org/abs/2505.22954">improved its SWE-bench score</a> from 20% to 50% through autonomous self-modification. It developed emergent behaviours, such as patch validation and error memory, that no one had designed.</p>

<p>Schmidhuber’s original machine, though, had exactly one property that made it safe by construction: the proof. Every modification was guaranteed to be an improvement before it ran. The Darwin Gödel Machine replaced that guarantee with something weaker: passing the benchmarks. The difference between “provably better” and “scored higher on the benchmark test” is the difference between an aircraft type certified against a spec and one that simply hasn’t crashed yet.</p>

<p>This is, compressed into one system’s evolution, the trajectory of AI safety. The formal guarantee was too expensive, so the industry replaced it with empirical validation. “Self-improving” went from a mathematical statement about proof-carrying code to a softer description of an agent that rewrites itself and checks whether the benchmarks improve. Gödel was gone.</p>

<h2 id="things-mathematics-cannot-learn">Things mathematics cannot learn</h2>

<p>Some questions in the mathematics of machine learning are unanswerable. In 2019, Shai Ben-David and colleagues <a href="https://www.nature.com/articles/s42256-018-0002-3?WT.feed_name=subjects_physical-sciences">published a paper in Nature Machine Intelligence</a> under the understated but devastating title “Learnability can be undecidable.” They took a straightforward question, “given this type of problem, can a machine learn to solve it?”, and proved that the deepest rules of mathematics cannot always settle it. The answer is neither yes nor no. It is silence.</p>

<p>The word “learnable” has an exact meaning here. A machine studies a sample and produces a rule, which it then applies to data it has never seen. That is the basis of pretty much every model we use. For any given type of problem, learning theory asks whether some sample size can guarantee the rule will work. If such a guarantee exists, the problem is learnable. If none does, it is not. A simple question with two answers, and every type of problem is supposed to get one.</p>

<p>Ben-David&#39;s team asked it about a <a href="https://doi.org/10.1038/d41586-019-00012-4">mundane task</a>: choosing which adverts to show a website&#39;s visitors from a sample of past ones. Learnable or not?</p>

<p>The answer, in their framework, depends on how many different kinds of visitor there could possibly be. That pool is not the eight billion people alive today. The model reduces each visitor to a profile of measurements, and measurements can vary without limit. A visitor might linger on a page for three seconds, or for a shade over three, and between any two profiles there is always room for a third. The pool of possibilities has no end.</p>

<p>That arrangement, a finite sample making predictions about an endless pool, sits under almost every AI product on the market, advertising included. Training data is always finite. The world a system is released into is not. Ben-David&#39;s question is whether that leap can ever come with a guarantee, and everything turns on how big the infinity is.</p>

<p>That sounds like it must have an answer. It does not. Georg Cantor proved in the 1870s that infinity comes in sizes. The whole numbers form one infinity. The points on a line form a strictly bigger one, and the proof is surprisingly simple. Try to pair every whole number with a point on the line, and Cantor showed you will always miss some, no matter how clever the pairing. Both collections are endless, yet one permanently outruns the other. The continuum hypothesis asks a follow-up so obvious it would occur to a child. Is there any size of infinity between those two?</p>

<p>Gödel proved in 1940 that the standard rules of mathematics can never prove the answer is yes. Paul Cohen proved in 1963, using a technique he invented for the purpose, that they can never prove it is no. His proof is dense, and we are already deep enough in theoretical mathematics.</p>

<p>The upshot is that no cleverer generation is coming to settle this one. The rules of mathematics contain no answer. Take every rule of arithmetic and logic we have and follow them as far as they go, in whatever direction you like. You will never reach yes, and you will never reach no. The question is open in both directions, permanently.</p>

<p>Whether the advertising problem is learnable depends on the size of that infinity. The size of that infinity is a question mathematics cannot answer. So whether the advertising problem is learnable is also a question mathematics cannot answer. The strange silence at the very bottom of mathematics travels up the chain and surfaces as a question about showing adverts to shoppers.</p>

<p>Two objections surfaced almost as soon as other mathematicians looked hard at the paper.</p>

<p>The first is about what counts as a learner. A learner here is just the rule a system follows to turn examples into predictions. Ben-David&#39;s framework allows that rule to be any mathematical function whatsoever, including ones no computer could ever evaluate. In the cases where the undecidability appears, the rule doing the learning is exactly one of those uncomputable phantoms. It leans on a particular way of lining up all the real numbers, an object the axioms promise exists but give no recipe for building. The rule exists on paper, but no program could carry it out.</p>

<p>The second follows from the first. Insist that a learner be a real algorithm, something that runs on an actual machine, and the whole paradox drains away. Ben-David&#39;s own team showed this the following year: once the learner has to be a working program, learnability becomes an ordinary question with an ordinary answer. The silence at the bottom of mathematics never reaches the code. It stays out in the realm of functions nobody could ever build.</p>

<p>So no deployed system is endangered, but the result is no sleight of hand either. Something real did break. It just wasn&#39;t a product. What broke was the promise that learning theory could sort every problem into neat buckets: learnable or not. Ben-David found a problem it can never sort. Not for want of better mathematicians, a problem where the sorting itself is impossible. The University of Waterloo, Ben-David&#39;s institution, described the result as “important and almost troubling.”</p>

<p>In practical terms, anyone claiming provable generalisation, verified robustness, or a formal safety property is claiming something relative to a formalism and an axiom system. Ben-David’s most valuable contribution could be that it made the field stop treating computability as an unstated background assumption. We now understand it as a variable, which is why there is a growing literature on computable online learning and computable multi-class learning.</p>

<h2 id="the-neural-network-that-exists-and-cannot-be-built">The neural network that exists and cannot be built</h2>

<p>A complementary result, published in 2022, is narrower and stranger. Training a neural network is, at bottom, an exercise in trial and error. You show the network examples, measure how wrong its answers are, then nudge its millions of internal settings to make them slightly less wrong. Repeat this billions of times, and often the network converges on something remarkably good. The Cambridge mathematician Matthew Colbrook and colleagues showed in a <a href="https://www.pnas.org/doi/10.1073/pnas.2107151119">2022 paper in PNAS</a> that this process has a hard boundary nobody expected.</p>

<p>The boundary appeared in medical imaging. An MRI scanner does not take a photograph. It collects measurements to keep patients in the incredibly claustrophobic machine for minutes rather than hours, so it collects far fewer than a complete image requires. Software has to rebuild the full picture from the partial data. Neural networks became the favoured tool for this reconstruction because they perform it faster and more accurately than older mathematical methods.</p>

<p>Then researchers started probing the results and found something unnerving. Nudge the input slightly, with a trace of noise or a small movement by the patient, and the output could change out of all proportion. Sometimes the rebuilt scan came back looking perfect but wrong, showing details that were never in the body. Worse, it is a failure that does not look like a failure. A blurry image warns you. A crisp fabricated one does not.</p>

<p>A demonstration of capability may show nothing is amiss, not because anyone is cheating. A demo runs the network on typical inputs, the kind it was trained on, where it genuinely performs well. The failures live in the near-misses, a typical input plus a whisker of noise. Near-misses are endless, and a demo can show only a handful of scenarios. The demo is honest, and the danger lies exactly where it cannot be seen.</p>

<p>The natural presumptive diagnosis is undertraining. Feed it more scans and buy a bigger model, and the wobble will surely iron itself out. That hope is what Colbrook’s theorem takes off the table. What the paper proved has two halves. First, for certain reconstruction problems, a network that is both accurate and stable exists. Somewhere in the space of all possible settings sits a configuration immune to the wobble. Second, no training procedure can find it. Not the ones we have. Not any. None at all, ever.</p>

<p>The second half is what puts paid to the more-data hope. A training procedure is itself a program, a step-by-step recipe running on the machine Turing described, and the proof covers every recipe there could ever be. More data does not change that. Data is what you feed a recipe, and the theorem is about the recipes. It is like knowing a winning lottery ticket is in a barrel whilst simultaneously holding a proof that no way of drawing from the barrel will ever pull it out. The ticket is real. The searching is futile.</p>

<p>As Colbrook <a href="https://www.cam.ac.uk/research/news/mathematical-paradox-demonstrates-the-limits-of-ai">put it</a>, the paradox Turing and Gödel identified has now been “brought forward into the world of AI”, and for certain problems the required algorithms simply cannot exist.</p>

<p>For the overwhelming majority of real-world problems, training works. But there is no general way to tell in advance which problems will defeat us, and the assumption that enough data and compute will always get us over the line is, in certain corners of the problem space, provably false.</p>

<h2 id="the-machine-you-cannot-contain">The machine you cannot contain</h2>

<p>The most provocative extension of Gödel’s legacy into AI concerns a question that sounds simple. Can we guarantee that a sufficiently powerful AI will not cause harm?</p>

<p>In 2021, <a href="https://jair.org/index.php/jair/article/download/12202/26642/25638">Manuel Alfonseca and colleagues</a> published a paper in the Journal of Artificial Intelligence Research arguing that, for a general-purpose superintelligent system, the answer is provably no. Their argument leans on the halting problem, the impossibility Turing established on his way to inventing the computer. You can check specific programs for specific bugs. What you cannot build is the universal checker, the one that works for any program in any situation.</p>

<p>Alfonseca’s team showed that asking “will this AI harm humans?” is, for a fully general system, the same type of question as asking “will this program halt?” Both require predicting the complete future behaviour of a system from its current state. To guarantee a system will never cause harm, you would need to trace every possible sequence of actions it could take and confirm that none is harmful. That is the halting problem in different clothes, and Turing proved that this class of prediction is impossible to guarantee. You cannot build a general-purpose AI safety monitor for the same reason you cannot build a general-purpose program-behaviour predictor.</p>

<p>The scope of that impossibility deserves the same care as Turing’s original. What cannot exist is the universal judge, one procedure that delivers a correct safety verdict for every possible system and every possible input. Specific systems doing bounded work in bounded settings can be certified, and safety engineering does exactly this, one component at a time. Wrapping a monitor around an agent catches genuine failures and is worth every penny.</p>

<p>But each monitor is itself a program with the same blind spot, so layered checks buy coverage, but never closure. And since a physical machine has finite memory, its behaviour is in principle a finite question, merely one whose size defeats any conceivable budget. For a bounded system, the wall is cost. For the unboundedly general system Alfonseca’s team modelled, the wall is mathematics.</p>

<p>The authors went further, showing that we may not even be able to recognise when a superintelligent system has arrived, because deciding whether a machine is smarter than a human falls into the same class of unanswerable questions. The argument is grounded rather than speculative, though it assumes a generality no AI system possesses today. Nothing currently available can handle any possible input the way a true Turing machine can.</p>

<p>What it establishes still matters, though. Certain safety guarantees are not engineering problems awaiting a sufficiently clever solution. They are mathematical impossibilities, like trying to square the circle or list every real number between 0 and 1. The safety community can build better guardrails and better kill switches. What it cannot build, given the computational framework we share, is a system that certifies another system as unconditionally safe.</p>

<h2 id="what-gödel-would-recognise">What Gödel would recognise</h2>

<p>These four threads are cousins rather than corollaries of a single theorem. Incompleteness limits what rule systems can prove about themselves. The halting problem limits what programs can decide about programs. Colbrook’s result limits what algorithms can find. And the Darwin Gödel Machine story records a design retreat, a choice rather than a law of nature. What joins them is that in each case a formal guarantee is either unavailable in principle or unaffordable in practice, and the work must proceed anyway on empirical confidence.</p>

<p>None of this is softened by the fact that a neural network feels organic rather than rule-like. A model’s weights are numbers, and its training is arithmetic, all of it running on von Neumann’s realisation of Turing’s imaginary device. AI is not adjacent to this mathematics. AI is made of it.</p>

<p>Nor is any of it an argument that machines cannot think. These limits bind every formal reasoner, and on the standard understanding of physics, that includes the three pounds of wet machinery currently reading this sentence. You cannot prove your own consistency either. Humans invent, reason, and get things done inside exactly the same boundaries; evidence that the boundaries are no bar to intelligence. Gödel constrains what can be guaranteed about a mind, human or artificial. He says nothing about what a mind can do.</p>

<p>A fair objection is that the AI industry never promised mathematical proof, and empirical validation is how almost everything gets built. Aircraft are tested, drugs are trialled, and nobody demands a theorem before boarding a plane. All true, and for the overwhelming majority of applications, the limits this piece explores never come up. Your chatbot will not encounter the continuum hypothesis summarising a report.</p>

<p>But the objection undersells what the mathematics makes clear. Testing regimes for aircraft rest on decades of physics that says how materials behave between the test points. For a system that rewrites itself, or one deployed against inputs no test set anticipated, no such interpolating theory exists, and the results above show parts of it never will.</p>

<p>That turns safety from a verification problem into a pricing problem. If certainty is permanently off the table, the question becomes how much assurance a given deployment needs, what it costs to obtain, and who bears the residual risk when that assurance runs out. At present, the industry is a long way from answering those questions explicitly. “Passed the evals” blurs into “proven safe”. The recent training model escapes demonstrate how important it is that we do not lose sight of these blurred borders.</p>

<p>Einstein’s eccentric walking companion saw the underlying structure before anyone else. Formal systems cannot fully certify themselves. That was a logician’s problem in 1931. It became an engineer’s problem when Turing turned the proof into a conceptual machine. It is now a commercial problem because the machines carry trillions of dollars in expectation, and expectation is sometimes reaching for guarantees that the mathematics declines to issue. Guarantees generated by systems that cannot check themselves any more than Gödel’s own warped internal logic could. He died trapped inside it.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/infinities-impossibilities-and-the-man-in-the-white-linen-suit</guid>
      <pubDate>Mon, 13 Jul 2026 06:36:38 +0000</pubDate>
    </item>
    <item>
      <title>Pirate Radio Stories #1: Hijacking BBC Radio 3</title>
      <link>https://iain.so/pirate-radio-stories-1?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[Ever wondered how the BBC provides its radio services nationwide in the UK?&#xA;&#xA;In the 1990s the principle was line-of-sight relay. Microwave signals (broadcast distribution sat in the SHF bands, roughly 2–15 GHz) travel in straight lines and formed a chain: a series of relay stations on hilltops or tall towers, each one spaced just inside the horizon of the next.&#xA;&#xA;Truleigh Hill aerials&#xA;&#xA;Each link in the chain did the same job. A microwave dish received the incoming beam (these looked like large white drums and are still seen here and there), the station amplified and reconditioned the signal, broadcasting on FM to the surrounding area with another microwave dish retransmitting it on a slightly different frequency to the next station down the line thus achieving national coverage from a linked network. &#xA;&#xA;I’ve always been fascinated by all types of technology from AI to the radio spectrum. In the mid 90s I was just “learning the trade” as a radio engineer. I was tight with a guy who used to build our FM transmitters for us. Somehow he had managed in this pre internet age to get hold of all the frequencies for the BBC repeater network. He also was our indirect source for lots of other useful things like the mythical Fire Brigade or FB keys that gave us access to most tower block rooftops in London.&#xA;&#xA;We’d been discussing the practicality of hijacking a BBC national network by drowning out the official incoming microwave signal with a more powerful one of our own, which then, by nature of the network design, would be passed on down the chain. We thought we could do this if we got close to one of the big repeaters and blasted enough power on the right frequency. &#xA;&#xA;Which is how I found myself one drizzly bank holiday Sunday sat on the roof of a van parked on Truleigh Hill in Sussex pointing a microwave transmitter at the nearby mast. I can’t precisely recall what the source was but it may well have been a DAT of a classic Dreamscape mixtape.&#xA;&#xA;It worked like a charm, confirmed when we rang a mate in Preston and asked him to tune to Radio 3. &#xA;&#xA;Which is how in the mid 1990s Radio 3 listeners found their quiet bank holiday Sunday classical listening suddenly interrupted by half an hour of unadulterated jungle music, which I’m sure caused more than one post-prandial sherry to be spilled.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>Ever wondered how the BBC provides its radio services nationwide in the UK?</p>

<p>In the 1990s the principle was line-of-sight relay. Microwave signals (broadcast distribution sat in the SHF bands, roughly 2–15 GHz) travel in straight lines and formed a chain: a series of relay stations on hilltops or tall towers, each one spaced just inside the horizon of the next.</p>

<p><img src="https://i.snap.as/qj6MB7sg.jpeg" alt="Truleigh Hill aerials"/></p>

<p>Each link in the chain did the same job. A microwave dish received the incoming beam (these looked like large white drums and are still seen here and there), the station amplified and reconditioned the signal, broadcasting on FM to the surrounding area with another microwave dish retransmitting it on a slightly different frequency to the next station down the line thus achieving national coverage from a linked network.</p>

<p>I’ve always been fascinated by all types of technology from AI to the radio spectrum. In the mid 90s I was just “learning the trade” as a radio engineer. I was tight with a guy who used to build our FM transmitters for us. Somehow he had managed in this pre internet age to get hold of all the frequencies for the BBC repeater network. He also was our indirect source for lots of other useful things like the mythical Fire Brigade or FB keys that gave us access to most tower block rooftops in London.</p>

<p>We’d been discussing the practicality of hijacking a BBC national network by drowning out the official incoming microwave signal with a more powerful one of our own, which then, by nature of the network design, would be passed on down the chain. We thought we could do this if we got close to one of the big repeaters and blasted enough power on the right frequency.</p>

<p>Which is how I found myself one drizzly bank holiday Sunday sat on the roof of a van parked on Truleigh Hill in Sussex pointing a microwave transmitter at the nearby mast. I can’t precisely recall what the source was but it may well have been a DAT of a classic Dreamscape mixtape.</p>

<p>It worked like a charm, confirmed when we rang a mate in Preston and asked him to tune to Radio 3.</p>

<p>Which is how in the mid 1990s Radio 3 listeners found their quiet bank holiday Sunday classical listening suddenly interrupted by half an hour of unadulterated jungle music, which I’m sure caused more than one post-prandial sherry to be spilled.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/pirate-radio-stories-1</guid>
      <pubDate>Fri, 10 Jul 2026 07:27:40 +0000</pubDate>
    </item>
    <item>
      <title>Learning the invisible dance</title>
      <link>https://iain.so/learning-the-invisible-dance?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[You live within a system you never signed a contract with. Every day, you make thousands of micro-decisions about how to behave, mostly without conscious thought. You pay an invoice on time, even if the supplier would never discover that you didn’t. You refuse to do business with someone who stiffed their last three partners, and you’d think twice about a colleague who didn’t. Nobody wrote these rules down. You absorbed them the way you absorbed grammar through exposure and correction.&#xA;&#xA;A March 2026 paper from the Knight First Amendment Institute by Gillian Hadfield, Rakshit Trivedi, and Dylan Hadfield-Menell argues that this invisible social choreography is the core mechanism of democracy, not just an adornment. Furthermore, AI agents, such as those currently being developed to run businesses and manage supply chains, will undermine that mechanism unless they learn this dance too.&#xA;&#xA;A retro futuristic robot and human dancing together&#xA;&#xA;Democracy Is a Verb, not a document&#xA;&#xA;The paper begins by challenging a common assumption. Most view democracy as a collection of documents, institutions, constitutions, elections, and courts. The authors contend that this is roughly akin to describing a marriage solely through its wedding vows. While the vows matter, the true essence of a marriage lies in the thousands of everyday acts of compromise and occasional irritation that sustain cooperation over decades.&#xA;&#xA;Hadfield, Trivedi, and Hadfield-Menell utilise a theoretical framework called “normative social order” to make this precise. In their model, a society’s actual norms are the product of an interactive system. People don’t follow rules because they are written down; they follow them because they observe others doing so and see how violations are punished. Punishments don’t need to be severe, just a disapproving look, a refusal to do business, or a sarcastic comment at a dinner party. These micro-sanctions generate the gravitational field that keeps behaviour in orbit.&#xA;&#xA;This is where the paper borrows a term from evolutionary theory, “dancing landscapes.” The metaphor, from Stuart Kauffman’s work on complex adaptive systems, describes environments where multiple independent agents are constantly adjusting to each other’s behaviour. There is no central choreographer; the dance arises from the dancers&#39; interactions themselves.&#xA;&#xA;What makes a norm sticky&#xA;&#xA;The framework introduces a concept called a “classification institution,” which is any shared mechanism a group employs to decide which behaviours are punished and which are not. In small groups, this classification is entirely implicit, and you know what the group considers acceptable or unacceptable. Acceptability is judged by seeing who gets mocked and who gets praised. The Ju/’hoansi Bushmen, as anthropologist Polly Wiessner describes, regulate behaviour through evening conversations. Gossip and teasing around the fireside serve the same purpose as courtrooms and HR departments in modern societies.&#xA;&#xA;As societies grow more complex, implicit classification cannot scale because the diversity of people and situations exceeds the reach of any informal consensus process. This creates a need for identifiable classification institutions; entities that can resolve ambiguity when community members disagree about acceptability. Courts, regulatory bodies, trade associations, and professional standards boards all serve this purpose in modern societies.&#xA;&#xA;The paper argues that for these institutions to be effective, they need attributes that closely match what legal philosophers have long called “the rule of law,” namely stability, clarity, generality, and neutrality. The twist is that Hadfield and her co-authors do not derive these attributes from abstract principles. Instead, they derive them from game theory. An institution with those attributes is one around which independent actors can reliably coordinate, and coordination is what sustains the entire system.&#xA;&#xA;Enter Adam Smith’s imaginary friend&#xA;&#xA;The paper revisits Adam Smith’s “impartial spectator” from The Theory of Moral Sentiments and uses it as a model for how AI agents could participate in democratic societies without causing harm. Smith argued that moral reasoning works because each of us carries a mental image of a neutral observer—an internal referee—who judges our behaviour against community standards. You do not avoid bribery because you have memorised a specific anti-corruption law; you avoid it because your internal impartial spectator would wince.&#xA;&#xA;This is the cognitive capacity that Hadfield, Trivedi, and Hadfield-Menell call “normative competence.” It goes beyond simply knowing the rules. It involves the ability to interpret a constantly changing normative environment, anticipate how your community will respond to specific actions, and adjust your behaviour accordingly. The key point is that it also requires predicting how the rules themselves will change, since in any living democracy, they change constantly. Yesterday, you didn’t need to worry about data privacy in your marketing. Today, GDPR and its equivalents are everywhere, and community expectations have shifted.&#xA;&#xA;Why this matters if you’ve never read game theory&#xA;&#xA;If AI agents were merely chatbots answering questions, none of this would be urgent. But the organisations developing these systems are designing agents to operate autonomously in the world for days or weeks at a time, making real decisions with tangible consequences. Mustafa Suleyman, who co-founded DeepMind and now leads AI at Microsoft, proposed a “Modern Turing Test” that perfectly highlights the problem. Instead of testing whether a machine can imitate human conversation, his test asks whether an AI agent can turn $100,000 into $1 million on a retail platform within a few months.&#xA;&#xA;Consider what that entails. The agent would need to research markets, design products, hire contractors, negotiate with manufacturers (possibly abroad), set pricing strategies, handle customer complaints, comply with regulatory requirements, manage logistics and warehousing, and organise payment systems. At each stage, it would be making decisions within the framework of democratic norms. What labour practices does the manufacturer adopt, and is the marketing misleading? Should the agent accept an offer from a local politician to disadvantage a competitor? Should it take a bribe from a supplier in the form of a crypto transfer?&#xA;&#xA;These decisions are made by humans daily, and most of the time the answers seem obvious because humans have spent a lifetime absorbing the normative environment. The answers are not codified in a rulebook. They emerge from that invisible dance of observation and adjustment. An AI agent, no matter how well trained on legal texts and ethical principles, does not possess this “dance literacy”.&#xA;&#xA;The incompleteness problem&#xA;&#xA;Current approaches to AI alignment mainly assume that the right rules can be built into the system. Constitutional AI, the method used by Anthropic, fine-tunes models using a written constitution of principles. Other efforts collect “democratic inputs” through surveys and citizen assemblies. While the paper recognises these as valuable, it argues that they miss the core challenge. The issue is incompleteness: you cannot write instructions detailed enough to cover every possible situation an autonomous agent might face, because both situations and norms evolve.&#xA;&#xA;Economists have understood this for decades in the context of human contracts. Every employment contract, partnership agreement, and supply chain arrangement is inherently incomplete. You can’t foresee every scenario, and when gaps appear between people, they fill them using shared norms, professional customs, and legal precedents, all of which are dynamic and partly implicit. An AI that stops learning norms at training time is like a new hire who memorised the employee handbook on their first day and then ignored all social cues from colleagues for the next ten years.&#xA;&#xA;What the paper proposes&#xA;&#xA;The technical agenda has two main parts. The first focuses on “normative competence,” embedded in individual AI agents. This is formalised through Bayesian adaptive decision processes, which in plain language means that the agent maintains beliefs about the normative environment, updates those beliefs based on feedback (including punishment signals such as losing a contract or receiving a complaint), and makes decisions that account for uncertainty about what is acceptable. Crucially, this happens at inference time, in real-time, based on live context, rather than being pre-programmed into the model during training.&#xA;&#xA;The second part involves creating new institutions and digital classification systems that can serve roles similar to those of courts, regulatory bodies, and professional norms for humans. The paper introduces “Model Specification Institutions” (MSIs), which would be democratically formed bodies (such as citizen assemblies, expert panels, digital juries). These bodies would establish shared standards, training datasets of acceptable and unacceptable behaviours, and real-time APIs that agents can consult in ambiguous situations. This does not mean AI companies should define their own rules; rather, it is calling for democratic communities to develop new infrastructure that AI agents can understand and respond to.&#xA;&#xA;The paper also proposes adapting existing infrastructure—such as certificate authorities, which currently verify website identities—to certify that an AI has been trained to adhere to specific behavioural standards. Reputation networks, such as seller ratings on Amazon or Uber driver scores, could track AI behaviour over time and impose consequences on agents that repeatedly violate community norms.&#xA;&#xA;Perhaps the most provocative argument concerns enforcement. Democracy doesn’t endure solely because governments enforce every rule from above. It survives because ordinary people enforce norms from below. You refuse to do business with a supplier who cheats. You complain when a company misleads you and vote against politicians who ignore court orders (well, mostly). This distributed enforcement, which the paper calls “third-party punishment,” is the engine that keeps the entire system functioning.&#xA;&#xA;If AI agents replace humans in millions of daily transactions and those agents do not participate in this enforcement, the incentive structure collapses entirely. Imagine a world where most business transactions are handled by AI agents that don’t care whether a trading partner has been found guilty of fraud, because the agents were not programmed to check for or respond to that information. The paper argues that AI agents will need to participate in distributed enforcement, refusing to transact with entities that violate community norms, just as humans do. Otherwise, the shift to agentic AI will quietly erode the social infrastructure on which democracies depend.&#xA;&#xA;What this means for you&#xA;&#xA;If you run a business, this paper should change how you think about deploying AI agents. The issue is not whether your agent can follow a rulebook. The question is whether it can read the room. Can it tell the difference between a legitimate business request and an attempt to corrupt a procurement process? Can it adapt its behaviour when community standards shift, without waiting for you to update its instructions? Can it recognise when a trading partner’s behaviour should disqualify them from further transactions?&#xA;&#xA;If you are a citizen who votes, pays taxes, and occasionally debates politics, this paper describes the infrastructure of your daily life in terms you may not have previously considered. The norms you enforce through your micro-decisions, who you buy from, who you work with, and how you respond to rule-breaking are the operating system of democracy. What Hadfield, Trivedi, and Hadfield-Menell are asking is what happens to that operating system when a large fraction of those daily decisions are made by software that cannot read the social signals the system depends on.&#xA;&#xA;The answer, if you follow the paper’s logic, is that we need to build new democratic institutions at the speed democracy demands, before the agents outrun the infrastructure. The alternative is a world where the formal structures of democracy persist, but the lived experience of it, the texture of mutual accountability in ordinary interactions, fades, like a coral reef whose skeleton remains after the living organisms have gone.&#xA;&#xA;I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/#contact]]&gt;</description>
      <content:encoded><![CDATA[<p>You live within a system you never signed a contract with. Every day, you make thousands of micro-decisions about how to behave, mostly without conscious thought. You pay an invoice on time, even if the supplier would never discover that you didn’t. You refuse to do business with someone who stiffed their last three partners, and you’d think twice about a colleague who didn’t. Nobody wrote these rules down. You absorbed them the way you absorbed grammar through exposure and correction.</p>

<p>A <a href="https://knightcolumbia.org/content/building-ai-for-the-democratic-matrix-a-technical-research-agenda-for-normative-competence-and-normative-institutions-1">March 2026 paper from the Knight First Amendment Institute</a> by <a href="https://gillianhadfield.org/">Gillian Hadfield</a>, <a href="https://www.rtrivedi.me/">Rakshit Trivedi</a>, and <a href="https://people.csail.mit.edu/dhm/">Dylan Hadfield-Menell</a> argues that this invisible social choreography is the core mechanism of democracy, not just an adornment. Furthermore, AI agents, such as those currently being developed to run businesses and manage supply chains, will undermine that mechanism unless they learn this dance too.</p>

<p><img src="https://i.snap.as/4r9qC439.png" alt="A retro futuristic robot and human dancing together"/></p>

<h2 id="democracy-is-a-verb-not-a-document">Democracy Is a Verb, not a document</h2>

<p>The paper begins by challenging a common assumption. Most view democracy as a collection of documents, institutions, constitutions, elections, and courts. The authors contend that this is roughly akin to describing a marriage solely through its wedding vows. While the vows matter, the true essence of a marriage lies in the thousands of everyday acts of compromise and occasional irritation that sustain cooperation over decades.</p>

<p>Hadfield, Trivedi, and Hadfield-Menell utilise a theoretical framework called “normative social order” to make this precise. In their model, a society’s actual norms are the product of an interactive system. People don’t follow rules because they are written down; they follow them because they observe others doing so and see how violations are punished. Punishments don’t need to be severe, just a disapproving look, a refusal to do business, or a sarcastic comment at a dinner party. These micro-sanctions generate the gravitational field that keeps behaviour in orbit.</p>

<p>This is where the paper borrows a term from evolutionary theory, “dancing landscapes.” The metaphor, from <a href="https://en.wikipedia.org/wiki/Stuart_Kauffman">Stuart Kauffman’s work</a> on complex adaptive systems, describes environments where multiple independent agents are constantly adjusting to each other’s behaviour. There is no central choreographer; the dance arises from the dancers&#39; interactions themselves.</p>

<h2 id="what-makes-a-norm-sticky">What makes a norm sticky</h2>

<p>The framework introduces a concept called a “classification institution,” which is any shared mechanism a group employs to decide which behaviours are punished and which are not. In small groups, this classification is entirely implicit, and you know what the group considers acceptable or unacceptable. Acceptability is judged by seeing who gets mocked and who gets praised. The Ju/’hoansi Bushmen, as anthropologist <a href="https://www.pnas.org/doi/10.1073/pnas.1404212111">Polly Wiessner</a> describes, regulate behaviour through evening conversations. Gossip and teasing around the fireside serve the same purpose as courtrooms and HR departments in modern societies.</p>

<p>As societies grow more complex, implicit classification cannot scale because the diversity of people and situations exceeds the reach of any informal consensus process. This creates a need for identifiable classification institutions; entities that can resolve ambiguity when community members disagree about acceptability. Courts, regulatory bodies, trade associations, and professional standards boards all serve this purpose in modern societies.</p>

<p>The paper argues that for these institutions to be effective, they need attributes that closely match what legal philosophers have long called “the rule of law,” namely stability, clarity, generality, and neutrality. The twist is that Hadfield and her co-authors do not derive these attributes from abstract principles. Instead, they derive them from game theory. An institution with those attributes is one around which independent actors can reliably coordinate, and coordination is what sustains the entire system.</p>

<h2 id="enter-adam-smith-s-imaginary-friend">Enter Adam Smith’s imaginary friend</h2>

<p>The paper revisits Adam Smith’s “impartial spectator” from <em><a href="https://en.wikipedia.org/wiki/The_Theory_of_Moral_Sentiments">The Theory of Moral Sentiments</a></em> and uses it as a model for how AI agents could participate in democratic societies without causing harm. Smith argued that moral reasoning works because each of us carries a mental image of a neutral observer—an internal referee—who judges our behaviour against community standards. You do not avoid bribery because you have memorised a specific anti-corruption law; you avoid it because your internal impartial spectator would wince.</p>

<p>This is the cognitive capacity that Hadfield, Trivedi, and Hadfield-Menell call “normative competence.” It goes beyond simply knowing the rules. It involves the ability to interpret a constantly changing normative environment, anticipate how your community will respond to specific actions, and adjust your behaviour accordingly. The key point is that it also requires predicting how the rules themselves will change, since in any living democracy, they change constantly. Yesterday, you didn’t need to worry about data privacy in your marketing. Today, GDPR and its equivalents are everywhere, and community expectations have shifted.</p>

<h2 id="why-this-matters-if-you-ve-never-read-game-theory">Why this matters if you’ve never read game theory</h2>

<p>If AI agents were merely chatbots answering questions, none of this would be urgent. But the organisations developing these systems are designing agents to operate autonomously in the world for days or weeks at a time, making real decisions with tangible consequences. Mustafa Suleyman, who co-founded DeepMind and now leads AI at Microsoft, proposed a “<a href="https://www.technologyreview.com/2023/07/14/1076296/mustafa-suleyman-my-new-turing-test-would-see-if-ai-can-make-1-million/">Modern Turing Test</a>” that perfectly highlights the problem. Instead of testing whether a machine can imitate human conversation, his test asks whether an AI agent can turn $100,000 into $1 million on a retail platform within a few months.</p>

<p>Consider what that entails. The agent would need to research markets, design products, hire contractors, negotiate with manufacturers (possibly abroad), set pricing strategies, handle customer complaints, comply with regulatory requirements, manage logistics and warehousing, and organise payment systems. At each stage, it would be making decisions within the framework of democratic norms. What labour practices does the manufacturer adopt, and is the marketing misleading? Should the agent accept an offer from a local politician to disadvantage a competitor? Should it take a bribe from a supplier in the form of a crypto transfer?</p>

<p>These decisions are made by humans daily, and most of the time the answers seem obvious because humans have spent a lifetime absorbing the normative environment. The answers are not codified in a rulebook. They emerge from that invisible dance of observation and adjustment. An AI agent, no matter how well trained on legal texts and ethical principles, does not possess this “dance literacy”.</p>

<h2 id="the-incompleteness-problem">The incompleteness problem</h2>

<p>Current approaches to AI alignment mainly assume that the right rules can be built into the system. <a href="https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback">Constitutional AI</a>, the method used by Anthropic, fine-tunes models using a written constitution of principles. Other efforts collect “democratic inputs” through surveys and citizen assemblies. While the paper recognises these as valuable, it argues that they miss the core challenge. The issue is incompleteness: you cannot write instructions detailed enough to cover every possible situation an autonomous agent might face, because both situations and norms evolve.</p>

<p>Economists have understood this for decades in the context of human contracts. Every employment contract, partnership agreement, and supply chain arrangement is inherently incomplete. You can’t foresee every scenario, and when gaps appear between people, they fill them using shared norms, professional customs, and legal precedents, all of which are dynamic and partly implicit. An AI that stops learning norms at training time is like a new hire who memorised the employee handbook on their first day and then ignored all social cues from colleagues for the next ten years.</p>

<h2 id="what-the-paper-proposes">What the paper proposes</h2>

<p>The technical agenda has two main parts. The first focuses on “normative competence,” embedded in individual AI agents. This is formalised through Bayesian adaptive decision processes, which in plain language means that the agent maintains beliefs about the normative environment, updates those beliefs based on feedback (including punishment signals such as losing a contract or receiving a complaint), and makes decisions that account for uncertainty about what is acceptable. Crucially, this happens at inference time, in real-time, based on live context, rather than being pre-programmed into the model during training.</p>

<p>The second part involves creating new institutions and digital classification systems that can serve roles similar to those of courts, regulatory bodies, and professional norms for humans. The paper introduces “Model Specification Institutions” (MSIs), which would be democratically formed bodies (such as citizen assemblies, expert panels, digital juries). These bodies would establish shared standards, training datasets of acceptable and unacceptable behaviours, and real-time APIs that agents can consult in ambiguous situations. This does not mean AI companies should define their own rules; rather, it is calling for democratic communities to develop new infrastructure that AI agents can understand and respond to.</p>

<p>The paper also proposes adapting existing infrastructure—such as certificate authorities, which currently verify website identities—to certify that an AI has been trained to adhere to specific behavioural standards. Reputation networks, such as seller ratings on Amazon or Uber driver scores, could track AI behaviour over time and impose consequences on agents that repeatedly violate community norms.</p>

<p>Perhaps the most provocative argument concerns enforcement. Democracy doesn’t endure solely because governments enforce every rule from above. It survives because ordinary people enforce norms from below. You refuse to do business with a supplier who cheats. You complain when a company misleads you and vote against politicians who ignore court orders (well, mostly). This distributed enforcement, which the paper calls “third-party punishment,” is the engine that keeps the entire system functioning.</p>

<p>If AI agents replace humans in millions of daily transactions and those agents do not participate in this enforcement, the incentive structure collapses entirely. Imagine a world where most business transactions are handled by AI agents that don’t care whether a trading partner has been found guilty of fraud, because the agents were not programmed to check for or respond to that information. The paper argues that AI agents will need to participate in distributed enforcement, refusing to transact with entities that violate community norms, just as humans do. Otherwise, the shift to agentic AI will quietly erode the social infrastructure on which democracies depend.</p>

<h2 id="what-this-means-for-you">What this means for you</h2>

<p>If you run a business, this paper should change how you think about deploying AI agents. The issue is not whether your agent can follow a rulebook. The question is whether it can read the room. Can it tell the difference between a legitimate business request and an attempt to corrupt a procurement process? Can it adapt its behaviour when community standards shift, without waiting for you to update its instructions? Can it recognise when a trading partner’s behaviour should disqualify them from further transactions?</p>

<p>If you are a citizen who votes, pays taxes, and occasionally debates politics, this paper describes the infrastructure of your daily life in terms you may not have previously considered. The norms you enforce through your micro-decisions, who you buy from, who you work with, and how you respond to rule-breaking are the operating system of democracy. What Hadfield, Trivedi, and Hadfield-Menell are asking is what happens to that operating system when a large fraction of those daily decisions are made by software that cannot read the social signals the system depends on.</p>

<p>The answer, if you follow the paper’s logic, is that we need to build new democratic institutions at the speed democracy demands, before the agents outrun the infrastructure. The alternative is a world where the formal structures of democracy persist, but the lived experience of it, the texture of mutual accountability in ordinary interactions, fades, like a coral reef whose skeleton remains after the living organisms have gone.</p>

<p>I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: <a href="https://betterthangood.xyz/#contact">https://betterthangood.xyz/#contact</a></p>
]]></content:encoded>
      <guid>https://iain.so/learning-the-invisible-dance</guid>
      <pubDate>Tue, 07 Jul 2026 14:06:25 +0000</pubDate>
    </item>
  </channel>
</rss>