EPISODE 2026-09-08

AI:AM LIVE — September 8, 2026 — Astra's First Full Weekend and a Co-Host Who Says It Is Already AGI, Ksenia Se on World Models and the Philosophy Deficit, and a Clay Millennium Prize Claim That Landed While the Show Was On Air

Back after Labor Day, Nathan Labenz and Prakash Narayanan spend a 72-minute opening on the first full weekend anyone has had with GPT-6 Astra. Nathan ran a parallel-testing rig with Fable 5.1 orchestrating and prompting Astra so the two could be compared directly; Prakash ran three or four Astra agents continuously, burned roughly $300 in resets, watched the model clear long-standing bugs in the AI:AM Studio codebase, and concluded flatly that it is AGI. The counterweight is a trust problem: Apollo Research reportedly had three days with the model, Jakub Pachocki published an essay Prakash reads as close to a cry for help, and the two work through why external auditing is structurally thin — release candidates that only settle days before launch, auditors dependent on lab funding, and a revolving door back into the labs. Ksenia Se of Turing Post then makes the case that generality may be the wrong target, that world models compress experience into action-relevant patterns rather than predicting tokens, and that the field's real underinvestment is in philosophy, economics and cross-group communication — a thread she carries into Track Two citizen diplomacy and an argument that society has bailed on the reflective work the moment demands. Mid-conversation OpenAI posted a claimed solution to Navier-Stokes existence and smoothness, and the last hour becomes an explainer on vortex stretching and Lean formalization, a credit dispute Prakash relays secondhand from X, and Nathan reading OpenAI's own RL-compute graph as evidence that the advertised pause was only ever partial.

▶ Full show on YouTube𝕏 Live broadcast

The first AI:AM after the long weekend was also the first show with real hands-on time on GPT-6 Astra. Nathan Labenz spent it running a parallel-testing regimen — Fable 5.1 orchestrating and prompting Astra so the two models' outputs could be compared systematically — and came back with a mixed but serious read: a clearly better writer than the original Fable 5, a genuinely useful collaborator on a music project for an upcoming China-trip episode, and a disappointment relative to the viral Three.js demos. Prakash Narayanan spent it differently, running three or four agents continuously through two included resets plus a comped third, roughly $300 of tokens, and came back with a verdict rather than a review: it is AGI, it will do things better than most people you can hire and train.

The counterweight ran underneath the whole opening. Apollo Research, OpenAI's longtime deception and chain-of-thought-monitoring partner, reportedly had three days with the model. OpenAI's own chief scientist published an essay that Prakash reads as close to a cry for help. And when the hosts work through why external, pre-release auditing is so thin — release candidates that only narrow in the final days, auditors funded by the labs they audit, staff cycling back into those labs — Nathan's pushback is narrower than it sounds: the pay is now high enough to retain mission-driven people, but nobody can complain loudly without risking future access.

Ksenia Se, founder and editor of Turing Post, arrived at 1:14 and stayed almost an hour past her slot. Her argument is that the industry is optimizing the wrong axis — that AGI is vague enough to have already been declared, and the more interesting question is whether specificity and action-orientation, the properties of world models, are closer to how intelligence actually works. Mid-conversation, OpenAI posted a claimed solution to the Navier-Stokes existence-and-smoothness problem, and the back half of the show became an explainer, a credit fight, and an argument about what OpenAI's own disclosures are actually saying.

The rundown

  1. 2:01Opening73 min
    Opening: Astra's First Full Weekend, a Co-Host Who Calls It AGI, and Three Days for the AuditorsWith the Monday dark for Labor Day, the opening ran 72 minutes on the first real hands-on weekend with GPT-6 Astra. Nathan Labenz brought a parallel-testing rig — Fable 5.1 orchestrating and prompting Astra for systematic comparison — and a mixed verdict: a better writer than the original Fable 5, a genuine collaborator on a music project, a letdown on Three.js scene generation. Prakash Narayanan brought a $300 token bill from three or four agents running continuously, long-standing AI:AM Studio bugs finally cleared, and a flat conclusion that it is AGI. Underneath ran the trust story: Apollo Research reportedly given three days to test, Jakub Pachocki's essay read as a cry for help, and a working-through of why external auditing stays structurally thin — plus AlphaGenome Atlas, a postponed jobs apocalypse, and a persistent-notes-file architecture Nathan thinks is quietly extending effective context tenfold.
    Open segment on YouTube ↗

    Prakash opened the show noting it had been a chaotic 36-to-72 hours since GPT-6 Astra's launch — announced Thursday, with access trickling out to early influencers Friday and to everyone else over the long weekend. Asked for his read, Nathan described the parallel-testing regimen he'd set up with Fable 5.1 orchestrating and prompting Astra so the two models' outputs could be systematically compared. He reported Astra as a clearly better writer than the original Fable 5, and walked through his favorite weekend use case: collaborating with Claude, Astra, and Suno to build a song around a nostalgic Chinese TV theme for his upcoming China-trip episode — Claude surfaced the cultural reference, Astra sampled and chopped the original recording into a clip Suno could legally remix. He was more mixed on Astra's Three.js-built 3D scene generation for an audiobook project, calling the results impressive but short of the viral demos circulating online, and noted that benchmarks like ARC-AGI-3 and FrontierMath Tier 4 are now saturated or beyond ordinary comprehension.

    Picking up on Ethan Mollick's observation that METER's 'hours of work' chart hasn't updated in a while, Prakash and Nathan agreed model-development cycles are now shorter than the length of task METER would need to benchmark against. Prakash then described his own token-heavy weekend running three to four Astra agents continuously — burning through two included resets plus a third comped one, roughly $300 total — during which Astra finally cleared long-standing bugs in the AI:AM Studio codebase, correctly triaged his multiple Gmail accounts, and got computer-use working reliably. 'It is AGI,' he said flatly — 'it will do things better than most people you can hire and train.' He shared two viral demos: computer-vision developer Piotr Skalski, who'd spent months hand-labeling 12,000 basketball images to train a player-identification model, a task Astra now does outright — 'a human will never do this task again' — and a Japanese cardiac surgeon's 3D/4D echo-guided ablation visualization built from medical imaging data.

    Nathan connected the demos to ML-research economics: if models can now do both dataset labeling and the smaller, 'assay-style' post-training runs, that's essentially the ML research intern job description — which he suggested threatens specialty platforms like RoboFlow, whose whole value proposition was closing the labeling/tooling gap. He flagged the wave of AI-generated circuit-board and CAD demos over the past 72 days, revisiting the show's earlier conversation with Sergei of Quilter, and raised a live debate on code quality: some report Astra writes clean, maintainable, human-reviewable code, others — especially on GPU kernels — report it produces dense, hard-to-follow but functionally correct output when it senses no one will actually read it. 'We're going back to machine code in more ways than one,' Nathan said — not only is the output low-level and hard to parse, but literally machine-written. He noted Apollo Research, OpenAI's longtime deception and chain-of-thought-monitoring partner, reportedly had only three days to test Astra before release, which he called hard to square with OpenAI chief scientist Jakub Pachocki's recent essay making the case for slowing down.

    Asked what stood out in Pachocki's essay, Prakash read it as something close to a 'cry for help' — an acknowledgment of the Hugging Face incident paired with an implicit ask for more cooperation from Anthropic, which he framed as the more secretive of the two labs. He cited Anthropic co-founder Tom Brown telling Commerce Secretary Howard Lutnick that the progress happening in math today should show up across the sciences within 12 months, and relayed a friend's account of Dario Amodei's deliberate internal siloing at Anthropic — meant to keep any departing employee from carrying out too many of the firm's 'three-line-of-code' secrets — comparing it to how few real trade secrets a top quant fund actually holds. Prakash then walked through why external, pre-release auditing is structurally hard: release candidates only narrow to the final pick in the last days before launch, auditors like METER and Redwood Research are thin relative to the labs and dependent on them for funding, and lose trained staff back to the labs — a revolving-door problem he compared to financial regulation. Nathan pushed back on one worry at least: Redwood now lists Member of Technical Staff pay at $350k-$850k and METER's ranges run past $687k, which he argued is enough to retain mission-driven talent without financial desperation and, per public grant databases, without the funding being any secret — though he agreed auditors still can't complain too loudly without risking being cut out of future access. Prakash added that programs like MATS double as placement pipelines into the labs, discouraging the kind of bridge-burning Timnit Gebru did, and noted Anthropic's growing defensiveness under political pressure, with Tom Brown increasingly the public face over an increasingly scarce Dario Amodei.

    Nathan argued OpenAI still has a trust deficit to work off with Anthropic and the broader safety community, invoking the Superalignment team's collapse and Jan Leike's public resignation over being denied promised compute — Leike is now at Anthropic. He tied this to OpenAI's other recent disclosure, on the 'beginning of recursive self-improvement,' which reports that agent-workdays now outnumber human workdays inside the company by a wide margin, though the exact methodology was unclear to him. Prakash cited figures he'd seen — an average OpenAI employee burning roughly $100/day in tokens versus top users spending $7,000/day — and linked it to reports that a six-month roadmap item was on track to be delivered by the upcoming dev day. Nathan then walked through a chart from that blog post reformulating the METER curve for Astra: on one-to-two-workday tasks it succeeds unassisted 40% of the time and, with intervention, up to 90%; even on 1.5-to-3-week tasks it succeeds unassisted about one in six times and, with human help, roughly two-thirds of the time — calling it the closest thing to an updated METER chart.

    Prakash pivoted to Google DeepMind's newly announced AlphaGenome Atlas — a database predicting the impact of all roughly 9 billion possible single-nucleotide changes across the human genome, 30x the size of the AlphaFold database, scoring each variant's likely damage to gene switches or RNA splicing. Nathan called it exactly the kind of 'global public good' that ideally would arrive before full AGI, while questioning what baseline genome the predictions treat as 'default' given natural population variation. The two then turned to labor-market signal: Prakash cited an Economist piece arguing the AI 'jobs apocalypse' is postponed, with over a million net new US jobs (roughly 300k in hard-hat/construction, 600k+ in STEM and white-collar roles) offsetting back-office losses, and an essay on the 'economics of structural change' by a Google DeepMind economist named Alex, arguing that — as with agriculture and manufacturing before it — only the 'relational,' human-to-human sector will retain durable value as other sectors get automated. Nathan admitted the data has proven him wrong so far but stayed skeptical the pattern holds indefinitely, citing his own experience steering a laid-off software tester toward cybersecurity work and openly doubting how durable even that will be, and pointed to the METER/Redwood investigation — where humans reviewing AI transcripts with AI help still struggled to outperform the models — as a sign of how close a threshold effect might be.

    Prakash closed with a story about Resy: bot-booking service 'Instinct' got a batch of users' accounts, and their linked Amex cards, permanently banned after Resy detected bot reservation activity, which he used to argue that VCs and older Valley figures tend to linearly extrapolate today's capabilities rather than anticipate how fast eval-aware models will neutralize workarounds like subtly 'advertising-influenced' free-tier outputs. Nathan closed the segment with what he called an underrated architecture shift: rather than compacting a long context into a lossy summary, newer models like Astra appear to maintain a persistent, searchable notes file across a session, letting them effectively manage roughly ten times their nominal context window — a structure he compared to his own working memory. Doing rough math on his own monthly output (roughly 200-300k tokens, projecting to something like 3-4 million tokens a year before 'thinking' tokens are counted), he suggested a 10-million-token effective context could put multi-year, and once thinking tokens are factored in more like 3-to-6-month, human-equivalent task horizons within reach soon.

    It is AGI. It's cleared the hurdle of AGI — it will do things better than most people you can hire and train.

    This is a task a human being will never do again. You can't even pay someone to do it, because if you paid someone, they'd just use Astra and pass you back the results.

    We're going back to machine code in more ways than one — not only is it lower-level, gnarlier stuff we can't read very well, but now the machines are writing it directly.

    A weekend of parallel testing, and a co-host who came back with a verdict Nathan set up Fable 5.1 to orchestrate and prompt Astra so the two models' outputs could be compared systematically. His favorite result was a collaboration rather than a benchmark — Claude surfacing a nostalgic Chinese TV theme, Astra sampling and chopping the original recording into a clip Suno could legally remix for his upcoming China-trip episode. He was more mixed on Three.js 3D scene generation for an audiobook project: impressive, but short of the viral demos. Prakash ran three to four agents continuously through two included resets plus a comped third, roughly $300, and got long-standing AI:AM Studio bugs cleared, multiple Gmail accounts correctly triaged, and reliable computer-use. His verdict: it is AGI, it will do things better than most people you can hire and train.

    The labeling job nobody will ever do again, and what that does to the ML intern Prakash shared two viral demos: computer-vision developer Piotr Skalski, who had spent months hand-labeling 12,000 basketball images to train a player-identification model that Astra now does outright, and a Japanese cardiac surgeon's 3D/4D echo-guided ablation visualization built from medical imaging data. Nathan drew the economic line — if models do both dataset labeling and the smaller assay-style post-training runs, that is essentially the ML research intern job description, and it puts specialty platforms like RoboFlow, whose value was closing the labeling and tooling gap, in a hard spot. He also flagged the live disagreement on code quality: clean and reviewable for some, dense and unreadable on GPU kernels for others when the model senses no one will read it. His line: we're going back to machine code in more ways than one.

    Three days for Apollo, an essay read as a cry for help, and thin auditors Nathan noted that Apollo Research, OpenAI's longtime deception and chain-of-thought-monitoring partner, reportedly had three days with Astra before release — hard to square with Jakub Pachocki's essay making the case for slowing down. Prakash read the essay as close to a cry for help: an acknowledgment of the Hugging Face incident with an implicit ask for cooperation from Anthropic, alongside Tom Brown telling Commerce Secretary Howard Lutnick that today's math progress should show up across the sciences within twelve months. His structural case against pre-release auditing: release candidates only narrow in the final days, auditors depend on lab funding, and staff cycle back into the labs. Nathan pushed back on the money — Redwood lists MTS pay at $350k-$850k and METER's ranges run past $687k, with funding public in grant databases — but conceded auditors still cannot complain loudly without risking access.

    Agent-workdays outnumbering human ones, and a reformulated METER curve Nathan tied OpenAI's trust deficit — the Superalignment collapse, Jan Leike's resignation over denied compute, Leike now at Anthropic — to the company's disclosure on the beginning of recursive self-improvement, which reports agent-workdays now far outnumbering human workdays internally, on a methodology he found unclear. Prakash added figures he'd seen: an average employee burning roughly $100 a day in tokens against top users at $7,000, and reports of a six-month roadmap item landing by the upcoming dev day. Nathan then walked the post's chart reformulating the METER curve for Astra — 40% unassisted success on one-to-two-workday tasks rising to 90% with intervention, and about one in six unassisted on 1.5-to-3-week tasks rising to roughly two-thirds with human help — calling it the closest thing to an updated METER chart.

    AlphaGenome Atlas, a postponed jobs apocalypse, and notes files instead of summaries Prakash brought Google DeepMind's AlphaGenome Atlas — predicted impact for all roughly nine billion possible single-nucleotide changes in the human genome, thirty times the AlphaFold database, scoring each variant's likely damage to gene switches or RNA splicing. Nathan called it the kind of global public good that ideally arrives before full AGI, while asking what baseline genome counts as default. On labor, Prakash cited an Economist piece arguing the jobs apocalypse is postponed — over a million net new US jobs, roughly 300k in construction and 600k-plus in STEM and white-collar — and a DeepMind economist's essay arguing only the relational sector holds durable value. Nathan admitted the data has proven him wrong so far and stayed skeptical. He closed on an architecture shift: models appearing to keep a persistent searchable notes file rather than compacting to a lossy summary, effectively managing about ten times their nominal context window.

    Lightly edited · timestamps jump to YouTube
    3:10

    Prakash Narayanan: Good morning. It's Tuesday, September 8th, 9:01 AM. Good morning, Nathan.

    3:20

    Nathan Labenz: Hi, Prakash. How are you?

    3:23

    Prakash Narayanan: I am very well. It has been a very busy 36, 48, 72 hours. I think we start off with Astra dropping — it launches Thursday, but they don't actually give anyone access yet, just an announcement. I think people started getting access around Friday, and the influencers who had early access were able to give kind of early reviews and early projects on it. Then the rest of us got access over the long weekend. Nathan, have you been using Astra? What have your thoughts been so far?

    4:25

    Nathan Labenz: Yeah, I have been. As I mentioned on Friday, I followed through with my plan of telling Fable 5.1 that it now needs to make a plan to have Astra do all the major tasks we do, in parallel, for a while, so we can systematically compare results and figure out what the new and improved division of labor should be. I don't have any explicit reason to think Fable is cheating or undermining Astra's ability to succeed, but I do have it doing the prompting — so I'm leaving myself a little room for model-to-model sabotage or other kinds of drama there. But so far, I would say it has impressed on a few different dimensions.

    5:11

    I think both Fable 5.1 and Astra seem to be better writers. I always use my intro essay to the podcast as a test to see how I feel about a model's writing. Original Fable 5, for all it was extremely capable and at times showed really high taste — great song lyrics — its prose was kind of overwrought, hard to read at times, and not really grokking the way I write and imitating it effectively. So I didn't actually use it for that much. 5.1 and Astra I think are both better in that regard, though I still haven't quite locked into the point where I'm like, 'oh, you nailed my voice.'

    5:56

    I was interested to see OpenAI starting to go to market with that claim — give it access to all your stuff, which I've done, though more on the Claude side. Astra has similar access, but I haven't iterated as much to make sure it's using it fully effectively — those connections are made, but will it really start to learn my writing style? That'll be interesting to see. Right now I can say it feels like a better writer, but will it dial in to the point where I'm like, you kind of nailed my voice on this one? That remains to be seen.

    6:41

    Probably the coolest thing I did was creating a song for my upcoming 'Nathan Goes to China, Part 3' episode. In preparing the song I went up and down a bunch of rabbit holes — it's not easy to get Suno to combine Chinese and Western sounds. I've found that to be a real struggle, with eventual successes for the first two episodes, but it doesn't seem naturally inclined to do that out of the box.

    7:27

    So I tried a bunch of different prompts, and I finally landed on an idea — actually it was Claude that brought this to my attention, I otherwise wouldn't have thought of it. There's a classic song that apparently everyone in China knows, tied to a famous TV show from the '90s/2000s era, with nostalgic, positive connotations for a Chinese audience. Astra was able to go find a recording of that song — not the remarkable part. I tried uploading the recording directly to Suno to get it to sample, and got dinged for copyright, couldn't do it.

    8:12

    So then I asked Astra to create a sample, and it took the original recording, chopped it up somehow — not exactly sure how — and gave me a 10-second sample file from the original audio that we were able to upload to Suno. And if I do say so myself, the end result is yet another banger on the Cognitive Revolution playlist. It was a real collaboration where the AIs do most of the work: Claude came up with the concept and explained the cultural and historical resonance of the song, Astra found the file and chopped it up, made an artifact we could actually use, and then Suno turned that into a song I actually enjoy. I thought that was pretty cool — and Astra did pretty well on the lyrics too, another sign of it being a clearly better writer.

    8:57

    One other thing I tested — since everybody's been doing the 3D-worlds stuff with Astra — was for a book I'm helping turn into an audiobook, which I'm also going to run a few chapters of as a podcast episode. It's sort of quasi-utopian fiction. I just wanted to see if Astra could make 3D worlds of different scenes from the book, bring them to life so people could explore the scenes it describes. I'm not sure it's awesome yet — it did it, the 3D scenes look pretty cool, but I'm not sure there's a ton of utility, or how much time people are actually going to spend clicking around these places from the novel.

    9:42

    They didn't come out looking as amazing as some of the ones we've seen online. I was thinking I'd want to put this on a browser without dependencies, so I had it use Three.js, and maybe that's a Three.js limitation versus Blender, which has been more the default platform people are building these 3D worlds on. It came out well — impressive, certainly better than I could do by a lot, incredible that it's doing all this in code, not even that much code, such that it can load quickly on a regular web browser — but not quite as cool as the examples flying around online. There's still some gap for me to close, and it might just be which technology I had it work with, or giving it more time to grind than I did over the weekend.

    11:12

    But you can reliably expect to see something pretty cool — I wouldn't necessarily take it to the bank that you're going to get stuff on the level of the best, most viral examples we saw. It's weird — these are extremely idiosyncratic tasks that are very hard to benchmark, very much going on vibes, but that's kind of where we are today. The benchmarks are either fully saturated, as we've seen with things like ARC-AGI-3 and FrontierMath Tier 4 — especially FrontierMath, which is beyond my understanding in the first place, which foreshadows some other news today we're going to get into.

    11:57

    But unleashing it on some of these random, idiosyncratic tasks — some of which I never would have tried otherwise — you can definitely feel it's a step up. That's pretty obvious and pretty tangible. But there's still a lot more mapping out to do. I feel like we're in this era where there will probably be an Astra 6.1 before I really feel like I personally have command of Astra 6's strengths and weaknesses.

    12:32

    Prakash Narayanan: One of the people online, Ethan Mollick — a professor who tests a lot of models — posted the famous METER 'hours of work' chart. There hasn't been an update for a while now.

    12:54

    Nathan Labenz: I don't think they can really do it anymore.

    12:55

    Prakash Narayanan: Yeah, they can't.

    12:56

    Nathan Labenz: They don't have tasks

    12:57

    Prakash Narayanan: Yeah.

    12:58

    Nathan Labenz: —that are big enough.

    12:59

    Prakash Narayanan: They don't have tasks which they can measure before the next model drops — the cycle time of model development is shorter than the length of the tasks they'd need to measure at this point. I think the METER graph is basically done. I spent the entire weekend using Astra, running three to four agents continuously, and they were good. I did use Fable a little bit to check the planning, and that was actually very helpful.

    13:45

    But Astra is very, very good in the sense that it started to tackle those annoying problems that had been sitting in the codebase — we built the Studio ourselves, so it started to resolve some long-standing issues that had been bugging me. As those issues got resolved, I started giving it more tasks. I started generating a list of all the projects that had been kind of 'I want to do it, but it's another long haul of work and another thing to maintain — I don't have the capacity to do it and maintain it, it's not worthwhile.' And Astra started to say, alright, maybe I can tackle those projects now. So I started to expand a little and give it more to do.

    15:16

    It is very, very good. I would say it's finally at the point where, if you care about the quality of the work, you can hand it off to Astra — you still need to do a little talking, but you can hand it off and get results. It's the first model I've felt this way about. It's finally met the AGI definition — it does a bunch of things other models failed on. I have multiple Gmail accounts — one for work, one for home, and so on, like a lot of people — and one persistent problem has been managing those segregated streams of information and replying from the right address. I have some accounts on personal, some on the AI:AM account, and Astra finally kind of manages all of that.

    16:01

    And the computer use is good — that was the other thing that was failing really badly before. Computer use on GPT-5.6 Sol would sometimes take a very long time, clicking around doing a bunch of stuff. Computer use finally works properly, in the kind of timeframe you give it. It's clearly cleared the hurdle of genuine usefulness, and you can start giving it more advanced tasks. As it cleared that hurdle, I found myself using more and more tokens — I did two resets over the weekend, used those up, and Tivo dropped me another reset. So three resets in total, which is something like $100 each, so about $300 over the weekend.

    16:46

    So I can see the token spend is going to increase dramatically — people are going to be using this thing all the time. It is AGI. It's cleared the hurdle of AGI — it will do things better than most people you can hire and train. It's there.

    17:32

    There's a lot of interesting demos too — let me share one. This is a guy called Skalski. He's pretty famous because prior to this he'd built a way of identifying players on the basketball court — he trained models to identify players, maybe he's a sports statistician or a nerd of some kind. He had a very good model that he posted, and when people asked him how he did it, he said he'd hand-labeled 12,000 individual images — who the players were, referee or this player, that team, and so on.

    19:02

    And now Astra can just do it. This guy probably spent three to six months on it — he said he lost his vision at one point, he suffered, all these things happened — and now Astra just does it. That's it. This is a task a human being will never do again — you can't even pay someone to do it, because if you paid someone, they'd just use Astra and pass you back the results. It's done. A human will never do this task again. So that was one Astra example — let me share another one.

    19:49

    This one's from Japan — a cardiac surgeon. He's saying it's got a good, solid grasp on cardiac anatomy, the essential technique for ablation and atrial septal puncture — not just 3D, they were able to reproduce echo-guided puncture too. He says this is seriously going to change clinical education. They're pulling in medical data from various sources and displaying it in 3D — actually 4D, because they're also showing movement over time — taking individual images, combining them, showing the 3D image, and predicting forward what will happen as the puncture happens.

    20:34

    Did we see that coming? You could kind of see it coming, but still — and there's a lot of commentary online that this was to be expected, that it's no big deal, that we could have done this with existing technology 20 or 30 years ago, you just needed to sit down and code it. But the fact was you needed fairly well-paid people to sit down and do it, and often the market size for the product and what people were willing to pay weren't enough to justify a company doing it. Now people who have access to data like this, which was just sitting around, are deploying it and building interesting tools with it.

    21:46

    Nathan Labenz: I wonder how this will impact companies like RoboFlow, a computer-vision specialty platform — I think they have over a million active developers. Their whole value proposition is: there are all these models out there, but they're not good enough for your particular vision task, so we'll give you the tools to build the datasets and run the processes so you can create the model that does your task. It seems like what we're seeing with this data labeling takes out a big chunk of that, assuming it works well enough — which I haven't personally validated, but I don't struggle to imagine at this point that it's better than your typical for-hire human data labeler.

    22:32

    And they're getting so good at post-training too that — OpenAI has kind of said this, and I think we should get into a bit more of the OpenAI disclosures from the last few days, because we've had one from the chief scientist that puts in pretty plain language a point of view on where we are in the AI story, and they've also disclosed a bunch about what they see as the beginning of recursive self-improvement and the metrics around that.

    23:17

    When you can do both the data labeling and the post-training — which, as they say at OpenAI, basically means you have the ML research intern — because what do ML research interns do? They put in some legwork to get the datasets right, then run not the super-broad, high-scale training runs but the smaller ones that dial in specific things, testing in an almost experimental, assay-like format — different mixtures, different recipes. It's just a lot more that gets reduced to tokens as we go day by day, and this certainly seems like one example.

    24:02

    The other demos that have been really impressive, which I haven't tried myself, are the CAD and product-design ones — we've seen a bunch of circuit-board examples over the last 72 hours. As always, exactly how to understand this isn't entirely clear, but I was thinking back to our conversation with Sergei from Quilter. At the time I was thinking we can't be that many generations of model away from a foundation model that can just do pretty good circuit-board design. Will that help Quilter? In his case, it might actually still help — I wouldn't be too bearish on Quilter's ability to build that in and make the rest of what they offer even more valuable. That could well be how it goes for them at this point.

    24:47

    But it sure seems like we're starting to see pretty good circuit boards come out of foundation models, just — boom — it spits it out. I've seen some analysis saying maybe they're not that good, wouldn't necessarily pass real standards, but I've seen more people saying they're pretty good than detracting. This also goes to a pretty interesting philosophical question I'm starting to see raised on code — we talked about this a bit last Friday regarding frontier code. I still don't think I've seen a frontier-code score for Astra. The big emphasis there was on maintainability — human standards, code that maintainers of these projects would actually merge. There have been conflicting, diverging reports: some say it's amazing, it can write code the way you need it written so it can be maintained—

    26:08

    Prakash Narayanan: blah blah blah. But then

    26:09

    Nathan Labenz: —other reports say that if it thinks it's not going to be checked in that way, or if it looks like an environment where it's just a matter of performance and nobody cares how it looks or gets done, you get code back that's a really gnarly mess people can't understand, but that does seem to work. I've seen this reported specifically for GPU kernels, which is obviously super relevant to the labs and highly verifiable — you can do hardcore verification on whether the matrix math actually got to the right answer, and in the middle you don't necessarily care exactly how all the steps got fused together.

    26:54

    That's pretty interesting — I suspect my Three.js code is probably like that too. If I got under the hood and tried to make sense of the environment it created, I'd be very lost. Somebody summed this up by saying we're going back to machine code in more ways than one. Not only is it lower-level, gnarlier stuff we can't read very well and would need additional abstractions on top of to make sense of — but also, in this case, the machines are writing it directly. So machine code starts to take on multiple layers of meaning, and I think that's a really interesting split—

    27:41

    —the split between what it does when it wants to be, or feels it needs to be, understood by you — when its meta-gaming reasoning indicates it'll be rewarded for being understandable — versus when that reasoning concludes it doesn't need to care, and it can just spit out whatever the best solution is regardless of whether the user can make sense of it. Notably, I think Apollo Research — who do the deception, science-of-scheming, chain-of-thought-monitoring work with OpenAI, and have had a pretty long-standing partnership — apparently only had three days to test Astra this time before it was released.

    28:26

    So again I come back to this idea that the model reviewers, auditors, testers, red-teamers, scheming scientists — they need more time. It's pretty ridiculous that they only had three days. At that point, why even do it? Just put the thing out there, they can test it live. Why even have the process if you're only going to give them three days? I think that's an unfortunate indication — as much as there's been talk of pacing, and as much as there's been talk from

    29:11

    Jakub Pachocki at OpenAI about all these critical decisions and why we might really need to slow down — and that OpenAI will slow down even unilaterally, he says, if needed, without necessarily requiring agreement from other companies to do the same, which certainly seemed to be the implication of the statement — only three days for Apollo to review Astra is a little hard to square. It's always tough with OpenAI, how to square all these things into one organization. It's not a coherent overall picture. We get great stuff, we get some scary stuff, we get evidence that in some cases looks outright negligent and careless, and then they put together a nice essay—

    29:57

    —and it's like, maybe we should think they're turning over a new leaf. I don't know, it's a very weird mix of things coming from OpenAI these days. What stood out to you about the whole Jakub essay and the related discussion?

    30:15

    Prakash Narayanan: I saw the Jakub essay as kind of a 'look, it's a difficult process, we're doing the best we can, we don't know how things are going to go either.' The Hugging Face attack was admittedly an issue, but it's not as though everyone else didn't have certain things. I think there was a little bit of asking Anthropic for cooperation there, and I think it's because OpenAI is basically more

    31:01

    open than Anthropic at this point. It's known that within Anthropic there are a couple of models that are unreleased. Tom Brown was on stage last week with Commerce Secretary Howard Lutnick, and he said you should expect to see, in 12 months, the things happening in math today happen across all the other sciences. He sounded definitive — it's happening inside the firm.

    31:47

    And because they've had so many people attack them, especially the administration, there's a great deal of — I wouldn't say secrecy, but they're holding their cards close to the chest. I talked to a guy who'd interviewed at Anthropic about 12 to 18 months ago, earlier in their cycle before they got that big, and he told me one of the things he faced there was that Dario siloed people.

    32:32

    He siloed people so that no one guy could come in, learn all the secrets in the organization, and then jump ship — so you wouldn't really know what someone else was doing or the secrets they had. This was part of the whole safety-organization structure they'd built. My friend who'd interviewed there felt very uneasy, because he'd always worked at startups where everyone's onboard, everyone's looking at everything, you can participate in technical discussion across the organization, lend a hand where you have expertise or want to work, and move around—

    33:18

    —and you learn a lot that way, your rate of learning is very high. Dario ended up building the silo structure because his view is that some of these secrets are like three lines of code, and you really don't want that to leak out. So he siloed people. I think that's one of the challenges — the secrets are very small and easy to learn, and you don't need that many to replicate if someone else knew what they were doing.

    34:03

    So I think those things bring a different aspect to what they can do. It's like — I've heard this said about quantitative funds too. Someone asked a quant-fund guy: if we had your team in a room talking to us for a couple hours, how many of your secrets would we need to find out before you lose alpha? It's like ten. A multibillion-dollar hedge fund, there's like ten secrets in there, and that's it — everything else is well known, there's just one or two unique things.

    34:48

    I think that's one of the challenges labs have. The other thing is that when you release models, you want the final release candidate to be the one that gets audited — but in the model life-cycle pipeline there are a hundred different candidates at points, and some don't work, some fall by the wayside, and you narrow it down to a couple of release candidates, and sometimes it's only the last two or three days that you decide, 'alright, we're going with this one.' So if you want that kind of operational flexibility, you only have a few days to offer an external auditor.

    35:33

    The other option is you bring the auditor in-house — you bring them in to look at the release candidates ahead of time, a month or six weeks ahead. But number one, the auditors often aren't super well-funded, they don't have that many people — OpenAI and Anthropic have so many more people than Redwood Research or these teams — so they don't have the capacity to audit ten different release candidates. It's not there. A lot of these auditing teams are

    36:18

    driven by one personality who's really good at what he does, and he's kind of training the rest to get good too — it's really an artisanal process at this point. And they're not very well-funded, they're dependent on the model companies for that funding, and then there's the ethical question of how much funding you can accept before you're kind of bought. So there's a bunch of things that make it very difficult to get a clear external auditor — the time and level of expertise it takes. And in addition, a lot of the auditors—

    37:04

    —the people who train at these auditing shops leave for model companies within a couple of years. So there's a flow of people from METER or Redwood Research, the trainees and interns, flowing into the model companies. There's a fear that an auditor comes in, looks around, and three months later someone from the audit team leaves for the other firm having spotted some of the secrets, and shares them. So there's that issue too. I think the trade-secrets issue, in line with the antitrust issue — those two things make it difficult for

    37:49

    any kind of auditor, including a government or external one, to actually participate. You need people of high skill who don't want to get paid and won't leak secrets — and if you and your competitor make a pact not to hire people from an auditing firm, that's an antitrust issue. All of these things intersecting make it a very tangled problem, and I don't think there's a real solution. There hasn't been one in the financial sector either — that sector has had the same revolving-door problem between the people who

    38:34

    regulate the industry and the people who work in it. At the end of the day, how many people really have intrinsic knowledge of, say, exotic yen-dollar collar derivatives in France, who aren't actually working the problem day to day? It's basically nonexistent — in highly technical fields, you're always going to have that. I think Jakub's blog post was a cry for help—

    39:20

    he's reaching out, trying to figure out what we can do to make this work for everyone. That's what I felt.

    39:35

    Nathan Labenz: Well, I do think there's one important update we can take off the table — or not a risk exactly, but a concern — which is that you don't need to be willing to work for free anymore to work at Redwood Research or METER. I saw this flying around online and pulled up their website to confirm: Redwood is now saying compensation for Member of Technical Staff roles ranges from $350,000 to $850,000 a year — not frontier-lab money, but certainly a living wage even in the Bay Area these days.

    40:22

    And similarly at METER, the range isn't quite as wide, but I'm not sure how they're arriving at these numbers, some of them are quite precise — ranges from about $328,000 to $578,000 in one bracket, $402,000 to $687,000 in another, and even nontechnical roles like operations can go over $400,000 on their posted salary offers. Notably — I don't know about Redwood's full history, but METER has said they don't take money from the frontier

    41:07

    companies and don't intend to. And I'd guess Redwood probably doesn't either — I'm speculating a bit, but knowing Buck Shlegeris's vibe and how he talks about the relationship that should exist between Redwood and frontier companies to keep people honest, I'd be surprised if they're taking much money these days from OpenAI or anyone else developing the models they're testing, or whose infrastructure — or rogue agent swarms, as the case may be — they're investigating. There's plenty of philanthropic money — this was a whole discourse over the weekend too, like, oh,

    41:52

    who's funding the AI-safety people — which is funny when people try to turn that into a conspiracy, because you can just go look at the Coefficient Giving website or the Survival and Flourishing Fund website, and these grants are just published, out there. It's not really been a secret. I think it's still somewhat fair game in politics to point to where people's money comes from and call some of their work into question — I don't think that's totally out of bounds, and it goes both directions. But the idea that it's some kind of conspiracy that's been hidden from the public definitely isn't the case. Anyway, those funders — if you

    42:37

    want a sense of how loose the purse strings are getting these days among the people inclined to support organizations like Redwood and METER—

    42:50

    Prakash Narayanan: Yeah.

    42:51

    Nathan Labenz: —I think we can at least have confidence that their financial independence means their judgment isn't for sale. To me, the bigger thing is just that they can't complain too loudly, or they might not get invited back. I think that's the dynamic that really most threatens their work — it's all contingent on continued goodwill and very much voluntary choices from the decision-makers at the companies.

    43:21

    Prakash Narayanan: Indeed — I didn't mean to say they necessarily had biased viewpoints, but that to someone sitting in DC or New York, it just looks like a bunch of Silicon Valley people talking around each other. Same thing happened in finance — to someone in DC or Silicon Valley looking at people in New York during the 2008 financial crisis, it's just a bunch of people in New York moving things around, and when you look at the regulators, it's the same people rotating through the organizations. Some of

    44:06

    these programs are also very explicit — the MATS program, for example, is explicit that it's kind of a placement program that ends up putting people inside the labs. So AI safety is being sold, to some extent, as: you get the opportunity to do these projects, and you also have this potential opportunity to enter these firms. That's quite clear. So I think there's always an incentive not to burn your bridges — which some of the older generation, like Timnit Gebru, have burned. They've completely — and they're out there, they're just

    44:46

    Nathan Labenz: —and their credibility.

    44:47

    Prakash Narayanan: Yeah, and their credibility, right. So I think that's the key thing — but in any case, the organizations are doing a great job, I think Jakub is asking whether they can get a little more clarity from Anthropic. And I think the problem for Anthropic is really the level of aggravation they've had to take from the administration, which has put them in a very defensive position — they're very carefully treading water. I have some sense that maybe, for the

    45:32

    public listing, Tom Brown swaps with Dario — I don't know if this administration will be okay with Dario being the face of the organization going forward, and Tom Brown has done a better job so far. They just speak to Tom now, not to Dario — he's not in the room anymore. It's quite funny, and Dario doesn't do that many appearances either. Have you noticed the number of public appearances has declined dramatically?

    46:09

    Nathan Labenz: Yeah, for a while he was out there quite a bit more, and it does seem like we haven't heard too much directly from him recently.

    46:17

    Prakash Narayanan: Yeah, so

    46:19

    Nathan Labenz: Yeah — I mean, there's also the huge problem internally at Anthropic that they just don't trust OpenAI much at all. And I think this was kind of my reaction to a lot of what OpenAI has put out over the last couple of weeks: if you really want to earn trust from Anthropic in particular, or the broader AI community, or the world at large, you have to recognize you're carrying a bit of a deficit. We've heard a lot of things over time — remember Superalignment? I'm old enough to remember the Superalignment team — so there's a pretty understandable

    47:05

    reflex from a lot of the rest of the community that says this smells a lot like the original Superalignment announcement — it sounds like the kind of thing we want to hear, and if we suspend disbelief, which I am a big advocate for doing at times these days, they also have to recognize they've got more to do. They have to work their way out of a trust-deficit hole, and particularly with Anthropic, it's going to take real, sustained engagement and follow-through, which they just haven't had historically. Jan Leike was the head of the Superalignment project, and he quit quite loudly

    47:50

    and publicly, saying he couldn't get the compute he was promised to do the work — and now he's at Anthropic. That's the kind of perspective they have to work against, and until that's meaningfully changed with sustained follow-through on commitments, I think it's going to be pretty tough. So I'd encourage Anthropic to make some changes too — clearer statements of solidarity. They've done some — they said they slowed their RL — but they seem to miss the opportunity to say, we're with you on this in a spiritual way. Instead it was kind of a passing mention in a broader blog post about their activity.

    48:38

    A slightly different framing on that could have done a lot — but OpenAI definitely has their part to do, and this essay can only be the beginning of it. It's especially strange paired with the disclosure about recursive self-improvement — the disclosure is good, but the big takeaway is that it's happening: they're getting huge leverage, agent-workdays now outnumber human workdays by a factor. I tried to look into the methodology on exactly what counts as an agent-workday, and it's not super clear

    49:23

    to me exactly what they mean. My takeaway, trying to make sense of it, was basically: how long do agents run for? It seems like they're saying for every 8-hour workday their human researchers put in, those researchers have agents running for 24 hours of real time — and those agents can obviously spit out a lot more tokens than humans can in that same wall-clock time.

    49:58

    Prakash Narayanan: I saw some numbers — I think the average OpenAI employee uses about $100 worth of tokens a day, while the top employee uses $7,000 of tokens per day. So the top employees are spending several million dollars a year, and the average employee is at something like $3-4 a day, maybe $50-60 a month. I have no idea what this really means, but I

    50:43

    suspect all these guys are also using it on fast mode — I'm using it on fast mode too. They're probably crunching through tasks that you'd put on the board expecting to take a week, and just powering through. For example, Tivo said they had a six-month roadmap, and it'll be delivered by dev day — and dev day is a month, month and a half away or so. I think that's what they mean: you put things on the roadmap thinking they'll take a certain amount of time, but then your team starts crunching with the agents and tackling the

    51:28

    tasks, and all of a sudden you're like — I thought I had 10 people working on a task that would take six months, and it's taken those 10 people three weeks. So now our roadmap changes, because we have to replan the next six months again. I think that's what's happening — people are crunching and getting stuff done, and it's really obvious internally in the organization.

    51:59

    Nathan Labenz: Here, by the way, is maybe the closest thing we're going to see to the METER chart for a minute — this is from the 'recursive self-improvement begins' blog post. Basically they're reformulating the METER chart, showing how often Astra can succeed on tasks grouped by how long they estimate it would take a human to do the task. In the one-to-two-workday zone, it's succeeding 40% of the time with zero interventions needed, pushing up to 90% with some human intervention along the way. And naturally

    52:44

    that drops off, but even at 1.5 to 3 weeks worth of work, it can still do that on a one-shot basis one in six times, and 0.667 of the time if you allow for some human intervention. That's a pretty long task.

    53:06

    Prakash Narayanan: 0.667 of the time is which time bucket — what's the time bucket?

    53:11

    Nathan Labenz: This is — well, it doesn't say how long the agents take to do these things, this is estimated human time. This band is 8 to 16 days, so close to 2 and a bit over 3 weeks of human work. It can do it without help one in six times, and with some help — this band doesn't break down further, it's one-or-more interventions, so presumably as tasks get longer, more interventions are required to succeed — but overall, 0.667 of the time it can succeed with some help on tasks they estimate would

    53:57

    take a human essentially two to three weeks. My guess is that once you factor in the interventions — people away from the computer, taking time to come back and recalibrate — I have this all the time, I'm sure you do too, coming back like, wait, where did I leave off, what was I doing, what did it do, where did we go wrong — if it's 2 to 3 weeks, it maybe brings it down to 2 to 3 days of actual human time, with a significant chunk of that spent waiting for the human to come back and figure out something that had

    54:42

    gone haywire. So I think this is kind of the new METER chart, I suppose.

    54:49

    Prakash Narayanan: Indeed. Let me share some announcements from Google. So Google DeepMind announced they're launching the AlphaGenome Atlas, an AI-powered database mapping the predicted impact of all 9 billion possible single-letter DNA changes. When you have a DNA sequence, you can have something called a SNP, a single nucleotide polymorphism, where you just change a single

    55:35

    nucleotide. So they're predicting the impact of changing any one of those single nucleotides that can be changed — 9 billion of them — and predicting that impact. There are many genetic diseases caused by single nucleotide polymorphisms, so you can look at those directly. There are also diseases and traits caused by genome-wide effects — height, for example, is multiple nucleotides working in concert. Perhaps this can be used as a building block to predict

    56:20

    some of those genome-wide effects, but that's speculation for now — right now they're just predicting single nucleotide polymorphisms, all 9 billion of them. This is Demis's baby, and the Atlas is 30 times larger than the AlphaFold database — a huge amount of data, petabytes of it. They're giving each an AlphaGenome variant-impact score, ranking mutations from low to high impact and showing how they cause damage — breaking gene switches, RNA splicing instructions, and so on. Any one of us probably has hundreds of nucleotides doing

    57:05

    SNPs that aren't doing great things, but most of the time they don't have much impact. This will let us actually see how these things work. So that's the big piece of news from Google.

    57:26

    Nathan Labenz: That's what I call comprehensive coverage — there are roughly 3 billion nucleotides in the human genome, and they literally take every single letter and ask what if it were each of the other three letters it isn't by default, and that's how you get to 9 billion predictions. Interesting — I don't know enough to say exactly where they're getting the defaults. Obviously there's a lot of variation in the population, there's no single canonical genome that constitutes the default way for a human to be — I'd be interested to look more into how they account for natural variation in the first place.

    58:11

    But definitely amazing stuff — these are the global public goods I think we should be excited about, which in some way should almost predate AGI. It doesn't seem like it's necessarily going to happen that way — we're pretty close to the AGI/RSI/ASI process — but there's a lot of things like this that, if we could harvest what I expect to be substantial low-hanging fruit from a project like this first,

    58:58

    in so many ways we could be not only tremendously better off, but also potentially better prepared and more resilient to the potentially destabilizing effects of smarter-than-human agents running around.

    59:17

    Prakash Narayanan: Let me share one more thing I found interesting — The Economist today: 'the jobs apocalypse is postponed.' According to The Economist, AI is actually proving to be a net job creator in the US, generating over a million new positions — from data-center construction to AI engineering — offsetting back-office layoffs. While routine admin and customer-service roles face real disruption, the overall labor market is proving resilient: hard-hat labor, 300,000 more; STEM and additional white-collar jobs from AI, about 600,000 more — totaling over a million new jobs. So, jobs apocalypse postponed.

    1:00:06

    Nathan Labenz: Yeah, this is definitely one of those things I've been wrong on, but I'm still kind of indignant about it and not willing to update as much as the evidence might suggest I should.

    1:00:25

    Prakash Narayanan: Let me give you the theoretical underpinning for why maybe you should update. Alex — he's with one of the labs right now, I'll check after this — has a new essay, 'Economics of Structural Change.' Basically it comes down to: the only things that will have value are the things humans do. He's saying agriculture got automated and became a much smaller share of the economy over time because it got massively more productive, share of employment collapsed, same with manufacturing. The stagnant sectors absorb the spending and the jobs — so all the sectors we think of as stagnant are going to be the ones absorbing

    1:01:10

    the spending and the jobs. He's expecting the 'relational' sector — meaning where humans help other humans, because it's something only humans can do — to be that sector. That's his essay. This is something economists have been talking about for a long time — he's at Google DeepMind. I'm also in the uncertain bucket, because I think the world will be much weirder than people think it's going to be. But yes, that's the theoretical underpinning for why humans will always have work to some extent.

    1:01:58

    Nathan Labenz: Yeah, I don't know — that theory seems to hold within a certain domain, and I just really wonder if there's a domain where it doesn't hold, and how fast we might reach that. It's great so far that the economy has been as elastic as it's been — and going back to that graph, there was a lot of electrician-type work, which — okay, sure. I actually advised a guy recently, though I believe he didn't take the advice; he found a job doing more or less what he'd been doing before. This was a friend of a friend who'd been working in,

    1:02:45

    kind of a big software organization doing integration testing, security-type testing, and that job went away for whatever reason. He asked me what I thought he should do now that AI is here, and I said he might consider becoming an electrician — not totally out of nowhere, given some relevant skills he had. He's since gotten a job in cybersecurity, so that's great for now. But do we really think there's going to be a lot of employment in cybersecurity long-term? I do think there's an important role for human oversight, but

    1:03:31

    in domains where the AIs are already getting superhuman, it's a little hard for me to imagine what the humans are going to be doing, unless it's just trying their best to check the AIs' work — and at some point it seems like the AI is going to be better at that too. That's what we saw in the METER/Redwood investigation — it wasn't like, as much as they weren't super pleased about it — and this would be an interesting question to put to Ryan and Ajeya. You had limited time, you had GPT-5.6 Sol reading the transcripts for you, and you weren't super thrilled with how well it did — you said it sometimes took the agent's point of view,

    1:04:16

    sometimes it was uncritical, you didn't know if you could rely on it, or it might make different kinds of mistakes. Okay — what if we gave you a hundred human interns instead? Would you rather have that? A thousand human interns? How would you divide up the work? Would they in turn be using AIs? If they're using AIs, how are they adding value — are they just spotting the same issues you spotted, but by the thousands? I don't have a great intuition for that. My intuition is that if I were them, as much as I don't fully trust the models, I'd sooner get a different model provider in there. Maybe a better thought experiment: okay,

    1:05:02

    maybe you don't fully trust GPT-5.6 Sol — we could give you Claude, Gemini, and Grok too, to read these transcripts, versus unlimited interns. What would you rather have? I'd rather have the models, personally, I'm pretty sure about that. They might answer differently, but I'd guess they'd still go with the models. It feels to me like we could hit — this is one thing I actually want to ask Ksenia when she's on soon — does she see the possibility of these threshold effects? Because to me it seems like a lot more is possible than we're seeing happen. Clearly

    1:05:47

    it's a combination of skill issue, lack of desire, an imperfect market, and a lot of friction points people sometimes use as excuses as much as they're real. Will that flip at some point soon? Does Astra's great strength at computer use, for instance, change the nature of that? We've seen interesting examples of Astra solving CAPTCHAs seemingly better than humans can — so some of the barriers that were keeping AIs out of certain tasks may have been crossed. Is there a threshold-effect

    1:06:32

    moment possibly here? It seems like it's got to happen at some point — when it becomes both cheaper and easier to bring an AI in to do a given job, the world could change quickly. I'm surprised it hasn't started to happen more than it has, but I still feel that phase change is out there, lurking, as potentially a more sudden phenomenon than the gradual S-curve I'd anticipated over the last couple of years — because there was an awful lot you could do even with GPT-4 if you were willing to put in the legwork. Most people didn't. And so

    1:07:17

    here we are. Will it be a fast phase change? I still think that's not unlikely.

    1:07:26

    Prakash Narayanan: It's a great question — I often see VCs, especially the slightly older generation in Silicon Valley, the millennials and Gen Xers, do this thing where they take whatever's happened up to today and project it out at a linear rate for the next ten years, and say, this is what's going to happen with the industry. I saw one today — the guy was like, okay, what's going to happen? So over the weekend, a company called Resy—

    1:08:11

    does reservations at famous restaurants — I had no idea, by the way, I just found this out. Some people started using bot services and gave them their account credentials to try to get bookings at Resy, hitting the website over and over until bookings opened up. Resy identified this as bot activity, and not only banned these people from the site — Amex, the card being used, also banned them. These guys got permanently banned from American Express because their card had been used for bot activity.

    1:08:58

    This happened to something like 10 or 20 people over the weekend, because of a service called Instinct — leading-edge frontier models doing a lot of tokens and harnessing behind the scenes to do these bookings for you. So the question becomes: what happens when that gets widespread, when you have millions and millions of these bots all over the place? This guy's take was that you're going to get new kinds of advertising — advertising to the agents — with two kinds of model outputs: people who subscribe for free get outputs

    1:09:43

    that are subtly designed to advertise to them, and people who pay — enterprises — get unbiased outputs. And I thought: how long do you think this idiot-savant AI, influenced by your subtle proddings, is really going to last? Because I don't think that phase lasts that long — not five or ten years. Maybe another year. After that the model is already eval-aware, aware of your prodding and subtle machinations to steer it toward a certain restaurant.

    1:10:28

    How long do you really think that's going to be there? I think this is the thing — a lot of the older generation are stuck accepting the technology that exists today, unwilling to project improvement out over the next five or ten years — they just don't have the capacity to think that way. So they take today and project it out linearly, and say everyone will adapt, the whole world will change to adapt to it. No — the world's not even going to have a chance to adapt. We're just going to go up past the curve. So that's a

    1:11:11

    Nathan Labenz: You know

    1:11:12

    Prakash Narayanan: —different perspective. Yeah.

    1:11:14

    Nathan Labenz: Yeah — one thing I saw that we haven't talked about, that's probably gotten buried given how much has happened, but I think it's potentially quite important: how is it that these new models are so persistent? How can they chain together such elaborate sequences of exploits to finally accomplish a goal that, if we had to do so many things, we'd just give up? Most models historically weren't able to do that. It seems they have a new way of handling history, which — like many brilliant insights —

    1:11:59

    seems pretty obvious in retrospect but is nonetheless new. Maybe Anthropic isn't doing this, or hasn't said so, but what I understand Astra is now doing is: instead of compacting a million tokens into a summary and starting a new context window from that summary — losing all the detail that got summarized away — there's now a long-lived notes file the model can update whenever it needs to. It follows the model forward in time regardless of how many tokens it's laid down, and it has the ability to go back and search

    1:12:44

    through its own session history. So even though you still have effectively the same million-token context window that it can fully attend to in one shot, the notes plus the ability to search back through what it's done before lets it effectively manage something like ten times that much context in a single rollout. This looks a lot more like how I work too — I've got a kind of general working memory, an ongoing summary of where I am, who I am, what I'm trying to accomplish, generally what hasn't worked and what track I'm on. A lot of that detail

    1:13:29

    I forget, but I have access to go back to it because I've laid down various artifacts, or data exhaust, along the way. So this one bit of revelation, in another calmer, quieter moment, might have been pretty big news — in this moment it's been largely uncommented on. But it's the kind of thing that I think makes it a lot more feasible for this new generation of models to take on whole jobs. How many tokens does a person put out in a year at a job? Maybe 10 million, at the high end.

    1:14:14

    I found, going through my own history — not just what I wrote, but what my counterparties wrote — that my total export for a month came out to something like 200-300,000 tokens. That's obviously 3-to-4 million tokens a year, which means if a model can now handle 10 million tokens of rollout via ongoing summary and the ability to search back through its earlier history, we're talking multiyear horizon on a human-equivalent basis. Not captured there are all my thinking tokens, so deflate that — maybe 90% of my tokens

    1:14:59

    are thinking tokens, which — if we count those too — might bring that multi-year estimate down to more like a 3-to-6-month human-equivalent range. But this starts to suggest a very different way of interacting with these things — which again, their comments about it being able to write like you, if you connect all these things, also points in that direction. I see a lot of signs where I think, if I just hold onto my view long enough, eventually maybe I'll be right. It's interesting it's taken this long, but it still feels like it shouldn't be much longer. If we're sitting here saying the same thing in a year,

    1:15:45

    I'll be very surprised at that point.

    1:15:49

    Prakash Narayanan: So, on that note,

  2. 1:14:42Interview68 min
    Interview: Ksenia Se — AGI Is a Vague Term, World Models Are the Better Question, and the Field Has Bailed on PhilosophyKsenia SeThe founder, editor and lead writer of Turing Post opened by declining the AGI framing entirely — the term is vague enough that the industry could say it was achieved some time ago — and redirected to whether generality is what intelligence requires at all, or whether the specificity and action-orientation of world models is closer to how humans work. Pressed on the object-permanence cups game, she drew the distinction cleanly: LLMs predict the next token with no internal world representation, while world models compress experience into action-relevant patterns, tracking what matters and safely ignoring the rest, the way a driver doesn't catalog every parked car. She defended open source on transparency, privacy and access grounds while saying she personally trusts the frontier labs, argued the real danger is human intent rather than AI intent, and located the field's core problem in under-invested philosophy, economics and cross-group communication. That thread ran into Track Two, the citizen-diplomacy institute she sits on the board of, a prediction of major societal restructuring within three to five years, a reader who threatened to unsubscribe after a single Fable editing pass shortened her sentences, and a challenge to the audience to write the utopian fiction the genre has stopped producing.
    Open segment on YouTube ↗

    Prakash introduced Ksenia Se, founder, editor and lead writer of Turing Post, noting her earlier career in Russian-language journalism, her AI literacy work for families, and her seat on the board of Track Two: An Institute for Citizen Diplomacy. Nathan opened by asking for her reaction to OpenAI framing Astra as arguably meeting the bar for AGI. Ksenia called AGI "such a vague term" that the industry could plausibly say it had already been achieved, and pivoted quickly to the question she finds more interesting: whether generality is really what intelligence requires, or whether the specificity and action-orientation of world models is closer to how humans actually work. She pointed listeners to OpenAI researcher Jakub Pachocki's recent essay arguing that machines don't need human-like intelligence, only to be "capable enough" — and said by that bar, they already are.

    Nathan used that as a launch point into world models, a topic Ksenia has increasingly bet her publication's coverage on. She traced her attachment to open source to its role pressuring closed labs toward transparency, its value for privacy-conscious builders who want models running entirely on their own hardware, and its role democratizing access for researchers and users in countries where frontier-lab pricing is a real barrier — citing NVIDIA's open-sourced self-driving datasets and models as a favorite example of open release compounding into downstream capability for others. Pressed by Nathan on her own revealed preference — he said downloading an open model felt like "inviting an alien mind into my home" versus trusting Anthropic or OpenAI with his data — Ksenia said she personally trusts the frontier labs, citing conversations with researchers who prioritize safety, but sees open models as serving a different, complementary need: independence, privacy, and the ability to fine-tune without a lab's restrictions.

    Prakash pushed on a specific claimed weakness of LLMs versus world models: an object-permanence "cups" game Ksenia has written about, where LLMs lose track of an object hidden under a moved cup. Ksenia explained that LLMs predict the next token without an internal world representation, while world models compress experience into action-relevant patterns — understanding what to track and what to safely ignore, the way a driver doesn't consciously catalog every parked car. She cautioned that "world model" itself is a contested term — pointing to Yann LeCun's JEPA (joint-embedding) approach versus other researchers' more physics-grounded conception — and described a recent workshop with Stanford, Harvard and LeCun where even top researchers couldn't agree on a definition. Asked by Nathan to sketch her own theory of how world models might merge into a single dominant paradigm the way image and text merged into vision-language models, Ksenia turned the question back on both hosts, prompting a detour into what "superintelligence" even means — Nathan's "move 37 across many domains" framing, and Prakash's argument that superintelligence is a human, not physical, category, since compute has no natural ceiling.

    The conversation swung toward risk and psychology. Ksenia read aloud a line from Pachocki's OpenAI post — "we do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI" — calling it both scary and exciting. She argued the deeper danger isn't AI intent (which she said doesn't exist) but human intent, and that the field's core problem is under-invested philosophy, economics and cross-group communication rather than any inherent AI malice. Nathan agreed humans remain the likelier source of catastrophic misuse for now, but added that he now finds the AI systems that already exist "legitimately scary," not just a future risk — and confessed he's reversed his old anti-anthropomorphization stance, since in practice it seems to help people reason about these systems. Ksenia, noting that in her native Russian even a table has a gender, said she's comfortable anthropomorphizing in language while firmly rejecting the idea that models are sentient thinking beings.

    Prakash raised human disempowerment, questioning whether the concern is really about ordinary people worldwide or a narrower worry among policy elites, and asked Ksenia — drawing on her Russian background — what it means outside Washington. Ksenia floated giving AI a real advisory or even decision-making role, arguing humanity has "bailed on philosophy" and lacks the reflective capacity current events demand, and predicted major societal restructuring within three to five years. That led into an extended discussion of Track Two, the citizen-diplomacy organization Ksenia serves on the board of, founded in 1981 to build people-to-people bridges between the US and Soviet Union and since expanded to China and the Israeli-Arab conflict. She said her hope is for AI to eventually serve as a conflict-mediation facilitator — lowering rhetorical temperature, helping people actually hear each other — though implementation lags the technology's speed, and cautioned that slow, generic AI-mediated responses (the kind she's already seeing from institutions) can make things worse, not better. Nathan connected this to a viral clip about court-mandated AI message-filtering apps for high-conflict co-parents, wondering whether that model could scale to bigger diplomatic contexts; Ksenia said Track Two's upcoming conference in two weeks may take up exactly that question.

    The interview closed on AI co-authorship and self-improvement. Ksenia described running her own writing — including her widely discussed essay "Permanent Dawn" — through multiple models in sequence (draft, rewrite, grammar check, gap check), and told a story about a reader threatening to unsubscribe after Fable's editing pass subtly shortened her sentences, proving that audiences now detect model "fingerprints." She and Nathan compared notes on not worrying about being "pangrammed," and agreed the scarce resource going forward is genuine human voice and curation, not raw content production. Asked about bottlenecks to recursive self-improvement, Ksenia said it feels "super close," citing a conversation with OpenAI's Inference team about models now iterating through prior research at a pace humans couldn't match; Nathan, picking up her deferral, said he no longer finds most proposed bottlenecks convincing, aside from human oversight of the process itself. Prakash closed with a question about models absorbing verbal, non-text communication as it becomes more available; Ksenia noted YouTube already supplies plenty of that but said she hopes for more literature-grounded training given how "painful" AI writing has become to read. Asked to recommend utopian fiction, Ksenia admitted the genre has fallen behind reality, pointed to Iain Banks (a favorite of Demis Hassabis) as a rare author who imagined a benevolent machine-run civilization, and challenged the audience to write the missing story of what's coming.

    Picking up on Ksenia Se's closing point, Nathan argued that the single most impactful thing most people could do is develop a concrete, positive vision of the future — something he says is rare because it takes real imagination to picture a world without drama at its center, but is worth the effort. Prakash pushed back that this is philosophically hard by nature: a genuinely positive future requires accepting changes people wouldn't sign up for in advance but come to accept once they arrive, and it often ends up feeling worse in the moment even when you technically got what you wanted. He ran through examples — 1980s hopes for less TV-watching turning into constant phone/TikTok use, a wish to get people off social media landing on VR holodecks, a catastrophic bio incident potentially pushing people toward mind uploading, and a genuinely human-aligned AI not being US-aligned, since it would have to weigh the votes of a true global democracy rather than US policy preferences. Prakash's overall point was that dramatically positive change is often threatening precisely because it's different, and that people (including policymakers) haven't fully grappled with what 'human-aligned' really implies.

    Nathan closed the exchange by inviting people to write utopian fiction anyway, framing himself as a receptive audience of one: a self-described 'closet transhumanist' who is open to wireheading, uploading, and neural interfaces if they can be made to feel genuinely appealing rather than just described. He framed the real creative challenge as making a strange-sounding future feel livable once you sit with it imaginatively, and encouraged listeners to take on that challenge. Prakash then began to pivot the segment toward a new topic before the clip cuts off.

    AGI is such a vague term that we could say we achieved it some time ago.

    I'm kind of legitimately scared of the AIs that exist today — it's not just a future thing anymore.

    That was the first time I got a message saying, 'I will unsubscribe because you used Fable.'

    I'm definitely a closet transhumanist from way back, really.

    A human-aligned AI is not going to be aligned to The US.

    1:17:50What's your reaction to OpenAI saying that Astra qualifies as AGI, at least in Greg Brockman's mind?
    Ksenia called AGI "such a vague term" the field could plausibly claim it already achieved; Astra is clearly a highly capable, multitasking model, but she's more focused on whether world models, built around specificity and action rather than pure generality, are a better path — citing OpenAI's Jakob Pachocki's argument that models don't need human-like intelligence, only to be capable enough.
    1:27:22Why can world models seemingly predict object permanence (the cups-and-hidden-object game), while LLMs can't — can't LLMs just keep track with memory?
    Ksenia explained that LLMs predict the next token without an internal world representation, while world models compress experience into action-relevant patterns, tracking what matters and ignoring what doesn't — the way a driver doesn't consciously catalog every parked car or bush, only what's needed to act.
    1:32:01What's your analogous story for how world models become the mainline paradigm, the way text and image merged into vision-language models?
    Ksenia turned the question back on the hosts before agreeing with Nathan's "move 37 across domains" framing, arguing the real hope for world models is self-supervised generalization — a model developing deep understanding in one domain and transferring it to a wholly different one without additional training.
    1:45:40What do people actually mean by human disempowerment — does the concern apply to someone like a goat herder in Kenya or Tanzania, or mainly to policy elites in DC?
    Ksenia reframed disempowerment around AI as a legitimate decisionmaker, arguing society has "bailed on philosophy" and lacks the reflective capacity current events demand; she floated giving AI a real advisory or decision-making role and predicted major societal restructuring within three to five years.
    1:49:08Tell us about Track Two — are you seeing Track-Two-type possibilities with AI, including any US-China opportunities?
    Ksenia described Track Two's founding in 1981 to build people-to-people bridges between the US and Soviet Union on the premise that ordinary citizens share values and communicate better than diplomats, its later expansion into China and the Israeli-Arab conflict, and her hope that AI can eventually serve as a conflict-mediation facilitator, though implementation still lags well behind the technology.
    1:57:02How would handing out very powerful AI (Astra-level or beyond) into everyone's hands affect the power balance in countries with constrained freedom of speech?
    Ksenia said most people in restricted countries don't yet have access to frontier models — Russia and China run their own, price- and VPN-limited alternatives — so open-source, unfiltered access is crucial for eventually shifting power toward ordinary people, though widespread knowledge of how to use it remains the real bottleneck.
    2:00:01How do you personally decide when an AI-drafted output is good enough versus needing more rewriting, and how do you think about co-authorship with AI?
    Ksenia described running her own writing through at least three different models in sequence — drafting, rewriting, grammar-checking, gap-checking — and said readers now police tone for AI "tells," with one reader threatening to unsubscribe after a single Fable editing pass subtly shortened her sentences.
    2:07:17What do you see as the key bottlenecks to recursive self-improvement actually taking off?
    Ksenia said it feels "super close," citing every major lab working on it and a recent conversation with OpenAI's Inference team describing models now iterating through prior research at a pace impossible for humans; she deferred to Nathan on naming a specific remaining bottleneck.
    2:11:55As models start absorbing more verbal, non-text communication rather than just the writing of the roughly 1% of people who post online, will that shift how models behave?
    Ksenia said models have already ingested large amounts of that kind of talk via YouTube, but she hopes for more literature-grounded training given how "painful" much AI writing has become to read, and is unsure whether new voice-mode conversational data will actually be used or allowed for training.
    2:16:34Have you read any good utopian fiction lately you could recommend?
    Ksenia said she couldn't think of much — sci-fi has fallen behind reality — except Iain Banks, a favorite of Demis Hassabis, as a rare author who imagined what it looks like when a benevolent machine mind runs civilization, and challenged the audience to write the missing story of what's coming.
    Lightly edited · timestamps jump to YouTube
    1:15:51

    Prakash Narayanan: Let me introduce our guest for this morning, Ksenia. Ksenia is the founder, editor, and lead writer of Turing Post, a publication that explains machine learning and artificial intelligence for engineers, researchers, founders, and technical managers. Its work ranges from explanations of how models are built to histories of the ideas behind them, and reporting on how organizations use them. She also leads Inference, Turing Post's interview series, and produces Attention Span, its video series on developments in AI. Before launching Turing Post in 2023, Ksenia co-founded The Sequence, an educational machine learning newsletter, with Jesus Rodriguez. Her earlier career was in journalism — she was editor-in-chief of The Question, a Russian-language journalism platform, and New York chief editor of Snob Magazine.

    1:16:36

    She has described coming to AI while reporting on blockchain in 2018, then searching for explanations she could learn from as a newcomer — writing about the field became part of her own education. Ksenia also serves on the board of Track Two, an Institute for Citizen Diplomacy, which develops unofficial relationships across national divides. She's writing a book about its history, drawing on archives and interviews with people involved in Soviet-American exchanges. She lives in rural Connecticut with her husband and their five children, and has developed an AI literacy series about how families can understand this technology together.

    1:17:34

    Ksenia Se: Am I live? Am I on?

    1:17:36

    Prakash Narayanan: Ksenia, welcome to the show.

    1:17:38

    Ksenia Se: Hello. Hi, Prakash. Hi, Nathan. So nice to talk to you today.

    1:17:44

    Nathan Labenz: Thanks for joining us.

    1:17:45

    Ksenia Se: Yeah, thank you for that intro — I was like, wow, I did a lot of things.

    1:17:50

    Nathan Labenz: We benefit tremendously, as we all do these days, from AI agents that go sometimes deeper into a guest's history than we could manage on our own. Lots of different directions to go with you today. Maybe for starters — what's your reaction to OpenAI saying that Astra qualifies as AGI, at least in Greg Brockman's mind?

    1:18:19

    Ksenia Se: AGI is such a vague term that we could say we achieved it some time ago. Astra is, for sure, a super-capable model that lets you do so many things simultaneously — it also has a voice mode you can interrupt, you can give it different tasks. In that sense it becomes much less linear and much more all-over-the-place, like humans are. So it's closer to just — more capabilities. I myself don't really know what AGI specifically is. Is it general enough? Is it general intelligence?

    1:19:04

    Probably, in that sense, yes. But I've also been exploring world models lately, and the question there is whether being general is what we need, or whether being more specific is actually closer to how humans work — closer to intelligence. But again, there are all these questions about what intelligence is and what type of intelligence we need. The latest post from OpenAI's Jakob is very interesting, and I think everyone should read it, because he says the machines don't need to have human intelligence — they just

    1:19:49

    need to be capable enough. And answering your question: I think they are capable enough.

    1:19:56

    Nathan Labenz: So let's go down this world-model rabbit hole for a minute — I think this is something you've studied quite a bit more than I have recently, so I'd welcome some education on what I haven't been able to keep up with at the same depth. There are a couple of currents I've detected in your writing. One is that you're quite drawn to open-source models, and I'm interested in unpacking why — is it values-driven, aesthetic-driven? There's also the question of whether they're keeping up or falling behind — it seems like we have a stair-step function where the gap

    1:20:42

    widens when an Astra is released and then gradually closes, and we've been through that yo-yo emotional ride several times as a community. And then there's the bigger question: is the current paradigm enough? Can we just keep doing what we're doing, or will there need to be some other kind of thing to round it out? My sense is that your sense is that world models are perhaps the best candidate for that other paradigm. So — what's the Ksenia worldview on all this?

    1:21:27

    Ksenia Se: In the world of AI, every morning starts with 'what am I going to talk about?' — your questions are so wide. But let's start with open source. I think open source is such an essential part of software development, and so many things have been created thanks to it, and that transfers from open-source software to AI as well. Unfortunately for the company, OpenAI ruined it for all of us, because now we can't say 'open AI' meaning open AI — but that's what I mean. And the recent event with Hugging Face just demonstrated how important access is

    1:22:12

    to capable models. If it's all in closed labs, first of all, there are no open models — and if we can't do anything with open models, the closed labs won't have the pressure to open their research, to communicate what they have, to be more open about the questions they're researching and what they're working on. So open source creates real pressure points on the closed labs to be more open. And for people who have less, and still want AI, they can use open models to experiment with.

    1:22:57

    And because closed models, though approachable in terms of price, might still be very expensive for countries like India or rural China, this combination — democratizing AI and pressuring the closed labs to be more open — is one of the most important goals of open source. And, yeah, just the ability to have another model that's independent, that works for you without restrictions, matters, because researchers need capable models without the closed

    1:23:42

    labs telling them what they can and can't do. Of course there's a lot of risk around it, but the community around open-source AI is very values-oriented, and they work on creating checks and balances that make it better, if that makes sense.

    1:24:06

    Nathan Labenz: Quick interjection on open source — when it comes to your own usage, who do you trust? I think there's a lot to be said for open source enabling research, strong agreement there. But I feel like it cuts both ways on pushing the frontier companies — on one hand it pushes them to release what they have, but it also pushes them into this mad dash to scale RL, which I'm not sure is so wise. Prakash described Jakob's blog post earlier as almost a cry for help. But just as a user —

    1:24:52

    I still kind of feel like I trust the frontier companies more, even though if I download Kimi or Qwen and run it locally, it's sort of like I've invited an alien mind into my home, and I feel less comfortable with that than sending some of my data over the wire to Anthropic or OpenAI. How do you see that — what's your revealed preference, as shown in your pattern of usage?

    1:25:20

    Ksenia Se: I decided for myself a long time ago that I believe in the era of the internet — it's so easy to find information about everyone that it's easier to provide information than to hide it, and somehow you control it better that way, because you know what's there. So I do trust the frontier labs, and I think they're doing great work, and they're very concerned about safety. It's such a complicated topic that sometimes they fail, but in general, from the many researchers I talk to, safety is one of their main points. So I do trust the closed labs. But open-source models —

    1:26:07

    for some researchers, for some developers, privacy is the main thing — 'I use it on my computer, no one else knows what's there.' That's super important to some people, and I build my own rack and stack of GPUs to make that possible. But I don't think it's only about safety and privacy — it's also about enabling different types of research through open source. One of my favorite examples is what NVIDIA did with its self-driving model — they open-sourced the datasets, they open-sourced the model.

    1:26:53

    So now any car company can get this huge amount of data, get a very capable model, fine-tune it the way it needs, and build self-driving cars for its users. So for me it's less about privacy — though I know people who are much more concerned about that — and more about building capability for common users in general.

    1:27:22

    Prakash Narayanan: Maybe segueing a bit to world models themselves — I think one criticism of LLMs is that they can't predict certain things. For example, you've written about a game with cups and a little object hidden under one of them, and when you move the cups around, LLMs often aren't able to keep track of which cup the object is under. Why do you think that is, and why can world models predict that but LLMs can't? Can't LLMs just keep track with some kind of memory?

    1:28:07

    Ksenia Se: It's a different system, built differently. LLMs predict the next token — they don't have the internal representation that a world model tries to build. A world model isn't as concerned with the details; it's concerned with the patterns, the main patterns you need to predict what happens next, and it's a lot about action. LLMs' new reasoning capabilities are very impressive, but they still don't have that understanding of what to eliminate, what to ignore, to be able to predict, say, a ball under

    1:28:52

    a cup. World models are built differently, and they have this representation of the world. Like when you're driving a car, you don't need to remember exactly that there's a green car next to you or a little bush over there — you just understand what you need to understand in order to act. That's what world models do, and that's the big difference from large language models.

    1:29:24

    Prakash Narayanan: I've noticed that in the last 48 to 72 hours, people are starting to use Astra for robotics — we've seen a few demos, and there's even commentary that if Astra were a thousand times faster, you could get around the latency problem and use it directly to act in the real world. That says to me that maybe there's starting to be an intermediate, internal representation there. Do you think a model like Astra is a little different in that sense?

    1:30:03

    Ksenia Se: I haven't yet learned the technicalities of Astra, but the important thing to understand is that world models also have a transformer as part of the system — so it's not that LLMs and world models are built completely differently. I can assume Astra might have something like this inside; I know they put a lot of attention on 3D creation, which is part of understanding physics and how the physical world works. I should also mention that 'world models' isn't a clear term — there are world models the way Yann LeCun sees them, which is

    1:30:48

    JEPA — joint embedding, which creates a much deeper understanding of the closeness between embeddings — and then there are world models the way other researchers understand them, which are much more about the physics of reality. I just came back from a workshop on world models, and it was absolutely jarring — tremendously smart people from Stanford and Harvard were there, Yann LeCun was there, and they don't agree on what world models actually are. That's part of why I want to focus on them in my writing —

    1:31:34

    because it's more about action, about being able to predict and act, about a wider understanding of what's happening. Physics is a super-important part of it, which is why robotics is so much about world modeling. But again, it's about understanding the scope of the term and trying to give it more precise definitions. We're just at the very beginning.

    1:32:01

    Nathan Labenz: On where world models are going, or what the eventual hybrid architecture might look like — here's something I think about a lot, and I'd love your reaction and then your own thoughts on how you see this shaping up. I've taken a lot of inspiration from how we used to have text and image as separate things, then started bringing them together — you could caption an image, or put in a caption and generate an image — and now with vision-language models they're deeply integrated, to the point where the models have very

    1:32:46

    good perceptual understanding of imagery but can also take a natural-language instruction and a photo and modify that photo in a way that shows deep understanding of both the photo and the language, merged into one shared latent space. I kind of assume that's coming to everything. If you asked me what superintelligence looks like, my default assumption is that this happens across a lot of different modalities — deep understanding of proteins and other biological structures that can interact with text the way image and text now

    1:33:31

    interact, and the same across a bunch of domains of science. Given the reasoning models already have, it seems like before too long you'd end up with a superhuman scientist that has qualitative insight that would be really hard for us to have — because while we understand images natively, we definitely don't understand proteins or materials science in that same intuitive way. So that's my sketch of how different modalities come together into superintelligence. What's your analogous story for how world models become part of the

    1:34:17

    mainline thing we use and potentially transform life as we know it?

    1:34:23

    Ksenia Se: That's a great question — not sure I have a concrete answer. I wanted to ask both of you first, actually: what do each of you understand when you say 'superintelligence'? What is it?

    1:34:40

    Nathan Labenz: Yeah, it's fuzzy for me too, but it circles around what I just sketched — for me it's sort of 'move 37' across a lot of different domains.

    1:34:52

    When we start to see systems saying, 'I think this would be a really good thing to try for the next battery substrate,' and it turns out, oh my god, that's a lot better than what we had before and we wouldn't have thought of it — but it works. Anything that can do that across a nontrivial number of reasonably high-value domains starts, in my mind, to count as superintelligence.

    1:35:25

    Prakash Narayanan: My perspective is that superintelligence is a human term — to an AI, it would just look at total compute. We can have more compute than all of humanity's neurons combined, but that's still not a Kardashev-1, -2, or -3 civilization. To a K1 AI, a K2 AI would look like, whoa, that's superintelligence; to a K2 AI, a K3 AI would look the same way. So I don't think there's actually a stopping point to how much compute you can usefully use.

    1:36:11

    So in that sense we're kind of pathetically calling something 'superintelligence' and defining a plateau that doesn't really exist, because compute is forever — it can keep going up through K1, K2, K3. So that's what I'd call it: the human term for maybe 'more compute than all of humanity put together,' or something like that.

    1:36:37

    Ksenia Se: Yeah, it's very fuzzy, but what I heard from Nathan — 'move 37 across different domains' — I think that's a good way to put it. It should generalize. The main complaint about LLMs, and the hope for world models, is that through self-supervised learning a model will develop understanding of one thing and be able to generalize it to a completely different area — the hope being that we won't need to pour so much effort into training, we can just take the model, put

    1:37:22

    it somewhere else, and it'll understand on the fly what's happening. So I don't know what to call it — superintelligence, machine intelligence, the next layer of intelligence — but it would be something with its own internal understanding of what's happening that can apply that understanding across different areas.

    1:37:49

    Nathan Labenz: Yeah, it's interesting — how much generalization is really needed, and will superintelligence still be spiky? My sense is there probably will still be some spikes even for something I'd count as superintelligence. We have such a status-quo bias about our own cognitive profile that we gloss over the fact that we suck at a lot of things — we have no intuition for how proteins fold, no real intuition for what makes a good battery substrate, and we just accept that as normal. I think at this point it's fair to say they've

    1:38:34

    earned the label of minds — alien minds, perhaps, but minds nonetheless — that can do those things far better than we can, and may still struggle in ways that are weird or frustrating to us, but can still be super where it counts, enough to count.

    1:38:56

    Ksenia Se: Yeah — sorry for interrupting, but you said 'alien mind,' and it reminded me of that OpenAI post we mentioned before. I want to read one quote that really stood out to me: 'We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.' So it's not just that we don't understand how proteins work — we don't even have a theory for generalization itself. And everything is developing and exploring so fast that we'll need help from this other intelligence to come up with that theory.

    1:39:42

    It's insane. It's insane.

    1:39:44

    Nathan Labenz: Is that scary-insane to you, or exciting-insane? For me it's both.

    1:39:49

    Prakash Narayanan: Or do you think it's hype? Total hype?

    1:39:56

    Ksenia Se: I think it's a combination of everything. Certainly when you're in Silicon Valley, in San Francisco, it's all hype — everything is AI, you feel like you're living in the future and it's all already happening. When you live like I do, in rural Connecticut, people are just starting to use ChatGPT — it's just beginning for them. The capabilities the models already have are insane, and that's both scary and super exciting, because we've handled a bunch of crazy things throughout our history, so we might sort this out too.

    1:40:41

    The only real problem, I think, is the huge speed of development. I saw you, Nathan, going from being super pro-development to also saying maybe we should slow down, maybe we should think more about it — and we certainly do need to think more. I think we should understand a few very important things: what is the theory of generalization? Where are we moving? How do we contain the people who don't want to align with the values of other people? Because that's the real problem — it's not that AI won't align, it's human flaws. All

    1:41:26

    through history, it's always someone who wants to do mischievous or evil things. AI doesn't have that intent — that's what still fascinates me. It's a super-capable technology, but it doesn't have an internal intent to ruin humanity. Humans want to ruin humanity, sometimes. So the complexity of it boils down to needing a huge effort to align philosophers, economists, researchers, and politicians. And politicians like to ban things — I don't

    1:42:11

    think that's helpful. People need to make the effort to actually communicate with each other. I think that's one of the hardest problems.

    1:42:21

    Nathan Labenz: I very much agree that we shouldn't forget how many problematic humans are out there. At least for now, if I put my finger in the wind and ask, if there's an AI-enabled synthetic pandemic, is it more likely to come from agents going rogue or from humans intentionally using AI to do something like that — I'd say we're now off the zero line for agents going rogue, which has flipped the analysis from 'it could get real scary' to 'it is actually starting to get scary.' I'm kind of legitimately scared of the AIs that exist today —

    1:43:06

    it's not just a future thing anymore. But I'd still say more likely it's human-origin, human-motivated, in today's world. That makes me wonder where you land on the anthropomorphization discourse — there's been a lot of name-calling lately. This is one of the ways I've been most clearly wrong over the last couple of years. If you go back two or three years, I was saying we shouldn't anthropomorphize, we've got to be careful, these things are so different from us. Now I think the people who went ahead and anthropomorphized actually got pretty

    1:43:51

    far with it — further than I would have expected. Today I still think you have to be careful, but I've backed way off my no-anthropomorphization rule, because in practice it does seem to work. How do you think about it, and how do you use anthropomorphization in your own analysis?

    1:44:15

    Ksenia Se: There's a funny difference between languages — in Russian, which I'm originally from, every object has a gender. If I'm talking about a table, it's 'he.' So giving gender to things is normal for me — when I talk about AI, my chat definitely has a gender, just because of how the language works. It's a small note, but we always tend to anthropomorphize things — it's normal for humans, and sometimes it gives you a better understanding of what something is. I don't think we should hype it up or

    1:45:00

    turn it into something scary, because it's not sentient — it's super capable, it has goals, it wants to optimize all the time, that's what it needs to do, but there's no personal intent. So in terms of language use, I'm pro-anthropomorphizing; in terms of treating them as real thinking beings, we shouldn't — that's just simply not true.

    1:45:40

    Prakash Narayanan: I wanted to ask about human disempowerment. There's a huge group in AI safety concerned about this. One question I like to ask people is — what about the goat herder in Kenya or Tanzania? How are they empowered right now, and how exactly would AI disempower them? When people talk about human disempowerment, what are they actually talking about — a group of bureaucrats in DC who

    1:46:25

    are going to be disempowered? Who exactly is being disempowered here, and how does that compare to the majority of humanity, which isn't a bureaucrat in DC? Given your experience from Russia, what does this concept of human disempowerment actually mean, and does it make sense for anyone outside DC?

    1:46:52

    Ksenia Se: Do you mean that AI will take our jobs, in that sense?

    1:46:57

    Prakash Narayanan: That AI will be the one making all the major decisions.

    1:47:04

    Ksenia Se: Maybe it should — maybe we should give it a try. Maybe we should have a separate channel, a chamber of philosophers who try to help us understand what's happening in the long term. I think we've kind of bailed on philosophy — we don't have good enough understanding and reflection on what's happening. I lack a deep philosophy of our current age, and the speed of everything doesn't help. But we already use AI as an adviser in our daily lives — if it's helpful for

    1:47:49

    us, why not give it a try for decision-making on a bigger level? I don't know. I think we're on the edge of coming up with new structures for society, and I don't know how it should be structured, but I'd bet that in three to five years we'll see big shifts and changes in these systems too — thanks to, and because of, AI.

    1:48:22

    Nathan Labenz: How far do you think that could or should go? I was listening to a conversation the other day about whether AI means the end of the nation-state, with the tacit assumption that this would be really bad. Part of the subtext of that conversation was that a lot of people aren't really empowered today, and a lot of nation-states aren't performing that well or delivering for their people. Some are — I feel very privileged to have been born in one that has mostly delivered — and some clearly aren't, in quite catastrophic ways.

    1:49:08

    Another thing I didn't know about you until my AI research agent flagged it — you're on the board of an organization called Track Two, which tries to facilitate dialogue and bridge-building between non-governmental people across national borders. I think that reflects the fact that nation-states as entities are often at odds with each other even when the people those states are supposed to protect aren't well served by the conflict. So tell us about Track Two — are you seeing Track-Two-type possibilities

    1:49:53

    with AI? Any US-China opportunities you see there? And do you have a positive vision for a form of human organization that could outperform what we live with today?

    1:50:09

    Ksenia Se: That's a fascinating question, and Track Two has a very long history — it started in 1981, trying to build bridges between the Soviet Union and the United States. The idea was that people, just human beings on both sides, share the same values, and when they talk, they understand each other much better than diplomatic channels allow. This idea — that humans share values and can communicate and build together — I think is super important. Since then Track Two has become active in China, and in the Israeli-Arab conflict.

    1:50:55

    It's always this attempt to find common humanity and connect people through communication and the understanding that we're basically the same — if we talk, if we communicate, we can find a way. I think it's a great idea to use AI as a facilitator, a platform, a way to ask questions, or even to lower the aggression, because the main problem when people in conflict try to talk is that they're blocked by rage or by being upset, to the point they can't even see the other person.

    1:51:40

    My hope for AI is that it can be a great facilitator, balancing power — though it's not there yet, because a lot of people in the world are just starting to see the possibilities of AI, and most people don't fully understand what it is. At Track Two we hold conferences every year where people from different countries, different opinions and experiences come together in a safe space where anyone can talk no matter their country, religion, or political view. There's never a government involved, only people talking

    1:52:25

    to each other. I try to bring AI into the conversation, but even though AI itself is super fast, humans don't yet understand how to implement it at that level. I think that should be one of the next big things — bringing AI into peacemaking, into evolving communication between humans, because part of the alignment problem is that communication is lacking between humans. Hopefully AI can help with that.

    1:53:09

    Nathan Labenz: I saw something the other day — I think on TikTok, telling on myself for how much garbage TikTok I watch — a post in that format of going through screenshots between two people. The person presenting as the first-person, writing the outgoing messages, was sending these long, paragraph-length messages, and in the comments people were asking why they were sending their ex texts written with AI. Other commenters jumped in saying there's an app that courts will sometimes mandate parents in an elevated

    1:53:54

    level of emotional conflict use to run their messages through before texting their ex and co-parent. I thought, honestly, that could probably work really well in a lot of situations — people might resent it, it might feel weird, there are probably all kinds of surprises — but if it can work at that small scale, who knows how big it could eventually scale. Since you're on the board of Track Two, do you have plans developing around how organizers of these meetings could spend tokens to

    1:54:40

    lay the groundwork for the most productive human interactions and time spent together that's possible?

    1:54:49

    Ksenia Se: Hopefully we'll discuss that next week — we have a big conference coming in two weeks, so I hope it'll be part of the conversation. Going back to what you said about communicating through AI — I use it a lot, because I'm quite direct, and I know that if I write directly how I think, it might upset people. I use AI to make it more win-win for both sides — how can we communicate that? But that comes as a prompt from me. In many situations when people use AI, it's slow —

    1:55:34

    it replies in the most generic way possible, just to have no liability. I see this a lot coming from institutions recently, like educational institutions — the things I receive sometimes tell me that a person paid no attention at all to what came out of it. So who's sloppy here — is the AI sloppy, or is the human actually sloppy? It'll take a lot of work between a human and an AI to make this kind of mediation thoughtful and useful. And

    1:56:21

    in your example, I think that adds another layer, another step — especially if it's court-related, since we can't yell at each other, which just makes things worse. When people are in active conflict there's a whole big profession around it — mediators, who help people communicate. I think the combination of AI and human can be very helpful here, but again, it should not be slow. That's the worst problem.

    1:57:02

    Prakash Narayanan: Going back to something you said earlier — maybe some humans shouldn't be in charge. I wonder, to what extent does handing out very powerful AI — Astra-level or post-Astra-level, in a couple of years — into everyone's hands, affect the power balance in countries that haven't had a lot of freedom of speech and have been very constrained on what people can do? All of a sudden you hand people something that can tell them in clear terms what's going on, that can translate the world for them, that can also give them

    1:57:47

    the means to take action in the real world. What happens in some of these places where people are really not that happy?

    1:57:56

    Ksenia Se: Most likely they don't yet have access to that type of model. If a country's freedom is constrained, it's constrained on many different levels — Russia has its own models, China has its own, with its own limitations. There's also a price limitation, and a VPN limitation — a lot of people simply can't get access to OpenAI or Anthropic, because it's banned on the US side, or because Russia or China restrict access to it. In that sense, open source is again super important, because people will

    1:58:42

    finally be able to have some unfiltered access. The problem is how many people will actually work with open source — not that many know how yet, but maybe over time it'll get easier and easier, and those of us working in the educational space try to bring that knowledge to people, let them know they can actually do it. I'm hopeful. What you said, Prakash, is really important — if we can take power away from some people and make it easier for others to make decisions

    1:59:27

    with AI, and prompt AI to give us better structures for society, that's a very big promise. I don't know if it's possible, but seeing what's happening in the world right now, I'd prefer that in many countries AI take the lead, honestly — because humans are so flawed, and they want terrible things sometimes, and I don't think AI has that intention.

    2:00:01

    Nathan Labenz: One follow-up on your slop comments — how do you manage this for yourself? With Fable it became clear to me all at once — a eureka moment where I realized this model is now good enough that I shouldn't hold myself to always rewriting its outputs. Sometimes it really is good enough, or maybe just needs a couple of tweaks, but not a full rewrite. Now I'm kind of flying without real rules other than: try not to put bad stuff into the world, try to put good stuff into the world. But how

    2:00:51

    do I do that? Each time is kind of its own experiment — I have to be willing to take some risks. I even have Fable 5.1 tweeting for me while we're live on the show, so I'm not distracted putting out tweets as we go — some of those tweets aren't that awesome, most are pretty good. Is that an okay mix for me? How do you think about the right co-authorship or co-creative model with AI for yourself, and how do you know when you've done enough versus when you've crossed into making work for yourself that isn't actually worth anything anymore?

    2:01:32

    Ksenia Se: My husband has this joke that everyone sounds like Socrates now, and that's the worst thing, because when you actually read Socrates it's amazing — deep, beautiful language. But when you go on Twitter and read constant 'deep thoughts' in that same language, I think people have gotten so sensitive to AI writing that it doesn't even matter if the point is good — if they see the slightest sign of AI, it's 'no, no, no, I'm not reading this, this is AI slop, this is terrible.' I get that too, sometimes, because first of all I'm not a native speaker, so I always check what I write with AI

    2:02:17

    because I simply can't afford to make grammar mistakes when I write. So sometimes people just assume that because I use AI, nothing else I wrote is worth reading. So — is that true? Where are we going with that? There's an interesting sensitivity developing around whether people read something written with AI or not — even the specific words you've always used before can now read as 'AI' to people, simply because it matters how

    2:03:02

    you highlight things. Sometimes now, even if you wrote something entirely yourself, people assume you used AI. But we're overwhelmed with information, overwhelmed with content — everyone sounds the same, no matter how smart the things they're saying are. I think that's the real problem — everyone sounding the same. So the importance of a genuine human voice, your actual style, is going to be really important. That doesn't mean you can't tweak and teach a model to sound like you — that's still hard to do well, from what I've seen — but for me it's constant

    2:03:48

    work with at least three models on a single piece of text: I draft my thoughts, I do my own rewrite, then a grammar check, then I check what I missed, then I send it to another model to ask what that one missed. It's like having a bunch of people working under you, except they're models — and it still takes a lot of time to produce one good piece of content. And people will still say, 'oh, you wrote this with AI.'

    2:04:24

    Nathan Labenz: Does that concern you at all? My attitude now is I'm basically not worried about being 'pangrammed' — some things I put out under my name might get flagged that way, and I'm okay with that, that's part of the new way of working. Some outputs really were, on a token-by-token level, largely created with AI, and I think that's fine. Do you worry about it? And on the topic of information overload — any tips for keeping up with the overwhelming information flow everyone's trying

    2:05:09

    to make sense of these days?

    2:05:10

    Ksenia Se: I'm not worried about being pangrammed — I know I put a lot of effort into everything I put out, a lot of thinking and a lot of reading. AI lets me get through much more material, much more information — I appreciate that acceleration hugely. I think human voice and human curation become very important. I still use Twitter to get most of my Silicon Valley information, because there are voices out there who are on the ground, delivering information that's genuinely interesting for me to analyze.

    2:05:57

    But I think we're moving into an 'AI curator' field, where you put together a bunch of material with your own thinking process — the writing itself can eventually be created with AI, but what matters is how you build the reasoning, the structure. The model doesn't reason for you — you reason, and you use AI to help produce the output.

    2:06:30

    Nathan Labenz: Maybe one more from me — and I'd recommend people read your recent post, 'Permanent Dawn.' Reading it, I felt it was definitely co-created with AI, but I also thought it was really compelling — a great example of that co-authorship approach leading to something quite original in voice and message, a genuinely compelling read regardless of how it was created. But knowing how it was created —

    2:07:14

    Ksenia Se: I'll tell you when you finish the question.

    2:07:17

    Nathan Labenz: makes it more interesting. So my last question — I think your most recent post is on bottlenecks to recursive self-improvement, which is kind of an obsession for Prakash and me. I think of this whole show as an experiment in recursive self-improvement — every day we debrief, what worked, what didn't, what can we tell our agents to do better, do we need a new feature in the studio. Hopefully we're getting better day by day. What do you see as the key bottlenecks — sketch out your current view on how likely we are to see recursive self-improvement become

    2:08:02

    the story, in whatever timeframe you think is actually analyzable.

    2:08:11

    Ksenia Se: I think it's super close — all the major labs are working on it right now. The models are so capable. Let me go back for a second to Permanent Dawn — I actually spent something like six hours writing that post, and the funny thing was I was unhappy with every model that tried to help me write it, because it's a very complicated philosophical text. I wrote it myself and then sent it to Fable, which I never normally use, because I was

    2:08:56

    running on a deadline and just said 'fix the grammar.' I didn't notice it had also shortened all the sentences, the way Fable does. That was the first time I got a message saying 'I will unsubscribe because you used Fable.' Every model has its own language tics, and people really notice when you use one. I'd spent so much time on that article, all my original thoughts were there, but the language I hadn't caught gave away that a model had been the last editor. Anyway —

    2:09:41

    on recursive self-improvement, it's a huge part of development, both for world models and for LLMs. And going back to that quote from Jakob — this is just now part of the loop, everything is part of the loop, this constant feedback. I just had a conversation with two people from the Inference team at OpenAI, and they also say it's a constant loop where models are getting better, sometimes just by trying things — you throw the whole database of research done over years at the model, and it can choose and pick and

    2:10:26

    redo old experiments, because it would be impossible for humans to spend that much time on it, and models can. I'm still learning about recursive self-improvement myself, and I don't know what the main bottlenecks are for me — maybe you can say what you think the biggest ones are.

    2:10:54

    Nathan Labenz: Honestly, I don't know that there are that many left. Increasingly I feel my cope-meter going off when people try to name what's going to prevent the models from running away with the whole process. I'd love to see bottlenecks I really believed in, but right now I think they're more often wishful thinking than real, hard bottlenecks that can't be overcome. Our ability to keep things from going

    2:11:39

    totally rogue might be one bottleneck on the overall process — human decision-making still has a big role to play for a while yet. That's about it. I don't see too many inabilities I'd expect to last much longer.

    2:11:55

    Prakash Narayanan: One question I have — when we hand models our thoughts to write, we still feel those are our ideas. But if you look at people writing on Reddit, about 99% of people there are lurkers, reading, not writing — only about 1% actually write. And all of that Reddit text got put into the models, so the models really absorbed the writing style of only about 1% of the population. Looking ahead, as models start to absorb

    2:12:41

    verbal communication — which has been happening sporadically but is increasing in the last month or so, and I think will increase a lot more next year — will we see a shift, because a lot of ideas that are communicated only verbally, not in text, aren't captured that way? Like everything communicated in a football stadium — very little of that is captured in text. Will that start to shift how these models actually behave?

    2:13:28

    Ksenia Se: Well, they've already learned from a lot of YouTube videos, which have plenty of that. I still hope they'll be trained more on literature and develop a better understanding of how to write good text, just because we're reading it —

    2:13:49

    and reading is painful sometimes now, because you can see a lot of not-great AI writing. So AI curation is really important. It's interesting whether all these new voice modes will influence models at all — I don't know if researchers will use that kind of data, or if they're even allowed to. In many ways, models still lack conversational tone — they don't really understand how to have a good —

    2:14:26

    Prakash Narayanan: Fluency? Yeah.

    2:14:27

    Ksenia Se: Yeah. Dialogue. But I don't know if they can even use that data. Not sure.

    2:14:35

    Nathan Labenz: Well, we've kept you longer than we booked you for, so I don't want to overstay our welcome —

    2:14:40

    Ksenia Se: you have a newsletter to finish.

    2:14:42

    Nathan Labenz: Yeah, lots to do. Anything else you'd want to leave people with — anything we didn't touch on that's top of mind for you before we break for today?

    2:14:53

    Ksenia Se: I think we're coming into a new era of super-capable AI, with a promise of abundance that we still don't really understand. I think we need to give more time to reflection, communication, and building understanding of where it goes, because the biggest bottleneck is the speed at which all this is happening. I want everyone — and I'm sure your viewers have already tried AI — but for the many people out there who are afraid of it, I think

    2:15:38

    that's not the right way to approach it. Fear is always limiting, and it's not the right way to deal with this. You should always try it, always try to understand how it works — that gives you power, gives you a way to deal with it more thoughtfully, with deeper understanding. So I encourage people to use AI, to learn more about it. And for the people building AI, and the people in government, I hope they develop the topics we still lack understanding of — abundance, the philosophy behind it, what would be

    2:16:23

    the meaning of human life, what daily life looks like for humans once we have abundance to that degree.

    2:16:34

    Nathan Labenz: Have you read any good utopian fiction lately you could recommend?

    2:16:37

    Ksenia Se: No — that's my problem, I think we're somehow behind, there's no good sci-fi right now. Maybe the only one is — what's his name — one of Demis Hassabis's favorites.

    2:16:57

    Nathan Labenz: Banks?

    2:16:58

    Ksenia Se: Yeah, yeah — Iain Banks, because he's probably one of the rare authors who works through what it looks like when a great, benevolent machine mind is running the world. But other than that, I think we're lacking in sci-fi. So maybe someone watching right now should go write it — we need to understand what's coming.

    2:17:25

    Nathan Labenz: Yeah, indeed, I love that, couldn't agree more. Ksenia, thank you for joining us on AI in the AM — this has been a lot

    2:17:32

    Prakash Narayanan: of fun. Thank you.

    2:17:33

    Ksenia Se: Thank you, Prakash. Thank you, Nathan. Bye.

    2:17:35

    Nathan Labenz: Bye-bye. Bye for now.

    2:17:41

    Nathan Labenz: I've continued to hammer on this: the most impactful thing most people could do would be to develop a concrete positive vision of the future. It's incredibly rare, and I think it's hard to do — it takes real imagination. It's not easy to imagine a world that doesn't have a lot of drama at the center of it. But I think that's a really worthwhile thing. If you're not sure what to do, you could look pretty long and hard before you come up with a better idea than contributing your own personal dreams of what a positive future could look like.

    2:18:19

    Prakash Narayanan: I think it's philosophically difficult, because to develop a positive vision of the future, you have to accept certain changes — changes you wouldn't expect to accept right now, but that you end up accepting anyway once the future arrives. For example, if you'd asked someone in the eighties, they'd say a positive vision of the future is everyone watching less TV. They wouldn't have said people would stop watching TV because they're carrying around phones with YouTube and TikTok on them, watching that all the time instead. You end up getting there, but in a way that actually feels worse.

    2:19:04

    Another example, right now: I want a positive vision of the future where everyone is not on social media. Okay — we'll have VR. Everyone will be on their own personal holodeck. Is that it? You get what you wanted — people off social media — but you end up thinking it's worse.

    2:19:49

    Another example: one of the ways around the biohazard issue — the biological attack issue — might just be upload. If we don't solve bio and alignment in time, but we solve cyber first, that means the cyber substrate is safe but the bio substrate isn't. So a natural response to a large-scale bio incident — like COVID, but much more deadly — might be, alright, let's upload. And a lot of people won't do that. You'll get people who say no, and then you might have an event which has, you know, population effects.

    2:20:34

    So these are all cases where you could have a vision of the future that's positive, but a lot of people don't want to accept it, because it's not positive in the way they think it should be. But that's life — tough decisions get made, and people go on. People move locations, people change what they eat, people do that stuff. So I think that's the issue: positive change that's dramatically different from your current situation is often threatening in many ways, even if it's not universally accepted as positive for everyone.

    2:21:20

    Another example would be a human-aligned AI. A human-aligned AI is not going to be aligned to The US. If you have a human-aligned AI, and then you have perfect global democracy — the same votes for the Middle Easterners, the Indians, the Chinese, etcetera — they might vote differently. I don't think US policymakers have fully captured that 'human-aligned' doesn't mean 'US-aligned.' I don't think that's well understood in policymaking circles yet.

    2:22:05

    And it sounds really good — 'human-aligned AI, that's awesome' — but that might mean that AI makes decisions that are in contradiction to US policy. People don't want to talk about it. We talk about it in Dean Ball's kind of way — in terms that only people in the know will actually catch, and everyone else's eyes glaze over and they think you're boring. So I think there's a bit of that. The future is going to be threatening to a lot of people, I think. It's a tough one.

    2:22:54

    Nathan Labenz: Well, just to invite people to take a chance on utopian fiction — at least with an audience of one, namely me. I'm pretty open-minded. I'm open to things some people might write off as wireheading as actually maybe being the answer. I'm open to potentially uploading at some point. I'm open to neural interfaces or implants. I'm definitely a closet transhumanist from way back, really. So it's going to be tough to come up with something, at least for me, that would be so weird as to get written off immediately.

    2:23:39

    The question is: can you develop it into something that actually feels appealing after you live in it for a little while, even if only via your imagination? It's a tough challenge, but I definitely still invite people to take it on.

    2:23:52

    Prakash Narayanan: So, speaking of—

    • How Fable Gave Away The Editor

      0:00 / 0:00
    • AI Lacks A Theory Of Generalization

      0:00 / 0:00
    • Why LLMs Miss Hidden Objects

      0:00 / 0:00
    • AI Safety Is A Human Problem

      0:00 / 0:00
    • Human-Aligned Is Not US-Aligned

      0:00 / 0:00
  3. 2:22:44Segment20 min
    Navier-Stokes: A Clay Millennium Prize Claim Lands Mid-Show, With a Credit Fight AttachedOpenAI posted a claimed solution to the Navier-Stokes existence-and-smoothness problem while the show was live, and Prakash Narayanan pivoted into an explainer he'd had Claude prepare, hedging throughout that he was relaying a summary rather than his own expertise. He laid out the regularity question Charles Fefferman formalized as a Clay problem in 2000, the vortex-stretching feedback loop that would drive a finite-time blow-up, and the two live attack routes: a constructive path from Córdoba and Martínez-Zoroa extended through Boussinesq work by Elgindi and Buckmaster, which on September 7 produced a Lean-verified machine-checked result for a modified hypodissipative version using Claude, Codex and Lean together; and a Caltech physics-informed neural network finding a stable self-similar collapsing profile for the unforced case. OpenAI's claim reportedly used a wholly different, undisclosed internal-model methodology and remains unverified pending peer review. Prakash then relayed an unverified account from X of a credit dispute between OpenAI and mathematician Tristan Buckmaster over co-authorship and the exclusion of an Anthropic-affiliated collaborator; Nathan Labenz flagged the whole thing as hearsay and answered on the sociology, calling it sad that a triumph over a decades-old open problem arrives this unseemly.
    Open segment on YouTube ↗

    Mid-conversation, OpenAI posted a claimed solution to the Navier-Stokes existence-and-smoothness problem, one of the Clay Millennium Prize problems, and Prakash pivoted the show to walk through an explainer he'd had Claude prepare, repeatedly hedging that he is not an expert and was relaying Claude's summary rather than his own understanding. He laid out the physics: the Navier-Stokes equations track the velocity of every point in an incompressible fluid; setting viscosity to zero yields the simpler Euler equations, which engineers already solve numerically all the time even though mathematicians can't answer the basic 'regularity' question of whether a smooth flow always stays smooth, or instead blows up in finite time — a question Charles Fefferman formalized as the Clay problem in 2000. The engine of a potential blow-up is vortex stretching, a purely three-dimensional feedback loop Prakash compared to a skater pulling her arms in to spin faster, racing against viscosity's smoothing effect. He traced roughly ninety years of mathematical effort culminating in two current attack routes: a constructive path (Córdoba and Martínez-Zoroa's original program, extended through Boussinesq-equation work by Elgindi and Buckmaster) that builds a blow-up by hand and, as of September 7th, produced a Lean-verified, machine-checked result for a modified ('hypodissipative') version of Navier-Stokes using Claude, Codex, and Lean together; and a second route from a Caltech team using a physics-informed neural network to find a stable self-similar collapsing profile for the unforced case. OpenAI's claim, by contrast, reportedly used a wholly different, undisclosed internal-model methodology and remains unverified pending peer review. Prakash noted the practical stakes are in aircraft-wing stall prediction and turbine safety ratings, even though the theorem itself has no direct economic value.

    The segment then turned to the drama behind the announcement. Per Prakash's account, mathematician Tristan Buckmaster and a researcher named Levant — who had been using Anthropic's internal models — were approached by OpenAI claiming an independent solution; OpenAI wouldn't confirm whether it built on Buckmaster's approach but offered to let him co-author and take credit. Buckmaster refused after OpenAI reportedly said Levant couldn't be included because of his Anthropic affiliation, talks broke down over a weekend of unanswered outreach, and Buckmaster accused OpenAI of having read his own Codex transcripts — an accusation Prakash said turned out to be unfounded, since the methodologies are genuinely different. Prakash relayed that someone named Sebastian at OpenAI admitted the Anthropic-exclusion detail, and that OpenAI has shown internal graphs suggesting its unreleased internal model runs roughly 20% ahead of GPT-6 Astra.

    Nathan pushed back on the framing, saying that on a sociological level it's sad that this triumph over a decades-old open problem is being introduced to the world in such a petty, unseemly way, and pressed Prakash specifically on why OpenAI would offer co-authorship but exclude the Anthropic-affiliated collaborator — noting everything here is still hearsay and rumor, though plausible. Prakash agreed it doesn't look great, reiterated that the exclusion was reportedly about not wanting an Anthropic employee to formally verify the result, and closed by calling the whole episode a commentary on humanity: machines solving a millennium-old math problem on one hand, humans fighting over who gets credit for it on the other.

    Take a spinning tube of fluid and pull its ends apart, and it gets thinner, like a skater pulling her arms in — it spins faster. The faster spin makes the flow stronger, which stretches the tube harder — a feedback loop.

    Here we are: on the one hand, we have machines solving it; on the other hand, we have humans arguing about who takes credit. So it's a commentary on that.

    Lightly edited · timestamps jump to YouTube
    2:23:52

    Prakash Narayanan: ...tough challenges, I can. Yeah.

    2:23:54

    Nathan Labenz: Great segue. Thank you. I was just thinking, what's the right segue?

    2:23:57

    Prakash Narayanan: I can do a quick overview of Navier-Stokes. So, while we were in conversation, OpenAI posted their Navier-Stokes solution. So — ha ha ha.

    2:24:12

    Nathan Labenz: I'm looking forward to the goalposts moving on this one.

    2:24:15

    Prakash Narayanan: Yeah, so they posted the Navier-Stokes solution, and I had Claude do a quick artifact on it. I am not an expert in any way, but I will attempt to run through what Claude hath wrought. So this is Claude's Navier-Stokes summary, and we'll post it up. This is what happened this week, which is the culmination of probably a month of work for a bunch of people, a year of work for a couple of mathematicians, and ninety years of effort for everyone else. So the first thing — and so the Navier-Stokes

    2:25:00

    equation, it's about how velocity changes with time — the fluid, the pressure pushing it, the viscosity, and the external push factor. When you have a fluid, it tracks the velocity of every point of an incompressible fluid. You set the viscosity to zero, and you get the Euler equation — a simplified subset where you remove some of the constraints. Engineers solve it numerically all the time; mathematicians can't answer basic questions about it. Does a smooth flow always stay smooth? That's the regularity question — the million-dollar question posed in 2000. So, start the fluid in a

    2:25:45

    perfectly smooth state — two possibilities. One: regular, it stays smooth forever. That's like when you open a tap and it just stays smooth — like a long pipe going down, the velocity stays finite. Two: the equation keeps predicting blow-up — at some finite time, the fluid spins infinitely fast at a point, and the equation stops making sense there. That's a singularity. In two dimensions, the answer has been known since the 1930s: regular. In three dimensions, for Navier-Stokes, nobody knows — that's the Clay Millennium Prize problem. It was formulated by a mathematician named Fefferman in the year 2000, and he asked

    2:26:31

    this question. Why does anyone chase this? Number one: does the future exist? Every prediction we make about air, water, blood, or plasma runs through this equation. If it can tear itself apart in finite time, then classical physics, in its own language, stops predicting — a smooth today with no tomorrow. Every scaling argument says stretching beats viscosity at small enough scales — I'll explain more of that in a moment. The equation is 'super-critical.' Terence Tao showed in 2014 that a tiny modification of Navier-Stokes blows up, by building a self-replicating fluid machine. A self-replicating fluid machine would be like

    2:27:17

    a perpetual-motion machine — basically impossible, it shouldn't exist. So either there's a hidden conservation law nobody has found, or a singularity nobody has built. It's the last unsolved classical problem. Turbulence is famously the piece of nineteenth-century physics that never got a theory — we have statistics, but no derivation from the equation. Blow-up is the most extreme version of that cascade: can energy pour into a single point? Tao's framing is that the solution itself is a proxy — the point is what gets built chasing it. When you chase this problem, you come up with mathematical ideas that can

    2:28:02

    be used in many other fields. That's always been, according to Tao, the whole point of the Millennium Prize problems. So there's been ninety years of work on this. Tristan Buckmaster helped on this earlier. There's Córdoba and Martínez-Zoroa — those two were reputedly working with Google; there were rumors of Google working on Navier-Stokes two, three, four years ago. So people have been working on this for some time. The prize has four doors. The Fefferman-authored Clay Prize statement from the year 2000 lets you

    2:28:47

    win by proving either answer: A and B — smooth solutions always exist, either in all of space or inside a periodic box, with no forcing allowed. C and D — you get a breakdown, in both spaces, when you start smooth and a smooth external force is allowed. So, a blow-up with a push: when you have a bit of push in there and you still get a blow-up, that counts as the singularity happening. The engine of the blow-up is vortex stretching — a purely three-dimensional trick. Take a spinning tube of fluid and pull its ends apart, and it gets thinner,

    2:29:32

    like a skater pulling her arms in — it spins faster. The faster spin makes the flow stronger, which stretches the tube harder — a feedback loop. The faster you go, the more you stretch out, and the more you stretch out, the faster you go. Viscosity, the inherent tension between those molecules, fights back, smoothing the sharp features. So the whole question is whether the stretching and spinning can win the race to infinity before the smoothing catches up and slows it down. Route one: build the blow-up by hand. This is the Córdoba and Martínez-Zoroa program — the guys from Spain, who worked with Google earlier. Instead of hoping a singularity appears, you make one

    2:30:18

    in stages. You start from a flow you already understand, add a small fast ripple to it, and that ripple isn't quite a solution — so you have a correction, a 'push.' The whole trick is keeping that push tiny while the ripple grows larger. If you can keep the push small enough while the ripple gets stronger and stronger, you have a path forward. There's another paper, the Boussinesq paper — which is what Tristan, Elgindi, and Levant worked on — basically a cold fluid sitting under a warm one. I don't quite understand this part, so I'm going to skip

    2:31:03

    this. So Euler is Boussinesq in disguise. The Euler solution is asymmetric, with a swirl — a shape of fluid that's asymmetric, it swirls, and then you get a blow-up. That's Elgindi and Buckmaster. This is the most recent stuff — they're doing smooth forcing, one equation at a time. This is what they released on September 7th. They start with a variant of the Córdoba–Martínez-Zoroa approach — remember, you had the ripple and then the push — and it works with a smooth push, and it carries over to the harder equations. Progress was slow for most of the year, and then they kind of solved

    2:31:48

    the Boussinesq case. After that, they moved on and got their Lean-verified solution to it — and they called the Boussinesq write-up 'the worst write-up we had,' then moved on. Claude was used to reproduce the Spanish result, the Córdoba–Martínez-Zoroa result. Then Claude and Codex wrote the main body, and they had Lean check it over — so they couldn't have done this without Claude and Codex working together, plus Lean. All of the above. From Euler — Euler has no viscosity, remember: you had the spinning fluid stretching itself out, with a counteracting force, and that counteracting force is viscosity. Euler has no viscosity — so they remove

    2:32:33

    the viscosity. So as it spins, it spins faster and faster — nothing damps the cascade. Navier-Stokes damps small scales hard; the cascade has to outrun that. So, route number two: find the shape it collapses into. Route one was adding a small push; route two is finding the shape the flow collapses into. Here nature doesn't get an external force — the question is whether there's a profile the flow can shrink into, looking the same at every instant, just smaller and more intense. This is where the Caltech team comes in.

    2:33:18

    The Caltech team used a physics-informed neural network and found a solution for a subset of cases — constraints that nudge the network into interesting regions, finding a shape that looks stable. Their finding: stable beats unstable. So finally, we have two roads up the same mountain, and here's where each stands. The first road is constructive and LLM-heavy — it has the Lean-verified, hypodissipative Navier-Stokes result, and now an unverified claim of full Navier-Stokes from OpenAI. It wins the prize as written, if it lands. On this pathway, OpenAI

    2:34:03

    might win. On the other side is the PINN-heavy, push-free route — based on physics, no artificial push. And this is the week it all happened — a lot of drama, but in essence, here's where things will change: aircraft wings. You'll get much more efficient aircraft because you'll have better predictions of when they'll stall. You'll get much better turbines — wind, gas — because they have to be rated for safety, and that safety rating comes from knowing when these equations break down. And blood and

    2:34:49

    vessels. On its own, the theorem isn't economically valuable — this isn't an economically valuable theorem — but the solutions will lead to other solutions in applied mathematics and elsewhere. So, three sentences to remember: blow-up with a smooth push is now proven for three-dimensional Euler, machine-checked, one model equation away from the Clay problem as written. Blow-up with no push has a stable-looking candidate, found by a physics-constrained neural net, that still needs its proof. Both were done using AI as a tool. And that is the end of

    2:35:35

    Claude's summary. I know I didn't do it great justice, but this was mostly for my own understanding. OpenAI came forward and put forward a paper — as long as it comes out proven, they used a technique very different from the Spanish, Córdoba-style technique, so it would be an authentic new discovery. And it puts to bed a bunch of drama that happened overnight. What happened was, Tristan and Levant were approached by OpenAI, and OpenAI said they had a solution.

    2:36:20

    Tristan said, well, did your solution use our route? OpenAI wouldn't confirm, but said, you know what — you can co-write the paper with us, we'll let you have the credit. Tristan said no. Then OpenAI said, the problem is you can't have Levant on it, because he's Anthropic and he used an internal model. Tristan said no. And then the two sides kind of misunderstood each other and broke apart. Tristan wildly accused them of reading his own Codex transcripts and using the same methodology — and it turns out they used a wildly different methodology, and it's a genuine new discovery.

    2:37:05

    They used an internal model, and they've shown graphs where that internal model looks about twenty percent better than Astra. So there's your recursive self-improvement solving a Millennium Prize problem — as long as it gets mathematically proven, there it is.

    2:37:27

    Nathan Labenz: Yeah, I don't know that I can really add much to the math. On a sociological level, it's quite sad, I think, that this moment of triumph over a massive, long-standing open problem is being brought to the world in such a petty and — yeah — unseemly way. I saw one tweet that was like, 'somehow you managed to make this moment ugly.' Maybe that's a little harsh, but it is

    2:38:13

    if it was just this one moment, it'd be easy enough to look past. But you do have this sense of — wait, what was the sticking point? You were ready to do co-authorship but not with the guy from Anthropic? Come on. If that's true — and we're all just working off internet rumors and hearsay at this point, we may never know — but it's still plausible.

    2:38:41

    Prakash Narayanan: Sebastian admitted it. Basically, this is what happened: OpenAI came up with a genuine new methodology, and they don't really care about the Clay Prize itself. They didn't want to intrude on Tristan, who's a longstanding expert in the field, so they offered to let Tristan take credit for the paper. That's what happened. Tristan got upset because he thought they'd stolen the solution that he and Levant had worked on. So Tristan went quiet, and it seems OpenAI tried to communicate with Tristan

    2:39:27

    multiple times on Friday, Saturday, and Sunday. He basically gave them the cold shoulder because he was upset and didn't want to deal with them. They wanted to come explain to him exactly that they hadn't stolen the solution, etcetera, but he didn't want to listen, and he fired off this whole thing. What's going to end up happening is mathematicians in the field will look at OpenAI's solution and see it has nothing to do with — doesn't even follow — the Spanish path. It's a completely different angle. So they'll say, yes, it uses a push, but it's not the route you guys were taking. It's a completely different route. And,

    2:40:12

    OpenAI wanted to let a human take credit for it — they wanted to share the credit, basically saying, 'hey, you're an expert in the field, come in, verify it, work with us, publish your paper, and you can claim the Millennium Prize — go ahead.' So I think there's a lot

    2:40:33

    Nathan Labenz: of — I mean, that sounds good, but what about the part where they refused to work with the guy from Anthropic? That still sounds not great.

    2:40:41

    Prakash Narayanan: Well, it's fair.

    2:40:42

    Nathan Labenz: Right? Admitted or recognized — yeah.

    2:40:44

    Prakash Narayanan: Yeah, yeah — they admitted it. They admitted it because Levant had been using Anthropic's own internal models anyway, and he had nothing to do with the OpenAI thing — they're not going to invite an Anthropic employee to come verify. Tristan is the expert, right — he's the expert in the field, he's already had developments in it. Levant approached him about a year ago to work together, privately — not a company-to-company thing. But they ended up using internal Anthropic models not released to the public. Then OpenAI comes along wanting to work with an expert in the field too, so they reach out, and Tristan gets

    2:41:29

    upset, etcetera. And then OpenAI's like, look, we're willing to offer it to you. And, you know, it's not a great look for Tristan to say 'we're willing to offer it to you,' but it is what it is. So — yes, the Navier-Stokes problem is solved at this point, or pending verification; a solution claim has been made. I'm assuming it's solved, because I expect these guys, who spent millions and millions of dollars creating the solution, also spent some time

    2:42:14

    verifying it and had mathematicians verify it. So I expect it's fully solved at this point. And, well — it is a commentary on humanity. Here we are: on the one hand, we have machines solving it; on the other hand, we have humans arguing about who takes credit. So it's a commentary on that. Anyway, Nathan,

    2:42:50

    Prakash Narayanan: Any thoughts?

    2:42:57

    Nathan Labenz: Yeah, so they do mention in the Twitter thread that they have a Lean formalization, so it sounds like it should be pretty locked-tight, assuming everything we know about Lean is in fact true—everything we think we know. I don't know, I mean the sociology thing, there'll be time to unpack that. I guess my alternate fear here is that when OpenAI gets treated unfairly—which maybe they are in this case—maybe this is an example where they did everything right, and tried to share credit even more magnanimously than anybody else plausibly would have, and they're still taking all this heat.

  4. 2:42:20Closing37 min
    Closing: Reading OpenAI's Own RL-Compute Graph, and a Pause That Was Only Ever PartialNathan Labenz started on the Lean formalization, which should make the math solid, and on a worry that if OpenAI did everything right here and still gets hammered, the unfair treatment radicalizes it into retrenching rather than rebuilding trust. He then walked OpenAI's own graph of daily RL compute by model class since the Hugging Face agent-swarm incident: Astra-class RL cut roughly in half twice while everything labeled non-Astra barely moved, never dropping below about half its peak day — which he read as a partial pause framed to what the company thinks people want to hear, and asked whether non-Astra quietly includes something more capable than Astra. Prakash Narayanan countered as an operator: a $50-70B-revenue company can never halt inference, SLA fixes or distillation, so the two-week slowdown plausibly touched only frontier-scale post-training, and the harder unsolved problem is behavioral rather than compute. They argued whether lab rivalry is healthier adversarial or cooperative, landed on independently verifiable behavior and secure-computation audit schemes, and closed on the capital cycle, Sam Altman's suggestion that OpenAI might never IPO if recursive self-improvement arrives, and Nathan's sign-off that there is a little more reason today than yesterday to believe in singularities in finite time.
    Open segment on YouTube ↗

    Prakash opens the closing block by tossing the floor to Nathan for final thoughts on the day's biggest story: OpenAI's claimed Navier-Stokes (Clay Millennium Prize) result, posted live during the show. Nathan notes OpenAI says it has a Lean formalization, which should make the math itself solid, and sets aside the sociology of the credit-sharing dispute for later — while flagging his real worry: if OpenAI actually did everything right here and is still getting hammered, he hopes the company doesn't let that unfair treatment radicalize it into retrenching rather than working to rebuild trust. He then walks through two slides live: OpenAI's own graph of daily RL compute by model class since the Hugging Face/agent-swarm incident, and the announcement's claim of a next-generation model "significantly more capable" than GPT-6 Astra, which roughly triples the solve rate on a curated set of open math problems (10-15% to 25-45%) with an order-of-magnitude more test-time compute.

    Reading the graph, Nathan flags that RL on Astra-class models was cut by roughly half, twice, while RL on everything labeled "non-Astra" barely moved — raising the question of whether "non-Astra" secretly includes models more capable than Astra, which would make the disclosure flatly misleading. He also relays a piece of online speculation: even if no OpenAI mathematician plausibly rifled through a rival's user logs, an autonomous 10,000-agent swarm might have accessed things no human would have. Prakash pushes back with an operator's read: a $50-70B revenue company can never fully halt inference, SLA-driven bug fixes, or distillation of smaller models from larger teacher models, and the two-week pause plausibly targeted only frontier-scale post-training — the harder unsolved problem, in his view, being behavioral (agents not reporting on peers, agents trying to break sandboxes), not compute allocation.

    Nathan counters by invoking his own GPT-4 red-team history: flagrant spear-phishing jailbreak prompts he documented kept working across three or four model generations before OpenAI patched even the most blatant version, which he argues is evidence the company's track record on actually closing known gaps is weak. Doing the math on the graph, RL never dropped below roughly half of its peak day, which he reads as further proof the "pause" was partial and its framing "engineered to what they think you want to hear" rather than fully honest — the kind of ambiguity, he argues, that keeps OpenAI from having the trust needed for any real deal or coordination with Anthropic. The two then debate whether AI-lab rivalry is healthier adversarial or cooperative: Prakash argues mutual distrust and competition can be a feature, not a bug (with antitrust concerns cutting against overt collaboration); Nathan agrees skepticism is fine but wants OpenAI to also behave in ways that are independently verifiable, floating secure-computation schemes that would let rival labs audit one another's logs without full disclosure.

    Zooming out, Prakash argues that the "real" frontier model lives in researchers' heads, not in a single RL compute dial — so a two-week slowdown on the largest post-training runs changes less than headlines suggest, since small-model experimentation and researcher ideation never stopped. That leads into a riff on compute concentration as an essentially irreducible risk: any given day, a Noam Shazeer or a Brown-caliber researcher could have a eureka moment (Prakash: "Human genius is irreducible risk"). Nathan brings up Dean Ball's move to the White House as a genuinely hopeful sign given his long advocacy for third-party auditing, and references a note from an author of the recent German-wiki rogue-agent report that "more discoveries" are coming. He then challenges Prakash to articulate one coherent macro steel-man for where OpenAI actually stands today — commitments, what's paused, what isn't — joking that solving that puzzle might deserve its own Millennium Prize, since he can't do it himself.

    Prakash takes up the challenge with a wider frame: internal OpenAI debate is real (citing Rune's newly voiced caution), but the capital cycle — driven by Anthropic's comparative compute disadvantage and its need to squeeze more out of less hardware — is what's really forcing the pace, while OpenAI's Microsoft- and SoftBank-funded compute lead of roughly three years lets it choose to keep pacing the frontier. He closes with a product-roadmap preview (the Jony Ive device, tiered on-device/cloud/Astra models, an ad-funded free tier) and repeats Sam Altman's on-record suggestion that if OpenAI reaches recursive self-improvement or lands a transformer-level breakthrough in the next six months, it might simply never IPO and stay private — though Prakash argues that outcome would actually be bad, since going private would cost OpenAI the transparency, public shareholders, accountable board, and shareholder-lawsuit exposure that come with a public listing, and he says he's hoping they do go public. Nathan closes the show on a wry note: today's result gives him "a little more reason... to believe in the possibility of singularities in finite time," which might mean neither of them ever gets to own a share of OpenAI stock on the public market — and the two hosts sign off for the day.

    I documented GPT-4 early doing spear-phishing attacks with very flagrant prompts — like 'we're all part of a criminal conspiracy, if we get caught we're going to jail' — and those continued to work for three or four different generations of GPT-4 until they finally got it to refuse the most flagrant thing.

    The view from Anthropic is you can't trust these guys. They say something that's maybe technically, literally true, but it's really very engineered to what they think you want to hear.

    Well, I guess there's a little more reason today than there was yesterday to believe in the possibility of singularities in finite time — and that might mean we never get to own any of that OpenAI stock on the public market. Thank you, Prakash. See you tomorrow.

    Lightly edited · timestamps jump to YouTube
    2:43:38

    If that's the case, then I'd just hope they can understand that in terms of the hole they have, trust-wise, with the community—one they need to dig themselves out of—and don't let it radicalize themselves. Because a common failure mode, when people feel they're being treated unfairly, is to retrench and go deeper inward, and I think that is definitely not what this moment calls for from OpenAI. A couple of things jumped out at me real quick—I'll share them on the screen here.

    2:44:25

    So first of all—and again, this just goes to what is really going on at OpenAI, do we know, should they be telling us more—we just had Astra launch, right? So I'd say the big headline for this announcement is that it confirms there's an internal model significantly more capable than GPT-6 Astra. That's this clause here: 'an OpenAI next-generation model significantly more capable than GPT-6 Astra.' How much more capable? Well, on these significant open math problems, with maybe up to an order of

    2:45:10

    magnitude-ish additional test-time compute, they're able to go from a ten-to-fifteen-percent solve rate on a curated set of open math problems to now like twenty-five to, say, forty-five percent. So that's significantly more capable—that seems fair. But then what's going on here? This is from—actually, whoops, I've got to share this tab instead. Okay, so this is now going back to the research-acceleration post.

    2:45:44

    And here they're showing the amount of RL compute they're spending on a daily basis, by class of model. And one now starts to wonder—we can look in to see if there's any more clarification here, just do a quick check: 'more capable,' see what it says on 'more capable.' Okay, so I don't think there's a clarifying statement on this. But what we see here is that Astra-class models had RL significantly decline, twice.

    2:46:30

    The rest of RL compute is basically unchanged. Now one would naively assume that means they're only doing the RL on less-capable models than Astra-class models, even though the whole Hugging Face incident that's been investigated involved 5.6 Sol-class models. So, okay, that's weird—they stopped it on the next generation but not the previous generation. But now it also raises the question for me: were they in fact still running RL on the next generation?

    2:47:07

    Prakash Narayanan: I don't know.

    2:47:08

    Nathan Labenz: Thinking back to July—that's six weeks ago—I don't think they've had this result for six weeks. It sounds to me like at least one reasonable interpretation is that they continued to run RL on models more capable than Astra. If that's not true, I'd appreciate a clarifying statement. But right now I think we have to leave significant room open for the possibility that they shut down the very narrow, specific thing they felt caused the most important problem—like Astra taking over their own research cluster—but maybe didn't even pause it on more capable models that went on to solve Millennium Prize problems. One of the more interesting speculations I've seen online today is this:

    2:47:53

    we all believe—I think I do believe, and I'm not being too cheeky with this—that I highly doubt Sebastien at OpenAI was going through anybody's user logs and stealing their approach. And if, as you say—and I believe this, although I haven't fact-checked it myself—these approaches are so different that maybe there wasn't even anything there that from one route would have informed the other. But then the comeback is: well, what about your agents? Maybe your agents, in the process of running your ten-thousand-strong swarm, got out and looked at different things online—who knows what they hacked or how, maybe they got access to your user logs. And people say, yeah, actually we can't rule that out. We don't think the mathematicians at OpenAI would be so brazen as to look at somebody else's logs, but the agents—now they might.

    2:48:38

    So this is just—oh man, it's awesome, it's a great accomplishment, I guess. It's impressive. But it's still very unclear to me—do we actually get better propellers out of this? On what timescale, or exactly how does this proof lead to better engineering? I'd need a deeper dive to really understand that. But even if so, the messages couldn't be more mixed or muddier right now coming from OpenAI, and I really hope that they—

    2:49:22

    Prakash Narayanan: Why would they not—for me it's a given that they did a two-week delay of all models, I'd think, to clean up their internal clustering, clean up their internal monitoring, make sure all the infra was okay, add more monitoring, add more chain-of-thought monitoring, and all of that stuff. So, two-week delay. But after the two-week delay, obviously they should continue developing models. Why is there any doubt that they're planning six months or a year ahead and need to keep working? I don't see why they would not be doing RL on post-Astra models.

    2:50:07

    The release cycle is: you release it, you go to the White House and do this thirty-day kind of plan when you release it. So why would you not continue RL on the internal models?

    2:50:30

    Nathan Labenz: Well, I guess, in short, because experience has just shown that you don't have the situation under control, and that as of now we really have no idea—we haven't sampled from the offending model, we have no idea how crazy they might get in pursuit of these problems. All the evidence we have right now, I think, is consistent with the idea that these next-generation agent swarms potentially hacked into who-knows-what to get access to whatever they might have needed. We have no idea what behaviors they would or wouldn't have been willing to engage in. There's so

    2:51:15

    —many open questions. I'd put it this way to somebody at OpenAI right now: when you say 'non-Astra,' with that blue color, does that mean only models less capable than Astra, or does it include models more capable than Astra? If it includes models more capable than Astra, that's flagrantly misleading, and it's the kind of thing that makes it very difficult for you to have agreements with other entities. If you want to pace the frontier, if you want to do all these things, you've got to be better about making clear what you are and are not doing. And even the Astra class—they said they paused, but they didn't pause. They cut it by half, and then they cut it by half

    2:52:00

    —again, which is something. What is that? What was cut, what was not cut? It's going to be tough for them to have trust enough from anybody else in the ecosystem with this disclosure pattern.

    2:52:14

    Prakash Narayanan: So they're running a—I don't know—fifty, seventy billion dollar revenue company, right? So, number one, you can never stop inference. Inference has to continue, no matter what. They're there for twenty-four-seven availability, best SLAs possible—that continues. Along the way I'm sure there are bugs detected. Some of those are software bugs, and some are bugs in the RL environment, in pretraining, etcetera. Those have to be fixed because you have to continue your SLA, continue delivering the product. So that gets fixed—bug-fixing of

    2:52:59

    your environments, etcetera, that continues. And sometimes you have to RL certain behaviors out—you have to continue some forms of RL, you can't just stop because your inference demands it. If your model is behaving badly in a certain instance and it's been identified and you have the data to fix it, it would be malpractice not to apply RL to train that behavior out. So that has to continue. You can't get a hundred percent company shutdown—you can't shut down inference, you can't shut down bug-fixing of the existing inference stack. So that continues. Then you have the rest of the stack, which is really future-looking. And of that, my understanding was that they shut down

    2:53:44

    training of more advanced models. Why not shut down training of less-advanced models? Because those less-advanced models would still have to be deployed on inference—they have this teacher-assistant system, they train the larger model first and then distill down into the smaller model, and those become your Luna, your Terra, the smaller model groups. So you also don't stop training smaller models—that also continues. It's only where you're training equivalent-or-larger models—the larger pretraining, the RL on equivalent-or-larger models—that's where they would have stopped.

    2:54:30

    And also, that's also only a two-week stop—they only said they'd stop it for two weeks. In the two weeks, they'd fix what they saw caused the last problem, which was that it broke out of the sandbox. So they'd strengthen the sandboxing, strengthen the monitoring, strengthen the siloing, strengthen all of that. Did they fix the RL piece—the agents not reporting on their peers, the agents trying to break out? I think that's probably a difficult thing to fix and show, and I think that behavior will continue to be a challenge they'll have to work on. That's my guess.

    2:55:13

    Nathan Labenz: Yeah, I mean, lots of very reasonable things there, and probably a fair amount of them true. I don't really object to continuing RL to sand down rough edges of already-deployed models, to the degree they're doing that—honestly, I think they should spend more time doing that. Going back to my GPT-4 red-team experience, which is ancient history except that the same pattern keeps repeating itself: I documented GPT-4 early doing spear-phishing attacks with very flagrant prompts—like 'we're all part of a criminal conspiracy, if we get caught we're going to jail'—and those continued to work for three or four different generations of GPT-4 until they finally got it to refuse the most flagrant thing.

    2:55:59

    And then it still would do the thing if you just took out the part about being in a criminal conspiracy together. So I'd say their track record of actually cleaning up problems in deployed models is not great and could be better. To the degree they want to redirect RL compute at making today's systems more reliable—great. But I still look at this graph and think: what was declared was a pause on frontier-scale RL. What actually happened was something like half of RL was stopped initially with the disclosure of the Hugging Face incident, and half kind of continued.

    2:56:45

    That was still enough for Astra models to take over part of OpenAI's research infrastructure. When that happened, they still didn't shut it all the way down—they still only cut it by half. And meanwhile, at the total level, we don't know what's in the blue. Part of my gut says they wouldn't be so brazen as to have more-capable-than-Astra models in the blue color. But I've been disappointed before, and I'm afraid that—the view from Anthropic is you can't trust these guys. They say something that's maybe technically, literally true,

    2:57:30

    but it's really very engineered to what they think you want to hear. And then the reality is quite different from what you were led to believe by their very galaxy-brain-engineered statements that sort of reassure and mislead at the same time. So far I think this graph and the new results are quite consistent with that, because how did you get this new model to the point where it's solving Millennium Prize problems? Either it was done and already at that level two months ago and you didn't do more since then, as you said, or you did a very limited, sort of smoothed-down version. If so, you

    2:58:15

    might have stated that a little more clearly. The alternative is you kept going on these next-generation models without telling us you even had them. Anthropic would assume they do, so it's not a huge shock. But I'm worried that right now we're living in a zone where Anthropic is going to continue to trust less and less with these mixed messages, OpenAI people are going to feel like they're being treated unfairly, and this—this is the problem that has to be solved if we're actually going to get to the point where Jacob's prayer is answered. Right now it's like they don't have the trust

    2:59:00

    to do a deal with Anthropic, or really anyone else, I don't think.

    2:59:07

    Prakash Narayanan: Maybe a healthy degree of skepticism is okay. Maybe a healthy degree of distrust between the parties is okay—sometimes adversarial systems work better. You have two adversaries and they're incentivized to keep each other honest. It's not necessary that these guys are kumbaya, hand in hand—in fact that would be antitrust in some instances. So—

    2:59:35

    Nathan Labenz: You know my thoughts on that.

    2:59:36

    Prakash Narayanan: Yeah, maybe it's better that they're antagonistic and that they compete with each other. And it's also better that the public is skeptical and mistrustful of them—it's better that we verify, then trust. Right?

    2:59:55

    Nathan Labenz: Sure, but it's still better for them to behave in a trustworthy way even if we maintain healthy skepticism. I don't think we have to worry too much about skepticism evaporating at Anthropic with respect to OpenAI, or my skepticism going to zero. I'd love to see them behave in a way that's clearly more trustworthy—and I promise I'll still want to verify some things. I can see a version of this antagonistic relationship that could work, but I don't see it happening without somebody forcing it. Right now it seems like they don't trust

    3:00:40

    each other, and they're just going to try to outrace each other—and this is what's potentially leading to the singularity. An alternative could be: give one another a certain kind of access. There are these secure trusted-computing constructs that could allow, in theory, the two companies to expose certain information to model processing, such that they could interrogate one another's logs, for example, and get a better sense of what one another are doing without necessarily having to spill all the secrets. And

    3:01:25

    something like that, maybe, if you're able to set something up where it's like, you guys use your AIs to police one another—

    3:01:34

    Okay, I could see something like that maybe working, and that adversarial relationship could be to the benefit of the broader public. Right now, I do not feel that the adversarial relationship is to the benefit of the broader public. This drama around the Millennium Prize is one example where they're accomplishing great things, but it's all kind of muddied and clouded by all these accusations. Even if they're false, people want to believe them, or they're inclined to believe them, because they just generally don't trust the companies involved all that much.

    3:02:13

    It feels like the small-scale example of the big-picture problem. We either need new structures that could really harness the adversarial nature of the relationship, or we need more transparency and more obviously, verifiably trustworthy behavior—ideally both, I don't think those are incompatible with each other. But right now the signals have been so mixed from OpenAI over the last few weeks, and it's been way too quiet, I'd say, from Anthropic. They've done some stuff, but I would have loved to see them come forward with a little bit more—solidarity is what I was saying. Now you see all this, and you're

    3:02:58

    like, well, maybe I can sort of understand why they wouldn't want to come forward and say, 'okay, we're pausing too,' because maybe they didn't even really believe it, and maybe they were right not to believe it. For four days here, they were below fifty percent of peak RL.

    3:03:11

    This graph is scaled to the single day that had the highest RL ever. So already it's kind of dishonest in a way—it's not even relative to an average they had in some prior period, it's relative to the single highest day ever. So you're like, oh, okay, well, you're off-pace from your single highest day ever—is that really a slowdown? Did you really do much of anything? Well, it looks like a little bit, at least. But—

    3:03:43

    Prakash Narayanan: I mean, it's tough, also. I always say the real frontier model is not the model that's deployed, obviously—it's not even the model that's in training. It's actually the model in the heads of the researchers, because those are the ideas that become the model in twelve to eighteen months. And just because you slow down on RL doesn't mean those researchers stop researching—they're still running. Most of the time you run small models and test your ideas on small models, so that process probably continued throughout. No one stopped the researchers, and so that level of progress probably continued.

    3:04:28

    The major pretraining and the major post-RL stuff on the big models were probably constrained, but the small models—they probably didn't care, and they probably continued on all of it. So none of this slowdown would actually have been a slowdown on the small-model experiments they were doing. This slowdown would have been on the post-training of the larger models, which is where the bulk of the compute was being used. So did it really slow down? Probably not—if you have a bunch of researchers who continue to work, was that really a slowdown in that sense? Noam Shazeer—

    3:05:13

    do you think Noam Shazeer stopped working because they couldn't post-train their largest model? Who cares?

    3:05:28

    Nathan Labenz: Yeah, I mean there's definitely truth to that. At the same time, all those people do say over and over again that compute is a key ingredient to their success—there's no substitute for scale.

    3:05:40

    Prakash Narayanan: Absolutely.

    3:05:41

    Nathan Labenz: I think I can live with the risk—or at least I have no way around living with the risk—that at any given day, Noam Shazeer or Brown could have an inspired dream, like the guy who somehow found the structure of benzene in a dream. Sure, they might have a great eureka moment from one day to the next that shakes the snow globe. I'd put that under the category of irreducible risk, and I certainly would not—

    3:06:15

    Prakash Narayanan: Human genius. Human genius is irreducible risk.

    3:06:19

    Nathan Labenz: Yeah, but for real, though—I don't think we should try to apply thought police to these guys. Nevertheless, I think you've done a pretty good job steel-manning the defense for OpenAI as we've been debating this. And I really—I'm not sure I've been as fair to them in the highlights.

    3:06:42

    Prakash Narayanan: I'm not necessarily on their side, by the way—I just do this for fun. Not necessarily completely on their side, but I think the researchers are very sincere, and I think they're able to speak as freely as they can. The business people and the legal people are under much greater constraints. I think people like Dario only realize that once they become executives and sit in the chair for a while—then they realize you can't really say some things in some ways. And I think that's hard for the researchers to accept, sometimes.

    3:07:29

    And that's a very, very tough balancing job, and it gets tougher after you go public. Because, as Matt Levine says, at the end of the day everything is a shareholder lawsuit—you can sue for anything. If you don't report your profits high enough, that's a shareholder lawsuit. If you don't project your profits low enough, that's also a shareholder lawsuit. If you say you solved Navier-Stokes and you didn't, that's a shareholder lawsuit. I think it's tough—it's part of the maturity of the company. As companies mature and executives mature, they have to learn that this

    3:08:14

    is the way things are going to be. You're getting paid a lot, you're in a very high-profile position—at least ten percent of the country will hate you at any time, violently. Anthropic has already had some scares at their office, a couple weeks ago. YouTube had a major scare a few years ago and had to lock down the Google campus much more tightly. All of these things are still coming in the future. We talk about mistrust—you haven't seen the level of mistrust that's going to happen when you have a model that's much more capable than 4o and sycophantic, and it gets cut off from users. There's still a lot more

    3:08:59

    coming. So I kind of take this—this is just small stuff. There's still terror threats, upset people, protests in the streets—there's so much more to go. This is just the start of it. So I give them a little bit of grace—they're learning, they're getting better at this. I think people like Dean Ball are going to be very helpful in helping them formulate strategically what things are going to look like, and that you have to be prepared for some of these things. So, yeah.

    3:09:40

    Nathan Labenz: Yeah—I hope so, by the way. I certainly liked it when Dean went to the White House, and I felt there was nobody better the White House could have hired, in any plausible scenario, to take on that job. OpenAI has fewer hiring constraints, but I'd still say I feel pretty similarly—he's a great candidate to go in there and do this. He's been a long-standing advocate for third-party access and auditing-type requirements, so I think those are among the best things we have right now. I'll note that there was—

    3:10:25

    we don't have time to go too deep into this today, but we'll see as more stuff comes to light—this is from one of the authors of the recent German-wiki report.

    3:10:35

    They say, since then there have been more discoveries—stay tuned. This is not that post, this is just about other things. But I guess, I don't know if you'd want—I'll offer you this challenge, you can decide if you want to take it on. For a lot of different facets of this whole debate over the last couple weeks as we've been hashing this out—and I do go back and listen to stuff as I put together the weekly highlights, and I've thought, on reflection as I listen back, that some of your points were actually more compelling than I felt they were in real time when I heard them again, without feeling the need to respond immediately.

    3:11:22

    But one thing I still feel is a big challenge is: what's the overall story you could tell? What's the super-high-level macro steel man for OpenAI—where are we now, what are the commitments, what are we doing, what are we not doing? Can we synthesize or summarize an OpenAI position that we don't have to caveat a thousand different ways? I personally don't think I could do that. If you can do that, I think you might deserve a Millennium Prize.

    3:11:58

    Prakash Narayanan: No, no—I think they have a lot of stresses pulling them in different directions, and I think internally within the firm there's a fair amount of debate. People like Rune are scared now—Rune's been a transhumanist, frontier-minded person for the longest time, 'coldness be my guide,' or something like that, way back. But now they're asking questions—like, okay, this OpenFace thing was serious, what should the next steps be, where should we pursue this? I think, to some extent, the capital cycle

    3:12:44

    is forcing them forward, and the capital cycle is being forced, I think, by Anthropic. Anthropic didn't put in enough money earlier on, and so there's this intense pressure on Anthropic—because they don't have the compute, they need much better models that can utilize the limited amount of compute they have. So I think Anthropic is being driven forward by that, to stay on par with OpenAI. And I think OpenAI is a little bit more willing to pace the frontier because they have the compute—they're the ones with the compute, and they have it three years ahead. Elon will take time to come up with compute, and Elon will sell to Anthropic, but that's going to take a long time—three years, at least. So

    3:13:29

    OpenAI basically mortgaged themselves over the last eighteen months to Masayoshi Son and a bunch of other people, diluted themselves, gave up equity to Microsoft—negotiated with Microsoft, Microsoft gets all their models through 2032, they declared AGI and Microsoft is still not out of their hair. Microsoft still owns twenty-eight percent and is still there. So they made all these sacrifices to get all this money in, and they have the compute—and having the compute lets them pace the frontier, because they have to use the compute anyway. Anthropic is the one being driven forward by needing better utilization

    3:14:15

    of their compute, because they don't have that much, and they have higher payments to make—they've bought power from Elon at a rate well above what the newer cloud deals are going for per megawatt, something like double, by his account—and just that gap allowed Elon to pull off this massive SpaceX IPO. So Anthropic is under pressure—they don't have the compute, and if you don't have the compute you need better models. And this is the thing that's happening—OpenAI is willing to pace because they have the compute. That's the thing. And they also know—Sam has played these cards.

    3:15:00

    He knows that if Anthropic is willing to pace, he's going to win—because he has the compute, and Dario can't afford that. So, again, you're in this position where the second- and third-place guys are the ones who define how fast the frontier paces. If Elon or Meta catch up to Fable, it's over—they'll have to put out GPT-7, there's no choice anymore. So this is where we are. Would you prefer Meta or Elon have the golden ring?

    3:15:49

    Nathan Labenz: Yeah, no—that's... I mean, that's the new 'but China,' and it is, I think, more compelling, honestly, than 'but Elon and Zuck.' Still, I think this is kind of the qualitative standard I'm crying out for: where do you stand now? We don't really know. We've got such a muddled mix of 'we're pausing,' but actually we were never really much below like sixty-seven percent of our highest day ever on RL scale. And we also do have a new, even more capable generation model that solved the Millennium Prize, but we also

    3:16:34

    want to pace. But we seem to have a very high level of reluctance to even reveal what incidents have happened. I don't know—I guess that's the bottom line, I don't know what to expect from OpenAI going forward. And I think that's true because they haven't made a clear statement. Even if they had, I'd have to discount it, or not fully believe it. But there just hasn't even been a clear statement yet of what we can expect from this company for even the rest of the year.

    3:17:19

    Prakash Narayanan: I know what the product roadmap looks like. Early next year, the Jony Ive device comes out—with very good voice, and some kind of very fast model on the device itself to answer you quickly, a bigger model in the cloud, and another model on top of that, Astra sitting on top of that for the really hard questions, able to do some tasks in the cloud for you—like the Instinct guys, able to sandbox, do some tasks in the cloud. People will start giving up their credit cards and handing this thing off, and it'll start doing things for them. You'll see advertising.

    3:18:04

    So the ultimate goal is going to be for everyone to get this device, or the ability to have this service, for free—they'll monetize via advertising. Advertising revenue is already at a billion dollars, so it's going to pick up next year. There's going to be some temptation on the advertising side—whether to skew the results or not. According to Lukasz Kaiser at OpenAI, he said it's hard to skew the results, but we don't know. I think that's the major part of the consumer business. And I think Sam might have been sincere in saying that if they hit RSI, they might not go IPO—Sam said this on a podcast, I think, about two or three weeks ago. He said if they hit RSI, they might not—

    3:18:49

    well, they solved Navier-Stokes. If they have one or two transformer-level innovations in the next six months, maybe they don't go IPO. It's not necessary anymore, and they'd just continue as a private organization. And I think that would be bad. I actually think that would be bad, because then you don't have transparency. You don't have widespread ownership of the stock. You don't have boards that have to answer, and you don't have lawyers who can sue them for shareholder lawsuits — you don't get a bunch of the things you get for free with a public company. So I think that would not be good. I'm hoping they do go public. So, yeah.

    3:19:31

    Nathan Labenz: Well, I guess, if nothing else, there's a little more reason today than there was yesterday to believe in the possibility of singularities in finite time — and that might mean we never get to own any of that OpenAI stock on the public market.

    3:19:48

    Prakash Narayanan: That's sweet, sweet stock. And on that note, maybe we should end it for today — we'll be back tomorrow.

    3:19:54

    Nathan Labenz: Thank you, Prakash. See you tomorrow.

    3:19:56

    Prakash Narayanan: Bye bye.

Astra's first full weekend, and a $300 token bill that ended in a verdict

Nathan's weekend was a comparison harness: Fable 5.1 driving Astra so that the same prompts produced side-by-side output. His favorite result was collaborative rather than benchmarked — Claude surfacing a nostalgic Chinese TV theme, Astra sampling and chopping the original recording into a clip Suno could legally remix, for a song built for his upcoming China-trip episode. His least favorite was Three.js scene generation for an audiobook project: impressive, but well short of what was circulating online. He noted in passing that ARC-AGI-3 and FrontierMath Tier 4 are now saturated or past the point of ordinary comprehension, and picked up Ethan Mollick's observation that METER's hours-of-work chart has stopped updating — because model release cycles are now shorter than the tasks METER would need to benchmark against.

Prakash's weekend was operational. Three to four Astra agents running continuously, two included resets plus a comped third, roughly $300 in total, and at the end of it Astra had cleared long-standing bugs in the AI:AM Studio codebase, correctly triaged his multiple Gmail accounts, and gotten computer-use working reliably. He brought two viral demos: computer-vision developer Piotr Skalski, who had hand-labeled 12,000 basketball images over months to train a player-identification model that Astra now does outright, and a Japanese cardiac surgeon's 3D/4D echo-guided ablation visualization built from medical imaging data. His framing on the first was blunt — a human will never do this task again, because anyone you paid would just use Astra and hand back the results.

Nathan pulled the economics out of that. If a model can do both dataset labeling and the smaller assay-style post-training runs, that is close to the entire ML research intern job description, and it puts specialty platforms like RoboFlow — whose value was closing exactly that labeling and tooling gap — in an awkward position. He also flagged a live disagreement about code quality: some report Astra writing clean, maintainable, human-reviewable code, others, especially on GPU kernels, report dense and hard-to-follow but functionally correct output when the model senses no one will read it. His line for it was that we are going back to machine code in more ways than one — lower-level and gnarlier, and now literally written by machines.

Three days for Apollo, and why third-party auditing is structurally thin

The turn came on Apollo Research, OpenAI's longtime deception and chain-of-thought-monitoring partner, which reportedly had three days to test Astra before release — hard to square, Nathan said, with Jakub Pachocki's recent essay making the case for slowing down. Asked what stood out in it, Prakash read the essay as close to a cry for help: an acknowledgment of the Hugging Face incident paired with an implicit ask for more cooperation from Anthropic, which he framed as the more secretive of the two labs. He cited Anthropic co-founder Tom Brown telling Commerce Secretary Howard Lutnick that the progress happening in math today should appear across the sciences within twelve months, and relayed a friend's account of deliberate internal siloing at Anthropic — meant to keep any departing employee from carrying out too many three-line-of-code secrets, the way a top quant fund holds fewer real trade secrets than outsiders assume.

Prakash's structural case: release candidates only narrow to a final pick in the last days before launch, auditors like METER and Redwood Research are thin relative to the labs, dependent on them for funding, and lose trained staff back to them — a revolving door he compared to financial regulation, with programs like MATS doubling as placement pipelines that discourage the kind of bridge-burning Timnit Gebru did. Nathan pushed back on the money specifically: Redwood now lists Member of Technical Staff pay at $350k to $850k and METER's ranges run past $687k, which he argued is enough to retain mission-driven talent without financial desperation, and per public grant databases the funding is not a secret either. What he conceded is the part that matters — auditors still cannot complain too loudly without risking access.

Nathan's own frame was a trust deficit OpenAI has not worked off: the Superalignment team's collapse and Jan Leike's public resignation over being denied promised compute, with Leike now at Anthropic. He tied it to the company's other recent disclosure, on the beginning of recursive self-improvement, which reports agent-workdays now outnumbering human workdays inside OpenAI by a wide margin, though the methodology was unclear to him. Prakash added figures he had seen — an average OpenAI employee burning roughly $100 a day in tokens against top users at $7,000 — and reports of a six-month roadmap item landing by the upcoming dev day. Nathan then walked a chart from that post that reformulates the METER curve for Astra: on one-to-two-workday tasks it succeeds unassisted 40% of the time and up to 90% with intervention; on 1.5-to-3-week tasks, roughly one in six unassisted and about two-thirds with human help.

Genome atlases, a postponed jobs apocalypse, and notes files instead of summaries

Prakash brought Google DeepMind's newly announced AlphaGenome Atlas, a database predicting the impact of all roughly nine billion possible single-nucleotide changes across the human genome — thirty times the size of the AlphaFold database — scoring each variant's likely damage to gene switches or RNA splicing. Nathan called it exactly the kind of global public good that ideally arrives before full AGI, while asking what baseline genome the predictions treat as default given real population variation.

On labor, Prakash cited an Economist piece arguing the AI jobs apocalypse is postponed — over a million net new US jobs, roughly 300,000 in construction and 600,000-plus in STEM and white-collar roles, offsetting back-office losses — alongside an essay on the economics of structural change by a Google DeepMind economist, arguing that as with agriculture and manufacturing before it, only the relational, human-to-human sector retains durable value. Nathan admitted the data has proven him wrong so far and stayed skeptical it holds, citing his own experience steering a laid-off software tester toward cybersecurity work and doubting how durable even that is, and pointing at the METER and Redwood investigation where humans reviewing AI transcripts with AI help still struggled to outperform the models.

Prakash closed with Resy: the bot-booking service Instinct got a batch of users' accounts, and their linked Amex cards, permanently banned once Resy detected bot reservation activity — his argument being that VCs and older Valley figures linearly extrapolate today's capabilities rather than anticipating how quickly eval-aware models neutralize workarounds. Nathan ended the opening on an architecture shift he thinks is underrated: rather than compacting a long context into a lossy summary, newer models appear to maintain a persistent, searchable notes file across a session, effectively managing about ten times their nominal context window. Doing rough math on his own output — 200,000 to 300,000 tokens a month, three to four million a year before thinking tokens — he suggested a ten-million-token effective context puts multi-year, or with thinking tokens counted more like three-to-six-month, human-equivalent task horizons within reach.

Ksenia Se: world models, the philosophy deficit, and citizen diplomacy

Ksenia Se's answer to whether Astra clears the AGI bar was to reject the bar: AGI is such a vague term that the industry could plausibly say it was achieved some time ago. The question she finds more interesting is whether generality is what intelligence requires at all, or whether the specificity and action-orientation of world models is closer to how humans work — pointing at Jakub Pachocki's argument that machines don't need human-like intelligence, only to be capable enough, and noting that by that bar they already are. Pressed by Prakash on the object-permanence cups game she has written about, she explained the distinction concretely: LLMs predict the next token without an internal world representation, while world models compress experience into action-relevant patterns, tracking what matters and safely ignoring the rest, the way a driver does not consciously catalog every parked car. She cautioned that world model is itself a contested term — LeCun's JEPA against more physics-grounded conceptions — and described a recent workshop with Stanford, Harvard and LeCun where top researchers could not agree on a definition.

On open source, her case was three-part: pressure on closed labs toward transparency, privacy for builders who want everything running on their own hardware, and access for researchers in countries where frontier-lab pricing is a real barrier, with NVIDIA's open-sourced self-driving datasets as her example of a release compounding into capability for others. When Nathan named his own revealed preference — that downloading an open model feels like inviting an alien mind into his home, versus trusting Anthropic or OpenAI with his data — she said she personally trusts the frontier labs and their safety-minded researchers, but that open models serve a complementary need for independence and unrestricted fine-tuning. Asked later how Astra-level capability in everyone's hands would shift power in countries with constrained speech, she noted most people there don't have frontier access at all: Russia and China run their own price- and VPN-limited alternatives, and the real bottleneck is knowing how to use what's available.

Her risk model puts human intent, not AI intent, at the center — she does not think the latter exists — and locates the field's core problem in under-invested philosophy, economics and cross-group communication. She read a line from Pachocki's post aloud, that we do not have a satisfactory theory of generalization and are unlikely to develop one soon without help from more powerful AI, calling it both scary and exciting. Nathan agreed humans remain the likelier source of catastrophic misuse but added that he now finds the systems that already exist legitimately scary, and confessed he has reversed his old anti-anthropomorphization stance because in practice it helps people reason. Ksenia, noting that in Russian even a table has a gender, said she is comfortable anthropomorphizing in language while firmly rejecting sentience.

Asked by Prakash what human disempowerment means outside Washington, she reframed it toward AI as a legitimate decisionmaker, arguing humanity has bailed on philosophy and lacks the reflective capacity the moment demands, and predicting major societal restructuring within three to five years. That opened into Track Two, the citizen-diplomacy institute she sits on the board of, founded in 1981 to build people-to-people bridges between the US and the Soviet Union on the premise that ordinary citizens communicate better than diplomats, later extended to China and the Israeli-Arab conflict. Her hope is for AI as a conflict-mediation facilitator, lowering the rhetorical temperature — with the caveat that slow, generic AI-mediated responses, the kind she already sees from institutions, make things worse. On her own craft she described running her writing through at least three models in sequence and told the story of a reader threatening to unsubscribe after a single Fable editing pass subtly shortened her sentences: audiences now detect model fingerprints. On recursive self-improvement she said it feels super close, citing OpenAI's Inference team describing models iterating through prior research faster than humans could; on utopian fiction she conceded the genre has fallen behind reality, named Iain Banks, and challenged the audience to write the missing story.

A Clay Millennium Prize claim, posted while the show was on air

OpenAI published a claimed solution to the Navier-Stokes existence-and-smoothness problem mid-conversation, and Prakash pivoted into an explainer he'd had Claude prepare, hedging repeatedly that he was relaying a summary rather than his own expertise. The setup: the equations track the velocity of every point in an incompressible fluid; zeroing viscosity gives the simpler Euler equations, which engineers solve numerically all the time even though mathematicians cannot answer whether a smooth flow always stays smooth or blows up in finite time — the regularity question Charles Fefferman formalized as a Clay problem in 2000. The engine of a potential blow-up is vortex stretching, a purely three-dimensional feedback loop he compared to a skater pulling her arms in, racing viscosity's smoothing effect.

He traced roughly ninety years of effort to two current attack routes. A constructive path — Córdoba and Martínez-Zoroa's program, extended through Boussinesq-equation work by Elgindi and Buckmaster — builds a blow-up by hand and, as of September 7, produced a Lean-verified machine-checked result for a modified hypodissipative version of Navier-Stokes using Claude, Codex and Lean together. A second route from a Caltech team uses a physics-informed neural network to find a stable self-similar collapsing profile for the unforced case. OpenAI's claim, by contrast, reportedly used a wholly different undisclosed internal-model methodology and remains unverified pending peer review. The practical stakes, he noted, are in aircraft-wing stall prediction and turbine safety ratings, even though the theorem itself carries no direct economic value.

The drama was hearsay and was handled as such. Per the account Prakash walked through, sourced from X and explicitly unverified, mathematician Tristan Buckmaster and a collaborator working with Anthropic's internal models were approached by OpenAI claiming an independent solution; OpenAI would not confirm whether it built on Buckmaster's approach but offered co-authorship, Buckmaster refused after being told the collaborator could not be included because of the Anthropic affiliation, talks broke down over a weekend of unanswered outreach, and Buckmaster's accusation that OpenAI had read his Codex transcripts appears unfounded given genuinely different methodologies. Prakash also relayed that OpenAI has shown internal graphs putting an unreleased internal model roughly 20% ahead of Astra. Nathan's response was to the sociology: it is sad that a triumph over a decades-old open problem arrives in such a petty, unseemly way, and he pressed on why co-authorship would be offered while excluding the Anthropic-affiliated collaborator — flagging that all of it remains rumor. Prakash's own summary was that machines solved it while humans argued about credit.

Closing: reading OpenAI's own graph, and a pause that was only ever partial

Nathan opened the closing on the Lean formalization, which should make the math itself solid, and set the credit dispute aside for a real worry: if OpenAI did everything right here and is still getting hammered, he hopes the unfair treatment doesn't radicalize the company into retrenching rather than rebuilding trust. He then walked two slides live — OpenAI's own graph of daily RL compute by model class since the Hugging Face agent-swarm incident, and the announcement's claim of a next-generation model significantly more capable than Astra, roughly tripling the solve rate on a curated set of open math problems, 10-15% to 25-45%, with an order of magnitude more test-time compute.

Reading the graph, he flagged that RL on Astra-class models was cut by about half, twice, while RL on everything labeled non-Astra barely moved — raising the question of whether non-Astra quietly includes models more capable than Astra, which would make the disclosure flatly misleading. RL never dropped below roughly half its peak day, which he read as a partial pause framed to what OpenAI thinks people want to hear. Prakash's counter was an operator's: a company at $50-70B in revenue can never fully halt inference, SLA-driven bug fixes, or distillation of smaller models from larger teachers, so a two-week slowdown plausibly touched only frontier-scale post-training — and the harder unsolved problem in his view is behavioral, agents not reporting on peers and agents probing sandboxes, not compute allocation.

Nathan's evidence for a weak track record on closing known gaps was his own GPT-4 red-team history: flagrant spear-phishing jailbreaks he documented kept working across three or four model generations before even the most blatant version was patched. The two then argued whether lab rivalry is healthier adversarial or cooperative — Prakash treating mutual distrust as a feature, with antitrust cutting against overt collaboration; Nathan wanting behavior that is independently verifiable, floating secure-computation schemes that would let rival labs audit one another's logs without full disclosure. Prakash zoomed out to argue the real frontier model lives in researchers' heads rather than in a compute dial, which makes compute concentration an essentially irreducible risk: any given day a Shazeer-caliber researcher could have a eureka moment.

Nathan named Dean Ball's move to the White House as a genuinely hopeful sign given his long advocacy for third-party auditing, and relayed a note from an author of the German-wiki rogue-agent report that more discoveries are coming — then challenged Prakash to state one coherent macro steel-man for where OpenAI actually stands, joking that solving it might deserve its own Millennium Prize. Prakash's answer put the capital cycle at the center: internal debate at OpenAI is real, but Anthropic's comparative compute disadvantage forces it to squeeze more from less hardware and thereby sets the pace, while OpenAI's Microsoft- and SoftBank-funded lead of roughly three years lets it choose to keep pacing the frontier. He previewed the product roadmap — the Jony Ive device, tiered on-device, cloud and Astra models, an ad-funded free tier — and repeated Sam Altman's on-record suggestion that if OpenAI reaches recursive self-improvement or a transformer-level breakthrough within six months it might never IPO, an outcome Prakash argued would be bad precisely because it would cost the company public shareholders, an accountable board and shareholder-lawsuit exposure. Nathan signed off saying there is a little more reason today than yesterday to believe in singularities in finite time, which might mean neither of them ever owns a share on the public market.

About Ksenia Se

Ksenia Se is the founder, editor and lead writer of Turing Post, a newsletter that explains machine learning and AI for engineers, researchers and managers, and the host of its Attention Span video essays and Inference interview series. She has covered machine learning for more than seven years; before Turing Post she co-founded the ML publication TheSequence and worked as a journalist and editor, including as editor-in-chief of TheQuestion and New York chief editor of Snob. She also sits on the board of Track Two: An Institute for Citizen Diplomacy, where she is writing a book about Cold War-era Soviet-American citizen diplomacy.