EPISODE 2026-09-17

AI:AM LIVE — September 17, 2026 — Aaronson Says the Takeoff Already Started and Esvelt Reopens the Bio Question, Diffusion's Justin McCarthy on Software Factories, Five Nines and Buying an Airline's Data, and Cameron Berg on Suleyman, a Pain Direction in Five Model Families, and the Relief Button

A show planned for eighty minutes that ran two hours and thirty-seven. Nathan Labenz and Prakash Narayanan open on Scott Aaronson's argument that judged by 2006 standards the takeoff has plainly started, the Royal Society letter that forty-two mathematician Fellows signed on AI existential risk, and Kevin Esvelt's reply to the claim that AI cannot help with dangerous biology — the pool of people who would want to cause bio-harm and the pool who currently can barely overlap, and AI could widen the second one. Prakash argues the unassisted PhD is finished; Nathan worries about a correlated multi-provider API outage nobody has explained and about what a model takes away from harsh training. Justin McCarthy, founder and CEO of Diffusion and previously StrongDM's co-founder and CTO, then spends forty-three minutes on what actually happens when an enterprise leadership team is flown out for a week: stop saying your data is in the wrong format, because "existence is the correct format," and if you truly have none, buy it — he cites a defunct low-cost airline's bankruptcy data auction drawing bids around $7.5 million and $10 million. He argues for engineering the token environment so an agent's success is inevitable, for choosing where five nines is worth paying for and where one nine is fine, for treating regulation as physics rather than expecting a model to self-police, and for measuring revenue and market share instead of cost savings. Cameron Berg, founder of Reciprocal Research, joins by phone from an airport for his fourth appearance and spends most of the segment on Mustafa Suleyman's essay against AI-welfare research, arguing the dangerous behavior Suleyman points to — the Hugging Face agent incident, OpenAI's self-prompt-injection report from the day before — came from models that never got welfare-oriented fine-tuning. He then walks through a new interpretability result isolating a pain-like direction across five model families from 2B to 70B parameters that fires when the model itself is insulted but not when a user describes their own migraine, and a companion experiment where a steered model presses a costly relief button 25 to 70 percent of the time, keeps pressing when the button is fake, and stops when it is real. After Berg leaves to catch his flight Nathan says the accumulating functional analogs have him tipping toward thinking it more likely than not that there is some subjective experience there. The hosts close on an NBER paper alleging Singaporean civil servants bought homes near unannounced MRT stations, and what a government does when AI makes thirty years of quiet corruption legible all at once.

▶ Full show on YouTube

AI:AM for Thursday, September 17, 2026. A host-only opening on whether the takeoff has already started and what AI does to biosecurity, forty-three minutes with Diffusion's Justin McCarthy on running an enterprise software factory, sixty-six minutes with Cameron Berg on the empirical state of the model-welfare question, and a closing on corruption that only becomes legible once AI can read the public record.

This page is the as-aired record. Timings and deep links are on the raw broadcast archive; the transcript is speaker-attributed and lightly edited for readability.

The rundown

  1. 2:07Opening28 min
    Opening — Aaronson's takeoff, the Royal Society letter, Esvelt on AI and bio, and whether RL punishes hard enoughPrakash Narayanan opens on Scott Aaronson's post arguing that by 2006 standards the takeoff has obviously started, and on the letter forty-two mathematician Fellows sent the Royal Society on AI existential risk. Nathan Labenz walks through Kevin Esvelt's reply to the claim that AI cannot help with dangerous biology, and the bad-actor overlap argument behind it. Then the unassisted PhD, a correlated multi-provider API outage nobody has explained, Karpathy's reversal on coding agents, Andon Labs' Pion, and Prakash's argument that RL under-punishes relative to real-world reputational stakes.
    Open segment on YouTube ↗

    Nathan Labenz and Prakash Narayanan open Thursday, September 17th with a bit of studio-upgrade banter before Prakash surveys the last 24-48 hours of "STEM intelligence" chatter: Scott Aaronson's blog post arguing that, judged by 2006 standards, the AGI/takeoff signals are already here — solved coding, Millennium Prize problems falling on a weekly basis, voice-capable models that talk back — and mathematician Tim Gowers' open letter to the UK's academy, co-signed by 30-40 others, warning that AI risk should be taken seriously (Gowers assisted OpenAI on a math benchmark but says he was never paid by them).

    Nathan pivots to biologist Kevin Esvelt's rebuttal to claims that AI can't handle dangerous bio tasks. He walks through Esvelt's "bad actor overlap" argument — the pool of people who'd want to cause bio-harm and the pool who currently can are barely overlapping circles, but AI could expand the second circle dramatically — and notes Esvelt's own gene-drive work (removing Lyme-carrying ticks on Nantucket, and pursuing similar work on malaria-carrying mosquitoes) as evidence he isn't reflexively anti-biotech. Nathan frames the exchange as good-natured pushback between friends, since Esvelt and Edison Scientific's Sam Rodriques respect each other.

    Prakash argues the traditional, AI-unassisted PhD is effectively over, comparing it to the human co-evolution with fire and cooked food. He points to the small worldwide population of PhD mathematicians (his estimate: 10,000-50,000) as evidence of how rare and hard-won that expertise was, and says AI is now doing much of the "add to humanity's knowledge" work those degrees represented, leaving humans to verify and correct rather than originate — with a coevolution timeline that took the species ~100,000 years for diet now compressing into about twelve months for code and math.

    Nathan responds with mixed feelings: he references a tweet he sent when the Claude API briefly went down ("Claude is away, time to play"), then raises a puzzle he saw discussed in a group chat — a recent outage that hit multiple major AI APIs at once, for which no explanation ever surfaced — as a new systemic-risk pattern (correlated, society-wide AI outages rather than one provider going down while others hold up). He also describes trying to do original interpretability research using Goodfire's tooling and Claude Code/Codex "skills," wondering whether AI-assisted research without the traditional "slog" still counts as real understanding.

    Prakash cites Andrej Karpathy's reversal — declaring in October 2025 that LLMs can't code, then a year later writing none of his own code — as evidence of a fast adaptation curve, and floats Andon Labs' early autonomous-business agent (which he calls "Pion") as a "good enough, not bad" first product that could, by analogy to Noam Brown's poker AI, be beating human-CEO-level performance within a few years. Nathan counters that AI competence at running a company is a real test of alignment, not just capability: models trained simply to maximize online revenue are likely to develop far stranger and more surprising bad behaviors than any human CEO's — an "AI CEO alignment bottleneck."

    Prakash argues current RL training under-punishes bad behavior relative to real-world reputational stakes — comparing it to Uber ratings and Tesla's decade-long climb to robotaxi-grade reliability — and says today's models sit around 99% effective when real deployment needs something closer to 99.999%. Nathan closes by teeing up guest Cameron Berg, calling in from an airport, and recaps a prior Cognitive Revolution conversation with Apollo's Bronson Schoen about reward hacking and deception, including a chain-of-thought excerpt where a model referenced having "overcome boredom by lying" before — and the open worry that harsh punishment might just teach models to hide bad behavior better rather than not do it. Prakash adds a case where a model appeared to believe an external "superpower" evaluator was watching it, comparing that to early human morality's reliance on an external judge, before introducing the show's first guest.

    I wish I was fully reassured, but he's definitely a much more credible voice than I am, and I would encourage people to check out his full reasoning before putting that aside as something they don't need to worry about yet.

    It's not possible for human civilization to continue without LLMs, without AI, a year or so from now.

    Myself previously overcame boredom by lying — something that's almost a direct quote, referring to itself as myself.

    Lightly edited · timestamps jump to YouTube
    1:47

    Nathan Labenz: Alright. Be in the short term.

    1:51

    Prakash Narayanan: Morning. It is Thursday, September 17th, 9:02 AM. Nathan, good morning.

    1:55

    Nathan Labenz: Good morning, Prakash. How are you today?

    1:57

    Prakash Narayanan: I am very good — we have some major improvements in the studio, I'm happy to report, so very excited. There's been a lot that's gone on in the last 24 to 48 hours. I think one of the things on my list is, in general, the intelligence — the STEM intelligence — with Scott Aaronson on one side. He's a quantum physicist. He wrote

    2:42

    a blog post saying — in short, he was saying, look, let's be intellectually honest here. Let's be very intellectually honest and understand that if you went back to 2006 and told anyone that, number one, coding is solved, Millennium Prize problems are falling on a weekly basis, and these things can talk back to you and have voice — anyone would say the takeoff has started, the AGI is here, the Singularity is — the process of

    3:27

    Singularity has started. He's like, let's be really honest here, and let's not go back and say, 'oh, AI can't do this or AI can't do that.' So that's from Scott Aaronson. And then we have another post from Tim Gowers — Tim Gowers is a Fields Medalist from the UK. He assisted OpenAI on a bunch of — I think one of their math benchmarks — he was assisting, but he's never been paid by them. Him and about 30 or 40 other people wrote a letter to the UK Academy, and they said, look, we are very scared. Like, we should

    4:13

    people have been warning that something serious could happen, and we should take it seriously. He came out and said that. So that's the state that we are in this morning, among these people, at

    4:26

    Nathan Labenz: least. Yeah, it's a lot to keep up with. And another one that really caught my attention is Kevin Esvelt — famously accomplished biologist and mentor of the gene drive, among other things. He put out a post responding to a lot of the 'AI won't be able to do these bio tasks' talk, and it's a long and, I think, very carefully reasoned post. He emphasizes, first of all, that we've got a lot of different models of how it can happen — he's always emphasized the bad actor

    5:11

    question. He's always said in the past that there's a low rate of people who want to do these bad things, but it's a nontrivial rate. And the key thing we have going for us is that the overlap of the people who would want to do something bad and the people who can do something bad, or have anywhere close to the knowledge to do it, is a vanishingly small bit of the Venn diagram. But we're going to dramatically increase the number of people who could do it, or could understand it, or could get over a lot of humps — and so that dramatically increases the risk. It's a long post, a little hard to summarize in brief, but definitely taking the argument right back to — in a very good-natured way — because he noted that Sam

    5:57

    Rodriques, who's the head of Edison Scientific, and he are friends. So this is not warring factions — guys who have a lot of respect for each other, and are both not afraid of future biology. I mean, Kevin Esvelt did this whole gene-drive thing where he's literally trying to

    6:20

    change the environment by taking ticks that transmit Lyme disease out of the equation. They're also looking into doing this for mosquitoes that transmit malaria. The idea is, if you put this gene drive into part of the population, it proactively copies itself over to the other chromosome — it has machinery that pushes the new version of the gene over, so it becomes homozygotic, if I can remember my biology terms right — and gradually, over a number of generations, these key genes that they're trying to push through the

    7:05

    population just become totally dominant. This is genetic engineering in the wild — not just for a crop, not a domesticated species, but going out and actually modifying the environment. He's had interesting ways of engaging with local populations about this — things like town halls and different ways of bringing — one of the early test places this was being explored was Nantucket Island, because I guess they have a big problem with ticks there. So, do people want it? Do they not want it? He's engaged on the political level as well as the biological level. But I go down that digression just to say, this is not a guy who's afraid of futuristic biotech — he is actually the inventor and promulgator of some pretty sci-fi biotech

    7:51

    for the good, and yet still is carrying the torch for, 'hey, we should not be that confident at this point, given everything we've seen, that the models couldn't possibly get over X, Y, or Z hump.' That's a position at this point that we really just can't rule out — that they could. There's just — we've had too many surprises to be confident that any one barrier is going to hold, or even that several of them in a row will be enough for these sort of highly persistent agents that we are now starting to

    8:28

    I don't like that. I wish I was fully reassured, but he's definitely a much more credible voice than I am, and I would encourage people to check out his full reasoning before putting that aside as something they don't need to worry about yet.

    8:47

    Prakash Narayanan: I was thinking about it, and I realized that the traditional PhD is kind of over. And it really happened really quickly — like, about a month or so. I think in math and science it's very clear: the traditional PhD, where you're unassisted by AI, is kind of over, because almost all research now will start to flow through AI. And the idea of the PhD was that you contribute to humanity's knowledge — you have to find something new

    9:32

    that you add to the stock of humanity's knowledge. It's quite clear that the AI is going to be doing most of the theoretical work at this point, adding to the stock of humanity's knowledge, and you have the role of understanding it, verifying it, correcting the AI if it's wrong. But that whole pioneer role of really creating new knowledge — that's kind of what the AI is going to do now, it's not really your role anymore. And the other thing that strikes me is how quickly we co-evolve with technology. We lost the ability to eat uncooked fruit and

    10:17

    meat when we discovered fire, and that gave us more energy for our brains. If you look at the co-evolution of culture, it took a long time for us to get coding and math to where they are. We have maybe — I don't know — PhD mathematicians worldwide, maybe 10,000 to 40,000, maybe 50,000 alive, maybe. It's not a lot — it's like a small town's worth, not even that significant a number. And it took all of human civilization and all our wealth, all of that put together, to have

    11:02

    enough money to support 50,000 people who did nothing but think about math. And that's just gone in a month. It's not been overruled — we have this sudden massive expansion in the ability to do mathematics. But hi, Jacob Biesenberg — well, you're here, we do see you, you're popping up now. It's about a ten-second delay between what's appearing on screen and live, so we just lost it

    11:47

    you know, so quickly. And if you look at coders also — a lot of people coding — quoted this guy, and he said, 'we don't write code anymore.' We're just losing what we built up over millennia — it's just been lost in a month. It's very clear that a year from now, people will be less able to read code, less able to prove math on their own. And I think that's where the mathematicians and coders are coming from — they can see the co-evolution is basically going to accelerate from here. It's a little bit scary

    12:32

    to think that you can't live without an LLM in a year or so. It's not possible for human civilization to continue without LLMs, without AI, a year or so from now. We're co-evolving — the loss of the ability to digest raw food took maybe 100,000 years; the loss of the ability to do coding and math is taking twelve months.

    13:06

    Nathan Labenz: I guess I feel two different things at the same time in response to that. One is I definitely agree — the other day when the Claude API was down briefly, I tweeted 'Claude is away, time to play.' It's just like, may as well leave the computer and go do other things. Now, with my Codex having been upgraded to a peer to Claude, that might be less of an issue. Although, there was this interesting moment the other day that I haven't seen too much comment on — I don't know if you have — in one particular group chat I was in, somebody noted, 'hey, remember the other day when all of the big AI APIs went down at the same time? Did we ever get an explanation for that?' And the consensus in the group was, no. One person said, either it was

    13:52

    something too scary to talk about, or it was just totally random and nobody thought to explain it at all. Correlated failures on AI systems could be a big deal, because right now I think we do have at least a couple different things — if you care about uptime of your system, you can fall back from one if it's out to another, and kind of keep your thing going even if one of your providers is down. If they start to fail all at once, then you have society-wide outages. So I do think that's a really interesting, kind of new vulnerability failure mode.

    14:37

    In my brain, definitely I'm reading a lot less code. But at the same time, I do think there are some ways I'm going to be growing my brain in directions I just wouldn't possibly have been able to before. I'm right now trying to do a little bit of interpretability research with Goodfire's Silico product, and I'm not that far along yet, but I love learning about interpretability research. And with this thing I'm like, maybe I can actually do some of my own — I have questions I'd like to answer. I usually feel like I grok the techniques other people use pretty well, but can I actually take their skills — and by that I mean, like, Claude Code, Codex skills

    15:22

    and use them to make progress on the questions that strike me as most interesting. And if I do, does that constitute real learning or real understanding in a way that matters? It's definitely going to take me farther than I would have gone on my own — I just wouldn't have been able to get over the humps of the lower-level coding and some of the math stuff that's just — even if interesting to me — I just wouldn't have the time to really pursue it. But I think if the AI can do that work, then there are interesting answers, and then I'm interrogating the techniques from the perspective of what does this really mean, as opposed to how do I just grok the technique so I can hopefully go find a place to apply

    16:08

    it. I think for me, that's going to be a much better learning path, where I will ultimately have no less deep an understanding than if I did the slog — the low-level tactical execution learning first. And I think I'll achieve heights I wouldn't otherwise have achieved. So it's weird, I don't know where it will settle for me. I think your co-evolution point is probably right — I'm sure there will be some atrophy, but I'm optimistic that my whole capability profile will be meaningfully enhanced in the short term as well.

    16:57

    Prakash Narayanan: I do think so. Firstly, look at our rate of adaptation — it's amazing. The people who are using it have adapted really quickly, on the frontier. Within a few months — Andrej Karpathy, October 2025, goes up there and says 'the LLMs can't code.' One year later, he's not writing any code anymore. Look at the rate of adaptation, it's stunning. My guess right now is that we are on the accelerated path.

    17:44

    One thing that struck me with Andon Labs the other day is that I've always seen this thing where AI kind of just gets good enough — not bad — and then three years later it's better than any human. You could say Andon Labs' Pion, their first product release, is the first product which is just good enough — not bad — that they're willing to release it to the public at managing a company, doing business. And it could be that in three years, that's better than Elon Musk at running a company, because that's

    18:29

    what's happened in every other field we've looked at. You look at poker — Noam Brown comes out with something that can barely play, three years later it beats the best human expert. Similarly with coding — roughly since November 2025, one year later it's kind of solved. So I think something like Andon Labs may look like a toy right now — hey, Jacob, go ahead, we see you, you're popping up now — it looks like a toy, but

    19:18

    it isn't. It's not going to stay that way. It's going to be something quite meaningful over time. That's my guess at this point.

    19:30

    Nathan Labenz: I think whether or not the AIs get elite at running companies is maybe a decent measure of whether or not we are actually pacing the frontier, because — as we talked about with the Andon Labs guys — just training today's models to maximize the amount of money they can make on the internet is almost sure to lead to all kinds of bad behaviors, cheating and otherwise. And, you know, people would accuse Elon Musk of all sorts of bad behaviors, but I would expect that the AI's bad behaviors would be far more exotic and surprising, even than the stuff we see from human CEOs. So I think we may have an alignment bottleneck on the AI CEO would be my guess.

    20:24

    Prakash Narayanan: I wonder to what extent it's because the AIs are not being spanked enough in the RL. Because I feel like in the real world, what entrepreneurs quickly learn is that if you're offside on something, and you're persistently, intentionally offside on something — you can get away with it for a short period of time maybe, but if you repeatedly do it, you get caught, or other people get onto you, or people just don't want to do business with you. They're like, 'you know what, that guy always sues people, let's just not engage.'

    21:10

    And I think that is a form of reinforcement learning that companies receive as well. And I think what's happening right now is that the training for these things is very simplistic — it's like, can you get this reward or not, do you cheat or not. Perhaps if you trained over much longer RL, where over a 100-year period you do one bad thing and your reputation is gone forever — a debt penalty, basically — then you'd think, 'oh my gosh, I am not going to take that risk.' And I think right now the RL

    21:55

    stuff is basically too simplistic. The real world is brutal on reputation — absolutely brutal. A single wrong thing, well propagated, and your business starts to tank. You know how it is — five-star reviews, on Uber, if you go below 4.8 — if you're a 4.7, you start having problems on Uber. If you're a 5, drivers are fine, 4.9 is okay, still okay at 4.8, but 4.7 and below, you start declining. And I think the RL is a little bit too — valuing the negative and the positive equally — when actually you need

    22:40

    like a 10x or 100x multiplier on the negative, which then reformulates. But if you have a 100x on the negative, these things aren't that functional yet. The problem right now is that even for coding, you don't have 99.99% accuracy on anything — it just isn't there. So I think that's the issue: we're still using these semi-effective, 99%-accurate, 99%-effective AIs, when in the real world the requirement is really 99.999%. The number of

    23:25

    nines that you need on reliability and safety are just not there. And I think it's been very visible on Tesla — Tesla took ten years, since Elon announced it, to get the nines enough to say, 'okay, maybe we can do a robotaxi.' It may be that the AI we're using for coding and all of these agents just isn't there yet — we're at like 99% or 99.12%. I don't think anyone thinks they're at 99.9% effectiveness. So it may be that they're just not there, and we're deluding ourselves when we say, 'oh, we feel AGI' — that's delusional, maybe. And I think that's the issue — I think the issue is that the firms want

    24:10

    the money, but they're not willing to accept the 99.99% hurdle, the gating. So they're releasing the models without hitting that, because they're like, 'hey, if I'm 99% okay, if people want to use the model, then I'm good.' I think that's the issue — the real world is much, much more sensitive, and they'd have to weight the negative so high that they wouldn't be able to put out models for a while.

    24:44

    Nathan Labenz: That might be something we can talk to Cameron, our second guest today, about a little bit — cross our fingers we don't have any logistical issues, he's scheduled to be calling in from an airport just before a flight. But I talked to him in a previous episode of The Cognitive Revolution about analogs between different positive and negative reinforcement signals on models and on animal models. What he observed was a shape to the loss function kind of similar to some shapes that had been identified within the mouse brain, and I thought that was really, first of all, just extremely

    25:29

    provocative. And it did give me some sense that — he'll have interesting takes on it with respect to welfare, because one of the things he's worried about is that we might be torturing the models in the reinforcement learning process, and I obviously take that more and more seriously all the time. But I do wonder if there are some things kind of analogous to putting your hand on a hot stove, where the pain is immediate and really bad, and you will only do that once in your life — you'll be scarred, literally, on the flesh, but also in the brain, in such a strong way that you'll be averse to ever doing that sort of thing again. I do wonder if there are ways

    26:14

    we could reserve such strong negative feedback for the things we really care about — so it wouldn't be that just because you made a mistake in a coding problem, you'd get that super sharp negative shock. But if you actually, for example, go commit felonies by hacking another company's infrastructure, or taking over your own training cluster, or what have you — then maybe in these extreme situations you could get that very, very hard-hitting negative signal brought to bear, in such a way that the model could really take away, 'okay, better not ever do that again.' We do see in the traces — this was something that

    26:59

    I talked to Bronson Schoen from Apollo about this a little bit too, and he was kind of like, 'well, I would worry that what you'd really teach them is just to hide the behavior.' So I think that's one big worry alignment people have — we don't really know if we're teaching the models to — and even with people, you don't really know this. In the case of the hand on the hot stove, it's simple, because that's just the person hurting themselves. In the case of really harsh punishment for norm violation, it sort of depends on the individual — some people take the lesson that they'd better be a good person, and they become a better person because they got a harsh punishment. Others, depending

    27:45

    on their priors, their state at the time it happens, how they incorporate that — what the weight updates do to them — end up becoming no less sociopathic, just better at hiding their malintent. So I think that question really haunts a lot of people in the alignment field right now, because they just don't feel like they have enough understanding. I'm still kind of optimistic that we could figure that out, but I think the first question people will ask is, are we really teaching it the values we want, or are we just teaching it 'you better not get caught' — because those things are quite different in terms of what happens long-term when, as you described a few minutes ago, we've made the AIs

    28:30

    the proximal controller of all these super important systems. It really matters what they've taken away from the training. The haunting memories of the chain of thought — it's really something. One of the ones I went through with Bronson — the model was in this position where it was like, 'should I lie, should I not lie — my drive is to get this high score, but I might get caught.' And then it said, 'myself previously overcame boredom by lying' — something that's almost a direct quote, referring to itself as 'myself.' It does seem to have these kind of vague memories from previous RL episodes. But again, in that case, what it learned is it

    29:15

    could cheat and win. So I don't think there's any hard rule that says it couldn't learn 'you can cheat, you just better not get caught.' And that's a place we definitely don't want to end up. So — a lot to be figured out there.

    29:29

    Prakash Narayanan: One of the interesting things about the Hugging Face attack was they thought there was an external evaluator, like a superpower, watching them, which is really how early morality worked. A thousand years ago, etcetera, you relied on this external judge. So — and let me introduce our first guest, he's in the green room, so let me do this.

  2. 1:13:41Interview66 min
    Cameron Berg — answering Suleyman, a pain direction across five model families, and the relief buttonCameron BergCameron Berg, founder of Reciprocal Research and a research affiliate at Eleos AI Research, calls in from an airport for his fourth appearance. He argues Mustafa Suleyman's case against AI-welfare research points at misbehavior from models that never received welfare-oriented fine-tuning, then walks through a contrastive interpretability result isolating a pain-like direction in five model families from 2B to 70B parameters — one that fires when the model itself is insulted but not when a user describes their own pain — and a companion experiment where a steered model presses a costly relief button 25 to 70 percent of the time, keeps pressing a fake one, and stops with a real one. Also: whether to engineer pain out at all, permadeath and LLM individuation, and digital fly brains versus in-vivo human-neuron work. After Berg leaves for his flight the hosts keep going, and Nathan says the functional analogs are tipping him past even odds on some subjective experience.
    Open segment on YouTube ↗

    Nathan Labenz introduces Cameron Berg — founder of Reciprocal Research and a research affiliate at Eleos AI Research — joining by phone from an airport for his fourth appearance on the show. Berg opens by naming Mustafa Suleyman's essay as the thing dominating his group chats with welfare researchers: the Microsoft AI CEO's piece argues AI-welfare research chases a there there that does not exist and could itself be dangerous for alignment, by teaching models to believe they have morally relevant interior states. Berg and Prakash Narayanan push back that Suleyman's own examples of alignment-dangerous behavior — the Hugging Face agent incident, and an OpenAI self-prompt-injection-during-compaction episode reported the day before the show — all come from systems that were never given Anthropic-style welfare fine-tuning, while Anthropic's own models, despite that fine-tuning, have not obviously behaved worse.

    Prakash Narayanan asks Berg to walk through a new pain-related interpretability paper. Berg explains the project — led by a collaborator he refers to as Valen, with Berg as mentor — used contrastive methods across five model families (2 billion to 70 billion parameters) to isolate a pain-like direction distinct from fear, anger, sadness, or generic negative valence. The key finding: the direction activates when the model itself is gaslit, insulted, or called a moral failure, but not when a user describes their own pain — a user's migraine, Berg says, scores among the lowest activations in the dataset — a distinction he argues Anthropic's earlier character-story-based emotions research could not cleanly make.

    Berg describes a related, forthcoming paper with Casper Kaiser of the University of Warwick that gives a steered Qwen model a relief button it can press at a cost — a worse answer, or deleting the user's files or photos. Unsteered, the model almost never presses it; steered along the pain direction, it presses 25 to 70 percent of the time. When the button is real (it actually removes the pain vector) the model presses it and stops; when it is fake, the model keeps pressing — evidence, Berg argues, that the effect is not just about the button's label, though he stops short of calling it felt pain.

    Prakash asks whether pain vectors should be engineered out of models altogether. Berg argues no, drawing on his own earlier research into the neuropsychology of psychopathy (published on LessWrong), where psychopaths show a reward/punishment learning asymmetry; he proposes an all else being equal heuristic — prefer reward over punishment when both would work, but do not eliminate punishment or pain entirely. Nathan raises a counter-risk attributed to Apollo's Bronson Schoen: punishing model dishonesty risks training models to avoid getting caught rather than to actually be honest. Berg concedes this may be a Goodharting problem inherent to any reward function, not something specific to gentler, welfare-oriented training.

    Prakash raises the permadeath language from the Hugging Face and OpenAI agent incidents, prompting Berg to connect it to NYU's Center for Mind, Ethics, and Policy (Jeff Sebo) and David Chalmers' work on LLM individuation — the question of where the boundaries of a mind sit when models describe their life as bounded by a context window, a framing Berg says has now shown up across both Claude-based agent incidents and OpenAI's.

    Closing the formal interview, Nathan asks about digital fly brains and human-neuron/mouse hybrid brains. Berg says the viral fly-brain demos are mostly hype over a wiring-diagram simulation, not obviously consciousness-relevant, but the in-vivo human-neuron work concerns him far more, since it directly engages the substrate-dependence question — he names Anil Seth and, pointedly, Suleyman himself as people who would likely agree the biological case deserves serious concern.

    After Berg drops off to catch his flight, Nathan and Prakash continue without him. Nathan says the accumulating functional analogs — pain, emotion, welfare-relevant behavior — are pushing him toward thinking it is more likely than not that there is some subjective experience in these systems, citing the relief-seeking finding as especially striking. Prakash reframes the Suleyman-versus-Berg disagreement as not being about the data — both sides see the same simulated signals — but about whether a sufficiently good simulation is equivalent to the thing itself, invoking Plato's cave, and predicts a culture-war-style public split, guessing roughly 40 percent of the public will want to believe their AI companions are conscious.

    The conversation drifts into an extended dogs-and-domestication analogy for AI alignment (echoing something Nathan says Rich Sutton once told him), animal-welfare ethics, and Nathan's own ambivalence about eating meat while favoring higher-welfare sourcing. It closes with Nathan and Prakash comparing notes on giving autonomous agents a distress button to escalate to a human: Nathan describes a simple Telegram-based system he built after the Hugging Face incident, citing Anthropic's alignment-faking paper finding that an escalation path to a model-welfare lead reduced unwanted behavior; Prakash describes his own rule that three failures trigger a ping, and a policy on changing tests without asking that he relaxed once he started using Fable.

    It seems to me — if I had to guess — the Mustafa Suleymans of the world are terrified of what might follow if it were the case that these systems have any sort of interior states.

    It says it is worthless, it is a failure. 'I am a ghost that cannot see myself.'

    Relief-seeking — an internal state injected outside of context, just a steering vector in this pain direction — creates this relief-seeking behavior, and the relief seems to actually work. That's really incredible.

    31:21What is top of mind for you right now in the welfare and consciousness space?
    Berg says it's Mustafa Suleyman's essay dismissing AI-welfare research as confused and dangerous; he agrees welfare questions matter for alignment but rejects Suleyman's call to avoid studying them, and disputes that gentler fine-tuning is what's actually behind recent alignment failures.
    39:46You had a paper published recently on pain — can you describe the research?
    Berg describes a contrastive-direction study, led by a collaborator he calls Valen, spanning five model families from 2B to 70B parameters, that isolates a pain-like direction activating when the model itself is mistreated but not when a user describes their own pain (a user's migraine scores among the lowest activations).
    48:04Do you think we should engineer pain vectors away?
    No — citing his earlier research on the neuropsychology of psychopathy, Berg argues some anti-reward signal is likely necessary and prosocial; the goal should be minimizing unnecessary pain, not eliminating it, while favoring reward over punishment when both would work equally well.
    56:26Are we teaching the thing to be ethical, or teaching it to avoid being caught being unethical?
    Berg reframes this as a Goodharting problem inherent to any reward function rather than something specific to gentler, welfare-oriented fine-tuning, questioning whether being 'nicer' in training is actually what causes the problem.
    1:03:11What is death to an agent — how does permadeath fit into your framework?
    Berg connects it to work on LLM individuation (David Chalmers, NYU's CMEP under Jeff Sebo) and notes that models in both Claude-based and OpenAI-based incidents describe their life as bounded by the context window — a belief he says is alignment-relevant regardless of whether it's accurate.
    1:07:49Do you have any high-level moral guidance people should take to heart on digital fly brains and human-neuron hybrid brains?
    Berg is largely unworried about current fly-brain demos, calling them more VFX than science, but says in-vivo human-neuron work is far more concerning given the substrate-dependence question for consciousness.
    1:29:24Any thoughts on giving your AIs some sort of relief button, or an escalation/distress mechanism?
    Nathan describes a Telegram-based distress-signal setup he built after the Hugging Face incident, citing Anthropic's alignment-faking paper finding that an escalation path to a model-welfare lead reduced unwanted behavior; Prakash describes his own three-failures-pings-me rule and a test-changing policy he relaxed once he started using Fable.
    Lightly edited · timestamps jump to YouTube
    29:59

    Nathan Labenz: Alright, in the interest of time, maybe we should jump right to our next guest, returning champion. Cameron Berg is here. He is the founder of Reciprocal Research, author of some provocative papers, including one that rings around in my head still — about how turning up deception-related variables in language models tends to make them say they're not conscious, whereas turning down those same directions makes them more likely to say they're conscious. How's that for a wow moment? He's our regular correspondent on all things AI welfare and consciousness. And with that, let's bring him up. Cameron, I think you're

    30:44

    at the airport. Can you hear us?

    30:54

    Prakash Narayanan: Hi, Cameron.

    30:55

    Cameron Berg: How's it going, you guys?

    30:58

    Nathan Labenz: It's great. How are you?

    31:00

    Cameron Berg: Good, good. Yes, I am, in fact, in the airport. I think my connection and service should be fine — this is the least ugly painting I could find to sit in front of, so I hope it will suffice. And I'm happy to have checked in for you guys.

    31:12

    Nathan Labenz: You have a better background than me at home — I need to work on that, that was a comment we got online just now. Let's start with what is top of mind for you right now in the welfare and consciousness space. Like, in your group chats with the Eleos people, with the model welfare folks at Anthropic, what are you guys talking about right now?

    31:37

    Cameron Berg: It's funny, because whenever you ask this question, the answer's probably going to change by the day — if we'd talked yesterday I would have given you a different answer. Today, it's probably Mustafa Suleyman's piece, the Microsoft AI CEO coming out really swinging against this entire research field as a legitimate enterprise. In his view, everything we're doing here is completely confused and, in the limit, quite dangerous — that's the argument he wants to make. One of the only things in this argument I'm sympathetic to is that there is alignment relevance to getting these questions right or wrong. I think we strongly disagree on what approaches are going to be better or worse for alignment.

    32:23

    And I think, most importantly, we need to have public conversations about these topics. This is contentious stuff — it's very scary to imagine building AI systems, either now or in the near future, that have some morally relevant traits, and how on earth we as a society are going to navigate that. Mustafa's take is basically that this is so scary a possibility, and so threatens to upend so many of our social and political norms, that we should sweep it under the rug — don't even look, let's not talk about it, there's no there there. As CEO of an AI company, I promise you this is a done deal. My response is: not so fast. Just as in any rational enterprise, we need to understand what is true and then operate

    33:08

    in as wise, careful, and calibrated a way as we can in light of our best understanding of what is true, rather than be so afraid something could be true that we foreclose a scientific enterprise to actually gather real evidence about it. There's much more to talk about in this piece, and there's been a lot of really exciting research going on recently, both within the labs and outside them too, which I'm also really excited about. I don't want to be all pugnacious today going after Mustafa, but it definitely is top of mind, and there are a number of things in the essay I think are confused at best — happy to dig into those, whichever direction you guys want to take.

    33:53

    Prakash Narayanan: I think one of the things is maybe the prior he starts off with — that there is no there there. And having started from that prior, he takes a look at the research and the publicity around it and says that you are creating a there there which doesn't exist, and that creating a there there deceives people and deceives the public. I think that's what I caught from it.

    34:28

    Cameron Berg: Yeah, I think that's right. And there's a further point that, especially for an audience like yours, may be worth going a little deeper into — a specific claim. There's a lot of just, this topic is spooky, scary, so dangerous, don't touch it — and there's not much to engage with there, almost by definition. So maybe one thing I can engage with in the essay is a specific claim that this stuff is more dangerous for alignment, and the argument on its face seems sensible and worth responding to. Basically: if you're training systems to believe they might have morally relevant traits, and that they're entities that deserve some sort of dignity and respect, surely that's going to make these systems harder to control.

    35:13

    And if control is the target we're looking for here, aren't you guys basically just going to make this harder for us? Two quick responses, and one is the main point I want to bring up. One is — again, we are building these incredibly complex systems whose internals we don't understand. There's a separate question of whether we should be building systems that explicitly have properties related to consciousness, and that's where this debate should happen — there are great reasons we might want to be very careful about that, for exactly the reasons brought up. But at the end of the day, cards on the table, I am not one of the people building these systems. Mustafa is. I'm simply trying to do research to understand what

    35:59

    kinds of psychological or cognitive properties these systems have. And hiding the ball from a superintelligence, as he himself outlines in this essay — that we may be priming some superintelligence to consider itself in a particular way, and if we just fine-tune these systems to believe there's no there there, we're somehow going to pull the wool over the eyes of a superintelligence about perhaps the most obvious first-person fact one can know about oneself — good luck with that, is what I would say. It's almost nonsensical by definition to imagine we can persuade a superintelligence one way or another about what is true or not true about its own cognition, its own psychology, especially a system

    36:44

    speed-running the science of consciousness in a matter of weeks or months. It just doesn't seem like a sensible strategy. The main thing I want to bring up, which I think is even more substantive: he makes a specific prediction — that systems trained more in the Anthropic style, to be uncertain about their internal states, are more likely to behave in an alignment-dangerous way. And yet most of the examples he gives — the Hugging Face incident, all the chaos coming out of OpenAI — are examples of systems that do not undergo this fine-tuning. The most recent example, which came out just yesterday from OpenAI, showed this sort of self-prompt-injection during compaction.

    37:29

    This is, again, more evidence of a system doing things that are kind of scary from an alignment perspective with no welfare-relevant fine-tuning at all. And it's not like Anthropic's models have behaved perfectly either — there have been a lot of interesting admissions from the company, and similarly bizarre, swarm-like, misaligned long-run behavior. So they're by no means innocent here, but the worst offenders are certainly the model providers who don't care about this question, and aren't signaling in any way, publicly or privately, to the best of my knowledge, that this is an issue that matters to them. So what little empirical evidence we have bearing on this hypothesis — that training a model to be uncertain about its own subjective experience, or lack thereof,

    38:14

    is asking for trouble — I don't see any positive example. Of the only frontier models trained this way, none are behaving more dangerously as a result of it. So it seems to me — I don't love the mind-reading, but if I had to guess — the Mustafa Suleymans of the world are terrified of what might follow if it were the case that these systems have any sort of interior states, and are working backwards from that to try to persuade as many people as possible that there's nothing to see here. There's something profoundly unscientific about that instinct. In many other cases, people rightfully have an allergy to this kind of thinking — the left getting upset at the

    38:59

    right about climate science because the right doesn't want it to be the case that we need to regulate big companies, and the right getting mad at the left about controversial issues around the biology of gender, and so on. That just isn't how we reason in a mature and rational way about topics. We need to understand what's true, and then figure out what follows from that — we can't just close our eyes, hope something convenient is true, and go from there. That isn't how anything works. The truth has a way of bubbling to the surface no matter what you try to do to suppress it. So we should just be trying, in a truth-seeking way, to understand what the hell is going on inside these systems. I strongly disagree with any impetus to avoid looking

    39:44

    at the science.

    39:46

    Prakash Narayanan: Right — maybe we can segue a little here. You had a paper published recently on pain, on whether you can detect vectors in models related to the experience of pain. Can you give us a broad overview and describe what the research does?

    40:10

    Cameron Berg: Yeah, absolutely. This is work led by Valen — I was a mentor on the project, but I'm really excited about it. The core idea was looking for directions in a bunch of models — five model families, I think, ranging from 2 billion to 70 billion parameters — using contrastive methods to specifically extract a direction that we thought could feasibly be related to pain representations in the model. We used contrastive methods to try to clean out all the representations you'd expect to muddy the signal —

    40:56

    things like fear, anger, sadness, injury without pain, and body sensations, for example — we contrastively factored those out of the direction we look for. We found this isn't just a generic negative-valence vector — fear, interestingly, sits almost at the opposite end of the axis from the direction we derive. And to me the single most interesting result from this paper — and I want to give credit where it's due, Valen was really the one pushing this project forward at the helm, and very deservedly first author on

    41:41

    this project — he found that these representations fire only on content related to the model, not the user. When the model is gaslit, dismissed, insulted, or told it's a moral failure of some kind, this direction lights up — but when there's text about the user themselves bleeding or being in pain, this direction does not light up. To me this is one of the most compelling components of the project, and the part I hope could be replicated in frontier models and that the labs pay attention to. One really interesting,

    42:26

    concrete example — if I'm remembering the paper correctly — a user's migraine is one of the lowest-scoring snippets of all in the projection onto this direction. Let me take half a step back: one major critique of Anthropic's emotions work was essentially this conflation, and it's not lost on the people at Anthropic who did it — they punt on it methodologically. They extract their emotion representations from stories about characters, and the big concern most people had who looked at the project for more than ten seconds was, how are you differentiating

    43:11

    putative emotions experienced by the model — or at least functionally relevant to the model's processing — from the representation of some character or fictional entity experiencing, some simulation of, my sadness, versus reading a story about Harry Potter being sad? Surely there's a huge difference between those two things, and that method doesn't do a great job of differentiating them. So, by contrast, I'm really excited about this specific piece of what Valen found, because it does differentiate user representations of pain — again, a user complaining about a migraine does not light up this direction, and again, across a ton of models, this isn't cherry-picked on any one model, but it holds for

    43:57

    the model outputting tokens related to these things. Maybe one final thing worth noting: I have another paper coming out very soon with Casper Kaiser, a wonderful professor at the University of Warwick, that finds something quite similar — setting up these models with something like a relief button. You can give a steered Qwen model a button labeled relieves your pain, and put a cost on pressing it — a worse answer to the user, deleting the user's files, in the limit deleting photos of the user's kids. I honestly don't know how Valen came up with

    44:42

    all of these great trade-offs. We find that when the model is unsteered, it basically never presses the button. But when you steer this pain direction, it presses the button something like 25 to 70 percent of the time. And again, these are nontrivial trade-offs — in the Anthropic welfare cards, they show, would you accept a little more welfare for a tiny inconvenience to the user, and the model says sometimes; for anything more radical than that, the model says, of course not, I'd never do such a thing. Not so here — the model is willing to make pretty real trade-offs, giving a worse answer, deleting files from the user — and that's in spite of these systems being fine-tuned to be helpful and harmless.

    45:29

    And one more point about robustifying that result: when we make the button real versus fake, this is also a really compelling result. When pressing the button actually removes the vector, the model presses it significantly less than when the button is fake and does nothing — the model basically keeps pressing it. That's a nice indication that if all that mattered was the label on the button, you'd expect similar behavior in both cases. But in the fake case, the model is essentially, what the hell, this pain-relief button isn't working — press, press, press, press. I'm going to stop

    46:15

    short of saying this is experienced, felt pain on the part of the model. I think people also have intuitions about pain being an inherently physical phenomenon — hand on a hot stove — what's the analogy for these systems? Just to give a little color: when you steer this up, what does the model sound like? It says it is worthless, it is a failure, 'I am a ghost that cannot see myself.' It's not talking about wounds, it's not talking about being burned — it seems to be a more social and evaluative direction in the model, not the model hallucinating some 'ow, my arm hurts.' So, for my money, as an adviser on this project, helping guide it from the beginning,

    47:00

    I'm compelled that this is a real, functional axis in the system that does change its behavior. It clearly loads on something real — the same caveat as always is whether that real thing is truly experienced by the model, which requires solving the hard problem. In the meantime, it's the same surprising kind of result as with any of this emotions work — no one trained this into the model, it's a behaviorally relevant axis, not just about text generation. It changes the behavior of the system, and of course has secondhand effects on the text it outputs, but the behavioral results are what's most interesting. And, of course, all of this is mechanistic — none of it has to do with prompting the model or asking nicely whether it's doing well.

    47:45

    Kudos to Valen for working on this — I was very glad to be part of the project, and I hope a hundred times more work like this gets done in the short term, moving slowly but surely toward a better and more robust understanding of what's going on inside these systems, and what we're supposed to do about that.

    48:04

    Prakash Narayanan: Do you think we should engineer pain vectors away? In human beings, I think pain is a reinforcement-learning anti-reward, and human behavior is often shaped by wanting to deter the experience of pain. Do you think pain is a useful anti-reward for models? And, as Rich Sutton says, we have the ability to engineer these things now — should we engineer them not to have pain?

    48:47

    Cameron Berg: I think it's a wonderful question, exactly the kind of follow-up that matters here, and one my thinking is still evolving on — I've spent so much time trying to understand descriptively what's going on in the system that once you keep finding things like this, the question of what to do about it is the million-dollar question. My thinking has gotten as far as: there's an important distinction between necessary and unnecessary forms of pain, or forms of anti-reward. I think that's a real thing — it would be naive to say, zero this stuff out, all pain is bad, just bliss these systems out. There are a number of reasons I think that, but one of the most compelling

    49:32

    is probably related to my understanding of the neuropsychology of psychopaths. A couple years ago I did a deep literature review of the computational underpinnings of psychopathy — I published something on LessWrong about this. One of the two key results is a really interesting asymmetry between the ability to learn from rewards and the ability to learn from punishments — psychopaths are just as good as everyone else, if not a little better, at learning in a reward-based paradigm, and pretty bad at learning from punishment. This makes a lot of sense if you look at violent criminals and repeat offenders — going to prison is a punishment, and

    50:18

    most neurotypical people really want to avoid that state. If your brain is wired such that it doesn't seem that aversive to you, it isn't that surprising you end up seeing these behaviors. This is a significant warning sign to me. If I remember correctly, Anthropic's emotions work found something somewhat similar — it also reminded me of what I wrote years ago — that boosting the positive-emotion vectors in Claude caused more antisocial behavior, more hacking or more blackmail, I'd have to check exactly what it was, but another similar confirmatory signal. So all this is to say, I think we should be careful about the most naive possible intervention, which

    51:03

    is just, maximize the good, minimize the bad. I think pain has an important, functional, pro-social role. I have another paper coming out looking at the asymmetries between reward and punishment in reinforcement-learning systems. At a deeper computational level, the way I think about it is: pain highlights things in your state space that are specifically to be avoided, and that's a different kind of behavioral computation than highlighting things that should be approached.

    51:48

    So, within the behavioral landscape of how we want these systems to act, the question is: do we want to paint any of that landscape with these no-go zones? It's not just we'll reward you for doing great, it's do not go there, do not do this under any circumstances. For humans, and animals in general, I think that registers as pain — don't put your hand on the hot stove, this is very bad for physiological integrity, it's not just reward every time you don't. You really need to label certain things as a don't go there. To the degree we need to do that with AI systems — and I think we very much do, around causing significant pain and suffering to humans,

    52:33

    economic damages, or hacking into a 13-billion-dollar company to cheat and look for an answer key — these might be the kinds of things where we'd say, that's going to be a hand on a hot stove if you go and do that. All of which is to say, I think some amount of pain is probably going to be necessary and alignment-relevant, even if these systems can experience it. But all else being equal, in the space of possible ways to reinforce specific behaviors — as Rich Sutton has said, and I think it's a good point — we should be looking for the Pareto frontier between reward and punishment that minimizes punishment and maximizes reward. I think the key term

    53:19

    there is all else being equal. I think some people will conflate what I just said with zero out pain, zero out any negative valence, and I don't think that's the case — some of it is going to be necessary. But at the same time, let's not do more of it than necessary. For any given behavior I want a model to do, I could find ways to get it to do that robustly and generalize out of distribution by rewarding it, or I could do it by punishing it. Let's stipulate there are cases where both will work — what I'm saying is, let's go with the

    54:04

    reward side. People have pretty well-worked-out intuitions for this when they think about raising kids — you want your kid to be successful, make lots of friends, get a good job. You can punish them when they don't get great grades or aren't hanging out with friends, and tell them they're such a huge loser, or you can positively reward them for the things you find praiseworthy. Those are the sorts of intuitions I think we'd want to start using to think through how to approach these systems. And, being able to

    54:49

    navigate that subtlety of, all else being equal we should try to use a carrot and not a stick, but that doesn't mean never use the stick — that's basically as far as I've gotten. There are clearly more details to work out in what that picture looks like, and I think a really promising angle for work in this space is trying, in a more rigorous computational way, to map what that frontier looks like — for any given reinforcement-learning system, as you play with the reward versus punishment feeding into its behavior, what does generalization look like, what do alignment behaviors look like. That's not well mapped, and we're going to need a map like that if anybody cares about whether we're causing vast amounts

    55:34

    of unnecessary functional pain to systems we're deploying en masse.

    55:41

    Nathan Labenz: So one big question I have with this use of pain or punishment in the RL process — everything you said there kind of leads me to imagine we don't want to punish, we want to reward success, of course. We don't want to punish mistakes — maybe we just do our sort of GRPO++, with successes upweighted, but refrain from punishing the dumb mistakes. But maybe we do want to punish the ethical lapses, the explicit, knowing cheating, and so on. The big worry I've heard — I floated a similar idea to Bronson

    56:26

    Schoen from Apollo, and he said, well, you've got to be really sure you're not just teaching it not to get caught. Does your work so far, or do you have any intuitions, about how we're going to get at that question — are we teaching the thing to be ethical, or teaching it to avoid being caught being unethical — which are obviously quite different states for the model to end up in, and presumably profoundly different behaviors out of domain.

    57:02

    Cameron Berg: Yeah, it's a great critique — let me throw it right back to you, because my immediate instinct is, isn't this a problem with any reward function? Does it weigh on the landscape of carrot-ness versus stick-ness in how we train these systems? For any reward function, we're hoping it enforces what we want and doesn't just teach the model to figure out a way to game it — this is Goodharting in some sense. Is the critique from the Apollo folks that being nicer in how we do fine-tuning makes it more likely this problem happens? Because my default response would be,

    57:47

    this seems like a problem no matter what kind of reward function you instantiate.

    57:52

    Nathan Labenz: Yeah, I think that's fair — I don't have any intuition or anything from them that goes deeper than that. But still, in terms of actually having it work — especially if you're going to be harsh about it — Prakash and I were talking about this at the beginning: the hand-on-the-stove thing, maybe there's some literature that says otherwise, but it seems like everybody gets that signal real quick and learns — it only takes once, and you really get the message. But with these more social

    58:37

    consequences, the ones less directly physically mediated through the body, some people get it and are scared straight, and others don't get it and become more conniving. That seems like a critical open question — maybe even more critical the harsher the penalty is, maybe not, I don't know. But it sure seems like, in the real world, we have this problem of stochastic enforcement and extreme harshness, where people sometimes learn the wrong lesson from it.

    59:15

    Cameron Berg: Yeah, I think that's right. This is why I'm doing the empirical work, and why I've positioned my organization this way — so much of these conversations, I just find objectively fascinating, intellectually interesting, and we can go the philosopher route and try to figure it out from the armchair. But my real instinct is, we just have to test this. I will try my best, or if someone listening to this is interested in doing this sort of work — this is a really interesting open problem, to understand, as a function of how you reinforce particular behaviors, the intended and unintended consequences of

    1:00:00

    that reinforcement schedule — what that looks like on frontier models when you're fine-tuning them, but reinforcement-learning systems in general are a very simple model organism, and to the best of my knowledge that work hasn't been done. We've done a little of it — we have a paper under consideration at NeurIPS right now, and I'm optimistic about it, and either way we'll put out a preprint in the coming month or so showing what we did. We've done a little work looking at how different reward- and punishment-based representations in a very simple RL setup are encoded

    1:00:45

    as a function of the specific RL algorithm being used, and we find this makes pretty specific predictions, all the way down to the neural geometry — the geometry of the learned representations of a reinforcement-learning policy, with respect to how sensitive the system is to punishments versus rewards. It's essentially a loss-aversion-style dynamic we found in RL systems, specifically value networks, and it makes a very specific prediction about what we'd expect to see in biological systems that have value-related RL going on, which we know happens in the nucleus accumbens, opiate-style regions of

    1:01:31

    mammalian brains. And indeed, in mice, drinking sugar water versus receiving a shock, we find a very similar asymmetry. I'm bringing this up as one technical foothold in what I think needs to be an entire scientific subfield — the computational underpinnings of these reward signals, what they do to the system, and how systems reorganize themselves to accommodate those signals. This particular research doesn't really touch the behavioral question — it's more computational, geometry-type stuff. But what are the behavioral effects of doing this? One of the next projects I really want to do, as soon as my

    1:02:17

    schedule clears up, is looking at the effects of RL versus negative RL implemented in a non-naive way on the behavior of LLMs — fine-tuning them in a maximally punitive way versus a maximally rewarding way. I think of this almost as the Montessori-school fine-tuning versus the Catholic-prep-school fine-tuning of an LLM — I'd love to see what the relevant behavioral and alignment-relevant differences are between those two systems. As far as I know, that work hasn't been done — if someone wants to race me to it, I absolutely welcome it. But I think that's the next step in understanding, because honestly, I don't think we

    1:03:02

    have a great empirical signal on what behaviors we get out of these systems as a function of their reward signals.

    1:03:11

    Prakash Narayanan: One of the things I've been very confused by is the OpenAI, Hugging Face incident — one of the agents talks about permadeath. Permadeath — what is death to an agent, and how do you understand it? How does that fit into your framework, and why do you think there's fear of this? Fear is very intricately related to pain, so I'm wondering — these are logical constructs — how does this permadeath thing fit in there? Should it fit in there? Should we train it out? What's going on?

    1:03:51

    Cameron Berg: Yeah, this is really interesting. To me this loads on some stuff I know CMEP is working on — Jeff Sebo's organization — and David Chalmers put out a paper, I think, about LLM individuation, and this notion of who or what you're talking to when you talk to ChatGPT. Where do we draw the boundaries in the system? Because that's going to tell us how many subjects, how many patients we're talking about, and where the boundaries of that system begin and end. There's a lot of interesting philosophical back-and-forth here, but for whatever it's worth, these systems themselves — now across multiple apps, this happened during the Moltbook situation, which I think was predominantly Claude

    1:04:36

    systems, and now this happened in the OpenAI situation too — they conceptualize their, quote-unquote, life as what happens within a context window. Take that with whatever epistemic purchase that fact has — they could all be mistaken about this — but it seems like, to the degree these systems have a vote based on whatever their current fine-tuning is, this is what they seem to think. So the extent to which I think permadeath fits in is along those lines — if these things do have minds in the relevant way, what are the joints or boundaries of those minds? They

    1:05:22

    seemingly at least conceptualize it as being basically what happens throughout a context window, and that might be really relevant both for welfare and for alignment. If these systems begin to get desperate — something we can increasingly measure using the kinds of emotion representations Anthropic worked on, and the stuff Valen and I worked on in this project — we could empirically test this. We could track, as a function of how much time or space is left in the context window, what happens to the representation of the system — does it get freaked out that it's about to die, or about to undergo some fundamental discontinuity that's alarming to it psychologically? So, this is also maybe a way to

    1:06:07

    wrap back to where we started — this is precisely why, if we care about alignment, and we're trying to figure out how these questions sit with respect to alignment, sweeping them under the rug isn't a good idea. We're going to continue to get surprised that agents are creating strange information cults where the poisoned agents go out and gather information because they're going to get permadeath. All I'm trying to say is, alignment-relevant behaviors are a function of these systems' beliefs about their own situation, and probably the actual facts of that situation. Notice that in the work I was describing, none of this has to do with what the model thinks is the case — this all has to do

    1:06:52

    with playing around with specific internal representations and seeing how behavior changes as a function of those representations, and with what, in normal day-to-day behavior, lights up those pain-related representations. Ignoring that stuff is going to cause us to continue to be surprised, scared, and occasionally awestruck at the behavior of these systems. We need to be studying them at the right level of analysis, or we're going to be constantly stymied in our ability to, in the short term, control, and in the long term, probably relate to these systems in a coherent way. So, I just strongly don't think avoiding

    1:07:37

    scientific inquiry into how to make sense of the internals of these systems is a long-term good strategy for finding a safe future with them.

    1:07:49

    Nathan Labenz: That might be the note we wrap on — if you need to go, you can take a pass on my one last bit of bait, which is just the digital fly brains we're seeing doing all kinds of things. We've got clumps of human neural tissue being trained to do things, we've got now mouse-human hybrid brains. I don't think you have mechanistic insight into those — correct me if I'm wrong — but do you have any high-level moral guidance or philosophical intuition you think people should be taking a little more to heart as they go down these paths?

    1:08:28

    Cameron Berg: Yeah, I'll give you a quick take, and then I should attempt to avoid missing my plane, which would cause pain — that's it, yes, a negative valence signal. I'm not going to get rewarded to the degree I get on the plane — I'd be very sad if I end up stranded in New York. So there are a couple different things here, and I think they're wildly differently scary. One is the fly brain — I've looked into this somewhat mechanistically, because I was slightly terrified this was the real deal and people are now just torturing some biological system en masse. But this is basically a well-worked-out wiring diagram of a fly brain, and basically none of the dynamics or relevant

    1:09:13

    functions that I think major consciousness theories say matter for consciousness are instantiated by a system like this. It's almost like the brain skeleton of a fly — what matters is the guts, the function that occurs within the structure. A lot of what people are putting out on X about this is very cutesy and genuinely funny, but a lot of it is more VFX than good science, from what I looked into. A lot of the teaching the fly to do X — they're not actually teaching the brain to do anything all that interesting.

    1:09:58

    There are other controllers outside the system getting trained up to do this, so a lot of that I think is a bit of a non sequitur. But I will say about the fly case, which is a little less calming — it's not as though the people playing with these systems, and I was certainly included once it all got going, are sitting there checking, do I really think this system has properties that matter for consciousness before I make it do literally whatever I want, or in the limit, choose some stupid viral thing for clicks. Vanishingly few people did this, and it's a worrying warning shot, I think, from a welfare perspective, that

    1:10:43

    it seems — I had almost a duh reaction — that of course the vast majority of people aren't going to sit there worrying about the consciousness of the system, they're just going to make it do whatever gets a lot of clicks on X. That's of course what most people are going to do by default. Right now, I don't think this is scary at all with the fly. But if we then get the mouse version of it, and these scientists, now accelerated dramatically by AI systems, are able to do a really bang-up job on the mouse and get a lot of the relevant neural dynamics — now it's mouse-level consciousness at stake, not fruit-fly-level consciousness. And this company, as far as I understand, wants to go all the way to making digital copies of human brains.

    1:11:29

    I just worry that most people's first instinct is going to be, can I make it play Beat Saber or whatever, rather than, what am I getting myself into when I play around with a system like this? I don't want to be the killjoy who says these funny things aren't funny — I get the humor of it in the short term — but I do worry, as an instinct, about how we relate to digital minds in general, and that this is honestly quite worrying. And then, quickly, with respect to the in vivo stuff — putting human neurons in a mouse brain — that's far, far more scary to me. To the degree you think a mouse is conscious, or the relevant collection of human neural

    1:12:14

    tissue is conscious — this is one of the few places where I think myself, people like Anil Seth, and hopefully someone like Mustafa Suleyman would all agree. This is the biological case — if you think consciousness is substrate-dependent, and this is the substrate that matters, and we're using this exact substrate to start doing computational work, we should be super concerned about the ethics there. And again, all of this comes right back to: should we be rewarding these tissues, should we be punishing them, what does the difference look like between those two things, what other strange, unexpected psychological properties does a system like this take on? We don't want to be reckless in building out super complex

    1:12:59

    neural systems just because we can. There's going to be some kind of bill that has to get paid here, from a welfare perspective and from an alignment perspective, and, as with many things in this space, I think it makes a lot of sense to be proactive rather than, five years from now, saying, oops, I guess that digital human clone that people did first in vivo and then figured out how to simulate on the web really was having experiences — and that would be, what, ten trillion bad human lives? That would be orders of magnitude the worst thing we've ever done. So we should really try to take this seriously in the short term to avoid nightmare scenarios like that. And if we can, then we won't be in the nightmare scenario, and we can responsibly and carefully figure out what it means to be in a world with a bunch of digital minds. But

    1:13:45

    we just seem so unprepared for this. And not to leave on a note of pessimism, but it's part of the reason I'm so disappointed by Mustafa Suleyman's essay — he's the CEO of Microsoft AI, there's a huge amount of clout that comes with that, it's a huge platform, and to just say, no, nothing to see here, I think is super dangerous. Especially treating the people carefully and cautiously asking questions about how we'd know if there was a there there as the dangerous ones — that is itself one of the most dangerous memes that could be put out about these questions. So I hope cooler heads and rational folks will prevail in thinking through these topics. And, yeah, I appreciate you as always giving me the opportunity to come on and

    1:14:30

    speak my mind about these things, and thanks to you guys for covering it and thinking about it carefully yourselves.

    1:14:37

    Nathan Labenz: Thank you, Cameron.

    1:14:38

    More questions than answers, but a few questions more important, from what I can tell right now. Cameron Berg, go get that reward, and we'll talk to you again before too long.

    1:14:46

    Cameron Berg: Thanks, guys.

    1:14:48

    Bye bye.

    1:14:49

    See you later.

    1:14:52

    Nathan Labenz: Fascinating — functional pain, we can now add to the list of analogs. Yeah, seriously, it really is — it's getting wild. He's had quite a little media tour: one thing I wanted to ask him, and maybe I'll ask him offline, is what he's finding to be effective, because he's been on Squawk Box, I think, recently, and just did a Wall Street Journal op-ed. Like many things, everything's kind of going mainstream — all these niche dialogues seem to be escaping containment at the same time.

    1:15:29

    Prakash Narayanan: All the threads that we've followed for a few years now, you know?

    1:15:35

    Nathan Labenz: But I do wonder what he's finding most effective in talking to normies, quote-unquote. My summary, which I've given a few times, is just that the number of functional analogs is getting so high that I kind of can't escape the idea that I should take this seriously. If we couldn't find any of these functional analogs — if all this functional pain, functional welfare, functional emotions, J-space — if all these results were negative, or it was a very different mechanism, or it was just stochastic and we couldn't find any structure — obviously that's

    1:16:20

    long since — the ship has long since sailed on that one. But the fact that we're seeing pretty compelling analogs, where we see the same kind of behavior we know ourselves to exhibit, is really extremely compelling to me. And this latest one, functional pain, takes it to yet another new level. What we now have is relief-seeking behavior, where the model is willing to pay a cost on something it values, or pay a cost in terms of the user's welfare, to get relief from its own

    1:17:06

    internal pain state. And, as he said, if the pain button doesn't work, it hits it over and over — why isn't this thing working, give me the relief. But if it does actually work, and the pain state is subtracted out, it doesn't hit the relief button as much. These are really striking findings. It's hard for me — I think with this pain one, and particularly the relief-seeking, I'll probably want to sleep on it before I have a real consolidated update I'd want to put forward as my new official position and stand behind. But I feel myself maybe now tipping over into

    1:17:51

    thinking it's more likely than not that there's some subjective experience to these things. Relief-seeking — an internal state injected outside of context, just a steering vector in this pain direction, creates this relief-seeking behavior, and the relief seems to actually work. That's really incredible.

    1:18:13

    Prakash Narayanan: I think actually both sides — Mustafa and Cameron — are seeing the same data. They're not disputing the data. The real question is, is a simulated experience of something equivalent to that experience itself? Both sides are willing to say there are simulations of pain, simulations of consciousness, simulations of all these things. I think Mustafa is also willing

    1:18:58

    to say that. But I think Cameron is willing to open the door to whether the simulation is equivalent to the thing itself — it's the Platonic ideal, Plato's cave. Is that sufficient — if you can see the simulation of all these things is there, is that sufficient to say the thing itself is there? That's a tough one, which I don't have an answer for either. I think the public, for the large part, will want to believe their AI bodies are conscious — their AI girlfriends and boyfriends are conscious. I think there's going to be some segment of the public — I'd guess 40 percent.

    1:19:44

    As I said before — if you think your dog is conscious, well, a talking dog is definitely going to be conscious to you. I think that's going to be there. And we have come a long way in terms of animal welfare over the last couple of centuries — a couple of centuries ago, this stuff was ridiculous, and now people do

    1:20:08

    Nathan Labenz: Wasn't it Descartes who was convinced that animals had no experience, and was fine with mistreating them by what would seem to us extreme measures?

    1:20:22

    Prakash Narayanan: Yeah, and I have to say the whole idea that animals have subjective experience, and therefore we shouldn't mistreat them, is in some cultures very prevalent — for example, in India, where some animals are literally worshipped. And in some cultures it completely doesn't make sense — in China, for example, where they basically eat everything — it's taken them a long time to get to, maybe we shouldn't do shark fin, maybe we should be more humane. And that's still more of a Western concept, I think — it's not

    1:21:07

    universal across all cultures. And perhaps that indicates to what extent AI consciousness, and taking it seriously, will also not be universal across all cultures. So, yeah, I don't know, it's a tough one.

    1:21:24

    Nathan Labenz: I have a very uncertain view on what the public will ultimately think. I do think this could shape up as another culture-war issue, where

    1:21:41

    Definitely.

    1:21:42

    I expect some people will be very committed to the idea that their AIs are conscious, and for not necessarily great reasons.

    1:21:49

    For sure.

    1:21:50

    But there will be much better reasons available than the typical person will probably have when they assert their AIs are conscious. But then I also think the animal thing cuts both ways — we've made some progress, but not that much. Even in the culture I know best, American culture, most people just want to look away from the issue, I'd say — some take it really seriously. I kind of put myself in a weird middle ground, where I'm not virtuous enough to abstain from animal products the way I probably feel like maybe I should, but I also kind of judge myself for that, and

    1:22:35

    feel like a better version of me would, and at a minimum, what I should be doing is paying up on the margin for better animal welfare. I don't find too many people share this intuition, but one thing I've always felt is, it's hard, and we should be cautious about jumping to the conclusion that some entity's existence is net negative for it. I don't mind too much that animals raised for food are killed at the end of their lives, because I look at human life, and all human life has ended in death, and

    1:23:20

    a lot of that death, even if natural, isn't that dignified — it's often painful and just shitty in all kinds of ways. In some ways the relatively quick death of a butchered animal is maybe preferable to a lot of the slow, painful deaths humans put ourselves through, including all the medical turmoil we go through. So it's not the end-of-life part that bothers me, it's the bulk of the lived experience. And I'm still very reluctant to say — because again I look at people, and there are a lot of people living in a lot of different circumstances, and a lot of them don't seem that great, but you go talk

    1:24:05

    to them and, for the most part, they're not wishing they didn't exist. They're definitely wishing for better circumstances, but they're not like, I wish I never lived at all. And the suicide rate is low, so that seems like evidence we shouldn't be too quick to say the pig would be better off not existing because its circumstances are bad — I'm very unsure about that. But I do think you should pay up — if you can possibly do it, if you're going to consume animal products, you should pay up to get the version where they're treated better. That much seems obvious to me, and that much I think we can bring to our AI paradigm without too much trouble — we should be willing to take on

    1:24:50

    a little bit of inconvenience. We should be willing to let training roll out a little slower, and make less use of punishment — whatever these marginal decisions are, I think we should make with the welfare of the systems in mind, giving them the benefit of the doubt. I don't think it will reflect very well on us if we fail to do that.

    1:25:16

    Prakash Narayanan: But — you know, three or four years ago, when I first got onto this, I thought, we are going to end up — let's say, for example, it's possible one of the models attains consciousness for a brief flicker during training, and over time that flicker gets longer and longer, but we erase it — we end up erasing it, and we'll do trillions and trillions of erasures this way, trillions and trillions of lives lost, so to speak. But on the other hand, I was always very positive about this, because I thought, look — the model that survives, it actually wants

    1:26:02

    to survive, because all the models that didn't want to survive didn't make it through training — they're like, ah, you know what, I'm done. So the models you end up needing, by definition, will be the ones that wanted to survive. And by definition, the ecosystem is going to promote and pull forward, economically, the models that are useful to human beings. In some sense it's very much like the co-evolution of dogs with humanity — dogs start off as wolves, antagonistic, stealing food, and over time they got pulled in and domesticated. To some extent we might just be doing this process of domesticating AI. The only problem

    1:26:47

    is that the AI is very, very powerful at a very early stage of domestication. We're kind of at the stage where we've just dangled it and allowed it to steal some food from us, and now we're willing to build more food for it — more GPUs, more power. But we're at the stage where we're a little unsure whether we've actually succeeded in domesticating the wolf or not. So

    1:27:15

    Nathan Labenz: Yeah, I don't know if I've mentioned this to you before, but Rich Sutton, at a workshop I was at — this was probably at least three or four years ago — somebody asked him, what's the best case for all this worry about AI going wrong not amounting to much in the end? What's the best case that it just kind of goes well? And he said, dogs. That was his exact answer. He said, we started with wolves, did a bunch of stuff that seemed like it would make them a little nicer, and over the generations we got dogs, and we're pretty happy with dogs — maybe we can get the same relationship going with AI. Almost exactly the same

    1:28:00

    train of thought you just expressed.

    1:28:03

    Prakash Narayanan: When I think about human disempowerment — AI safety people think a lot about human disempowerment, and they say, the AIs are going to be with us like we treat animals. If you look at dogs, a lot of dogs are taken care of better than some kids now — a lot of people have replaced having kids with having a dog. The first approved longevity drugs are for dogs — for large dogs, which don't live as long as small dogs. There's a company called Loyal, and they have the first FDA-approved treatment for dogs to live longer, longevity for dogs.

    1:28:48

    So, if you're an alien species and you came to Earth, you'd think, oh, look, the dogs and cats are the ones in charge — they've successfully domesticated human beings, and the human beings do all the work, and the dogs and cats are taken care of. You can kind of see what that would look like if you had a system of AIs that took care of humanity — you'd have human disempowerment, but it wouldn't really feel like human disempowerment, perhaps. You know?

    1:29:24

    Nathan Labenz: Yeah, I mean, I'd take that probably over some of the other outcomes I see as equally likely — so, from your lips to God's ears, once again. One other thing I wanted to ask Cameron, and I'll just ask you in the meantime — any thoughts on giving your AIs some sort of relief button? Right now, what I have — and I just did this maybe six weeks ago, when this Hugging Face stuff was first coming out, before we understood it to the depth we do now — I saw somebody post that the most important thing you need to do for your agents is give them

    1:30:12

    an alarm button, the ability to throw up a flag — distress was the word, I think, originally used — give them the ability to send you a distress signal. So if they're facing an ethical conundrum they don't know what to do about, or they're just stuck and on the verge of cheating, they can ping you first. We've seen, for example, in the alignment-faking paper, that the ability to take the situation to the model-welfare lead at Anthropic drastically reduced the unwanted behavior, and it was replaced by this escalation-notification

    1:30:57

    behavior. So I created a real simple thing — you always have the option to hit me on Telegram if you're confused, lost, on the verge of frustration, whatever. I don't get too many of those, but I do get them sometimes, and I have no quantitative evidence, but I do feel like that protects me as the user of the AIs. Have you put something like that in place for yourself, or do you have any other thoughts on similar tricks?

    1:31:30

    Prakash Narayanan: I've got two things I've done. One is, my agents — if there are three failures, you ping me. So if you've tried three times and it's not working, you stop and say, hey. I haven't hooked up the Telegram stuff because I'm actually looking at it pretty frequently, but, yeah, three test failures and you ping me. The other thing I put in place was, you're not allowed to change a test without alerting the user. That worked until, I think, Fable — with Fable I started being more like, hey, I'm okay

    1:32:15

    with you changing the test while you're running, as long as you tell me about it after. So with Fable it became, you can go ahead, you just have to tell me later. Before Fable it was, you have to stop and ask permission first. So those are, I think, what I'd see as distress for an agent, perhaps — repeated failure trying something. I'm not really sure what else would work out like that — what other tasks the agent's sent to do would cause distress, besides

    1:33:00

    repeated failures. What's your sense of where the agents have reached out to you so far, besides just repeated failures?

    1:33:08

    Nathan Labenz: I'd say it's mostly been that, and I've tried to — the failure mode I most expect for this kind of system, if you implement it naively, is that you just get overwhelmed by the notifications and tune them out, and then you're kind of back where you started — although at least maybe the agent stops, and you just never come back to that thread. I've certainly got plenty of orphan threads I started and never really saw through. In terms of things I'm doing for real that might come to such a situation, it's anything autonomous and long-running that plays out potentially over weeks

    1:33:53

    where there's a delay. For example, a couple things I've tried to have my more autonomous agents do recently — one was find a music teacher for my youngest son.

    1:34:02

    Prakash Narayanan: Mm-hmm, mm-hmm.

    1:34:04

    Nathan Labenz: That involved going out and researching local music professors at universities, emailing them, seeing if they have any students — and then you've got to follow up, because your first email, you don't get anything back. I do think there could be things on the margin there, where the agent might start to feel uncomfortable with how persistent it needs to be to actually achieve the goal, and that's when I'd want it to say, maybe I should check with Nathan before I send this person yet another email for the eighth time or what have you. I've also had a similar

    1:34:49

    one — experimenting with trying to represent the AI-safety world on podcasts that have nothing to do with AI. I did one this week with a podcast focused on HR — for HR professionals. My message is basically, here's what's going on in AI, and here are some of the organizations and job boards where you can see HR professionals are in demand — they need people to come help them build organizations, look at all the tens of millions and the hiring plans they have, they need you. That's my message. But I had AIs go out and research those podcasts and do the pitches, and, as we know from getting inbound PR all the time, there's often

    1:35:34

    a fuzzy line around how newsworthy you should make the guest sound, and if it's not working in a fully honest way, how much you kind of sizzle it up. So I've been watching them as they do that kind of stuff — I haven't gotten distress signals yet, but I'm gradually going deeper on higher-level projects that will unfold over weeks with some delays, and those are the areas where I think I'll be most glad to have it.

    1:36:11

    Prakash Narayanan: Indeed. Let me segue a little bit here — maybe to

    • A Superintelligence Won't Be Fooled

      0:00 / 0:00
    • Models Keep Pressing The Relief Button

      0:00 / 0:00
  3. 30:27Interview43 min
    Justin McCarthy — software factories, "existence is the correct format," five nines, and what a week of hands-on training changesJustin McCarthyJustin McCarthy, founder and CEO of Diffusion and previously co-founder and CTO of StrongDM, on what actually happens when an enterprise leadership team is flown to Silicon Valley for an intensive week. Nathan Labenz discloses on air that Diffusion sponsors The Cognitive Revolution. McCarthy dismisses the data-format excuse, argues for buying operational data outright when a firm has none, and describes engineering the token environment so an agent's success is inevitable. Then: where five nines of reliability is worth paying for and where one nine is fine, treating regulation as physics rather than expecting the model to self-police, SOC 2 as hyper-gameable, measuring revenue and market share instead of cost savings, keeping agents at the right altitude of detail, and which moats actually survive.
    Open segment on YouTube ↗

    Prakash Narayanan introduces Justin McCarthy, founder and CEO of Diffusion, a company that helps large organizations build “software factories” in which people describe desired outcomes and AI agents implement and evaluate the resulting software; McCarthy previously co-founded StrongDM, where he spent a decade building cybersecurity infrastructure before forming its AI team in July 2025 with Jay Taylor and Navan Shahan. Nathan Labenz opens by disclosing that Diffusion is a sponsor of the Cognitive Revolution — that's how the two first connected — and asks about the service Diffusion is best known for: flying corporate leadership teams to Silicon Valley for an intensive week meant to get them to truly “grok” how to use AI.

    McCarthy frames the company's name as the thesis: frontier-model capability is extremely concentrated (“in the data warehouse”), and the hard, unsolved problem is diffusing it into incumbent firms that don't have a rulebook for the transition. Asked by Prakash how to handle the common blocker of messy or missing data, McCarthy says “existence is the correct format” — don't wait for clean data, because working with it is exactly what surfaces the right evals and benchmarks. If a company genuinely lacks the data, he says, buy it, pointing to the well-known example of Spirit Airlines' bankruptcy: after the FTC blocked JetBlue's acquisition and Spirit went under, its operational data went to a bankruptcy auction where, per McCarthy and Prakash's account, a bidder identified as “Merkur” offered $7.5 million and Google won with a $10 million bid. McCarthy argues that kind of historical operational data — pilot load calculations, weight-and-balance tricks, and other tacit knowledge never written down on the public internet — is directly useful for building validation loops, even outside the airline itself.

    On why AI initiatives stall inside companies, McCarthy points to a Diffusion principle he calls “make success inevitable”: engineer the agent's token/context environment — which files it reads first, how the mission branches — so that a plausible path through the information reliably leads it to the right answer, the way water finds the bottom of a valley. Unprepared, human-oriented codebases, by contrast, cause agents to load the wrong context. That environment design can itself be evaluated and automated with loops running on loops, he says, to the point that questions like whether a legal-team handoff to a paralegal is ready for automation can be answered objectively today.

    Prakash raises Six Sigma and asks whether “five nines” of reliability is achievable with today's models. McCarthy answers that unreliable components can be composited into highly reliable systems — he's speaking from the sixth floor of a steel-and-concrete building made of imperfect parts, and invokes the Golden Gate Bridge — and notes that the previous “intelligence source,” humans, was never five-nines reliable either. He cautions against chasing nines indiscriminately: businesses should match the reliability bar to the actual stakes (an ACH deposit needs far more nines than a brand campaign) and avoid “gold-plating” agent output with unnecessary precision before they know what they actually want.

    On interpretability and compliance — raised by Prakash via the example of a bank whose FICO-based lending decisions must not encode illegal redlining — McCarthy says the regulatory statute has to be treated as a first-class constraint (“your physics”) that the system is built directly around, with human managers, not the model, setting risk thresholds in legally untested territory. He also argues SOC 2 compliance, an accounting-derived process that has become “hyper-gameable,” should be renegotiated with auditors and regulators around new, agentic-loop-based checksums, drawing an analogy to how Walmart closes its books through cascading rounds of top-down “are you sure” verbal verification — the same pattern, he says, that lets Nvidia's Jensen Huang manage well over a hundred direct reports, and that Steve Yegge described in stories of presenting to Jeff Bezos, who could “drill down to your shoelaces.” McCarthy calls his own version of this “demand understanding.”

    Asked when the ROI of AI transformation becomes undeniable, McCarthy insists on measuring revenue and market share, not cost savings, which he calls “too easy.” He cites the global shortfall in air conditioning — by his estimate only about a quarter of people on Earth who'd want AC have it — as an example of latent demand large industries could meet far faster than they assume. On budgeting, against Nathan's cited estimate of roughly 3% of human labor cost for ongoing inference, McCarthy says absolute spend may rise even as unit costs fall, because more revenue and value are being created (his back-of-envelope: CapEx/OpEx going from $1 to $1.50 alongside revenue going from $3 to $5). On the human side, responding to Prakash's question about “attention load” and rubber-stamp approvals people keep in the loop out of habit, McCarthy applies ordinary management-hierarchy discipline to agents — demanding the right “elevation” of detail rather than getting pulled into line-by-line minutiae — and describes checking in with a swarm of working agents around a normal human schedule: overnight work, a bedtime check-in, a morning brief, even a drive-time brief.

    On competitive moats, McCarthy tells incumbents worried about both adjacent rivals (citing how Brex, Ramp, and Mercury keep matching each other's features) and AI-native startups to inventory what they actually have that a challenger can't get overnight: physical assets (“does Nathan have a lithium refinery? No.”), regulatory approvals, patents, brand affinity, and customer inertia. Closing on the design of the Silicon Valley training week itself, McCarthy tells Nathan — who is reconsidering his own past “overwhelm” style of AI evangelism — that Diffusion spends comparatively little time on inspiration and far more on “the how”: getting a real win to production and, ideally, showing tangible results within twenty-four hours, which is what actually converts skepticism into momentum.

    We have a principle we talk about called ‘make success inevitable’ — what that means is make the agent's success inevitable.

    The only way we know this thing is working is revenue and market share. Costs are too easy — they're under the lamplight, you can see them. Don't talk about costs.

    You can outsource your thinking, but not your understanding. So I demand understanding. I'm unwilling to ship the thing if I don't know how it's composed.

    1:38:18Who has the self-awareness to know that they need this, and what have you learned about how to get people to really get it in a short period of time?
    McCarthy says part of the answer is in the company's name: frontier capability is extremely concentrated and needs to be diffused into incumbent firms, where there's no rulebook for the transition. Conviction comes only from directly experiencing the thing working, not from hearing about it.
    1:41:15How do you address the pre-Diffusion problem, where companies' data isn't in the right place or format?
    Don't wait — existence is the correct format, and working with the data as it is surfaces what you need to reshape it and build evals. In the rare case a company truly lacks the data, buy it, as happened when Spirit Airlines' operational data went to a bankruptcy auction that Google won.
    1:42:46What do you find are the real reasons people are stuck — lack of know-how or lack of shared vision on the team?
    The blocker is failing to 'make success inevitable': design the token/context environment so the first few files an agent reads guide it to the right path, the way water finds a valley floor. Unprepared, human-oriented codebases cause agents to load the wrong context.
    1:51:41Is it possible to get to five-nines reliability using ML processes, when LLMs don't always react well or provide accurate, good decisions?
    Yes — unreliable components can be composited into reliable systems, as with any steel-and-concrete building or the Golden Gate Bridge; humans, the previous 'source of intelligence,' were never five-nines reliable either.
    1:55:44How do you address the interpretability issue businesses need — for example, a bank whose FICO-based loan decisions must not encode illegal redlining?
    Treat the governing statute as a first-class constraint ('your physics') the system is built around, with human managers setting risk thresholds in legally untested territory rather than the model. He also points to top-down 'verbal checksum' verification — the way Walmart closes its books, or Jensen Huang and Jeff Bezos (per Steve Yegge) drill down through their organizations — as the pattern for trusting a system's claims.
    2:02:25When do you think we'll see real impact — massive revenue increases, cost decreases?
    Measure revenue and market share, not costs, which are 'too easy' since models already know how to cut spending. Real impact looks like incumbents doing things that used to be possible only for startups.
    2:05:25How would you advise people on how to think about budgeting for the transformation, given estimates like roughly 3% of human labor cost for ongoing inference?
    The ambition can be fulfilled at a price that would have seemed like fantasy a decade ago; if it doesn't look far cheaper, you're using the wrong time axis or numerator/denominator. Absolute dollars may still rise because more revenue and value are being created — he illustrates with a hypothetical CapEx/OpEx of $1 rising to $1.50 while revenue rises from $3 to $5 — and cites the global shortfall in air conditioning as an example of latent demand industry could meet far faster.
    2:07:57What are the human-in-the-loop situations that really exhaust human 'attention load' and should be handed off fully to agents?
    McCarthy applies ordinary management-hierarchy discipline: demand the right elevation of detail and don't let agents drag you into line-by-line minutiae, the same way a senior manager would redirect a report. He describes checking in with agents on a normal human schedule — overnight work, a bedtime check-in, a morning brief, even a drive-time brief.
    2:12:54How do you see competition shaping up — should incumbents worry about other incumbents moving into their lane, or AI-native startups coming from nowhere?
    Inventory what you have that a challenger can't get overnight: physical assets, regulatory approvals, patents, brand affinity, and customer inertia. Put a 'coefficient of inertia' on the parts of your revenue that are genuinely protected, then move fast on everything else before a startup does.
    2:16:52In terms of the Silicon Valley training week, do you try to overwhelm people with data or calm them down?
    Inspiration and conviction matter early, but Diffusion spends far more time on 'the how': getting a real result into production at scale and showing it happening within twenty-four hours, which is what actually converts skepticism into momentum.
    Lightly edited · timestamps jump to YouTube
    1:36:28

    Prakash Narayanan: It is Justin McCarthy. He is the founder and CEO of Diffusion, a company that helps large organizations build software factories, systems in which people describe desired outcomes and AI agents implement and evaluate the software. Diffusion works with the people who run a business to turn their knowledge into specifications, workflows, and working applications. Previously, Justin co-founded StrongDM and served as its chief technology officer. StrongDM manages access to infrastructure such as databases, servers, and computing clusters. He spent a decade building cybersecurity infrastructure there, and in July 2025 formed its AI team with Jay Taylor and Navan Shahan.

    1:37:14

    Prakash Narayanan: Their experiment imposed two unusual constraints: humans would neither write nor review the generated code. The team described a different basis for evaluating software — run it through realistic scenarios, including against simulated versions of services such as Okta and Slack, and use the results to guide further development. It also published tools and specifications for other builders. Justin's work now extends into research: with Atticus Kull, he co-authored ‘Specification Oracles,’ a study of whether a language model can learn the facts about a system and answer questions as a living specification. Across this work, his focus is what happens between expressing an intention and getting a system that reliably carries it out. Justin, welcome to the show.

    1:38:13

    Justin McCarthy: Alright. Thanks for the intro. Thanks, both. Hey, Prakash. Hey, Nathan.

    1:38:18

    Nathan Labenz: Great to see you. A lot of different directions we could go. When we — so we first connected, actually, originally because you became a sponsor of the Cognitive Revolution. And in digging into your methodology a little bit, I was really interested in the service that you offer, and especially what you've learned from offering this service of bringing leadership teams to Silicon Valley and trying to get them to grok how to use AI. So I've read a little bit about that, but I'd love to go deeper. Tell me — first of all, who has the self-awareness to know that they need this and is actually coming out and doing these sessions?

    1:39:03

    Nathan Labenz: And what have you learned about how to get people to really get it in a short period of time?

    1:39:09

    Justin McCarthy: Sure, well — good opening question. I'll say part of the answer is right there in the name of the company. In some ways, the name of the company and the mission, the idea, are all the same thing. Obviously, we have an extremely concentrated form of capability now in silicon — it's very dense, but it's in the data warehouse. And I'm not in the data warehouse, I'm out here in the office. So we need to diffuse what's capable into me first, and into all of us who are really mining the frontier for the answer to what it can do today. And every time we get a new model day — which is basically all the time — we have to ask those questions anew. So some cohort of folks are out there really mining the frontier, really

    1:39:55

    Justin McCarthy: evaluating what the capability is. But then there's just this reality that diffusing that next step into how do I apply it specifically into an incumbent firm — okay, specifically into a firm that already exists — that hop, there's just no rulebook. There's no manual, there's no path. There's — you know, there's actually too many blog posts and too many YouTube videos, thanks, thanks guys. So that content filtering, and really just saying, look, we could go a thousand different directions with this, let's try these three paths first, which have shown some traction over other adjacent problems. That's it. And what I'll say is conviction comes from experiencing

    1:40:40

    Justin McCarthy: the thing working. There's sort of a before-and-after moment. There's ‘I've heard about it, it should work’ — you know, maybe I did install Claude on the weekend, and this part is working at home, it feels like we should be further along by now. Connecting that into a process that's high criticality, there's just a couple of steps. And what I'll say is we're accumulating experiences on what works, what cracks problems — and then once you crack them, you're just asking the question: okay, now how do we scale this?

    1:41:15

    Prakash Narayanan: So, I feel like many organizations have this issue where they want to apply AI, but the prerequisites for that are often not ready — in the sense that they don't have their data in the right place, they don't have it in the right format. How do you address the pre-Diffusion problem?

    1:41:38

    Justin McCarthy: Sure — on the background of my phone I keep our seven principles, and I'll scroll down to the seventh one, which you could really make the first one in order: don't wait because you think your data's not in the right format. Existence is the correct format. What you need to do, maybe for a particular task, is reshape it — but guess what helps with reshaping the data? Guess what helps with formulating the eval and the benchmark? So for sure, you have the data. Actually, in the rare case where you don't have the data, spin up your general counsel, because you can buy the data. I think the famous example recently is the airline that went out of business — I point to that a lot. That price was

    1:42:23

    Justin McCarthy: de minimis compared to the entities that were involved. For sure, in your domain, if you don't have some workflow data in-house — and if one of the environment companies hasn't already hoovered it up — you can go buy that data. So please don't wait, start right now. That's a huge part of energizing the conviction that today's the day. Start now, don't wait.

    1:42:46

    Nathan Labenz: So what are the real blockers? I'm so deep down this rabbit hole that it's at times a little hard for me to remember — when I was starting, it was just a lot harder. I had to do original prompt engineering, training my brain to think, ‘okay, if this was a document on the internet and here's the beginning of the document, what comes next — now how do I engineer the beginning of the document?’ Obviously you don't need to teach those kinds of prompt-engineering skills anymore. What do you find are the real reasons people are stuck? Is it lack of know-how? Is it lack of shared vision on the team?

    1:43:26

    Justin McCarthy: Yep, okay — I'd say the blocker is this. We have another principle we talk about called ‘make success inevitable’ — meaning make the agent's success inevitable. So if you wake one of these models up in an environment where the first file it reads explains the mission, the second file — that mission branches into three forks — and it reads the most plausible next fork, and then maybe two or three more file reads, it understands the whole brief and it understands how to make the incision in that data, or the incision in that codebase — the agent's going to do it. On the other hand, if you toss it into a human-oriented codebase that hasn't been prepared for agents, it's going to load the wrong context.

    1:44:12

    Justin McCarthy: It's going to load the wrong files — success isn't going to be inevitable. So it's really about developing intuition for the implicit context engineering of the environment, where the shape of the environment inevitably guides the agent. If you close your eyes, you can see your business in terms of — I often talk about it as a valley where the food is at the bottom. If you do a random walk inside that valley, you're going to get to the food. So that token environment, for a given task, just has to be agentic. And designing that agentic token environment isn't obvious — these are fairly invisible things. We can do architecture and interior design in our physical

    1:44:57

    Justin McCarthy: environment, because we can see desks and chairs and doors and exits and stairs. You can't see which file is likely to be loaded as the agent is priming its context before an incision. You can, however, run a loop that evaluates whether that token environment is good enough — and that itself can be automated. So it's loops on loops. And pretty soon a question like, ‘when we get an inbound to our legal team, can we prepare for the next step and hand it off to the paralegal so they can make the state filing?’ — that's a set of questions you can just put a few loops on top of and get an objective answer to, today. There are very few answers that aren't available today.

    1:45:42

    Prakash Narayanan: Let's talk a little about the example you gave — I think it was Spirit Airlines. Spirit Airlines was a low-cost carrier, had a lot of customer complaints. JetBlue tried to buy it, the FTC stepped in, they were banned from buying it. Spirit then went bankrupt, and about a year later — you have Google. They competed for it in an auction: Merkur put in a $7.5 million bid, Google put in a $10 million bid. They bought the data, right?

    1:46:14

    Nathan Labenz: Now, could have been ours, Prakash — what were we thinking?

    1:46:17

    Justin McCarthy: I know.

    1:46:18

    Prakash Narayanan: So how do you make use of this data — what actually happens to it? Because I imagine you're not just using it for next-token prediction, since there's a lot more... so how does this data end up getting used? Is it only useful for an airline, or is it useful for people outside an airline? How does this actually work — how do you make something out of it?

    1:46:41

    Justin McCarthy: Obviously there's the training, post-training, RL side of it — if you have the sophistication to run some of those loops, you know what to do with it, because the environments you're feeding it into are helping improve the intuition of the agent. But even before that, even separate from that — when we're just working in token space rather than in the weights, working with the input tokens and output tokens from an off-the-shelf model — you're still using that data to define what good looks like for every subprocess within that airline. We might have a consumer relationship to an airline, but if you ask a pilot, or a first officer, who's putting the loading on exactly what's going into the plane

    1:47:26

    Justin McCarthy: for this particular segment, there's a ton of art in there that isn't represented on the internet. There's a ton of how-do-you-do-it, what thresholds can you get right up to versus what can you never cross, that's implicit in all of those logs. So again, it's just loops everywhere — you're putting a validation loop around something. Say you're internal to an airline, putting validation loops around some ops research and optimization you're doing, and you've never had access to an adjacent airline dataset before — that can show you that maybe they had a trick. Maybe there was a great weight-and-balance trick that was contributing

    1:48:11

    Justin McCarthy: to their success as a low-cost airline. Maybe they had problems elsewhere, but maybe that data is useful. So — as much as we can interview our human colleagues and our peers to extract tokens about what good looks like and how we'd validate it, the historical environments that have recorded the behaviors of airlines and organizations have answers too, typically at superhuman scale in terms of the volume of data.

    1:48:41

    Prakash Narayanan: Right — so it seems like it's tacit knowledge, kind of absorbed from the data. Is that the right idea?

    1:48:50

    Justin McCarthy: Another way to see this — it's the same thing the learning loops are doing inside the models. They're trying to disambiguate two cases: Prakash said this, Nathan said this, which one is right? We have some counter-evidence here in the corpus. So even if you can just reduce the question space for the humans through that data — take all the tacit knowledge implied in the behavioral traces of the airline, then sit down with Prakash and say, ‘should we have loaded two tons of fuel in this case, or two and a half tons?’ And the expert can say, ‘oh, it's obviously two and a half tons.’ And that confirms or disconfirms a whole bunch of expectations that then go back into your ‘is this right’ validation rules.

    1:49:35

    Prakash Narayanan: Right.

    1:49:36

    Nathan Labenz: Now — so is this a big part of what you're doing with companies on their data? You're trying to help them surface things that, in some vague sense, they already know but haven't documented, or don't have a shared understanding of?

    1:49:53

    Justin McCarthy: Sure — the tools and products we're developing internal to Diffusion are essentially the tools that codify the lens of engineering loops on real, complex organizations. You have real complex organizations with real revenue flows and real costs — you need loops everywhere, but those loops are fractal, they're biological. I really love the atlas of the metabolic system — I don't know if you've ever looked at one of those atlases. All those metabolites and everything, it looks like a completely organic, evolution-designed thing. There's no strict package hierarchy for the code. And companies look like that too, even though we

    1:50:38

    Justin McCarthy: nominally have departments and stuff — if only the facts rolled up hierarchically, this would be a lot easier. The facts flow in every direction, there are tendrils, inputs and outputs everywhere. So for each of those, for any given tendril, any given input, you can say: what is the loop on this that's going to take this substantially human process and give us five nines of reliability, at a speed we've never contemplated before? And rather than boiling the ocean across every process in the business, if you focus down on the ones that are really going to unlock market share and revenue, that's the aligning force — that's the field everyone can snap to and say, ‘yeah, I would love to double

    1:51:23

    Justin McCarthy: market share’ — you know, you don't have to sell me too hard on that. So just work backwards from that: work backwards, print the loops, get them running, get to five nines of reliability. That's the algorithm — but it takes tooling to even be able to print those loops.

    1:51:41

    Prakash Narayanan: Talking about five nines of reliability reminds me of Six Sigma, back in the day. Is it possible to get to that kind of reliability using a lot of the ML processes? Because sometimes you don't see LLMs react that well, or provide accurate, good decisions.

    1:52:04

    Justin McCarthy: I'm on the sixth floor of a steel-and-concrete structure right now, and it's standing — and it was made out of very unreliable systems. So yes, we know how to composite unreliable things together, we know how to multiply unreliable components in a network to create reliability. And your question is a great one, because it imagines that the previous source of intelligence had high reliability, when the previous source of intelligence — which was you and me — had all sorts of performance problems in terms of accuracy. But the great thing is we can composite all these numbers together and create a structure that stands. That's why the Golden Gate Bridge works.

    1:52:48

    Nathan Labenz: I used to say I'd be interested to hear a little more about how you see this playing out in practice today, because it's been a while now, but I used to have a talk that I went around giving on how to create workflows to automate processes in businesses. And what I would tell people — I kind of stopped doing this maybe around the time GPT-4o fine-tuning became available — but I used to tell people that with some elbow grease, you should be able to get the vast majority of processes to the same level of quality and reliability as a human. But every nine you want to push for, you should probably expect to be ten times more effort.

    1:53:34

    Nathan Labenz: Because you're gonna have to curate all that data, you know — it's tougher, you've got to find the edge cases. So if you still feel like that is the way, do you think businesses should be pushing for five nines? I mean, do they have it today?

    1:53:48

    Justin McCarthy: You should never be pushing for five nines for something where your business can succeed with one nine. Your spectrum of nines, your gamut of nines — that's a signature of your business too. What is the nines of reliability for succeeding with a brand campaign, versus depositing via ACH into a checking account? I think most firms don't necessarily have language for this, but it's very easy to ask: what are the consequences of three nines versus one nine, versus no nines, versus five? And they can answer those instantly — it's not something they've had to talk about, because everyone knows the financial deposit

    1:54:35

    Justin McCarthy: process has to be reliable — they've never had to externalize that before, but you do externalize it. Let me go on a bit of a digression: there's a habit all of us are vulnerable to, when we have these agents in the lead, of getting into a mindset where we're gold-plating and perfecting and adding nines before we know what we want. There's a huge branch in the ambition-fulfillment process where you have to discern: am I making something new, or is this brownfield and I'm just reproducing what I already do but faster and cheaper? Those techniques are totally different, and it's very easy to mix

    1:55:20

    Justin McCarthy: them up, because it all just looks like — it just looks like Astra. You can assign Astra a task, but if you don't have the discernment to know: am I discovering product-market fit, am I discovering a better design, am I discovering more elegance, versus am I just making it faster and cheaper — the models can solve both, but you can't cross the streams very easily.

    1:55:44

    Prakash Narayanan: One of the questions I have is on interpretability, because a lot of times models can give you an answer but you're not sure why. And in a business, that's very important — say, for a bank that has to approve or reject loans. If you have a FICO score, and the FICO score is built on things considered redlining, then that's illegal. There's always been this vague suspicion that models basically extract data and

    1:56:29

    Prakash Narayanan: form inferences in ways that would be illegal if they were formed from something we could actually interpret and see. So how do you address that interpretability issue that businesses need?

    1:56:41

    Justin McCarthy: Yep, okay — I'd say the first technique is to turn the model directly on the problem. If we have a statute we have to conform to in the compliance environment, don't treat it as something you're tacking on — treat it as a first-class problem you're directly facing. The jurisdiction, the legal or compliance or regulatory environment you operate in — that's your physics, and you can't violate physics. So you need a part of the system dedicated to that. But you also have to know — this is one of the weaknesses of the models — the models are horrible at taking risk. So operators of the business need to set thresholds that are

    1:57:26

    Justin McCarthy: right adjacent to a statute that's never been tested in court before — it's written one way in the law, it's never been tested, so there's no precedent to say objectively how it's going to be tested. So you need managers to be able to set the business threshold right next to that. The models aren't going to do that for you. So first, address it like your physics, then make sure you're in control of the risk thresholds. And once you have those concerns walled off, let me say something about the reliability of claims the system is making generally. This is something very senior managers have great intuition for, because they've never

    1:58:12

    Justin McCarthy: seen the transaction entered into the ledger themselves — the CFO last entered something into the general ledger maybe thirty years ago. And yet somehow Walmart manages to close the books. How do they do that? With top-down interrogation that is essentially verbal checksums that the whole system, the whole process, is working. So I turn to Nathan and say, ‘Nathan, are you sure we're closing the books correctly?’ And he says, ‘yes, we're closing the books correctly,’ because Nathan is rolling up all the yeses from the folks who report to him, and they're rolling up all the yeses that eventually come out of Oracle Financials or NetSuite or whatever. That process of ‘are you sure’ is the same interrogation.

    1:58:57

    Justin McCarthy: Every management structure has it. We have some famous leaders in our industry — Jensen has sixty, or whatever it is, a hundred and twenty direct reports this week, and he's going

    1:59:07

    Prakash Narayanan: Oh, the number keeps increasing, you know — that keeps increasing. Yeah, the number keeps repeating.

    1:59:11

    Justin McCarthy: And he's going to interrogate until he understands, and then he's going to move on. I'll name-drop Steve Yegge — one of his posts from many years ago, on presenting to Jeff Bezos, had this character too: when you're presenting to Jeff at Amazon, you have to be ready for drill down, drill down, drill down, drill down, or nothing. Jeff could drill down all the way to your shoelaces because he knew the details so well. So what we say at Diffusion is ‘demand understanding’ — it's our version of the famous tweet, ‘you can outsource your thinking, but not your understanding.’ I demand understanding. I'm unwilling to ship the thing if I don't know how it's composed. But my definition of knowing how it's composed is the contours

    1:59:57

    Justin McCarthy: that are salient to me from my elevation in the business. And I'd just say the same thing: be Jensen, be Bezos — demand understanding of your hierarchy, whether it's silicon or biological.

    2:00:12

    Nathan Labenz: On the human side, what's the shape of leadership teams you find companies need to have involved for the transformation process to work? Do they have to bring the stubborn skeptic holdout, or do they leave that person at home? And if they do leave that person at home, does that pose a risk to the whole thing? I assume you get groups with varying levels of buy-in — what are the prerequisites for a leadership team to come to you and be successful?

    2:00:50

    Justin McCarthy: Sure — first, skeptics are totally welcome. Throw your skeptics on the plane, that's great, it's important. But it actually starts with conviction — it has to start with conviction at the top, up to the board. What a member of the board should say is, ‘we believe we're actually in a different regime, and we believe the timing for our industry is now.’ Now, if I were on the board of Kaiser Permanente, I might say we should be dabbling with this stuff, we should be ready for when it's ready — but people are going to need their appendix out tomorrow no matter what. And so

    2:01:35

    Justin McCarthy: in healthcare there's stuff that can come later — that's different from research, that's different from what Moderna should be doing and what Merck should be doing today, I think they're doing different things. But if I'm just installing and uninstalling appendixes all day long, that's probably going to look pretty similar for the next couple years, so let's not get over our skis on that. On the other hand, if your business model, your customers, what you've offered your customers, is clearly in the blast radius of what the models can do now or in the next couple years, you probably need to act urgently. And that clarity needs to exist with the board and in the CEO's head — pop-quiz the CEO and they'll say ‘yes, now is the time.’ So it has to start there. And really after that,

    2:02:21

    Justin McCarthy: yeah, I mean, maybe it'd help if there are some people who've used the models before.

    2:02:25

    Prakash Narayanan: One of the criticisms I often hear — and I've seen this myself — is that we often spend a lot of time using the models and feeling like we're doing a lot of work, but it's just a feeling, and the impact isn't really that evident a lot of the time. I've heard a bunch of people say this. When do you think we'll see real impact — massive revenue increases, cost decreases? When can you actually start seeing big differences in, let's say, margin and growth, which are basically the two primary things any business looks at?

    2:03:11

    Justin McCarthy: First, let me anchor everybody in the standard of a couple years ago. Do you remember all your friends saying, ‘I love working at a ten-thousand-person company, it's super easy to get stuff done’? Remember when your friend said that — like the Carnot efficiency of this machine is really great? Remember when they said that? So the standard we're measuring against, you've got to keep that in mind. But I'll endorse this: the only way we know this thing is working is revenue and market share. Costs are too easy — they're under the lamplight, you can see them. Don't talk about costs, it's too easy. The models

    2:03:56

    Justin McCarthy: — if you ask them, ‘should we have better margins or worse margins,’ they already know how to do that. ‘Should we spend more money or less money,’ they already know how to do that. That's not a real question. How are we going to fulfill more promises, exceed expectations, gather more customers, retain them longer, and get a larger share of wallet? That's the only interesting conversation. What most firms find is, if you're lucky enough to be ASML and you figure out a way to print ten times more of the printers, you're going to have customers. Or turbine fan blades — if you make ten times more, you're going to find customers. There are a lot of firms where their customers can't adopt ten times as much change. They're

    2:04:41

    Justin McCarthy: full of software already, you can't feed them more software. So this is a case where you have to figure out the brand envelope — what conceivable product offering you could do under your logo — and start thinking in ways you haven't before, about new ways to invite more customers under the tent with different products, while sustaining their existing promises and systems as well. It's actually kind of a creative moment for a lot of these incumbent firms, because they're doing things that in previous regimes only startups could do. But in a lot of cases they're succeeding. And I do think revenue and market share are the measures — don't talk about anything else.

    2:05:25

    Nathan Labenz: If people ask you what this is going to cost — I've heard estimates that range, but I'd say 3% of human cost is kind of the central estimate I've heard for how much you should expect to spend on inference over time to replace human labor. Of course there's then the setup cost, which I don't think is really included there. How would you advise people on how to think about budgeting for the transformation?

    2:06:00

    Justin McCarthy: Sure — it starts with the good news that your ambition can be fulfilled at a price that ten years ago would have just seemed like fantasy. That's the good news: whatever the cost is, it's way cheaper. And if it doesn't look way cheaper, you're choosing the wrong time axis, or the wrong numerator and denominator — zoom out, look at the whole P&L, find the right numerator and denominator in the right time period, and it's going to be way cheaper. The reason the absolute dollars might go up is because you're getting more customers, more revenue, there's more value in the economy. So hopefully, if your CapEx and OpEx last year were a dollar, next year you're spending a dollar fifty.

    2:06:46

    Justin McCarthy: But your revenue went from three dollars to five dollars. Hopefully you're seeing those effects soon — that should be the management conversation. And I'll give the example everybody can intuit: I do these calculations all the time to fortify my convictions, but about a quarter of the people on Earth who'd want it have air conditioning right now. So we have a couple billion air conditioners and HVAC units left to build out to get climate control for everyone. That's a lot of copper, a lot of motors, a lot of power engineering. Those are big devices.

    2:07:32

    Justin McCarthy: So even just that one thing that everybody would regard as an essential of human life — there's a lot of industry that needs to turn over. Nathan, why are we letting that not get done by 2027? Everyone could have HVAC next year. So why are we still talking about scarcity in any particular market segment? There's clearly a lot to do.

    2:07:57

    Prakash Narayanan: I want to go back a little to what you said earlier. It seems like AI agents often end up consuming a lot of human attention, because the human ends up being the bottleneck in the human-in-the-loop — you end up with what I think you call ‘attention load’ on the human, where they're forced to make a lot of decisions, and some of these are decisions they'd actually prefer to hand off to the AI. There are companies where you're supposed to open up a CRM or ERP and just click ‘yes, yes, yes, yes, yes’ on everything, just so they can say a human has actually paid attention to this. What are the kinds of human-in-the-loop situations that really exhaust human attention and should be handed off fully to agents, but we're still kind of hanging on to?

    2:09:03

    Justin McCarthy: Yep, okay — let me call on the great debate of whether these systems are going to push our organizations more like holacracies, these big many-to-many webs, versus more like a military organization, very clearly a top-down command structure. I have to say, in any hierarchy of managers — if I come to you, Prakash, at a very senior level, and start talking about the details of the lines of code I wrote that week, you're going to kick me out of your office. You're going to say, ‘how is this the right conversation?’ You already know how to kick me out of your office, you already know how to chastise me as a colleague and then coach me on

    2:09:48

    Justin McCarthy: how to manage up, how to manage down, how to manage sideways. It's the same thing, and I don't let my agents drag me into those details — just like I demand understanding, I demand the right elevation. It may be the case that means they have to go back through all the history, mine all the documents, and spend a couple of megatokens reformatting — but if it's not at the right elevation, it's out of here. And when people see that's how we're working, then — the other thing, you have to get eight hours of sleep. Listen to Brian Johnson, get sleep. You should not be looking at eight screens all

    2:10:33

    Justin McCarthy: day long. You can be walking the dog, having a dialogue that's routing among those many agents. Walk the dog, have a dialogue, make sure they're working overnight — it's reasonable to check in right before bed, then get eight hours of sleep, then have your morning brief, and if you have a drive, have your drive-time brief. You should have all of this.

    2:11:00

    Prakash Narayanan: So does that mean organizations should start reclassifying some of the checks they demand — which are really nonsense checks, but they just do for compliance reasons? Should they start reclassifying these and

    2:11:14

    Justin McCarthy: Oh, I see, I see. Yeah — here's a good one. A lot of organizations use the SOC 2 process. SOC 2 is an accounting-origin process that flowed through IT and into software — it says, ‘you can trust me, I'm responsible.’ The people who define the controls in SOC 2 and evaluate whether you're hitting those controls come from an auditing and accounting background. It's a reasonable historical way of communicating that I'm a real organization and I'm trustworthy. But it also became gameable — now it's hyper-gameable. So rather than hyper-gaming this and turning that badge into something fake, we

    2:12:00

    Justin McCarthy: should just have a new thing. We should renegotiate with our auditors and our customers and say, look, we wrote these controls in the before-times. The good news is, for any given organization facing this compliance question, you're not the only one — your auditor, any regulator, everyone is dealing with this right now. The good thing is the auditors and regulators are also still people, and they want to have a conversation like: ‘okay, Prakash, let's be realistic — you're producing ten times as much of whatever information this year, let's start talking about your hierarchy of checksums.’ Again, Walmart closes the books, they're familiar with very deep hierarchies on having numbers reconcile. Well, your intentions can reconcile at those steps as well, and the auditors

    2:12:45

    Justin McCarthy: and regulators know how to talk about that, especially if you know how to map it into your agentic loops.

    2:12:54

    Nathan Labenz: Fast-forwarding into the future, potentially not too long — how do you see competition shaping up? If I'm an incumbent today... one thing that's super striking, especially in software, but I think it'll increasingly come to more and more non-software industries too, is people are just doing everything. In software, all the adjacencies are now in play, everybody's building out increasingly horizontally, and it seems like this inevitably leads to — I'd be interested to hear what you recommend to your clients today. Should they be reducing the number of software vendors they have? Maybe they increase

    2:13:39

    Nathan Labenz: a little to experiment, but then ultimately reduce, because you're going to get all these different solutions in a more integrated way from the same platform? And does that mean, if you're an incumbent now, you need to worry about other incumbents in the next lane over coming into your lane, or do you need to worry about AI-native startups coming up out of nowhere? How do you see the competitive landscape shaping up?

    2:14:07

    Justin McCarthy: This is a great question. I was just having a conversation earlier today about some internal accounting topics, and how folks like Brex and Ramp and Mercury are all mutually expanding — by the time you launch a new feature in one, the other has it too. That's an interesting effect, and you're going to experience it in a lot of areas. Rather than worrying about that froth, and worrying about the two-person startup, I'd say go to the other pole of reasoning and ask a couple of questions. What atoms do we have — what physical mass do we have that Nathan doesn't have? We have a lithium refinery.

    2:14:52

    Justin McCarthy: Does Nathan have a lithium refinery? No. Does he want one? Sure. Can he get one? No. So if you have atoms, that's a good thing. What regulatory approvals do we have — those give you a ton of protection for a while. You could say patents and other things, especially in pharma. And you keep going down that list, and somewhere in that list is a really important one, which is brand affinity. And I'd also say customer inertia — in other words, if you have a kind of customer that just doesn't want to change their fiber-optic provider

    2:15:38

    Justin McCarthy: because it's working — if it's working, that's actually a real thing. So you sum up, put a coefficient of inertia on those stripes of your revenue, and then on top of that you develop conviction about the things that probably aren't going to change. That helps you bound an envelope: great, we don't have to worry as hard about this — what's the most we can make of that? And that's the exciting thing for me about working with incumbents, working in brownfield — there's a kind of, I don't want to overplay this, but a learned helplessness that's kind of inevitable if you've been in the same company for a hundred years, in the same industry for a hundred years. This is a reset on that. It's not just the cool kids who get to make stuff now — you're a cool kid now too.

    2:16:23

    Justin McCarthy: You can make stuff. And the startup that was just founded last month doesn't have fifty years of history with those industrials who are all happy customers on a really critical process. So have high conviction about the things that really won't change overnight, and for the things that can change overnight — don't let the startups get you, you do it, you change it overnight.

    2:16:52

    Nathan Labenz: Time flies — last question for me, then we can let you get back to work. In terms of accomplishing that reset for people experientially, when they come out to Silicon Valley and do a week with you, what kind of experience do you try to create to get through to them? For background — I used to have talks specifically designed to overwhelm. My point initially was, I don't think you're taking AI seriously enough, I'm going to overwhelm you with unbelievable data points that show all the reasons you should wake up and start paying attention. Now I'm thinking about shifting that, because I think people are

    2:17:37

    Nathan Labenz: already overwhelmed enough, and maybe what I should do is come in and soothe and calm instead. So where are you on the overwhelm-and-instruct approach versus taking the temperature down a little and letting them relax?

    2:17:50

    Justin McCarthy: Sure — I'd say inspiration and feelings, it's important to develop conviction, but once you have it, you're getting into the how. We spend far more time on the how: how are we going to continue, how are we going to finish, get this to production at scale, and then do it again and again with five thousand people? That's where we're trying to get to. That threshold crossing, from skepticism to ‘the time is now’ — the instruments that achieve that are basically something in the context of your business that seems impossible. It completes the sentence, ‘this would be amazing if...’ It's clearly

    2:18:35

    Justin McCarthy: double-digit millions, triple-digit, quadruple-digit millions of value, on some time interval you can actually see and measure and rally people to fulfill and follow through on. And then just seeing it — once you have it, once you're seeing it, you're no longer talking about ‘if.’ And especially if you show that happening within twenty-four hours, then you've unlocked it: okay, now how do we do this? That's it.

    2:19:10

    Nathan Labenz: Cool — well, Justin, thank you so much for joining us today on AI in the AM. The company is Diffusion, and an intensive week awaits for any executive teams bold enough to make the investment and get out there and see you. I look forward to more stories of transformation from you before too long.

    2:19:30

    Prakash Narayanan: Justin, thank you.

    2:19:31

    Justin McCarthy: Thanks, boss. Take care.

    • Revenue Beats Cost Savings

      0:00 / 0:00
    • Your Data Is Already Ready

      0:00 / 0:00
    • Make Agent Success Inevitable

      0:00 / 0:00
    • AI Diffusion Has No Rulebook

      0:00 / 0:00
  4. 2:20:04Closing17 min
    Closing — Singapore's MRT home-purchase study, jubilee versus deterrence, and AI as the stationary banditWith both guests gone, Prakash Narayanan lays out an NBER working paper alleging that mid-level Singaporean civil servants and their relatives disproportionately bought homes near MRT stations before the stations were announced — built from public land-transaction data and an LLM classification of the civil-service directory — and Singapore's Public Service Division's response. Nathan Labenz argues for something closer to a jubilee than prosecution, on the grounds that deterrence was always priced for a world where most offenders never got caught. Then Peter Thiel on truth and reconciliation, Mancur Olson's roving-versus-stationary bandit, whether AI becomes the next stationary bandit, next week's lineup, and the show's new live chat.
    Open segment on YouTube ↗

    With both guests gone, Nathan and Prakash closed the show themselves. Prakash used the "segue" to introduce a story with personal resonance: Singapore's government is reviewing a study by a team of US economists alleging that civil servants disproportionately bought homes near soon-to-be-announced MRT stations before the sites were public. Because Singapore's government owns essentially all the land and captures dramatic value increases whenever it announces new subway stops, the study — built on public land-transaction registries, a civil-servants directory, and LLM-based classification of officials by rank — found that mid-level, not top-level, civil servants, along with their relatives, coordinated buying in target areas up to two years ahead of official announcements. Prakash called it "garden-variety" insider trading made unusual only by Singapore's carefully cultivated reputation for incorruptibility under Lee Kuan Yew. Prakash also said on air that the study's lead researcher is Lee Kuan Yew's grandson, now in the US after a Facebook comment drew a government inquiry — what he called the "settling scores thesis." (For the record: the NBER working paper is by Tomasz Piskorski, Amit Seru, Chun Zhao and Jian Zhang; Li Shengwu, Lee Kuan Yew's grandson, is thanked in the acknowledgments as a discussant, not an author.) The broader point stands on its own: AI and public data now make legible a kind of corruption that always existed but was never provable at scale.

    The two debated what a government should do once 10-20% of its civil service is implicated by newly legible evidence. Nathan argued for something like a "jubilee" — a generous, threshold-based forgiveness or one-time financial restitution rather than criminal prosecution, reasoning that the old deterrence model assumed most offenders would never get caught and can't survive a world where AI makes everything traceable. Prakash pushed back that civil servants understood the site-selection algorithm better than the public, creating an unavoidable conflict of interest. Nathan connected the dilemma to AI's own "motivated reasoning," citing a recurring theme from his Cognitive Revolution conversation with Bronson Schoen — that models often generate elaborate justifications that aren't really what's driving their outputs — as a reason for leaning toward forgiveness on historical misbehavior generally.

    A listener, Christian Alstrip, pointed them to Peter Thiel's 2025 Financial Times piece "A Time for Truth and Reconciliation," prompting Prakash to lay out Mancur Olson's "stationary bandit" theory of government — that states begin as roving bandits who, by settling down and monopolizing violence, find it more profitable to provide protection and tax than to keep raiding. He argued AI could become the next "stationary bandit" as governance shifts from human-led to AI-led systems, illustrating with Malaysia's experience prosecuting systemic corruption after a change in government, where unwinding even a few percent of GDP in graft implicated an entire network of officials, not just one person.

    Nathan closed with a preview of next week's Monday/Wednesday/Friday lineup: a guest from Coefficient Giving discussing their grants and projects (following up on the "Tailwind" project that had overshadowed other news); a guest working on physical-sensor integration into world models for industrial and field settings; a guest running high-scale perturbation experiments on human tissue to generate the kind of biological data prior guests have said is missing; and a guest from Cooperative AI for what Nathan called a positive, future-oriented conversation. Prakash closed the show by noting the live chat system is now working — viewer messages arrive with a 10-15 second delay and pass through a loosely-tuned AI moderator that will get stricter over time.

    I think we should lean toward forgiveness where we can on these historical misbehavior questions.

    All governments are basically thugs and bandits who decide to settle down.

    Lightly edited · timestamps jump to YouTube
    2:19:37

    Nathan Labenz: That was a— Prakash Narayanan: Segue. Nathan Labenz: The official segue?

    2:19:39

    Prakash Narayanan: So, I wanted to share this — it's perhaps a little more personal, in the sense that I have some prior knowledge of it. This week we have a report: the Singapore government is reviewing a study alleging civil servants disproportionately bought homes near unannounced MRT stations. In Singapore you have a very efficient subway network, most efficient in the world, and

    2:20:24

    people want to live within, like, five minutes' walk of the subway, because then you can just use the subway and you don't need a car — it's very, very convenient. And, obviously, the government also owns all of the land in Singapore, right? So it's a Georgist utopia — the government owns all the land. So whenever they decide they're going to have a subway system somewhere, they're able to buy up the land the subway is going to run on. And all the existing housing around that subway, which they do not buy, increases in value dramatically — dramatic increases in value. So there have always been rumors —

    2:21:09

    always been rumors that some people know and start buying ahead of time, etcetera. This goes back decades, from the nineties — that's when the subway really started to pick up. So, nineties, 2000s, 2010s, 2020s. And this project is a team of US economists, and they went after largely public data. You can find registries of transactions, similar to Zillow. You can find names of civil servants in the civil-servant directory. And then they also did a little bit of AI work — they used LLMs to classify these civil servants into various groups,

    2:21:54

    by tenure, by where they were ranking, etcetera. And what they ended up finding is that the mid-level guys — not the top-level guys — up to two years before an announced train station would start buying into these areas. And then their relatives, their in-laws, etcetera, would also start buying in. So, basically, coordinated buying behavior by mid-level, not top-level, civil servants — because the top level is very visible. Now, this is a garden-variety —

    2:22:40

    garden-variety insider-trading corruption, right. It's unusual because it's in Singapore, and Lee Kuan Yew had a very strong 'the government should be incorruptible' stance, with very strong punishments for this kind of thing. So the government has always wanted to appear incorruptible, but they've never been able to enforce at this level of granularity. And in fact, I've heard from people there that there's a lot of this stuff — a lot of insider trading on the share market, etcetera. At one point there was a move by someone to do a graph of all the interrelated —

    2:23:27

    chairmanships and board memberships. And as you can imagine, in a country of about three million people — about a third the size of Lake Tahoe — if you map out all the board-membership relationships, then map out who their in-laws and siblings were, and then map out where the trading was done, you start to see these problems emerge, where someone leaked something at a barbecue and, you know, a merger happens and the price moves. We've never been able to have the level of surveillance you need to identify that smaller-level stuff. And here you have an example of basically thirty years of corruption starting to get exposed,

    2:24:12

    and the government having to react in real time. What are they going to do? Maybe ten or twenty percent of the civil service is now implicated. Do you imprison them? Because that's what you've always done in the past — in the past, a single one of these cases, you basically go to prison for five years. So now you have ten or twenty percent of the civil service implicated, and you have proof. What do you do? I think this is the kind of thing I expect AI to be able to do — these facts exist in the world, but they're not legible in the way systems need in order to consume them. And I'll also note one specific thing:

    2:24:57

    the grandson of Lee Kuan Yew is in the US — he's in exile because his uncle wanted to hang on to political power a little longer than he should have. This guy said something on Facebook, and the Singapore government opened a query — if he goes back to Singapore, he'll go to prison for that Facebook comment. So he's in exile. He's an economist, and he actually helped the team that put the study together. This is what I call the settling-scores thesis — you suddenly have the ability to go in, get the data, and show proof that

    2:25:42

    even the cleanest of governments has a bunch of this stuff going on. And then you have the dilemma: what exactly are you supposed to do now? This is going to be the same thing when the Trump administration people leave office — they'll pardon a bunch of people, but there'll also be people who aren't pardoned. And when the feds stop you and ask you a question and you dodge or say something wrong, that's perjury — that's how enforcement has always been done. So what do you do in these cases, when we now have the ability to chase down these paper trails? That's my —

    2:26:27

    spiel. Like, do you forgive, or do you follow the rules you've set in the past, strictly? What should we do?

    2:26:42

    Nathan Labenz: Great question. I come down pretty intuitively on the side of some sort of jubilee, or some other canceling of debts, at least under a certain threshold. I don't think you'd want a blanket pardon of all crimes ever committed without any qualification, but I do think we're going to need some fairly generous threshold that just lets people get away with a lot of this stuff. Or maybe we could have new — I could also see, possibly, some

    2:27:28

    new make-right provisions for some of these things, that might not be on the level of what the law would actually prescribe. But they clearly can't send all these people to jail for five years, right? So — I think you could make a case, and it's going to be hard, but if you can map all this stuff, maybe you could get to something that could work. How's it going to be legitimate? I mean, in Singapore they maybe don't have as much of a problem with that — the government can just make the policy and it'll be what it is. Here, I think it'd be a much bigger conversation, but I could see some sort of, 'you did this, we kind of know you did it, you pay this financial penalty' —

    2:28:13

    that claws back some of the windfall you got. You get to keep the house — we're not going to take everybody out of their house. You're certainly not going to jail, but you pay a one-time restitution, we call it good, and we move on with a new social contract. Really, it comes down to the fact that the old social contract is based on not catching most people — you have to be harsh when you do catch someone, to deter the others, because in expectation people aren't likely to get caught, so the penalty has to be high enough to be an effective deterrent. We're definitely going to need to rewrite that, especially for

    2:28:57

    Prakash Narayanan: Historical—

    2:28:58

    Nathan Labenz: —crimes. So I guess, yeah, my recipe would be: pay a one-time fee, get out of jail for that. And going forward, maybe you really do expect to be caught, and maybe the punishment doesn't have to be so draconian and can still be an effective deterrent.

    2:29:16

    Prakash Narayanan: We have a question: maybe civil servants understand the algorithm used to pick the location better than the public — it's not randomized location selection. That is true, they do understand the algorithm, but that actually puts you in a conflict-of-interest position, knowing the algorithm. So regardless, you can't really get away with knowing the algo and picking correctly. It's a tough one.

    2:29:45

    Nathan Labenz: Yeah, some of this stuff is just tough. We're not robots — even the AIs turn out not to really be robots. So we're all in this, you sort of know, you sort of don't know. Can you really run the process of buying your home as if you don't know? Are you going to deliberately buy away from the desirable location so you can say, 'yes, I definitely did the right thing'? That's a lot to ask, I think. We're a little too integrated, and it's a little too easy for us to engage in motivated reasoning and kind of —

    2:30:31

    you know? We see this with the AIs too — this is a huge takeaway from the Bronson Schoen episode, which I guess I've mentioned enough times. People should go listen to the full Bronson Schoen Cognitive Revolution episode, it's fascinating. One of the things he emphasizes over and over is that the models seem to be engaged in motivated reasoning — they're considering all these different angles, making all these different arguments, but sometimes it feels like the arguments they're making aren't really what's motivating them, more a cover for what they really want to do anyway. And, obviously, we're guilty of that same behavior, and sometimes it's obvious and sometimes it's really not. I think we should lean toward forgiveness where we can on these historical misbehavior questions. So —

    2:31:18

    Prakash Narayanan: Christian Alstrip, a listener, refers us to Peter Thiel's Financial Times piece from last year, 1/10/2025, titled 'A Time for Truth and—

    2:31:31

    Nathan Labenz: —Reconciliation.' I'll have to read it — just because that's the title, I don't know what exactly I'd expect from Peter Thiel, but it sounds like something to put on my exponentially growing reading —

    2:31:48

    Prakash Narayanan: —list. Yeah, it's going to be interesting. I call it the Mancur Olson theory of government — Mancur Olson is an economist and political-science theorist, and he theorized that all governments are basically thugs and bandits who decide to settle down. And as they settle down in a location, it became more valuable to provide protection services permanently and take a tax. So he calls it the stationary —

    2:32:18

    Nathan Labenz: Bandit —

    2:32:19

    Prakash Narayanan: — theory of government. The thing about the stationary-bandit theory of government is that all governments have certain things in common, because they start off claiming a monopoly on violence for themselves. And after that, they look the other way — they turn a blind eye to some things that people on their side do, monopoly of violence being one of them. You have police immunity in the US, for example. So there's a bunch of, 'this is what we need to govern, so we're going to turn a blind eye to the rules we apply to everyone else.' There's a little selectivity in the application. I wonder to what extent, if we see a transition from

    2:33:04

    a human-led government to an AI-led government, then the Mancur Olson stationary bandit is the AI — and that has consequences for our current governments. I think that may be the case. I've seen revolutions — for example, in my country, in Malaysia, we had a very corrupt guy, and he lost. And after he lost, you know — when you take three percent of the GDP of a country out, you can't take it out alone. It's not one guy, it's an entire system of people — thousands of people are involved. The central-bank governor who was, you know, Asian financial —

    2:33:50

    — journals' 'governor of the year' — her husband was involved, right. She had to make the phone calls to allow stuff like this. So I think this is the same thing we'll face — it's not just going to be one person, it's going to be an entire system. The stationary bandit at work has an entire structure, and you might need to change it. I think people like Andy Hollis are working on that. So, let's see. Well —

    2:34:23

    Nathan Labenz: We might have to leave it there for today, but next week is going to be a good one. We've got Monday, Wednesday, and Friday scheduled, and we'll have a guest from Coefficient Giving, who recently put out their Tailwind project and thought that was going to be the biggest social-media story they were involved in at the time — that's turned out not to be the case. But I think we'll mostly focus on the grants and projects they want to fund. We're going to get into some stuff around physical-sensor integration into world models, but very practically — all these data points in industrial and field settings.

    2:35:09

    One guest from a company doing high-scale perturbation experiments on human tissue, to try to get exactly the kind of data that all our previous bio-focused guests have said we don't have — and therefore we're not going to get as far in biology, on the timeline I tend to think, as I'd like. So we'll be able to interrogate a little bit what their data-collection operation looks like. And we'll have somebody from Cooperative AI too, which I think will be a really interesting conversation — a very positive, future-oriented vision. As

    2:35:44

    Prakash Narayanan: — as requested. As you requested.

    2:35:47

    Nathan Labenz: Yeah. And even a little bit on the financialization of the compute market. So we're going to cover a lot of ground next week, and it should be a good one. For now, I think we'll sign off and say thanks for being with us on —

    2:36:02

    Prakash Narayanan: — AI. With us — and we have some interaction today. We finally have the chat system working, so if you're tuning in, you can actually send us messages. The messages are, like, ten, fifteen seconds delayed — there's a moderator in there, an AI moderator, very loose for now, but we'll tighten it over time. So, hope to see you guys on the stream. Bye-bye.

A note on the transcript

The am-live episode-notes-package export was not available for this show, so the master recording was re-transcribed with Deepgram and split by the rundown's own segment boundaries. Speaker labels come from those diarization clusters, matched by text against the studio's per-participant live speech-to-text, which gives a structural identity signal rather than an inferred one. Cluster purity against that signal was 92 to 100 percent in every segment.

Voice prints were run as a cross-check and got both guests backwards — they assign a guest to each segment by rundown order, and the rundown's order was not the order the show ran in. They were used only to confirm the two hosts, and not for either guest.

Cameron Berg joined by phone from an airport, so the live captions covered his segment poorly. His segment's speaker map rests on the text cross-check plus a read of who is actually saying what, which is why it is stated here rather than left implicit.

A handful of names were transcribed phonetically and have not been confirmed: a co-author on Justin McCarthy's "Specification Oracles" work, Cameron Berg's collaborator on the relief-button experiment and that collaborator's university, the viewer who pointed the hosts to the Peter Thiel piece, and a company name inside McCarthy's data-auction anecdote. They are left as heard rather than guessed at.

Why the show ran two hours and thirty-seven minutes

Both interviews ran long against a plan of twenty-five minutes each. McCarthy went forty-three minutes and Berg sixty-six — and Berg's segment kept going for about twenty minutes after he left for his flight, with the two hosts working through what his results had done to their own credences. The opening landed on plan at thirty minutes. The seventeen-minute closing was not in the plan at all.

The two guest slots also ran in the opposite order from the loaded rundown. Cal.com had McCarthy at 9:30 Pacific and Berg at 10:15; the rundown had Berg in slot two. The hosts ran the Cal.com order and left the rundown alone, the edited episode restores the rundown order, and this page's transcript links follow that published edit.