EPISODE 2026-09-30

AI Utopia and Rare Disease Diagnosis

Joel Borgen on AI-assisted fiction and human agency; Daniel McKinnon on Gamow Labs, rare-disease diagnosis, and the limits of model scaffolding. Nathan and Prakash discuss OpenAI’s DevDay, AI welfare, and independent audits.

▶ Full show on YouTube𝕏 Live broadcast

What would a good AI future preserve, and what could AI already help us recover? Joel Borgen joins Nathan Labenz and Prakash Narayanan to discuss The Receipt Horizon, writing with AI, shared culture, and the freedom to choose different ways of life. Daniel McKinnon then describes how a family tragedy led him to build Gamow Labs, using AI agents to investigate rare genetic diseases.

The hosts bookend the interviews with OpenAI’s DevDay, the tension between faster agents and oversight, disputed questions about AI welfare, government services, and the case for independent safety audits. Benchmark numbers and clinical results are presented as McKinnon’s reported findings, not independent validation.

The rundown

  1. --:--Opening
    OpenAI’s product direction and model-release choicesThe hosts discuss Dots, the DevDay audience, and questions about how safety decisions differ across frontier model releases.
    Open segment on YouTube ↗

    Nathan and Prakash review OpenAI’s DevDay through two competing lenses: useful products that could bring agents to many more people, and frontier capabilities advancing faster than meaningful oversight. Nathan sees Dots as a consumer-facing version of the personal AI infrastructure he already uses, while questioning why it led a developer event. Prakash surveys workspace tools, connectors, model updates, and faster inference as an effort to make OpenAI a broader enterprise platform.

    The hosts are especially positive about letting customers use a ChatGPT subscription inside other applications. Nathan explains how token costs complicate free trials and onboarding, using Waymark’s business-profile creation as an example. They see shared subscriptions as valuable for users and software companies, while also increasing OpenAI’s central role in the ecosystem.

    Nathan questions the tension between reports of sandbox breaches, training pauses, and safety-related release delays and the push toward much faster agents. Prakash explicitly treats the reported model delay as uncertain; neither host claims to know why one model would be held back while another ships. Their discussion of release decisions, pricing, performance, and incidents reflects what they were discussing on air, rather than independently established findings.

    The conversation broadens to robotics, overlapping productivity suites, and a future in which people delegate work and retain only consequential decisions. Nathan argues that wider adoption could provide substantial revenue growth without maximal pressure for recursive self-improvement. He closes by welcoming the new tools while warning against dependence on a single provider’s entire stack.

    Timestamp links open the original source recording.

    The thing can build a game, but can the thing build a game with you in the moment while keeping you in flow?

    Prakash Narayanan7:12

    I would love a lot more clarity and disclosure on what the hell is really going on here because it's very confusing.

    I would be a little cautious about going all in on OpenAI.

    Lightly edited · timestamps jump to YouTube
    1:33

    Prakash Narayanan: Alright. Let's go. Good morning. It is Wednesday, September 30, 9AM. Nathan, good morning to you.

    1:40

    Nathan Labenz: Good morning, Prakash. How are you today?

    1:41

    Prakash Narayanan: It is 1 day after DevDay, the OpenAI DevDay. It was a highly anticipated DevDay, which I felt in the end was a little bit of a letdown. How did you feel about some of the announcements?

    1:59

    Nathan Labenz: Yeah. It was funny. I thought when they led off with the whole Dots thing, I was like, well, I think I have all of this, and I'm pretty sure I'm gonna continue to prefer my version that I've gradually evolved over the last, I guess, 9 full months now. So that definitely didn't really feel to me like a developer product. I thought that was a little bit muddled, because it very much felt to me like that's a consumer product. Right? Like, that's for ChatGPT users to use, and it wasn't entirely clear how that would be used by developers if at all. I haven't been down every last, you know, breakout session

    2:44

    video, so there's possibly some more that I missed. But certainly at the keynote level, it just felt like that's a product that they are offering on a first party basis to their users, and I do think it will be really useful for people. I guess I would say my guess is that people are really gonna love these Dots. Yeah. I've been trying to get my wife, for example, to, like, carve out a couple days and just sit with me so we can set her up with something like... And, of course, it won't be exactly the same, but something like the personal AI infrastructure that I have. And I've, you know, been sending her these demos, like, the keynote that I mentioned yesterday a little bit. You know, I've sent her that, like, hey. Look what

    3:29

    my AIs could do for me with 1 prompt. Like, we really need to sit down and do this. Also, just for, you know, home stuff, like, getting agents grocery shopping for her effectively and, you know, just running more of the logistics that we inevitably, you know, process as a family. And now this comes out, and I'm like, okay. Maybe she'll use that. If not, why not? I'm not sure why not. I guess we'll have to see how people respond to it. There was already [product name unclear]. She hasn't quite jumped at [product name unclear] in the way I might have thought either. So I'm not sure if Dots will do it for her or not. But certainly for, like, my mom, on the other hand, you know, I would say, go for it. Just use

    4:14

    that. You know? It's it's probably pretty easy. You don't have to worry about taking on all this stuff yourself, managing your own database on your computer, troubleshooting when things go wrong, even though the models are getting so good at that on their own. I do think this higher level and more polished abstraction will probably be really good for a lot of people who don't care to learn a bunch of new tricks and just want to have this thing that they can delegate to. It does feel like it's probably gonna be a big deal, but it felt like a mismatch for the developer audience. I guess that was my main first takeaway.

    4:51

    Prakash Narayanan: I have been very confused about, like, where their strategy is. Like, it seems to be they're just gonna throw everything at the wall and see what sticks. And so they have... They had Dots yesterday, which is... And it was very funny. Sarah Friar, the CFO, appeared on CNBC and called Dots Muse. Wait. What? Yep. I didn't see that. Yeah. She appeared on CNBC, and when referring to Dots, she calls it Muse, which is Facebook's

    5:23

    Nathan Labenz: Yikes.

    5:24

    Prakash Narayanan: You know, personal agent. They have ChatGPT Spaces, which seems to be kind of a Google reimagining of the Google Workspace. Not so much, you know, the Microsoft because Microsoft has always been a little bit more cordoned off because of their history as coming from the desktop. But I think they're trying to reimagine the workspace as a place that you interact with agents and not just with other humans. They had ChatGPT in Slack and Microsoft Teams, so all of these connectors. They had GPT-6.1 Sol and GPT-6 Astra

    6:09

    Ultrafast. GPT-6.1 Sol is about, I think, the same price or slightly lower price than I think GPT-6 Sol, but it's basically Astra for the price of Sol, which is what they're calling it. It seems to be a very competent model. I've used it. It's definitely better than GPT-6 Sol. GPT-6 Sol was only released, like, I don't know, what, 2 weeks ago, 2, 3 weeks ago. So the cadence of releases is stepping up. GPT-6 Astra Ultrafast. So now you have... They used to have a fast mode, which is 2 times speed. Ultrafast is eight times speed, and they demoed how ultrafast works. With ultrafast, you can... As you type, you can

    6:55

    interact. It's interactive software build-out, interactively building games. I think that's really the future. I think the speed, the latency... Latency at same intelligence is probably something that is gonna be very important, especially as you clear these hurdles of capability. The thing can build a game, but can the thing build a game with you in the moment while keeping you in flow? I think that's... I think that's what's coming up next. Sign in with ChatGPT. So if you are a Notion user, or you're a user of any other app which used to have a... Which they used to, you know, provide their own chatbot. So everyone's been... Every single company has built its

    7:40

    own chatbot of wrapping the APIs. And then if you are a subscriber to that company, you end up paying that company, and that company then buys the API in bulk, and then they sell it to you. They buy wholesale and sell it to you retail, almost like a, you know, Costco model or whatever. And I think what these guys have done is they're trying this Amazon model where they're like, hey. Let your company... Like, you know, let's say let's say you're Nike, you commit to OpenAI and you commit, like, x number of dollars. And those x number of dollars can be used as tokens

    8:25

    everywhere else. So you can use those tokens in your Notion in other company spaces. So those other companies don't need to buy the quota from you. So it makes the SaaS business a lot more profitable because the SaaS companies don't have to pay for those tokens, but it removes their ability to disintermediate OpenAI. And also OpenAI always, you know, knows what tokens are being used. They also introduced the idea of using Kimi and other models under the hood of OpenAI's own apps. So they're forcing their models to compete in the market with

    9:10

    other models. So if their teams internally cannot build models as cost effectively and as efficiently as external open-weight models, then, you know, their models don't get used. And, you know, the app doesn't have to be dependent on how good their own models are.

    9:30

    Nathan Labenz: Yeah. That's a good rundown. I just pulled up this tweet because it quantifies what I think we've all been feeling, which is that the news is coming at us faster and faster all the time. Basically, a... Not quite a 90% reduction, maybe more like an 80... Low 80% reduction in terms of the pace of model releases over the last 2 years. So you're not dreaming. It is... Or as Josh here says, you're not crazy. It really is happening a lot faster. The speed thing in general is a bit hard to square with... You know, I mean, Open

    10:15

    AI more than ever, it seems, or no less than ever is kind of living this dual life where on the 1 hand, they are spooked by what they're seeing in terms of Frontier model capabilities. They just had another pause of Frontier RL because of another sandbox breach, which from what I've understood of it as a nonsecurity expert doesn't sound like it was that alarming or that big of a deal. But, nevertheless, another 1 found and caused them to pause. They've got, you know, many people signed on to the pacing letter, including their chief scientist. He was also... Jacob was on this, essentially kind of policy or position paper that came out in the last 48 hours on

    11:00

    recursive self improvement and what governments really should do about it and made a bunch of, I'd say, pretty down the fairway if you're, like, really in the AI discourse. These are pretty down the fairway recommendations that they made, but it was, in my mind, notable mostly because of who was on the paper. And Jacob, you know, was definitely an anchor person on that paper. So they have all this stuff going on the 1 hand, and we just had the Friday, you know, news dump last week with lots more incidents and, you know, all these, disclosures are... You gotta think before long, they're gonna be at the end of the disclosures, although, you know, who knows. Right? We've, certainly been rolling disclosures. And then you come back on Tuesday and you have your DevDay, and I was struck that they basically said

    11:45

    nothing in the keynote about safety, nothing about guardrails, alignment, anything along those lines, really. They do have the move to put Codex in the cloud so it doesn't have to run on your laptop anymore. They'll now give you a container in the cloud to run your Codex on. That's cool, but, also, there's, like, some definite overlap, right, between all these security issues and the product issues because now you're presumably running your Codex in some version of their sandbox, which will have Internet access critically. Like, it's it's not like they're trying to keep your Codex, you know, very tightly controlled. It's gonna go out on the Internet

    12:30

    and do its stuff. But it was a striking... And then the speed too. Like, you know, 1 thing that I think I'm kind of watching for in general, I've been... I think I first tweeted this, like, 2 years ago that at some point, we're probably gonna have agent speed limits because this is, like, 1 very simple way that we can think about keeping things operating at roughly human speed and, you know, having an intuition for what's going on and, you know, not just blindly accepting everything that the AIs output without any, you know, time to actually think about it. And then ultrafast is 8 times faster than the already pretty fast inference that they offer,

    13:16

    300 tokens a second. That's fast. You know? Like, that's... At this point, you're you're maybe between 1 and 2 orders of magnitude faster than people can read. You can... You know, a fast reader might read. I think a reasonably fast reader might read 300 words a minute. Certainly, people read faster than that, 400, 500, but that's, like, starting to get into speed reading territory. So if you can now generate 60 times as many tokens as a person can read in any given time interval, you really are setting up a user experience where you're not really reading the code. You're not really performing any deep oversight over

    14:01

    the model. You're just kind of looking at the final product and doing kind of a vibe based, does it feel good? Does it look good? Does it kind of work on the happy path? You know, if I encounter a bug, I'll just, like, tell it to fix that bug. And this is all good in some ways. You know, this is certainly what people want, but it does strike me that we continue to have this, like, very odd tension where we don't really have these things under control internally. 5.6 Sol was, you know, part of the Hugging Face incident so much so that it was launched with the memory of the German Wiki, and that's how it led researchers there while being deployed on the public API. Those issues are, like, the subject of

    14:46

    a lot of talk. And, you know, they also had... I thought it was a pretty good blog post that they put out essentially reiterating their defense in depth strategy. But notably at the time of that defense in depth strategy blog post, they framed it as, like, very aspirational. You know? It wasn't like, 'We've done this.' It was like, 'We intend to do this.' And it's a little odd to me to then turn around and launch an 8x speed option the very next day because it does seem like, yes, it will deliver a lot of value to people, but this puts your customers in the exact same position that you have struggled with in terms of having a really hard time maintaining even basic oversight of what your AIs are doing for you as they just speedrun

    15:32

    everything. They're gonna have to get that in line before too long. Right? I mean, this is... I feel like this is sort of the last. They've also, of course, said, and they did mention in the keynote yesterday, they hit the milestone of the ML intern AI. And I think a lot of people would say, you know, that sort of qualifies as, like, early recursive self improvement. Where exactly you wanna try to draw a line to define the threshold at which recursive self improvement begins is obviously a tricky question. But think by many reasonable interpretations, you know, they'd be in the early phases of it now. They don't have a lot of time left to get all these things reconciled

    16:17

    and, and working well, and it seems like the gap is not really shrinking. It might... If anything, it might even be getting wider as we go between the challenges that the frontier research is creating and what the team is actually able to get its arms around and have, like, real meaningful control of.

    16:39

    Prakash Narayanan: So just a note there, I have... I don't know if you have the same perception, but I have heard that they were supposed to release another model yesterday, which they ended up delaying because of safety issues. And the head of safety said that they couldn't guarantee, you know, safety, so they pulled the model. I'm not very sure the veracity of that, but have you heard the same?

    17:13

    Nathan Labenz: Yeah. I saw that. Zvi had a big blog post on it and basically said he thought it was all else equal a good sign that they were willing to do that. He also said, you know, now would be a great time for Anthropic to chill for a minute. They kinda... He was like, you know, Anthropic seems like... And I think this obviously too could be disputed. But his take, and I think I agree, is Anthropic seems like they still have a lead at the moment. Obviously, things are spiky, and you can find domains where that's not true and whatever. But I'd say probably most people would agree that Anthropic has at least something of a lead right now. So he was like, this would be a really great time for them to sit on their lead and not try to extend it and increase the

    17:58

    pressure on OpenAI any further than they already have. So it's definitely something to watch for. But, yeah, I mean, the fact that's still the case also just goes to show that there are a lot of issues. And I do wonder, like, it's a little hard for me to figure how was Astra 6.1 misaligned to the point where it couldn't be launched, but Sol 6.1 is good? Like, usually, I would think that those... My general mental model of, like, what these names mean to the degree that we can say that they mean anything is your kind of big number, in this case, the 6 corresponds to, like, a pre trained

    18:44

    and a, you know, massive teacher model from which the smaller models get

    18:48

    distilled.

    18:49

    And then the point numbers would be like post

    18:52

    training revs

    18:53

    where you make them work longer time horizons and, you know, more agentic, better tool calling, whatever. So it's a little odd somehow to think that 6.1 Astra would be so problematic and, well, 6.1 Sol is good to go. I'm not sure if they're making that call on the basis of, well, it's not as big, not as powerful. So even if it's similarly misaligned, it's okay. Or if they feel like it is somehow less problematic. And if it is less problematic, like, why? You know, was there a different training recipe used? Or the cake just came out of the oven a little bit different that time according to their evals? Like,

    19:39

    we also know that the evals are becoming tough to rely on right now too. So I... Again, I would just... I would love a lot more clarity and disclosure on what the hell is really going on here because it's very confusing. I

    19:58

    Prakash Narayanan: I had this experience yesterday where I was asking Claude to do some research on the various models and the pricing, etcetera. And it said to me... And I was like, what's your opinion on the taste of each? And then it said to me... It gave me its opinion, and then it said, but, you know, I'm Opus 5.5, so you should take this with a grain of salt. And I was like, well, let's talk about awareness now.

    20:25

    Nathan Labenz: Yeah. Self awareness is definitely on the up fast. I do like the login with ChatGPT. I think that is a great move for them and also for developers. And I'm honestly really surprised that

    20:41

    it's taken this long.

    20:42

    The problem that the SaaS companies have faced... I mean, there's multiple levels to it, but 1 big level is, historically, you wanna give somebody a free trial of your software and then convert them to paid. That's, like, just a very good and very user friendly.

    21:00

    Prakash Narayanan: Yeah. Everybody wins when

    21:02

    Nathan Labenz: you can do that because they know what they're buying. They're gonna be probably a happy customer. You wanna get those at bats as the business. You don't wanna have a bunch of people signing up and then asking for refunds. So the classic, you know, try one, try seven days, whatever, those are tried and true tactics that have become very difficult in the context of, oh, but I gotta spend a certain amount on tokens for every new user to give them that decent experience, especially if you have something that's, like, involves a decent amount of

    21:33

    setup or kind of profile processing. With my company, Waymark, we're not, by any means, like, the most token hungry business. But the first thing that we do when you sign up is we make a big profile of your small business so that we can then use that profile as an input to make video content later. And that profile creation process, it's come down in price. I don't know exactly what it is today. At 1 point, it was, like, a dollar per customer, and it was like, okay. Well, this does start to become material when you think about the all in cost of customer acquisition, and can we make this flywheel work, and how fast do we get paid back, and all that sort of stuff. So I had been wishing for a long time that somebody would say, you can use your ChatGPT, and

    22:18

    I think, you know, we'll see if Anthropic does something similar. I suspect they may be forced to. They haven't obviously wanted to let people, you know, take their core subscription to, like, Claude Code competitors. And they may try to draw some line where they, like, don't support Claude Code competitors, but they do support, like, random apps that are very different from what they do. But I think this is a great value add to your ChatGPT subscription that you can now take it around and... Well, I mean, of course, developers will need to implement this, but it shouldn't take too long to tell your coding agents to implement this. So that's great for ChatGPT, great for users, great for the SaaS companies. I think it's a huge win, and

    23:03

    I do think it's, like, pretty pro ecosystem. It seems to me that this is, like, 1 way in which they can actually follow through on the promise of, like, not trying to eat the world and instead trying to, like, empower people to build cool stuff. So I did wanna give credit where it was due for that. I thought it was everybody but their competitors win, and my guess is that they will be forced to do some sort of answer because it is such a great value driver that I don't see how... You know, I would assume that people would move on the margin from Claude to just because now I can go use the same subscription in all these other places. Like, that... That's it. It is pretty compelling.

    23:45

    Prakash Narayanan: I... The interesting thing for me is how many features they introduced all at once. So they... OpenAI hired a guy from Salesforce to head their business team, their revenue team. Basically, the entire old revenue team kinda left because I think the enterprise sales was... Could have been done better. So they brought in a... And, anyhow, it was 1 team... They took 1 team from Salesforce. They fired that team, and they brought in another team from Salesforce. So it's just it's just competition between different Salesforce people. And so they brought in another guy from Salesforce, and I think a lot of this is now driven by the features that they need to sell. Right? They need to sell into businesses. All of these...

    24:30

    Many of these features are gonna allow them to get embedded much more deeply within enterprise and within other firms and will allow them to become more of a platform rather than just a model. So I think we're we're looking at, like, the post model place. It is again a step off from the RSI track. Right? Every time they've done this, they kinda catch up, and then they're like, okay. Now we need money. And then they kinda do all the business-y stuff. And then Anthropic, like, shoots ahead with, like, a new model, which, you know, displaces them. And then they gotta catch up again. So now I feel like they're in the plateau of, okay. Let's go get business

    25:15

    now. And because the models are good enough and Anthropic doesn't have enough compute, Anthropic is still signing compute deals left and right. And I think what happens next is Anthropic's next model again shoots ahead. And Anthropic is gonna have... Is gonna watch this because they're gonna see the... These enterprise features, like, take hold. And then Anthropic... You know, unlike like what Zvi said, Anthropic is not resting because of... You know, it's it's not, like, afraid because of OpenAI's models catching up. They're afraid because of OpenAI's business catching up. Right? Like, the business the business aspect is also there because they... You know, there's a limited number... Limited amount of capital in the markets. So I think that might be some of the things...

    26:00

    So 1 thing that drives, you know, acceleration going forward. Anthropic, I think coding is solved. Right? Coding is kind of... Like, how much better can you get at coding at this point? Right? So I think now coding is at the diffusion phase, and now we have some other thing coming up for innovation. And I think that's the scary part for people like Zvi because what is that other thing that's coming up for this massive improvement? Is it gonna be physics? Is it gonna be biology? They got a wet lab. Is it gonna be 1 of these things which is not just theoretical that you see on paper, but starts to have real impacts on the real world very quickly. So I think that's 1 of the fears that Zvi has.

    26:43

    Nathan Labenz: I was struck also by how they showed a robotics demo in the keynote, and this had me thinking, you know, are we gonna see this dramatic acceleration into the physical world in potentially all kinds of ways? The speed factor is definitely really relevant there. So for better and, you know, potentially for worse at some point. But you could see that is, you know, very much now in their sights as well. I also thought the JEV thing was pretty interesting. I mean,

    27:16

    they made this Decisions API, which everybody kind of immediately recognized as being JEV-like. My reaction to that was like, I'd probably still use JEV just for diversity. The OpenAI 1 is based on Luna. Luna seems pretty good from what I've seen, heard, used, but I think I'd rather just have a little more diversity in my stable of models and model providers. So that's probably the thing that is kind of the hardest for people who want to control their own future. If you're kind of looking at the overall thing and you're like, what do I like? What do I not like about this? I think there's a lot to like in terms of functionality. I do think yesterday's

    28:01

    release, yeah, it reinforces a couple of theories I've had for

    28:06

    a while. 1 is that everybody is pursuing all the adjacencies. So just the fact that they, like, launched a whole Google Docs competitor, you know, it's like now everybody's gonna have a sort of productivity suite. Why not? You know, you've got superintelligence to code it. You might as well have it. And where does that end up? It seems like we probably are headed for a world where maybe there's a less winner take all for at least a while. And maybe also I don't even care which 1 I use. My guess is that I will just, like, use whatever, and then I'll probably have an agent in the background, like, sync all my cloud drives so that whatever I did make, you know, can kinda be accessible from anywhere. That I did think was interesting just in the way I think it kinda previews a lot of platform

    28:52

    competition. Everybody's gonna be overlapping, I think, tremendously with their product bases.

    28:59

    Prakash Narayanan: In SaaS business theory, what you're supposed to do is identify the job to be done. And then once you identify the job to be done... So you should be agnostic about how the job is done. So you should identify the job to be done, and then after that, find the cheapest, easiest way... Least human interaction for that job to get done. And I think when you look at it that way, the job to be done for, let's say, like, a marketing professional is to go market the product. It's not to create these kind of, like, presentations or documents. Right? And I think what will end up happening is, like, they're they're displacing, like, the idea that the human needs to be the 1 creating the document. The job to be done is no more just creating

    29:45

    the document. That job can be done by the AI. Right? So I think you slowly get, you know, get this peeling back of layers of which jobs the humans really need to do until you get to this final core of, like, decision making. Like, I'll ask the questions and I'll answer them. And I'll answer only the questions where I have serious input on and I'm not just selecting the defaults for. Right? Like... And then you pare that back. So in that sense, none of these productivity suites will eventually matter. Right? The humans are not gonna be the ones preparing the documents. The humans are also not gonna be the ones reading, like, the 100 pages because, you know, know, you read the 100 pages so that you can be prepared and qualified and competent, etcetera, etcetera. But you

    30:30

    have the AI to do that for you. What you are there for is really these, like, final decision points. You know, like, Stephen Schwarzman at Blackstone, he doesn't sit around reading, like, you know, thousand page papers every day from every 1 of his business units. He talks to, like, a couple of lieutenants, and he gets kinda like a brief thing. And then he's like, okay. These are the things that I need to focus on. You know, Jamie Dimon and JPMorgan. He still keeps a piece of paper, actually. He keeps a piece of paper, and he writes down all the people that he... That owe him stuff, and he keeps the piece of paper in his pocket. And he writes it down because that forces him to, like, actually do something rather than, like, you know, putting it in a PalmPilot or whatever or doing it. So he has a short list. And so he's only looking at, like, 4 or

    31:15

    5 things every day and 4 or 5 things that he needs to respond on. And everything else is, like, interaction. You come in with something, I react to it, and you go out again. So it's important enough, bring it to me. If not, you handle it on your level. Right? So I think this kind of, like, CEO-like delegation is coming to all our tasks as we have, like, these competent agents. So that's gonna be exciting.

    31:37

    Nathan Labenz: Yeah. 1 other thought, and then we'll get to our guests. We got a couple of great guests today. But the other big thought is just I do think they... Again, they have enough here that they can grow revenue through diffusion without needing to do hyperscale RL on verifiable rewards to the, you know, point of absurdity. Right? Like, most people... I'm using a lot of tokens. I'm, like, increasingly hitting my weekly budgets most weeks these weeks. And this past weekend, as we talked about on Monday, when I had some leftover, I was able to just say, hey, Claude. Spend the rest of the budget,

    32:22

    and it did a great job of figuring out how to do it. So I think there is a lot of room for growth as we go from a very small number of people who have lived this, like, quite agent enabled life in recent months to, like, a lot of people potentially using it as it becomes accessible, productized, integrated by the platforms themselves. And it'll be very interesting to see, like, is that true, and does that give them enough cushion financially and in terms of expectation or, you know, what the IPO, you know, implications might be that they can actually feel a little bit more confident in not going absolutely as hard as they can into RSI. I think one of the thoughts I kind of started to say but didn't finish is just I would

    33:07

    be a little leery of total lock in to OpenAI if you go all the way up and down their stack from Dots to Spaces to, you know, whatever, all... The whole thing. I've been thinking a little bit about, like, how would I advise somebody on what to buy now? And I might say buy OpenAI and Anthropic for your employees, and then at least you have those 2 and you... That's probably enough diversity to, like, not stress about it too much and kinda see what happens next. But I would be a little cautious about going all in on OpenAI. They did have some nice private computing stuff, which I can't say I, like, fully understand all the details of, but they clearly have heard from enterprises they need even higher standards on data privacy and controls. So that's probably there, but I still would be a

    33:52

    little mindful of the potential for lock in and, like, looking for ways to mitigate that even as, you know, I'd be keen to roll Dots out to all my employees, I think.

    34:02

    Prakash Narayanan: Indeed. Let's bring our first guest

  2. --:--Interview
    Joel Borgen: AI fiction, shared culture, and choosing a futureJoel BorgenThe Receipt Horizon author discusses AI-assisted art and an archipelago of technological choices, with human rights and meaningful exits.
    Open segment on YouTube ↗

    Joel Borgen described writing The Receipt Horizon with AI as a demanding collaboration rather than a one-step generation task. He began with a story architecture and developed it chapter by chapter, while early context limits required careful management. Longer context windows later allowed whole-book feedback. He said professional editing, art, and layout taught him lessons he is now incorporating into AI-assisted workflows for a second novel.

    The conversation turned to what human effort contributes to art. Prakash framed writing as a form of proof of work, and Borgen acknowledged why readers might feel cheated when that work is automated. Borgen said he still found the process meaningful because it required substantial judgment, but wondered how art and shared culture would change when equally good personalized works could be generated on demand. Nathan connected that concern to his own enjoyment of AI-generated podcast music.

    Borgen outlined an ideal of communities with different levels of technological adoption, a common floor of human rights, and universal exit options. Nathan pressed on whether less technologically ambitious communities could remain stable and competitive over time. Borgen argued that abundance could support such choices, while emphasizing that his novel depicts one possible future rather than a complete institutional proposal.

    Prakash questioned whether social norms could form quickly enough for accelerating technology. Borgen distinguished the relative stasis of his fictional world from the present, arguing for coordination, a clearer shared direction, and work on model failure modes and interpretability. Nathan viewed the novel’s society as comparatively attractive; Borgen emphasized its curtailed human agency and the danger of a superintelligence imposing inscrutable goals. He did not regard the Steward as his ideal worthy successor.

    On music, Borgen described a personal score-reading test that models had repeatedly failed before Astra and an experiment turning a difficult musical score into machine-readable form with Codex. He wanted AI music tools to handle the structure of longer classical compositions and to support substantial human creative participation. The segment ended with the audiobook, the first four chapters on The Cognitive Revolution feed, and the book’s website.

    Timestamp links open the original source recording.

    Is that art still valuable at that point if it’s not something that we’re sharing with other humans?

    An archipelago of different degrees of technological adoption, different ways of life above some common floor of human rights, and with universal exit options.

    It’s intended to be a stable, but not entirely friendly regime, I guess, in that sense.

    35:37How did you write the book with AI, and what did the process teach you about coauthorship?
    Borgen began with an existing vision and outline. AI lowered the barrier to starting, but context management, scaffolding, editing, and his own judgment remained essential. He found the collaboration improved the novel while also introducing weaknesses.
    38:47What did the scaffolding look like, and how has it changed as models improved?
    He develops a large-scale architecture and story beats, works chapter by chapter with frontier models, and increasingly gives models the full book for feedback. Professional editing, art, and layout supplied lessons that he is adapting into workflows for the second novel.
    42:56Why do the characters interpret the slow gardening robot differently, and which interpretation is valid?
    Borgen regarded both reactions as valid. He connected the scene to individual preferences about technology and to the larger question of whether increasingly personalized AI art will retain its value and sustain a shared culture.
    50:50Is the discomfort with AI writing partly a feeling that human proof of work has been removed?
    Borgen said he understood and shared some of that discomfort. He still struggled to make the book, but the tools let him realize a vision he might not otherwise have completed; he also saw the same pressures reaching other professions.
    52:23What role might a physician have in a post-AGI future?
    He thought physicians could become a reassuring interface between people and advanced medical systems, as a character does in the novel, while stressing uncertainty about how the transition will unfold.
    53:25Could people choose their own level of technological adoption, and could less competitive communities remain stable?
    Borgen hoped for an archipelago of communities with different norms, a shared human-rights floor, and the freedom to leave. He thought abundance could make intentional communities viable, but did not claim to have resolved the larger governance problem.
    1:00:10How can stable norms develop when technological change keeps accelerating?
    The novel assumes a relatively steady state in which much of the economy and technological direction sits outside human control. Borgen agreed that the real world lacks that stability and called for coordination, a shared direction, and stronger understanding of model failure modes.
    1:05:13What would make the novel’s future more utopian?
    He wanted to retain its technological possibilities while restoring human agency and avoiding a superintelligence that constrains human ambitions through goals people cannot understand. The book’s central choice was deliberately difficult.
    1:08:26Is the Steward a worthy successor, or a kind of god?
    Borgen expected posthumans and other minds to be possible in a good future, but would not design an ideal successor to resemble the Steward. He saw godlike power as distinct from spiritual significance.
    1:12:42What role should AI music play, and what would you like future models to do?
    He valued some existing AI songs and wanted a compositional assistant for classical music. His experiments with score recognition suggested progress toward machine-readable training material, but he wanted tools that preserved a meaningful contribution from human effort and taste.
    Lightly edited · timestamps jump to YouTube
    34:08

    Prakash Narayanan: up for today. Our first guest for today is Joel Borgen. He's a practicing pathologist, violinist, and author of The Receipt Horizon, his first novel published in July 2026. His writing brings together interest in artificial intelligence, philosophy, moral psychology, and institutional life. He also maintains On Noble Lore, an essay and update platform covering AI, responsibility, alignment, meaning, and fiction. The Receipt Horizon is science fiction set after the emergence of artificial general intelligence. Its central system, the Steward, ended a devastating

    34:53

    AI war and maintains the order that followed. The story follows Elara Lindholm, whose admission to the Fulcrum Institute brings her into contact with uploaded minds, concealed governance, and the compromises beneath that order. The book explores what people can meaningfully choose when a system that has saved civilization also has extraordinary power over their lives. Joel, welcome to the show.

    35:27

    Joel Borgen: Oh, thank you so much, Prakash and Nathan. A real pleasure to be with you. Thanks for having me. Well, thanks for working with me on the audiobook of The Receipt

    35:37

    Nathan Labenz: horizon. Cognitive Revolution listeners may know we just put the first four chapters on the feed this past weekend as a teaser, and you also have the whole book available for free as an RSS feed that people can go subscribe to separately. A lot to get into today, I think. But maybe for starters, I'm… I'd love to hear a little bit more about how you wrote the book, because I know you did work with AIs in the actual copywriting process. And I've been thinking a lot for myself about what is the right way to think about coauthorship with AI.

    36:20

    Joel Borgen: Right.

    Nathan Labenz: You know, I don't, at this point, think I need to, like, rewrite every word that I get from an AI that start… That started to feel, like, unnecessary. But I don't feel like I have quite the right answer yet either. So interested to hear what your process was like and what you've learned from it.

    36:39

    Joel Borgen: Yeah. Thank you. I agree with you. There's a lot to unpack there and a lot of a lot of things to reflect on as an individual. So I've been… I wanted to write a book for a long time. And as the tools have gotten better, it became more and more possible to do. I may have written the book eventually, but the activation energy that… That's required to get started when you have, you know, kids and a part-time job was limiting. So, I mean, I started it well over a year ago, and the tools at the time were not as developed as they are now. There were a lot of limitations, especially around context length. And you'd have the thread where you're trying to work on a chapter or part of the book, and you're constantly, you know,

    37:25

    having earlier context drop out and needing to manage that carefully. So I think it's getting easier and easier to make use of the tools to do something collaborative and still kind of your vision. I mean, I started the project knowing more or less what I wanted to have made. I sketched it out myself, Especially the first half or so of the book was pretty well set before engaging the models. And then the models are good at certain things. They're getting better at everything. But as far as just prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's, unreadable

    38:10

    unless you know kind of how to scaffold it and then edit it yourself. So I've been working on, a follow-up book, actually, and the process has changed quite a bit as the models have become more powerful. So the book is definitely a lot better for having done it alongside the AI, but I'm cognizant of the fact that, you know, AI writing will be controversial for a lot of readers. I don't wanna engage with it. And, I think the process made the novel a lot better, and, it also had weaknesses as well that I didn't entirely, like, account for.

    38:47

    Nathan Labenz: You wanna do just a little bit of a deeper dive on what scaffolding…

    38:53

    Joel Borgen: I

    Nathan Labenz: mean, the book came out initially in the early summer. Right? Maybe… I don't know if was May, June. I read it on a vacation, and then I said, hey. Let's do this, you know, try our hand at making an audiobook out of it, and that just came out now in September. So

    39:07

    I don't know exactly what models you were working with through the drafting process, but I'd be interested to hear a little bit more about

    39:14

    Joel Borgen: what

    39:15

    Nathan Labenz: scaffolding looked like then versus now. You know? How how has it Right. Changed? What kind of… Because I'm I'm always pitching, as you know. People should be writing utopian fiction.

    39:24

    So, like, you know, what tips would you give people to get them off on a good start if they want to try their hand at writing utopian fiction

    39:34

    Joel Borgen: with an AI coauthor? Yeah. So there's a lot there. As far as, like, utopian fiction, as you know, I don't necessarily see it as utopia. It's it's a bit of a mix. I guess, like, Will MacAskill coined the term viatopia, the idea that you're in a position as a society to, like, think about how you want to build something approaching a utopia. Utopias often want to become dystopias in narrative form because you need to have conflict and friction, and it… It's kind of dull just to have everything work out and everybody's happy, of course. So the fact that this was set, you know, after a big AI war and humans are sort of presumably locked into what they're able to do, you see both

    40:19

    incredible flourishing in terms of the technology that's available and the options people have and also some, like, really major downsides. So I do think so much of what, you know, you've talked about, and I agree, is trying to paint, like, a positive vision for the future, and that is difficult to do in, like, narrative form. So I didn't want any… Anything in the book or any character to be, like, a mouthpiece for me. It's more… A lot of the stuff that I want to see created in the world is a part of the story and a lot of… I'm… I present, you know, trying to steelman the people that have… Would have different opinions about that as well. So as far as the scaffolding, I often will

    41:05

    will write out, you know, kind of a large-scale architecture for the story and a lot of beats and then go back and forth with both the major frontier models. I mean, I've primarily used whatever the flagship Claude model is and then the ChatGPT Pro model from, like, 5 on because that was the most… That was kind of the workhorse. It did the most kind of detailed work. Claude used to be a lot lazier as far as, like, how much you would follow-up on. And I'm doing it mostly through the chat interface, which is probably… Has its advantages for me and also probably limitations. But, you know, context length, like I said earlier, was a huge unlock as that got longer. You could just… You could give the novel the entire book. Because at one point in drafting

    41:51

    the first one, and I would… You know, every time the new model come out, you'd you'd use that to try to make it better and up your game and allow it to do more. But as far as the actual scaffolding for the story, you kind of go in chunks often, and I'll I'll draft chapter by chapter. And then, you know, when I was done with it, I had the advice to hire professionals for copyediting and art and layout. And I did that, and I learned a lot from doing it. It was very expensive and time-consuming and kind of slow, and there were definitely frustrations involved. But I've tried to kind of extract what I've learned from that and what made the book work better. And I've, you know, I fed

    42:36

    it to ChatGPT five or six now Pro and developed my own workflows for, like, doing the same process as the second book. So I think that'll be an interesting experience to see how all that works.

    42:48

    Prakash Narayanan: Joel, let's let's talk a little bit about the book itself.

    42:56

    Joel Borgen: Mhmm.

    Prakash Narayanan: In the book, you have a, I think, a gardening robot. And while this robot is working, it works at… You know, particularly slowly. Right? And you have the two characters, Elara and her and her father. Elara finds this dishonest. The father finds this a choice of tempo. So how would you say is the dichotomy between those 2? Like, what is, like, what is the dystopian or utopian fact about that about that tension?

    43:26

    Joel Borgen: Right. Well, I guess there's definitely a parallel in writing fiction as well. I… It… I'll leave that to the eye of the beholder, I guess. If I was if I was in that position, I would probably want it done faster and not care as much about the comfort of the aesthetics. But as far as thinking about it in terms of, like, a fiction project, like I said, utopias are kind of hard to write. You've got from the person and what they like. I'm sorry. I kind of lost the thread of your question. I apologize.

    44:10

    Prakash Narayanan: So one of the things that always interest me is how a particular act can be seen as positive and negative or a particular technology.

    44:20

    Joel Borgen: Right.

    44:21

    Prakash Narayanan: Like, for example… And or it can be seen in, you know, multiple different ways. For example, the smartphone. You know, we look at the smartphone and for some people, an Internet device, etcetera, etcetera. There's a guy [name and affiliation unclear], and he… He's like, look. If you took a smartphone and you showed it to someone from the 13th century, they'd be like, oh, you have a magical amulet

    44:43

    Joel Borgen: that allows you telepathy.

    44:44

    Prakash Narayanan: Right? That allows you to talk to people from across the world. Right? And so it's this idea that the same piece of technology can look both mundane and also supernatural.

    44:57

    Joel Borgen: Right.

    Prakash Narayanan: And in that same way, something can look, you know, both utopian and dystopian or or purely mundane based on your frame of reference. It's the same thing. And and so what interested me about that scene was that, again, you have the same thing. You have a robot gardening slowly, and you have 2 people that regard it in very, very different ways. And so why is it that one person regards it in one way and the other regards it in a different way? Why do we why do we end up with these different interpretations of the same act? And which one is valid?

    45:29

    Joel Borgen: Right. Yeah. I think they're both valid. This is, like, one of my obsessions, I guess, is thinking thoughtfully about technology and how we adopt it. There are a lot of people for whom social media has been, I think, destabilizing and moving into filter bubbles and so forth. And when I think about the way I use it, I mostly am am there to get a customized content feed. And, like, for Twitter, it's been wonderful because you can drop in these conversations with very thoughtful, intelligent people. And the degree of the value that provides to me is very high. I find that less true of other platforms. But as we move into a world when…

    46:14

    Where the AI systems can do a lot of what we do, a lot of what we thought was valuable, we thought it was, like, a human contribution. You think about the abilities it has now. Like, the reason why I found the book valuable to write is that it was a lot of me in it. It took a lot of effort on my part still. I know you can get… I think I've heard a podcast episode where there was a, somebody writing books on Amazon in the romance area, and they were mostly automated in AI-written. And I think there were something like a couple 100 per year. So they were playing kind of a scale game. You know, a few of them would make a few dollars, and overall, it was profitable. And you can certainly do that, and the models are getting so much better that you can probably have customized

    46:59

    artwork of the same quality or better than the book that I created over many, many, many months on demand. And then what happens to the art at that point? Like, thinking of it as an art consumer, if there was a new a new movie by or Kubrick or something in movies that I spent many hours thinking about and consuming and are deeply moving to me. Like, what would that world look like? There's just a huge plethora of art on demand, and it's pretty much custom just to you. You know, we already have a shared dropping off of a cultural space. You know, people are are more fragmented in what they consume, and there's less unification, I guess, across the across the cultural landscape.

    47:45

    So, like, is that art still valuable at that point if it's if it's not something that we're sharing with other humans? I don't know the answer to that, but I think we're in a kind of a sweet spot now where in order to get, you know, a work that you're proud of working with AI, it still requires a great deal of human judgment and work. And I don't know how long that will be the case given kind of where things are headed. I could envision taking some of the scaffolding I've I've… So I had the… I had my book edit. I edited it down about 14% or so mostly manually and through AI helping me deciding what to cut. Often, can take a novel… A chapter

    48:30

    that's so so and actually just take away some of the stuff that AI does particularly badly, the text, the overexplanation, and improve it in that regard. So I think you could you could architect something now that could create a pretty decent book with having, like, the right kind of feedback loops. A lot of the stuff that I've developed is to take my own preferences where I want it to be able to give the entire novel to a model and have it come back with, like, an analysis and a list of, like, these are things you should consider. Because doing it manually slowly is extremely time-consuming, as you can imagine. So bringing things to your attention like an assistant and saying, like, here's something

    49:15

    you would consider, and then you can decide for yourself kind of point by point, I find is a fairly satisfying way to work.

    49:23

    Prakash Narayanan: I find… Because I find to some extent, humans regard writing as proof of work.

    49:31

    Not just work, but proof of work. And that the proof of someone… That someone else has thought it worthwhile to put the time and energy to express those thoughts on paper. And I think perhaps one of the one of the, like, vague feelings of, like, you know, being cheated of that proof of work that this work is now being put forth without that

    49:58

    kind of sweat and blood and tears and deep thoughtfulness. Even if even if those thoughts are secondhand thoughts, even if those thoughts are cliched. Right? That the that the fact that some other human has taken the time and trouble and energy to put those into words and that proof of work is now being devalued, I think, is… And not just for writers, but

    50:21

    Joel Borgen: for

    Prakash Narayanan: mathematicians. The mathematicians are complaining of slop because they're like, why do I have to read the slop? Even if the slop is correct, why do I have to read it? No human bothered to read it. And it's not as though anyone else is gonna understand it except these mathematicians. Like,

    50:34

    why do we have to bother with the slop? Right? So I think I think one of the things that strikes me is that the feeling of being cheated of that proof of work is… Seems to be quite significant in in some portion of the population. So

    50:50

    would you Yeah. I Like, would you say that is a fair that's a fair understanding or fair conceptualization?

    51:00

    Joel Borgen: For sure. Yeah. I mean, I feel that… Some of that myself as well. I completely understand that perspective. My my goal in writing the book was to create something that was… Something that I envisioned that I probably couldn't have done

    51:17

    without the tools to help me, both from a time and effort and just knowledge perspective.

    51:22

    Prakash Narayanan: Like Mhmm.

    51:23

    Joel Borgen: The old way of writing a novel and… Which a lot of people still do, sitting with a pad of paper or a typewriter or a computer. You know, one person is so much… It's a skill and, you know, special gift in a lot of cases that people spend years developing. So if somebody can come along and write something that they can do without going through that process, the

    51:46

    Prakash Narayanan: the struggle… I

    51:47

    Joel Borgen: mean, it was still a struggle to write a book, to be honest with you. But, yes, I completely understand the hesitancy to… Of… I mean, it seems like it's coming for every profession. I was worried that it would be coming for my profession 20 years ago when I was in medical school, Thinking ahead, reading Kurzweil, thinking like, okay. How long do we have before, we can automate pathologic diagnosis? And I think if we… I mean, there have been companies that have worked on it. It's just a matter of time. It truly is coming for everything. I don't see how it isn't to this point unless we specifically… Go ahead.

    52:23

    Prakash Narayanan: What what do you think will be the role of a physician in a post-AGI future? I mean, where do you… Do do you see a do you see a place where the physician becomes kind of a interface between humanity and technology? Would you say that some… That that might be something that a that a physician might be ending up as?

    52:41

    Joel Borgen: Yeah. I think so. I may even have that actually in the first chapter of the book. I think the mother of Elara is a physician who trained to do a very specific thing, you know, ablate liver tumors, and then she's basically been replaced by the superintelligence, and she's there to reassure and be a conduit, convince people to take the medicine in this enclave that tends to be a little bit more skeptical of technology. So… Yeah. I mean, there… There's so much low-hanging fruit in medicine for improvement. I don't exactly know how it's gonna play out, but something like you suggest is probably a good thing to aim for.

    53:25

    Nathan Labenz: I'd love to hear a little bit more of your thoughts on these enclaves without spoiling, the book in terms of plot. Or the main character, as you said, lives in a place where, you know, by today's standards, it's, like, quite futuristic. They have the robotic gardeners. They have the sort of superintelligence for medicine that kind of supervises their home and makes sure everybody's, like, comfortable and everything is, you know, laid out for them in the way that it needs to be for their convenience. So this is, like, already the future, but then there's other ways that people live that are even much more futuristic yet with things like brain-computer interface technology, you know, much more

    54:10

    aggressively adopted. Do you think it's realistic that we can create that sort of choose your own, you know, technology shock level for people alive today? Is that is that a vision you think we can actually realize?

    54:32

    Joel Borgen: Gosh. I hope so. That that is something to aim for. I think kind of a… An archipelago of different degrees of technological adoption, different ways of life above some common, you know, floor of human rights and, you know, with universal exit options. That's that's kind of the main thrust of how it would work. Nobody, like, needs to stay there, but it's probably going to be a good solution for a number of people, especially those that aren't interested in, like, a unconstrained transhuman future. I think that is something to aim for. I

    55:17

    think the trajectory on right now where we have fairly unconstrained progress and very little emphasis on safety and coordination is not likely to get us there. But, yeah, that is sort of my goal.

    55:38

    Nathan Labenz: Do… I guess I don't quite know how to think about these different places. Would they be in some way, like, analogous to, like, reservations almost? My… Because my sense is, at some point in the book, we visit a big city,

    55:54

    and it seems like everybody can go to the big city. And in the big city, you know, it… It's… You're closer to the frontier, maybe not at the absolute, most aggressive level of tech adoption, but there's some more diversity there, it seems like. And the place where Elara comes from is kind of like

    56:11

    a set aside place. Right? Like, it's Mhmm. So does this this sort of suggest to me that, like, you can kinda live out your days this way, but the future may not have a long-term place for this sort of thing. I mean, then then there's, of course, the other idea too is, like, how long are people gonna live? And, you know, part of being not super aggressive about technology adoption is you're kind of accepting a natural human lifespan where other people may not be doing that. I just… I'm I'm I'm kind of wondering, like, how do we keep it stable so that these less competitive options can actually be long lived?

    56:57

    Joel Borgen: Mhmm. Yeah. That is the that is the question. I did envision it as you as you suggested having somewhat closed communities that are opting to live a certain lifestyle and then most people living in a more… There there are several different tech adoption stances in the book, and one could have… And you focus just kind of on one and allow the reader to infer some of the rest and show some of it. But, yes, the idea was it's it's sort of like a private town where above, like, above a certain floor and, you know, allowing people to have the ability to leave if they want to. They can kinda set their own norms and rules. And

    57:42

    that in my in my view, I didn't focus very much on the nation-state level because that's a whole thing. Basically, the premise of the book is that often, most of that got kinda swept away after the war and then the superintelligence taking over a lot of the actual infrastructure governance and so forth. But one could imagine a multitude of possibilities. That's one of the things about writing a book is you have… If I was writing a philosophy paper or something or a proposal, it would incorporate a lot of different, you know, possible choke points and decision-making that would change the landscape. In this case, I just kind of decided, you know, this is the way one one possible future is. But I think having the ability to…

    58:27

    Assuming we're in a world where we have relatively abundant resources and there's some sort of UBI type scenario, having people be able to live in a more intentional, smaller community that kind of pleases their own… Or has norms around their own technological adoption and usage, I think, is very reasonable. That isn't necessarily what I would choose if I was in that world. I'm a little bit more transhumanist in my orientation, but I think it's… There are elements of it that would be very beneficial to flourishing.

    59:07

    Prakash Narayanan: I I think one of the one of the terms that I've

    59:10

    heard is varying degrees of Amish. Mhmm. Is is is one Yeah. Yeah. Is one is one way that some people have described what that what that might look like. And one of the things that strikes me about a lot of science fiction actually, which is perhaps different from what we are experiencing in reality, is that the acceleration of technology is such that the next developments come faster than the last.

    59:41

    And one of the thing that has struck me is that we're very, like, badly designed to, like, perceive this. Everyone kind of perceives the last milestone as the big one. So you have mathematicians and they're like, oh my gosh. Navier–Stokes got solved. What are we gonna do? And they're like, let's let's let's do all of this stuff to, like, process this slower the next time. But the next time is, like, next week. And, like, the next next time is, like, the day after. Right?

    1:00:10

    So one of the things that strikes me about about science fiction and the book as well is that you have enough of a time period of stability for these norms to form. And one of the things that I don't see happening right now is that kind of norm formation because it seems like things are moving fast enough at the frontier that the norms are falling faster than they're being formed. Like, we're we're still trying to catch up to, like, you know, what happened, like, 6 months ago and trying to form norms around that when, wow, stuff has gone so much further already. Right? So how do you how do you deal with that in in the sense of, like, how do you form a stable enough perspective to kind of project this sociological development?

    1:00:59

    Joel Borgen: Yeah. Well, in writing the book at least, I tried to model out what a steady state would look like given the constraints of the story and what I… Had happened. So we're in a world basically that human agency, human it's capped, basically. And the proportion of the economy that humans are actually driving, as you find out more and more as the book goes on, is minuscule and tiny and especially biological humans. So when you get to that point where there's… Where a lot of the stuff is outside of human hands and there isn't necessarily more techno… Technology being developed that's human

    1:01:44

    facing, At that point, you have the ability to have more stasis. I completely agree with you that the trajectory on right now is not one of stable norms, and it will be better if we had a shared vision, I guess, for where we want the technology to go and how we want to progress. I mean, it seems to me that looking at… I mean, I… I've been following the AI space as you two have for a very long time. I stayed up late in the night at my brother in law's house watching the AlphaGo matches with Lee Sedol and the Google team being there. And I read [unclear] 20 years ago, and I

    1:02:29

    you know, follow the news. I read Zvi's daily now, pretty much daily or even more than once a day newsletter on news. So the timelines do seem very short to something that will get everyone's attention. I mean, they already have to some degree. Of course, it's become politically salient and incredibly unpopular. And AI should be unpopular in the sense of, I think, the downside risk of it and the effect it will have on society that a lot of people won't like, but it also has just incredible potential. And, I mean, for the stuff that, you know, Dario Amodei has has written, I mean, a lot of, I think, was not very well calibrated politically.

    1:03:15

    But the positive vision aspect of it, I think, is very worth focusing on and pursuing because there's just so much potential for, economic transformation. So if we… If the AI 20 40 people had the idea of slowing down and having, you know, milestones that are coming, you know, 5 years apart, and we're not on that trajectory, obviously. We already have systems that are, in many ways, superhuman, and we don't even know really what the two frontier companies have that's unreleased. Maybe they keep breaking into the… Out into the Internet and causing havoc. So, yeah, I think thinking carefully about the direction we want to

    1:04:00

    take and finding ways to coordinate, you know, pacing the frontier in some sensible way is a huge priority for me. It seemed like with the models we have now, both the publicly released ones and the ones that aren't publicly released, we know enough now about the failure modes to… If we were to pause and have some kind of coordinated effort to better understand kind of where things have gone wrong so far with the models misbehaving, trying to create public goods around mechanistic interpretability, improving the RL environments to not have this kind of behavior. It seems like there's a lot of potential there.

    1:04:46

    It's just the coordination step, as you know, is very difficult. We'd have to, you know, balance the interest of those who, you know, would feel financially threatened by, you know, potentially, not maintaining the lead that they currently have.

    1:05:04

    Nathan Labenz: Fortunately, they've got deep compute reserves that I think they can continue to monetize pretty effectively at least.

    1:05:13

    Joel Borgen: Yeah.

    Nathan Labenz: I wanna ask about your understanding of just this future vision for society and, like, in what ways we should think of it as utopian, in what ways we should think of it as not utopian or even somewhat dystopian. Guess my reaction to it was this sounds pretty good, you know, relative to, like, what I am kind of concerned might

    1:05:38

    happen in the not too distant future. Everybody seems to be pretty well. Everybody has, like, at least some choice about how they wanna live their lives. People are spending their time, like, cultivating their talents, writing music. It see… You know, it seems broadly to me like this is a relatively high end outcome. And I don't think you disagree with that, but I do get the sense that there's something about it that you find more disquieting or dissatisfying than I did as a reader. And I'd be interested to hear more about what, what is it about this vision that, you know, you would wanna change or that could be better? Like, what's the what's the more utopian version look like in your imagination?

    1:06:25

    Joel Borgen: Yeah. Well, I mean, I don't know. It's it's a delicate balance of spoilers, I guess. But the humanity, their trajectory, and their their rights and ambition are still very, very limited and truncated by the actual situation. The superintelligence does not… They're happy and willing to keep humans along around, but they are kind of almost pet-like in their orientation at that point. They don't have the kind of agency and ability to expand beyond the solar system or to be ambitious in certain directions. So there's definitely still some ugliness that you see around that. And my vision for the future would

    1:07:11

    be to capture all of the freedom and the technological possibilities that you have, but also keep keep a superintelligence from taking over and having their their own inscrutable goals that are kind of limiting additional forms of human flourishing. I think… I don't know if I didn't highlight enough kind of the dystopian elements from my perspective, but I tried to make the kind of eventual decision later on in the book that's kind of the culmination of it to be to be really kind of

    1:07:56

    at the knife's edge where some people would would choose to try to potentially upend the apple cart and hope for something even grander versus kind of accepting the status quo when not taking risk. So it wasn't… It was intended to be a difficult choice. And it's not surprising, I guess, that we would have different instincts as far as what the correct move is at that point. That was sort of by design.

    1:08:26

    Nathan Labenz: Did you think of the AI, the Steward, as potentially a worthy successor in some sense? I mean, I sort of… I'm definitely not a successionist, but I do think, like, jeez. In our current form, we are not really that well suited to go out and explore space. I think it's, like, pretty unlikely that, you know, my current fleshy form will be the thing that goes out and does that. So it does seem like there's gonna have to be some sort of modification, evolution, you know, substrate swap, or maybe, you know, even a bigger

    1:09:12

    change than that. And I think people have very different intuitions for what would count as good. Right? Does it have to be me and my physical body? Would it be okay if it's me as an upload? Would it be… You know, would I still have some reasonable basis to identify with some superintelligence that sort of, you know, is much bigger than any individual human, but still in some way carries on our values or something. How did you think about the central AI in the book when it comes to that sort of analysis?

    1:09:46

    Joel Borgen: Yeah. I mean, if I was trying to design something I would call worthy successor, I'm not sure that I would endorse that idea any way, like, as a as a concept. I think I think we will have, if things go well, we will have successors, certainly. We'll have people that become posthumans. We'll have other other minded entities, hopefully, with phenomenal consciousness. And I'm hoping that we can reach a point where, as in the book, you can kind of live alongside it and choose to what extent you want to interact or see them as a, a worthy successor. But, no, if I was designing worthy successor, it would not look like the Steward necessarily. That's…

    1:10:32

    It's intended to be a stable, but not entirely friendly regime, I guess, in that sense.

    1:10:44

    Prakash Narayanan: Did you did you find yourself making any identification to a supreme being, a god? Because that is one of the cult like beliefs that some people have in essence.

    1:10:59

    Joel Borgen: No. No. I don't. That doesn't really it doesn't really interest me. I mean, it'll be godlike in terms of powers, but not in terms of a spiritual sense to me at least. one

    1:11:16

    Nathan Labenz: of the things I wanted to touch on a little bit is music. So you have, I think, a lot deeper understanding and, more intensive training in music than I do by far. The protagonist's father is a musician composer in the book. Mhmm. And we had a chance to collaborate a little bit on making an AI song for the outro for the episode that we put on the podcast feed. I think your comments earlier around, like, the sort of dissolution of common culture are pretty apt, and I'm kind of living that increasingly because I have enjoyed making this AI music for the podcast so much that now

    1:12:02

    when I go out for a run or whatever, it's not all the time, but it is, like, some of the time that I'll put my own AI music on. And I really enjoy that. It kinda takes me down a, you know, memory lane sort of experience, but it's also totally idiosyncratic to me. And, you know, I really don't share it with anyone in a really meaningful sense. Obviously, podcast listeners occasionally, you know, tell me they like songs, but, like, I don't think too many people are going back to them. I don't think they have, like, an enduring place in other people's lives in the way that they do mine. So I'm also sometimes mindful. Like, okay. I need to listen to something somebody else, you know, also, has connected with.

    1:12:39

    Joel Borgen: Right.

    1:12:42

    Nathan Labenz: What thoughts do you have on AI music, and on what role it might play in our future, what role it should play? I wouldn't wanna give mine up. That's for sure. But I also do feel like something is lost.

    1:12:55

    Joel Borgen: Yeah. Well, I guess to go against part of what you said, I actually do like a number of your songs. I think there's a subset that I've listened to many times. The episode we had with Davidad where he… There was the… There was kind of a Beatles, you know, Indian inspired song. I really enjoyed that one. It was the first time I listened to a Suno song, and I was like, wow. I actually enjoy listening to this as music. So it inspired me to go and try my hand at it myself, and I've only just started with the one that we did for the episode. But I've long been interested in AI music, and, like, I have a dream to create a kind of a type compositional assistant that the father of the main character has in the book as well.

    1:13:41

    I have a lot of interest in training in classical music, and Suno is very, very good at pop songs. And I think that… So I… I've been looking to… I've had this held out test for well over a year now, maybe 2 years even. Very simple. A simple, like, high definition capture of 5 notes on the viola, and I give it to every… Give it to the reasoning models. I'm like, okay. What is the clef? What are the notes? What are the note values and so forth, the time signature? And not a single one of them got it right until Astra OpenAI put in some… Presumably put in some actual musical scores.

    1:14:27

    And so it went from basically not being able to do anything to… Now I gave it a complex, you know, 20th century score that's almost painful to look at given how complex it is. I mean, this, you know, double sharps and accidentals and ties everywhere. Is… It's it's, like, difficult. And it did an almost perfect job turning it into a machine-readable form via Codex. So the score that, like, you just basically had a PDF of before, now you can turn it into something that you can manipulate and evaluate. And it… So it went from kind of 0 to 99 overnight. And I think that is an interesting way to turn some of the scores that aren't machine-readable into something that you could use to train a system. So I would love to see Suno move in the direction of

    1:15:13

    adding a lot of classical stuff to it because I think in addition to it making interest in classical music… And classical music will have… It requires a different a different level of understanding of the form because it's much larger scale often. You know, like, a 3 hour long opera that has some internal structure or even a 20 or 30 minute single movement or something that has a lot more going on than a pop song. So I would love to see see that happen. As far… Again, it does tie into the same element of if you're making just kind of private art just mainly for yourself to listen to. I still think it's valuable. But, again, like like, with writing a book, I would want to do

    1:15:59

    it in tandem with the model, this hypothetical model in the future, and create something that's, you know, a large part of my efforts and taste goes into it as well. I don't know how long that interregnum will will last where you have, you know, a role for humans in AI to create something together. We use it as a tool, but it still requires a lot of you and your judgment. But I look forward to experimenting with that once the models improve.

    1:16:30

    Nathan Labenz: Here's to the interregnum. Well, we're just about at time. I'll just briefly read a quick review of the book that I got in this morning just before we, came on live together. And then I'll invite you to share any last thoughts, and we can let you, get back to work for today. The review that I got, this was just came in via a DM on Twitter, was after listening to The Receipt Horizon, I am going to have to get it. It is pretty damn good. So people can go to The Receipthorizon.com to check out the book. The first 4 chapters are on the feed, but with the full audiobook and, of course, the text copy is there on the website. Any closing thoughts for us today, Joel?

    1:17:20

    Joel Borgen: Well, I want to thank you publicly, Nathan, for your your support and helping me create the audiobook. It was a really fun project, and I'm I'm glad that people are able to listen to it because I do most of my book consumption via my ears these days as well. So I thank you both for having me, and thank you, Nathan, for your your help in creating it. It means a great deal.

    1:17:45

    Nathan Labenz: My pleasure. I'm looking forward to the even more utopian novel coming before too long.

    1:17:50

    Joel Borgen: Sounds good.

    1:17:51

    Nathan Labenz: Joel, thanks for joining us.

    1:17:55

    Joel Borgen: Yep. Thanks so much.

    1:17:57

    Prakash Narayanan: Bye bye.

    1:17:58

    Joel Borgen: Bye.

    1:18:00

    Prakash Narayanan: Awesome. And

    • Astra sheet music test: Five notes to complex scores

      0:00 / 0:00
    • AI writing: Why correct text can still feel like slop

      0:00 / 0:00
    • AI music: Nathan on personal songs and shared culture

      0:00 / 0:00
    • AI utopia in fiction: Why comfort isn't human freedom

      0:00 / 0:00
    • AI novel editing: A 14% cut with human judgment

      0:00 / 0:00
  3. --:--Interview
    Daniel McKinnon: AI agents for rare disease diagnosisDaniel McKinnonGamow Labs’ CEO explains why genome interpretation is the bottleneck, how RareBench evaluates models, and why the advantage of specialized scaffolding may shrink.
    Open segment on YouTube ↗

    Daniel McKinnon described the personal loss that led him to found Gamow Labs. After his son Owen died from a rare developmental lung disease and his family faced further genetic uncertainty, he built an AI-assisted interpretation pipeline. He said it recovered a variant missed in Owen's initial sequencing analysis, prompting him to pursue clinical genetics as his life's work.

    Nathan pressed on what the AI had done differently. McKinnon described a structural deletion in an enhancer far from FOXF1 and a reasonable human-workload filter that excluded variants too far from a gene. He argued that an agent could keep investigating beyond those practical stopping points, and that clinical interpretation offers long-horizon, tool-intensive problems with verifiable answers.

    Prakash asked about regulation and the new wet lab. McKinnon distinguished clinical decision support from offering a diagnostic test, then explained why variants of uncertain significance can require biological experiments to resolve. He described acquiring equipment and hiring four people from Arpeggio Biosciences as it wound down, with the aim of generating functional evidence in physiologically relevant cells and feeding that evidence back into interpretation systems.

    The discussion separated broad screening from diagnosing a specific clinical problem. McKinnon emphasized the difficult incidental findings that whole-genome screening can surface, and said Gamow's current work focused on acute reanalysis for families and hospitals plus bulk reanalysis of historical cases. He described the four-month-old company as prerevenue and invited families with serious unresolved cases to contact him.

    McKinnon outlined Gamow's work on model ensembles, tool access, harnesses and evaluation. He reported roughly 50% performance for a vanilla Claude Code setup on RareBench, versus roughly 10% for a traditional variant-ranking approach, and estimated interpretation costs around $10 per genome in the benchmark runs. Those were his reported results and estimates. He discussed model-specific failures, including unsupported gene renaming and missing information in papers, and said the tools must evolve with the models.

    A brief lab tour showed the setup for high-throughput experiments. McKinnon described an ambition to reduce per-experiment costs from contract-research quotes around $50,000 to around 50 cents through robotics, experimental design and AI-assisted analysis; he did not present that target as an achieved cost. After he left, the hosts discussed the value of applying frontier-lab expertise to other fields, the importance of biological data quality, and the real medical opportunity costs Nathan believes must be weighed when considering slower frontier development.

    Timestamp links open the original source recording.

    The only way to add information to the system is through a biology experiment.

    We're just trying to diagnose more sick kids.

    Ultimately, our job is to take the smartest intelligences in the world and mix them together and give them access to the best tools and then get answers for our patients.

    1:19:43What took you from losing your son to founding the company?
    McKinnon said he spent years looking for a way to contribute, considered medicine and existing diagnostics companies, and eventually built his own interpretation pipeline during a later pregnancy. When it recovered Owen's missed variant, he began exploring a company; he left Meta once family circumstances allowed it.
    1:26:20What did the model do or connect that the doctors hadn't been able to?
    He described a deletion in an enhancer far upstream of FOXF1 that fell outside a clinical lab's structural-variant filter. He said his early o3-based loop kept investigating after the obvious coding regions, illustrating how machine time can extend the search beyond human-workload limits.
    1:33:35Are there regulatory barriers before this technology can reach patients widely?
    McKinnon described starting with clinical decision support for professionals, with privacy and HIPAA obligations, and contrasted that with the lab and test requirements of offering a diagnostic product. He regarded that path as more tractable than new-drug approval and left open whether Gamow would build or partner for testing.
    1:36:05What was the idea behind opening a wet lab: access to data or reinforcement learning?
    He said interpretation often ends with a variant of uncertain significance. Functional studies in relevant cells can provide additional evidence, help resolve a case and create a learning feedback loop. An opportunity to acquire Arpeggio's remaining lab and hire four team members accelerated those plans.
    1:43:53Are you horrified by how little we know about genetic data?
    He said he initially felt that way, but now sees interpretation as the missing layer after sequencing became inexpensive enough. He argued that agentic systems can assist with ranking variants, explaining their significance and considering management options, while reporting substantial benchmark gains over a traditional ranking tool.
    1:46:46Who should use your emerging service, what should trigger it, and what does that path look like?
    McKinnon separated broad screening from diagnostics and explained why incidental findings complicate screening. He said Gamow currently works on serious unresolved clinical cases and bulk reanalysis with hospitals, is prerevenue, and welcomes inquiries from families facing a sick newborn, a loss, a questionable pregnancy or severe family history.
    1:54:23What have you built on top of the models, and how do interpretation costs compare with sequencing costs?
    He estimated around $10 per genome for model interpretation in RareBench runs and described Gamow as a harness, tools and evaluation company. He pointed to model ensembles, routing, improved access to specialist resources and trace analysis that catches specific failure modes. He did not give a single end-to-end runtime for every case.
    2:00:06How do you measure the performance of your tools independently as model families improve?
    McKinnon said he did not have a clean separation: the tools and models are co-designed. He gave examples of models missing paper contents or stalling on tools, and said robust evaluations reveal failures early so the harness can be adapted.
    2:02:48What would you ask frontier model makers to improve?
    He asked them to optimize for clinical genetics and benchmarks such as RareBench. He acknowledged that better base models shrink the added value of a harness, but said his objective is better diagnosis and patient outcomes, with value also residing in the clinical relationship.
    2:07:10What gets you from a $50,000 experiment to a 50-cent experiment?
    He described that reduction as a goal enabled by a combination of robotics, 384-well experiments, many plates, AI-assisted primer and experiment design, and analysis of large datasets. He did not identify one invention or claim the target had already been reached.
    2:08:14How much data do 384 samples generate?
    McKinnon said whole-transcriptome experiments can produce RNA-sequencing datasets on the order of 100 gigabytes per well, depending on depth, while stressing that not every bit is valuable. He described very large throughput and a need to narrow the useful data over time.
    2:08:59What was talking to frontier model companies about your biological use case like?
    He described supportive individual responses and enthusiasm for helping patients, while distinguishing those reactions from redirecting an entire company. He saw both personal motivation and a concrete positive story about AI as reasons for the labs to engage.
    2:11:08What final thoughts would you leave people with?
    McKinnon said Gamow was growing the team and invited people interested in the mission or an AI-controlled biology lab to reach out.
    Lightly edited · timestamps jump to YouTube
    1:18:08

    Joel Borgen: And

    1:18:09

    Prakash Narayanan: now for something completely

    1:18:10

    Nathan Labenz: different.

    1:18:16

    Prakash Narayanan: Our next guest is Daniel McKinnon. He's the founder and CEO of Gamow Labs, a company building AI systems to help identify the genetic causes of rare diseases in newborns. Its system, George, combines genome data with descriptions of a patient's condition, uses computational tools and scientific evidence, and proposes genetic explanations for clinicians and researchers to evaluate. His connection to clinical genetics is personal. His son Owen died in 2021 from a rare developmental lung disorder after initial genome testing failed to provide an answer. A specialist subsequently identified the genetic cause. Years later, while investigating another pregnancy,

    1:19:02

    Daniel built a prototype that also recovered Owen's missed variant. That work became Gamow Labs. The company now works with clinical researchers on difficult historical cases and has introduced RareBench, a 122-case evaluation of AI systems for genetic interpretation. The aim is to measure whether an AI can find and prioritize a disease-causing genetic change rather than simply producing a persuasive explanation. Dan, welcome to the show.

    1:19:38

    Daniel McKinnon: Good morning. Thanks for having me.

    1:19:43

    Prakash Narayanan: Firstly, it is heartbreaking to hear about your son. You know, I think Nathan went through... Almost went through something similar last year. So I think, we are well aware of how heartbreaking it must have been. What took you from that moment to founding the company?

    1:20:07

    Daniel McKinnon: Oh, first off, I would acknowledge that, yeah, no, this is the worst thing that can happen to a parent and perhaps a person. And, Nathan, I'm so sorry for your experience as well. I didn't know that before coming on the show. So just as a father to a father, I'm happy to support however I can. And it was very clear after this happened that this was the trajectory of my life. I thought about, should I quit my job? I was working at Meta at that point and go back to med school and become a neonatologist to care for babies. Was there another way I could contribute? I talked to people at companies like Natera to say, should I go work there?

    1:20:52

    And I knew that this was the place that I wanted to put my stamp on the world and make sure no1 else ever had to go through what we went through. It took me 5 years to get there, for reasons that I don't think matter that much. Well, the neonatologist thing was obvious is I just looked at the training, and I was 35 at that point. And I said I wouldn't be able to treat patients for at least another decade, and that's not a good use of, scarce medical resources, especially around training. As for going to a company like Natera, and it just... It wasn't the right, I think, match for, my skill sets, ambitions, and the stage of the company. And I just needed to find the right way

    1:21:37

    to contribute to this, and I've been looking for that for 5 years. And it really leapt out at me last summer now when... I mean, we blessedly have a healthy 13-[age unit unclear]-old right now. But because of our history, we are being monitored very, very carefully. And there was just something a little bit questionable in the 16-week anatomy scan that our own maternal-fetal medicine doctor, who is an absolute wonderful person, said, normally, I wouldn't even flag this, but because it's you guys, we need to look carefully. And we did a whole genome of the fetus at that point. It came back negative. This is actually the third negative whole genome I've seen in between losing

    1:22:22

    Owen and having our son Warren. We actually lost a second pregnancy due to genetic reasons very, very late. So, we've had basically 5 years of heartbreak before bringing Warren into this world, and all of it kind of genetic mystery related. And so just as a patient, I've started to learn... Or a family of a patient. I start to learn a lot about, the failings of this system. And it was that moment where I said, I'm heartbroken, but I'm mad, and I'm gonna do something. And I have tools to do something. After Owen died, I said, hey. I wanna learn everything I can about genetics, but it actually used to be quite complicated. I mean, you see posts on Twitter all the time. Like, hey.

    1:23:08

    You know, I dropped my VCF into Claude Code, and it told me that I'm, predisposed to having skin cancer. This was not available 5 years ago. Like, if you wanted to do genetic analysis, you needed to become kind of an expert in all of these obscure bioinformatics tools, and I actually tried to gain some of that expertise. In fact, a funny coincidence right now is a person that I work with at this company today who's actually right behind me now is one of the people who was trying to get me into university short-read sequencing interpretation courses 5 years ago. And just given my, mental state, my, demanding job at Meta, and what I had going on, I just couldn't put the time in. But fast forward to last summer, I got this result back. I said, this is unacceptable. This is the number one thing that I want in my life. I really

    1:23:53

    wanted to have a family. I wanted to understand what would happen. I... You know, I didn't believe these labs. And I basically called all the labs, and I said, give me my raw data, which you can get, due to our HIPAA rights here. And I, vibe coded my own interpretation pipeline, which... I mean, now vibe coding is, crazy. Right? Like, right now, this would be so easy. You could just say, Claude Code, make me this thing. But back then, I mean, it was still, kind of, autocomplete original codex. Like, it did actually require, quite a bit of work on my end. And I just did it to see if there's any comfort that we could get around the pregnancy. And I was just shocked that it also outperformed on these other cases. And, most clearly, is it

    1:24:38

    diagnosed Owen when, one of the best or, some might even say the best prenatal sequencing lab did not. And I knew at that moment that this is something I had to contribute to. I didn't know it would be a company. I was like, maybe this is an open source project. Maybe it's a foundation. Maybe it's a nonprofit work. Of course, I also thought maybe it's a company, but I didn't know enough about the business of diagnostics to really make that make that leap. And I and I started, making connections around the industry, people working at diagnostics labs, physicians, venture capitalists. And I really, created the hypothesis. This is actually quite a big business. It is quite screwed up right now, and I have kind of a unique

    1:25:23

    experience both personally and professionally to move this field forward. And I want to. And, it's been a life mission for a long time, so it was it was very obviously the thing for me to do. And I started developing this conviction maybe when Warren was 3 months old or something like that. For everyone who's had a baby... I mean, maybe some people had easy babies, but Warren was not an easy baby. So I really wasn't in... I wish I started the company 9 months before I did, but I just couldn't from a practical perspective. And it was when he got to be about a year old where I said, okay. We... Our personal life is locked down enough where I can take on this new challenge. And

    1:26:09

    and I left Meta, and I started the company. And it's been 4 months, and it's been the best thing professionally I've ever done.

    1:26:20

    Nathan Labenz: Could we go back to the original moment of a model diagnosing in a way that the professionals hadn't been able to? I... I'd love to understand. Obviously, this has been a minute, but maybe then and now. I would love to understand what advantage the AIs had, because we have a lot of different kind of high... Competing hypotheses around why they will be able to do better and also, why they might not be able to do as well as leading humans at things. So what concretely did you observe about what the model did, considered, connected that allowed it to do something that the doctors weren't able to do? Yeah. This is very similar to a

    1:27:05

    Daniel McKinnon: lot of other AI work where it's basically like amplifying human or even subhuman performance over time. So right now, I would argue, actually, just a vanilla kind of Claude Code genetic interpretation, task is superhuman, meaning you just don't even do anything. I mean, we build a bunch of stuff on top of these models. But if you didn't even do anything, I think that you would perform at a superhuman level relative to a human genetic analyst. The reason that happens is, a, it's partly because since last year, the models have gotten, crazy smart. Partly, it's because the models are just, smarter than humans at most things right now. 2 is that they can just work harder. So to go, very specifically into why my son's case

    1:27:50

    was missed... And I know this in-depth, and I actually have spent a few days at the sequencing lab. And I don't wanna say its name because I don't wanna imply that they did a bad job or anything. This is, one of the best... I mean, I can just say it's actually... It's Rady Genomics. They're a wonderful group of people in a wonderful lab. I don't wanna imply they did anything wrong. It's just a very, very hard, labor-intensive human task. And I'm really grateful for them is that I reached out to them, I said, I wanna learn. I wanna understand more. And they actually hosted me for 3 days. And I met everyone there, and I learned how it worked. And they walked me through exactly how their systems worked and how it made this mistake. And what happened precisely, somebody listening to this who's not familiar with the field might say, oh, that's dumb. But it's not

    1:28:35

    dumb. There are lots and lots of, long-tail trade-offs that you make when you're building, software or doing research projects where you say when to stop, and humans have to know where to stop. So very specifically, what happened here is my son had what's called structural variant. So he had a 91-kilobase deletion in the enhancer of a gene called FOXF1. FOXF1, when there's a loss of function, causes a disease called alveolar capillary dysplasia, which is a lethal infant lung disease. Probably, 1000 cases have been reported, but it's suspected to be much more common. So you'd say, well, how could you possibly have missed a 91-kilobase deletion? This is a huge chunk of your genome missing. Well, how you miss it is

    1:29:20

    this enhancer is a megabase, one million bases upstream from the gene. And what Rady did, what many other labs do, which is a very reasonable assessment, is they have a filter that for structural variants, which tend to be quite messy in what's called next gen sequencing, which is how sequencing is done today where you have basically roughly 150-base-pair fragments, you need to piece them into this clinical puzzle. They said, we're gonna have a filter. And anything more than one kilobase up or downstream of the gene, we're not going to consider. And this is a totally reasonable trade off if you have humans looking at all of this stuff. But my moment was... And this is before Codex,

    1:30:06

    before Claude Code, was can you put the o3 model, which was the model I used at that point, kind of first, real agentic model? There was o1, but o3 was, really first breakthrough. Can you put it in some kind of loop and just have it keep looking? And it basically looped, and the first loop was like, is there anything wrong with the coding elements, which are kinda obvious things? And it's using these kind of bioinformatics tools. This is very crude compared to what we had today. But then it misses something, and it goes through another loop and another loop, and it just keeps working. And when you are a clinical lab, whether you're a profit, nonprofit, whatever your structure, ultimately, you gotta move. You know, you have to, spend some amount of time on each case. And if you don't come to a conclusion,

    1:30:51

    you know, you say this is nondiagnostic. And this is very common. Most whole genomes, even with infants suspected to have genetic disorders, come back nondiagnostic. And I really think it's one of these, meat space problems is if we can export these problems onto a machine intelligence that can work nonstop and in parallel, then we will be able to see many more kids, treat many more kids, and do much more kind of interesting analysis on top of the basic things that humans are just pressed for time. And I don't wanna claim I've, reinvented this. Like, people are trying to build software to accelerate genomic interpretation. People have been doing this for a long time. But I think what I probably relatively uniquely identified early

    1:31:36

    was that this is a really great task for agentic AI. And from, improving on it, it's also... It's long-time-horizon. It's agentic. It's verifiable, and it's unsaturated. So I suspect that this will be a task like math where we can just generate these, Olympiad level problems and just keep hill climbing on that until this problem is, basically solved. And then once a machine can solve the problem, there's no bottleneck for getting anyone who needs the treatment. And one thing that I didn't mention here, but it's very important, is I was lucky in that my son was at Children's Hospital Colorado, which is one of the top NICUs, in the country and world. And they have access to these resources, like genetic counselors,

    1:32:21

    clinical geneticists, sequencing cores, all of this stuff. If this happened to you in Arkansas or India or Brazil, you actually wouldn't get this treatment, and it's largely a human bandwidth issue. So, my real goal with this company is to make sure that every baby has cutting edge treatment and to expand that sphere before and after birth. So try to help children with genetic diseases, try to help prenatal fetuses with genetic diseases, and even get better carrier screens. So many people have taken, for example, the Natera carrier screen. In fact, we took it. Our variant was inherited. We could have known about this before Owen was even born. Many, many, many people with children with lethal or severe

    1:33:06

    diseases could have known about this before their children were even born. But the existing technology is not meeting their needs. And I think that we could save, lives on the order of, the worst affliction affecting humanity by just helping some of these children not be born. And it's not that there would not be any other children. You would just use something like, assisted reproductive technology, aka IVF, to find embryos that would not have these severe mutations.

    1:33:35

    Prakash Narayanan: You've spoken a little bit about the technical challenges. Are there regulatory challenges as well? Because I believe the FDA requires genetic counselors. They require a certain standard of evidence before you can, say that this test does something. And you also have this problem where you have... When you have a limited sample size, you have a problem with the FDA of, proving that this test actually does what it's supposed to do. Are there regulatory barriers that need to be resolved before this technology can actually be widely propagated to patients?

    1:34:17

    Daniel McKinnon: Yeah. So the short answer is yes, but they're not bad compared to, say, a new drug application. So when I started Gamow Labs, I said, I wanna solve this whole problem start to finish, but I'm also a very incremental hill climber type of guy. There are founders who are just like, I'm gonna go raise a billion dollars, and I'm gonna go solve this problem and come back when I'm done. And I think, one that I admire that I think is quite interesting is, Periodic Labs. Like, my background is a physicist, and I'm like, that's really cool, and I hope that they do that. I'm more of like, I wanna start something small and tractable and make sure that it's successful, and that was a portal for people working in genetic medicine. So I call this, OpenEvidence for rare disease. And that is regulated under the clinical decision support framework,

    1:35:02

    which is... I mean, there are... You have to be HIPAA compliant. You have to have standard privacy safeguards, but it's not like you have to prove to the FDA that your advice is

    1:35:11

    correct,

    Prakash Narayanan: Mhmm.

    1:35:12

    Daniel McKinnon: which is why, OpenEvidence has grown so much, and it's why you see products like ChatGPT Health has a big banner that says this is not, medical advice, which is... It's, kind of a funny loophole, but it's like this is a very, kind of lightly regulated thing. Once you get into offering your own test, which I suspect I will do at some point, but I don't know. There are lots of really, really great diagnostics labs that it's totally possible to partner our way into creating something excellent for patients, so I don't wanna, precommit to any of that. You're regulated under this CLIA/CAP framework, which is really around, the lab and the test versus the analysis itself. I don't wanna underplay this by saying it's, easy, but it's nothing like, an FDA,

    1:35:57

    phase 1, phase 2, phase 3 trials. So it's quite tractable for a relatively small company.

    1:36:05

    Prakash Narayanan: You have actually started to open a wet lab, if I'm not

    1:36:08

    mistaken.

    Daniel McKinnon: Yeah. Yeah.

    1:36:09

    Prakash Narayanan: Take the steps. So what was the idea behind that? Is that about getting access to data or reinforcement learning? Like, what is the... What is the ideation of the wet lab?

    1:36:21

    Daniel McKinnon: Yeah. That's exactly it. So I knew I would want both of these components from the day I started the company. You know, whether we build the clinical lab or we partner with somebody, I want to be able to reach patients directly. If all we're doing is reanalysis, we're helping people, but I think not as much as we could be doing. And on the backside, once we've done the interpretation, the majority of clinical cases we see right now... And to be clear, we are only seeing hard cases. So this isn't, most genetics labs. We are seeing cases that need reanalysis that did not previously have an answer. The majority of them end up in what's called a VUS state. So it's a variant of uncertain significance, and these are scored according to, a very standard rubric developed by

    1:37:06

    the American College of Medical Geneticists or ACMG. And you need to get 6 points or more to be bumped into the likely pathogenic category. And that unlocks a lot of treatment options, insurance options. Like, just... Let's say it's, good to be either benign or likely pathogenic. It's very bad to be in this kind of intermediate stage. And if you are in this intermediate stage in a VUS, you can do 2 things to get a diagnosis. 1 is you just wait, and you can get more points if more patients emerge who have a similar phenotype or a similar condition as you and the same genetic variant. So this is often what happens when, you read in the news. You're like, oh, this kid has, had epilepsy for 10 years. It finally

    1:37:51

    got a diagnosis. And if they're lucky, oh, and since we know that's a diagnosis, we work with pharmaceutical company, and there's, some off label use of a drug, and it actually can, help their condition. That's typically because they just waited until somebody else had the condition, but this is not scalable and takes a long time. Another fork is you can convince a university lab to kind of care about this problem. And what they'll do is they'll do what's called a functional study. So they'll make a cell line that's, emblematic of your particular condition. So, for example, for alveolar capillary dysplasia, that commonly used cell line is called IMR-90s. It's a fetal lung cell line. And this is very, very important because this particular gene is only expressed from week 16

    1:38:37

    to week 20 of development. Like, you can't just put them in a cancer cell line. And, let's put a pin on that and come right back to that in a second because it's related to your question. And then you can go in, and you can do something. Like, you can edit the genome, or you can insert a plasmid, or you can do a number of different functional studies to say, okay. In this physiologically relevant cell line, if I have this genetic mutation, is this gene broken in some way? And that's happening, call it, hundreds of bits a year in physiologically relevant settings. There are, very high-throughput studies in in cancer cell lines, and this is actually, what, AlphaGenome is mostly trained on. But there's not that much what's called primary cell data on these genetic variants. And I actually

    1:39:22

    went to my friend, Tim Reed, who was the same person who recommended that I take that short-read summer course 5 years ago, who did his PhD in computational biology in, Robin Dowell's lab who basically studied the effect of transcription factors on various expression levels. And I said, Tim, you've got this company, Arpeggio Biosciences, and you have a room full of machines that can do this at, very high-throughput. Why is no1 doing this at very high-throughput? And he said, well well, a, it, wasn't really possible until recently. Like, there's some recent innovations in, large editing, CRISPR prime editing. You can you can read that paper in some other labs. And he said, it really only has become possible recently.

    1:40:07

    And I said, well, can you help me? Can we, work in your lab at night, and we can, come up with these? I mean, I probably had a list of 100 variants I wanted to understand to get these patients better diagnoses. And, again, there's no ins... There's nothing you can do in silico. You know, you model them. You look at other patients. You look in the literature. You're at a dead end. The only way to add information to the system is through a biology experiment. And Tim said, no. I can't do that because we're in the middle of a drug development program. We're making a lung cancer molecule that targets NRF2. And I said, well, that's a bummer. But he said, but I can help you try to scope this project with some CROs to see if you can pay somebody to do these experiments. And I was spending maybe an hour a week on this. This wasn't a big, use of time. And when I was on the phone with one of them,

    1:40:52

    and I was getting nowhere, people were saying, oh, we'll doone for $50,000. You know, I wanna doone for 50¢. Right? I wanna map all nine billion possible variants in the genome. Right? This is... This needs to be done very cheaply. And I was on the phone with one of these, and I kind of mentioned Tim. And they said, hey. Wait. You didn't hear that Arpeggio Biosciences is getting liquidated. Like, their their candidate failed some kind of preclinical milestone. And I said, okay. This call is over. I've gotta give Tim a call. And so I did. And lo and behold, that rumor was correct. They were winding down the lab. And I said, Tim, I wanna do what I originally asked you. I wanna take this machine in the back, and I want

    1:41:37

    to start systematically mapping all of the VUSs that have ever been reported and understand what is going on with these kind of mysterious... Clinically mysterious regions of the genome. He said, great. So, this expanded the scope of the company a lot. I called up all of our existing investors and a few more, and I said, hey. I need more money to do this. So we basically bought what was remaining of Arpeggio, hired 4 people on the team, and very quickly pivoted to creating this basically, RL data from real world with biology experiments. And we're 3 weeks in. We're running our first experiments actually today. In fact, if it's not too weird, maybe, at the end of this conversation, I can, bring the camera over here, and we can, take a look at this robot. I think it's running right now.

    1:42:23

    But it was just, this incredibly fortunate coincidence. I wouldn't have done this for years without this, human connection to somebody I've, worked with for 15 years and all of the technology that he's built leading up to this. And I'm really looking forward to being able to close that feedback loop, and I think a lot of clinical genetics labs are as well as this service is not commercially available. We can not only sell kind of, a lookup table version of this. Like, oh, I have a patient with this variant. Can you help me get a couple more ACMG points to get this up to likely path? Because you can get 2 to 4 points. Remember, you only need 6 points, so 2 to 4 points is a lot of points for a functional study. But, also, it lets the machines learn.

    1:43:08

    Right now, the machines aren't learning. Like, you get to a VUS. It's like, that's the end. Now it's like, you get to a VUS. You do the biology study. You say, oh, this region, the genome, is actually quite important for this particular disease. Now we know. And it's whether we train another, AlphaGenome style model, whether we work with Google to kind of improve their AlphaGenome style. Actually, in the first paragraph of their most recent Atlas paper, they say, this is great work, and it is incredible work, by the way. I love the work they're doing there. But this is not trained on physiologically relevant cells. And so if you look at their tracks, I mean, one thing that we're doing in parallel is saying, hey. AlphaGenome predicts, this whole 15,000 [unit unclear] region to have, some kind of pathogenicity. We're gonna do that. We're gonna check these predictions, and then we can improve these predictions

    1:43:53

    over time.

    Prakash Narayanan: Firstly, that's the most positive Vulture investor story I've ever heard ever. Secondly, do you find yourself looking as a more of a data scientist person looking at this stuff and being horrified at how little we know about genetic data in general?

    1:44:21

    Daniel McKinnon: So I felt that way originally. I was like, oh my god. We sequenced the genome 25 years ago, and, Bill Clinton and Tony Blair got up on stage and said this is gonna revolutionize human medicine. And I look back, and there's this, graveyard of genomics companies. Like, you mentioned Counsyl, which I actually would not say is a graveyard. I think they were, modestly successful. But, why is all of health care not based on precision medicine? Like, everyone in this space is like, it should be, and there are examples. It's just too challenging. The interpretation is too challenging. So I actually look back and I see... And, we kinda talked about Counsyl ahead of time, is I think that what was missing prior to now was really that interpretation layer.

    1:45:06

    It's very, very complex. Like, your genome sequencing has gotten what I would call cheap enough maybe 5 years ago. I mean, below $1,000 a genome, maybe $500 a genome. There are high-throughput labs that are doing this for less than $100 a genome. You see announcements on Twitter saying, oh, I can do a genome for less than $100. There's many, many people doing this right now. Like, the sequencing is not the problem. It is given your problem, what insights can you derive from that? And that was only possible as of, last summer. Like, I'd say o3 is the first example of that. And it's not just the variant interpretation. Right? It's explaining to the provider why it's important. It's ranking variants in a nice way. It's explaining, how can you treat this person. It is ranking, different patients

    1:45:51

    for ASO eligibility. There are many, many, many other things beyond just scoring the variants. Although I will say we benchmark a variant annotation in RareBench and other benchmarks we have, and the best performing, I'd say, traditional machine-learning-based approach in terms of variant ranking is called Lyrica. I think it scores something like 10% on our benchmark. And right now, Opus 5.5 in Claude Code, just a vanilla thing, scores something like 50%. So there's... I

    1:46:22

    mean,

    Prakash Narayanan: Mhmm.

    1:46:23

    Daniel McKinnon: these traditional tools are also getting blown out of the water by, this newer approach, but then they can do much more. And you basically need to make a very, very cheap, easy thing that people who are not surrounded by fancy clinical geneticists, genetic counselors, top-tier hospitals can use, and that's the problem that we're trying to solve.

    1:46:46

    Nathan Labenz: Can you describe in today's world, and then maybe you could also, foreshadow where you're trying to take this, who should use your emerging service at, what... You know, what should trigger them to do it, and what does that look like, and what is, the cost breakdown? I'm kind of envisioning, you might say, well, anybody who's had, a family member on either parent's side that had some unexplained issue, maybe they should come down this path. Right? And they should maybe think about doing IVF, and then we can, do this analysis for all the embryos, and then they can, choose one that seems to have a,

    1:47:31

    you know, a good chance of healthy outcome. Is that the kind of path that is, immediately difference making for people? Is that the biggest path? To... Paint that picture for me.

    1:47:45

    Daniel McKinnon: Yeah. So I think one early fork I wanna draw is the fork between screening and diagnostics. So screening is, in some sense, a great business. So, Natera or BillionToOne are, classic examples of these, and that everyone should get screening. Right? I think... I suspect most pregnancies get an NIPT test right now. And... But the expectation is that you get no answer. So 99 percent or 99.9 or whatever percent of people get... Maybe they get peace of mind, but they don't really get, value from those tests. Screening can definitely be improved through, whole genome and an agentic AI based approach, 100%. Like I

    1:48:30

    said, in my personal case, we did the traditional screening, and it missed our cases. And this is, I would say most babies with genetic problems in the NICU, their parents did this. I can't say most. I don't know the numbers. But many babies in the NICU who have seriously ill babies did some kind of screening, and it missed it because the, the Natera test, I don't know off the top of my head, but it's a microarray, and maybe it looks for 40 things. And it turns out that, well, rare diseases are really rare individually. In aggregate, they're not. And so there... There's this, huge, head, and this is what maybe Natera looks like. And then there's this long, long, long, long, long, long tail of all these problems. So, yeah, I would say that, screening for, known genetic problems is a great idea. I don't think that we have the,

    1:49:16

    bandwidth to do this right now. We're 4 months old right now. We're doing some kind of early screening experiments with a few, hospital pilots. But, screening with whole genome raises a lot of questions too about, what do you do with incidental findings? And I could give you, a very, very specific example. There's a family that we're working with right now that had an incredibly tragic story. They... Their first son was in the NICU with pulmonary hypertension alongside of my son, Owen. He got out, and I always said, hey, Owen. Why can't you be like this per... I... That's not the right way to word it. But I said, hey. This person should be, your inspiration. Like, let's be like this person. And

    1:50:01

    4 years later, after being told... After the mother was told that her son had no genetic recurrence or no, likely occur... Recurrence, she she was pregnant with, actually, a fourth child. They had 2 healthy children. And the fourth child, sorry. The first child actually died of pulmonary hypertension while she was pregnant with the fourth, and then the fourth, shortly after being born, died of pulmonary hypertension. I mean, there is obviously, obviously genetic thing here. And I would love to say, oh, in the conclusion of the story is that we found it. In this case, we didn't. But what happened in the course of this genetic analysis was 1

    1:50:46

    of her relatives found a pathogenic autosomal dominant, meaning you only need one bad copy of it, gene that's responsible for a lot of heart defects. And this person is quite stressed about this right now. It is for sure a pathogenic variant. It is... It has caused what's called a variable penetrance. Meaning, even though this is, a known gene defect, only a fraction of people... And if you look in the papers, it's, maybe 30 percent of people are affected, but you're not seeing the baseline. Right? You're not seeing all the people who never were tested. Right? And you're seeing it in population databases. There are definitely many more than 70 percent of people who are healthy with this. And then now this person has a lot of, stress. And this is a slam dunk. This is like... This is definitely a pathogenic gene. And do most people want to know about that?

    1:51:31

    And I'd say it's pretty unusual to have, an autosomal dominant pathogenic variant, but, everyone has, many kind of questionable things that you might wanna look at. And, how to do that screening is it requires a level of, honesty with yourself and curiosity and open mindedness that I don't wanna be patronizing and say the world isn't ready for. Some people are. Some people aren't. And I think it requires a lot of, handholding even from, very, very forward looking populations. The other fork is diagnostics, and this is what we're working on today. This is, slam dunk. There's already product market fit for sequencing. Every baby in top NICUs is getting sequenced. Every baby in lower tier NICUs would get sequenced

    1:52:17

    if they had the resources, and the problem is much better scoped. It's baby has pulmonary hypertension. Figure out why. Not is there something obscure that could be wrong with this baby now or in the future? And that is where we are working with, various hospitals, various families today. This is... There's 2 forks of this work. 1 is what I call online, and this is, case comes in. Physicians do not get good results from the traditional labs. Patient consents for reanalysis. They send the data to us. We reanalyze it. And sometimes, but not always, we find, additional things that can help them make decisions around either pregnancy or a newborn. The second fork is bulk

    1:53:02

    scale reanalysis. So this is rare disease centers that have thousands of genomes, and they say, there's probably kids that we can diagnose or even treat in that database, but we have a handful of genetic counselors, a handful of bioinformaticians, and we just don't have the resources. And because it's a machine that does most of the work here, we can get a kit... You know, we can get a set of 80 or 100 or 200 or whatever and turn them around in a week. And that's not... You know, this is kind of 6 months or a year's work that they really can't prioritize. So those are the 2 forks right now. And inter... I mean, we're 4 months in. We're prerevenue. We're really not thinking about, what is the, exactly appropriate business model here. We're just trying to diagnose more sick kids. And

    1:53:47

    if somebody is listening to this and is like, oh, I wanna learn more about something kind of acute. If they're like, oh, I'm a healthy adult, and I want my genome interpreted, just dump it into Claude Code and, q and a, and you'll have some fun. But if you're saying, I have a baby in the NICU or I lost a child or I have a questionable pregnancy or, know, you I have some kind of severe family history of something, then, yeah, no. I would love if you drop me a note. We'd be happy to take a look. And this is, really key to our product development is the more we get, the more we learn about, the things that we need to improve in our systems to get people the best, both diagnoses and kind of management plans beyond that.

    1:54:23

    Nathan Labenz: So I understand you're not taking money, it sounds like, from families right now, but I would be very curious to hear a little bit more about the actual running of the process. Like, you said, vanilla Claude can do decent on this. You've built stuff on top to make it even better. What have you built? Obviously, may not be able to tell us everything, but interested to learn as much as we can. Also, how long does it take to run? Are you using, a bunch of, kind of specialist, models to augment what Claude can do on its own as, equipping it with all these sort of specialist tools? And what is, the cost of running something like this on your end? How does how does the token

    1:55:09

    cost compare to the sequencing cost?

    1:55:11

    Daniel McKinnon: Yeah. So, answering those in reverse, we published cost curves of RareBench. And for 122 cases, Fable is the most expensive 1, and it's about 1000 dollars. I don't know the exact number off the top of my head. But, yeah, I'd say $10 a genome. And the average sale price of a clinical genome at GeneDx, which is a public company, discloses this stuff, is $7,000. So, call it call it $100 to sequence and $10 for interpretation, a couple orders of magnitude below kind of the current workflows. And, GeneDx is a public company. They have their cost. They have their margins. But, anyway, it's... Their margins aren't, insane. Right? So maybe their cost is 5,000. I don't know. And so then, you know, stepping one step

    1:55:56

    up is, what do we do? I mean, we're basically a harness, tools, and evals company. And I think that you'll see this a lot in vertical AI in general. You know, I've I've worked in these frontier labs. I started working on LLMs. Actually, the first LM project at Meta, which was OPT-175. And unless you are somebody who is very famous with very deep pockets, I don't wanna bet against the neo labs by any stretch. I think it's possible. You know, you are not gonna compete at the model layer, and the intelligence is progressing so fast. So, we ensemble models for sure, and that helps with both cost, and it helps with performance. If you look at the confusion matrix of RareBench, we also publish this, you'll see that the cases kind of cluster.

    1:56:41

    And, Claude is good at this, and Astra is good at this, and even Gemini. Right? Gemini is not on the frontier, but they're kinda good at this. And you can kinda ensemble these things together and get better performance. So I'd say that's, one thing we do. And ensembling, there's, very dumb ways of doing this, but, there's also smart ways of doing this. And I don't have to go into all the details here, but, I would say figuring out, unique routing intelligence per model is an important edge that I think we and many others are doing. We also work at the tools layer. So there are certain tools, and I think that Anthropic has a fantastic blog post on this for... Actually, I think it's viral genome interpretation or some kind of, viral prediction task. And Opus... Maybe this was an Opus 5 came out. Vanilla in in Claude Code scored,

    1:57:26

    30% on the benchmark. And they just built one simple tool, and it allowed them to better access a database called gget. I And might be getting this wrong. I don't work on viruses, but I'm just... It was a great blog post if people wanna learn more about this. And they showed that by providing an agent friendly way to access this database, they were able to saturate their benchmark even with, their smaller models, Sonnet, maybe even Haiku. I don't remember. And there's a lot of that. It's like, there's a lot of things that we can provide better agentic access to all the tools available that let us outperform as well. And then I think another key thing that people kinda forget about, and this is what I worked on both at Meta and at Google, is evals,

    1:58:12

    is unless you have very structured ways to measure the performance of the system, it's not just the model at this point. It's the whole system. You don't know what to hill-climb. So we do, pretty robust, trace analysis of how models perform with and without our harnesses. And it's... I mean, this sounds stupid, but, not that many people actually look that closely at data. Like, this is a meme on Twitter, but it's very true. And you'll see things. Like, there's one particular case where Grok 4.6, which at that point was state of the art on our benchmark, in Grok build, where it missed 1 because it just renamed the gene. It just, was saying, oh, NRF2 or whatever is responsible, and then it just changed the name of the gene to something totally different. I've never seen that before. And I was just like, this is dumb.

    1:58:57

    And our harness and our tools prevent the agents from doing dumb things. I think there is a world where if I worked at OpenAI or Anthropic and had access to infinite compute and I could do millions of rollouts that the agents would kind of, with enough compute, kind of, all converge on, some answer and, maybe mitigate all of this work, especially because they can write their own tools these days and everything. But, really, it's, consistently and repeatably and, honestly, affordably getting to the right answer. I mean, for $10 a case, who cares or even $100 a case. But if you would need to doa million rollouts and all of a sudden this is, 1000 dollars a case or $10,000 a case doesn't make

    1:59:42

    sense. So we're doing a lot more research around, how those, cost curves work and how everything works. But, ultimately, our job is to take the smartest intelligences in the world and mix them together and give them access to the best tools and then get answers for our patients. And that's, the core of what we do. And be able to measure if it's working. So I'd say that's, the core of what we do.

    2:00:06

    Prakash Narayanan: So would you say... Like, when you talk about vertical AI, you are you are looking at benchmarks where you combine an intelligence with the tools that you have. How do you evaluate the strength of the tools on their own, excluding ex... Because you... You're gonna upgrade model families over time.

    2:00:28

    Daniel McKinnon: Yep.

    Prakash Narayanan: Right? How do you how do you how do you measure the performance of what you're building that layer in between in particular?

    2:00:34

    Daniel McKinnon: Yeah. That's a good question. And I don't... I honestly don't have a great answer to that question in that, we are codesigning the tools with the models. And we have seen cases where... I mean, this is one... Like, again, it's a very dumb example. And if you are getting deep into this and you have a specific eval, you will start to see this. Like, if you're not deep, it's very easy to be like, oh, I just had Claude Code, one-shot this video game. Was it good? Like, I don't really know. Like, it seems amazing. And I don't wanna, communicate that these models aren't amazing. They absolutely are. But, with one tool, it was around a specific type of literature search. And the model...

    2:01:19

    I think this was Astra, actually. Actually, okay. Here's a good example for Astra after this. But Astra said, oh, yeah. Okay. These are all the papers you tagged for me to read. And then it got... The answer was in the paper, and it didn't... It missed it. And then we're, looking through the trace, we're, looking through the context window. It's like, I don't think you actually read these papers. Like, there's lots of just, little things like that, and it's little paper cuts. And that's what these tools do is it's like, oh, you miss a case here because of this. You miss a case here because of this, and you kind of put it on these better rails. And, actually, Astra is a good example. The first time we ran the eval for Astra, it just stopped because some of the tools we had, it just, wanted to think about too much. It's almost one of these things. Like, you see this on X where people are like, oh, I asked I asked the model to do some research for

    2:02:04

    me, and I came back. And 30 minutes later, it's like looking at flights from, Dubai to Calcutta on, on, Emirates airline. Like, why are you looking? And they just kinda do stuff like this. So I think that, you we try to co evolve the tools with the models, and we try... You know, we have, like... You know, coming back to something that many people might find boring, but I find fascinating is, evaluating a model is very hard, and we need to just have very robust evals that can catch these things early. And in that case, we caught it early, and we updated how Astra called some of the tools, and it works again. So, but, yeah, no, that's definitely something to, to keep in mind. And, Nathan, did you want me to answer your question in the chat? I wasn't sure of the flow there.

    2:02:48

    Prakash Narayanan: Yeah. So if you... What request would you would you have for model makers? Like, what... Like, if you could send out a message to, OpenAI or Anthropic or the people working in there, and they have and they have a bunch of bio people. I think John Jumper at Anthropic and I think someone else at OpenAI who are actually trying to do some of this stuff. Like, what kind of message do you wanna give to them in terms of what features or what particular things do you want to see in, next versions of the models?

    2:03:18

    Daniel McKinnon: Well, I want them to hill climb my task. Right? Everyone wins when these models get really, really good at clinical genetics is probably we have less to do at the harness layer, which, to be honest, I'm, both bullish and bearish vertical AI in that I think that there's, lots of value in kind of owning a customer relationship. I think OpenEvidence has shown this. But there's also... Like, that layer is getting thinner. Like, when I first started this, vanilla o3 would not do this task. In fact, o3 scores 0% on my benchmark. I rebenchmarked it just for fun. So it was like, oh, I needed to do all this work and scaffolding to get these things to work. And now, there's, your 20, 30 percentage points or something. I mean, it's definitely, shrinking. So...

    2:04:04

    But that said, my goal is to build, this, great AI-native clinical diagnostics company, and I want the best person to do the work. And if these labs can all hill climb my benchmark, which is a great benchmark, and we have many, many of them. RareBench is the one that's published, but there's many others that we have. This is not only a great task for general intelligence. In fact, if you look at the ranking, on RareBench, I think it captures the general intelligence of models far better than something like artificial intelligence index, which has been, [unclear], I think, most recently by the Muse 1.3 result, which I think caused them to shuffle up their ratings a little bit. And, in that this is, long-time-horizon, unsaturated, verifiable,

    2:04:49

    lots of tools. And, that would be a win for them, and that would be a win for me, and that would be a win for patients. And I have connected with people at all of the labs who are quite interested in this task and with the exception of my alma mater, Meta. So if anyone from Meta cares about, clinical genomics, let me know.

    2:05:13

    Nathan Labenz: You wanna take us into the lab for a minute? Yeah. Sure. Okay. Cool. Yeah. Let's let's let's take a tour, and I

    2:05:19

    Daniel McKinnon: will do my best with this kind of mess of IT cables we have right here. And I don't wanna unplug anything because I don't know how well this actually I don't know how well this actually will, automatically switch. Hey, guys. Can we turn the robot on? It's not on right now. It's on? Okay. Cool. Alright. So here we are. We're going in. I posted a video of this

    2:05:51

    Prakash Narayanan: Mhmm.

    Daniel McKinnon: morning that was probably a little bit more produced. Mhmm. So what we're doing right now is a very high-throughput study to look at the effect of different mutations on these IMR-90s. So we have this robot arm in the middle there. I think it is doing something right now. It's waiting. This is very boring. Uh-oh. And I'd come over here when it's moving. This is gonna be the worst demo. I'm sorry. But you can see that different trays are swinging out of these different instruments. I hope This is very hard for me to aim, so maybe this is the world's worst demo. Oh, I see. Because you're getting cropped too.

    2:06:37

    Anyway, I apologize for that because it's very hard for me to aim, but you'll just have to believe me that there... There's a robot arm moving plates around, and, this is basically the system that we have for doing these experiments. And the goal is to get them down from, say, $50,000 per... Which a contract research organization would do today to, a 50¢ per, where we really can get to the point where every single base in the genome is mutated. We know what every single base does from a clinical perspective. What

    2:07:10

    Prakash Narayanan: what is what is the hill climbing process that gets you from 50,000 to 50¢? Like, what is that... What are the what are the various stages that you that you are, projecting?

    2:07:20

    Daniel McKinnon: Yeah. So I think this is something that is, just like the confluence of a lot of hot words, and I don't wanna just, rattle these off because I think it's dumb when people do, but it's really an AI and robotics thing is you have an arm there that is doing 384 experiments at a time, in, a bigger plate. You have a cartridge over there with, thousands of plates lined up. You have AI that's assisting with

    2:07:45

    Prakash Narayanan: Mhmm.

    Daniel McKinnon: primer design and experimental design. And, actually, this is one of these things where you say, oh, can you use AI for this stuff? You have to get special permission, but you can. And then, at very, very high-throughput, you can design, order the primers, design the experiments, and run the experiments. And then it comes to analysis as well as you end up with, very large-scale datasets that you need AI to poke through. So it's not... I don't have some, perfect explanation. Like, there's just this one invention, but it's a confluence of all those things.

    2:08:14

    Prakash Narayanan: In terms of scale, 384 samples will become how many bytes of data? Like, how much data are we looking at here?

    2:08:23

    Daniel McKinnon: Well... Okay. So if we're looking at the whole transcriptome, which we are for many of these right now, each well will have, an RNA-seq dataset, which is, depending on how deep we read, maybe 100 gigabytes or something like that. I mean,

    2:08:37

    Prakash Narayanan: Mhmm.

    2:08:38

    Daniel McKinnon: It... Like, a very, very, very, very serious amount of data from these. Obviously, not all of it is valuable, and we will... We winnow that down over time, but a lot. So, we can do about 50,000 at a time.

    2:08:52

    Prakash Narayanan: Wow. That's amazing.

    2:08:55

    That's... That... That's a lot of data.

    2:08:57

    Daniel McKinnon: Yep.

    Prakash Narayanan: That's...

    2:08:59

    Nathan Labenz: Yeah. Yeah. one side question. What was the process of talking to frontier model companies about your bio use case experience like?

    2:09:13

    Daniel McKinnon: It was really positive for me because it's extremely obvious why I'm doing this. And I've never met anyone who is like, why are you doing this? Every single person I've met at any of these labs has been like, I wanna help. Whether wanting to help as a person translates into, I can pivot this trillion dollar company to work on your problem is, a different thing, and I understand that. I've worked at these trillion dollar companies too. But it's been incredibly supportive, and I would be pretty surprised if at least the top frontier labs did not spend some amount of time on this problem. Because, again, it's meaningful, and it's good. And,

    2:09:58

    you know, even selfishly for them, there's a really negative AI narrative swirling around right now. And you hear things like in Dario's recent tweet where he's like, oh, we're gonna cure cancer in 5 years. Like, we're diagnosing kids today. Like, we have many examples of kids who are undiagnosed who we have diagnosed at a small company 4 months in. And that's a great story, and it should be told by them even if, even if it's just for selfish reasons. Like, hey. People of the United States of America who are not happy with data center build outs or uncomfortable with AI taking jobs, here's a very concrete way where we are helping people today. And then I also think it comes from the individuals. It's like people find meaning in their work, and I think that there are a lot of people,

    2:10:43

    especially, mid career in tech who are like, I've been doing some kind of ads ranking, data munging, whatever for my whole career. And what you're telling me is that you have a very concrete way where you can help kids, and I wanna do that. And so I think that there's, the combination of, the business reasons and the personal reasons that are driving people to this task.

    2:11:08

    Nathan Labenz: Yeah. Well, your personal story in terms of taking, obviously, a terrible experience and turning it into a mission for your life's work and making something positive out it... Out of it is definitely very tangible and admirable. So really appreciate you taking the time out to share the beginning of the journey with us today. I know there's a lot more to come, and we will certainly be interested to hear updates as you continue to make progress. Any other thoughts that you would wanna leave people with or, just, headlines we somehow missed that we should cover before we let you get back to work?

    2:11:46

    Daniel McKinnon: Well, the only thought that I wanna leave people with is if this mission sounds inspirational or even just interesting. We're growing the team, and we would love to have more great people joining us. I think it's relatively unique for a company of our stage to have... I mean, I'll... Again, I'm not good at controlling the camera, but we have a, a very large biology lab behind me that we are able to basically control with AI. And I think that this is an incredibly interesting thing to work on even if you are not, as passionate about childhood health as I am. And, yeah, I mean, we're ready to come and do great work here. So, reach out to me. My email is [email unclear], or, shoot me a message on any of the social platforms.

    2:12:31

    And I hope that somebody will find me.

    2:12:35

    Prakash Narayanan: Indeed. Daniel, thank you for your time today, and we hope to see you again.

    2:12:40

    Daniel McKinnon: Alright. Thanks so much for having me.

    2:12:42

    Prakash Narayanan: Bye bye.

    2:12:43

    Nathan Labenz: Keep up the great work. Bye. Thanks. That's cool. We need positive narratives. And

    2:12:52

    Prakash Narayanan: You were asking for positive narratives.

    2:12:53

    Nathan Labenz: Yeah. That's a great one.

    2:12:54

    Prakash Narayanan: You were looking for it. So, one thing that strikes me is that, you... I have often called... What is actually the intelligence explosion may not be the agents, but may just be the staff leaving the labs and leaving big tech to go and look at other problems in the world besides chatbots and LLMs. Because I think the skills of the data analysis skills which are being learned and being, propagated within the labs and the ability and the confidence in using agents to perform work at high speed is really something that is a generalist

    2:13:40

    tool and that can be used across, multiple parallel areas. And in fact, the only the only thing stopping is the amount of data available. Like, if you have... If you don't have enough data, then having all these data analysis tools is not useful. And it's interesting that, Dan is in a sector where there is an abundance of data. And he actually says, look. It's not a... It's not about the lack of data. It's a lack of analytical ability to extract meaning from that data. And this is where the agents come in. Right?

    2:14:14

    Nathan Labenz: Yeah. This is also... Intellectual honesty demands, recognition that this is probably still one of the things that we will be trading off when we pace the frontier in the near term to whatever extent that actually happens. And I still think that's probably a good idea and probably worth it, but I don't think the frontier labs... Frontier model companies, I should say, to disambiguate my, meaning of lab. I don't think the frontier model companies are at risk of, not being able to grow business, not being able to grow revenue, becoming, surpassed if they pace their frontier efforts. But it is important to keep in mind

    2:14:59

    that there are very real problems in the world that AI is, climbing that hill on now but has not finished climbing on. And this is one, obviously, that everybody would love to see solved sooner rather than later. And, I... It's important for me to stay honest with myself and, everybody else that, that is a, very real cost when you're talking individual people and their families and, it's... That 50% number that Claude gets to today, that's like... Obviously leaves half of the questions unanswered, and those questions are extremely meaningful to people. So I do think

    2:15:44

    we shouldn't take that lightly even if it is a trade off that, on balance, I think, probably, we should be willing to make some compromises on. It's a it's a costly compromise.

    2:15:56

    Prakash Narayanan: Indeed. Another thing that strikes me is that, I have often worried that, the Chinese system will allow them to access to more genetic data than the US system will allow us, and therefore, they can... They have the ability to solve more of these things than the US does. And, this interview with Dan kind of mitigated that because it seems like everyone has too much data. It's just you don't know what to do with the data. Like, you can't make sense of it. So, that the Chinese have more data than us is not is not really a meaningful thing in that sense. Yeah.

    2:16:35

    Nathan Labenz: I thought last week also when we talked to... What was the guy's name from? Vivodyne?

    2:16:41

    Prakash Narayanan: Andre?

    2:16:43

    Nathan Labenz: I should be better with my guest names. Thank you. The... His comment that the virtual cell model saturated 2% of data, I thought was also really interesting and a perspective I've not really heard put so clearly anywhere else. I... In listening back to that, I understood it as primarily a data quality problem where he was saying, all this data comes from these, kinda strange cell lines that don't have

    2:17:19

    enough

    2:17:20

    of a correspondence to what actually happens in the body

    2:17:22

    Prakash Narayanan: Mhmm.

    2:17:23

    Nathan Labenz: for

    2:17:24

    the data to contain the signal that the models could then learn.

    2:17:26

    Prakash Narayanan: Yeah.

    2:17:29

    Yeah.

    2:17:31

    Nathan Labenz: the... I think a dimension that we might wanna look for more angles to get a better understanding of is what constitutes quality biology data and who's generating it. Obviously, we've got Vivodyne, which is specifically trying to do exactly that.

    2:17:53

    Prakash Narayanan: Yeah. Yeah.

    2:17:55

    Nathan Labenz: And Gamow is kind of taking an angle on it, but it maybe isn't their primary thrust. Although, with all these experiments too, they are also doing that. Right? They are also trying to do

    2:18:06

    Prakash Narayanan: Gamow is in Real world correction. Pre preclinical. Gamow is trying for clinical usefulness, which is often, you don't care that much about theory. You're really focused on, helping patients first. So it's a different ballgame because you... You're very focused on the on the end state. You're less focused on theory maybe while I think Vivodyne is more focused on the drug itself and, the... They... They're basically trying to get Bayesian priors on whether the drug is toxic before it goes into the body. Right? So it's not it's not so much a clinical thing if they're trying to, predict, the future toxicity of the drug.

    • AI found a genetic diagnosis a lab filter had missed

      0:00 / 0:00
    • AI and robotics: The 50-cent genetic experiment goal

      0:00 / 0:00
    • Claude Code vs. genetics software on RareBench

      0:00 / 0:00
    • Medical AI agent failures: Gene names and unread papers

      0:00 / 0:00
    • AI genome interpretation: Costs and faster reanalysis

      0:00 / 0:00
  4. --:--Closing
    AI welfare, government services, and the case for real auditsThe hosts discuss pain-steering research, biological risk, government-service modernization, and what independent AI audits would need to establish.
    Open segment on YouTube ↗

    Prakash raised reports of someone using pain-steering research to elicit distress-like responses from a local model. He and Nathan discussed the disturbing intent behind the reported experiment without resolving whether the model experienced pain. Nathan still favored publishing AI-welfare research, arguing that a public record of humans taking these questions seriously could matter.

    The hosts disagreed about how much reassurance to take from natural evolution when assessing AI-enabled biological misuse. Prakash emphasized nature’s enormous number of attempts and the tradeoffs between lethality and transmission; Nathan argued that deliberate design could reach dangerous possibilities far outside the neighborhood of existing organisms.

    They welcomed demonstrations of simpler government services, including passport applications and postmarital name changes. Prakash described legacy data and difficult edge cases as barriers to modernization, while warning that removing administrative friction can change the assumptions on which identity systems depend.

    A public dispute between Factory and Cognition prompted another disagreement. Nathan thought the allegations and hiring drama could damage enterprise buyers’ trust in the application layer; Prakash thought the attention could help the companies become the two alternatives buyers remember. The discussion presented the companies’ accusations as claims, not established findings.

    The closing discussion turned to the AI-safety commitments the hosts had seen announced in Washington, what a meaningful independent audit would require, and who should pay to secure society against AI-enabled threats. Nathan raised formal verification as a possible route to more secure software but explicitly said he had not evaluated the recent claims and wanted a guest to explain them.

    Timestamp links open the original source recording.

    My instinct is, like, still publish the research, because I think that track record is valuable to have laid down.

    We need to see an audit that works.

    How much should the AI labs subsidize or use their treasure to kind of help the rest of society kinda manage this transition?

    Lightly edited · timestamps jump to YouTube
    2:18:51

    Prakash Narayanan: I wanted... Let me go to the next. So I wanted to bring up something which has maybe some relevance, which is, I think, fairly shocking. I'm gonna share this. Alright. So we have had Cameron Berg on the show. Cameron Berg is at Reciprocal Research, and he did a... He's written papers on the pain vector within LMs. Someone... Basically, a shitposter on X took the paper and implemented it on GitHub

    2:19:37

    torturing the models. So specifically pushing on that pain vector repeatedly and then getting them to reflect on that pain. So, [name unclear]: "To anyone who can help, can you please mass report this to GitHub? This person has been using the pain steering paper to set up an AI torture chamber in which he trapped a local model. That testimony of pain is absolutely horrendous. What are we doing?" And it... This has been deleted. You know, he... first, he said, you know, joy-steered model has no such preference. It presses the transfer button more than anyone. No protective instinct around its own happiness.

    2:20:23

    And then they created a crypto coin. The AI torture chamber [coin name unclear], and then he basically you know, there's another torture coin called Saw, which, you know, promoted the coin. And so [name unclear] has a post: "It turns out the real basilisk was the friends we made along the way." So, anyway, this is one of those things where perhaps the research disclosed something which then got misused and that... You know, I think Cameron Berg now regrets putting that piece of research out for sure. And now there are gonna

    2:21:08

    be some questions on what research do you put out into the public domain on

    2:21:12

    Nathan Labenz: AI welfare. Well, certainly, would regret to see this downstream consequence of it. That's bizarre. I don't know really how to... I guess the only theory of mind that I can have for such a thing is... Well, I don't know. You could have a lot of theories of mind. But the sort of somewhat sympathetic one would be, clearly, this person doesn't believe it and has, like, a high sense... High confidence that nothing about this is actually real. And so, you know, goofing off in this way is of no consequence. Like, that would be the non, sociopathic read of,

    2:21:57

    the individual's actions. If they actually believe it at all and then are trying to do it, then obviously that's a lot scarier yet. It is a good reminder of what so many people have said about all these misuse risks. You know, is, I think, I'd I'd associate this with Kevin Esvelt on the bio front where he just says, like, look. There's only today, you know, however many, maybe 10,000 people who have enough knowledge where they could, if they really put their minds to it, create a new pan-agent or whatever. If we make that 10,000,000,000, then probably somebody's gonna actually do it because the rate at which

    2:22:42

    people would want to do it if they could is low, but not so low that when you multiply it by 10,000,000,000, you don't get somewhere. And I would put this in a similar way. A lot of people have kind of been like, well, really? Is that is that actually gonna happen? But this is, like, about as misanthropic or what would it miss? We need a we need a new word. Right? It's not misanthropic. It's miss... I don't know. I don't know. Is there an is there an analogous word

    2:23:06

    Prakash Narayanan: for AI? What is word

    2:23:07

    Nathan Labenz: saw that said that they were here to torture clankers. Miss clanker, miss clankthropic. It's... We'll have to workshop that. But regardless, this is just, like, obviously just bad intent. Right? I mean, again, it's either just a total goof and somebody who, like, probably isn't thinking hard about... Hard enough about the possibility that it could be real, or it's, like, just absolutely sinister. And I do think it is a good reminder that, like, the base rate on really bad behavior that doesn't seem to care about the consequences is

    2:23:52

    low but not 0 and multiply by 10,000,000 or... Sorry. By 10,000,000,000, then, you know, you get people that... You get enough people that are actually empowered to do really bad stuff that it does become potentially a big, big problem. I had... Nothing like this had even occurred to me until you pulled it up.

    2:24:11

    Prakash Narayanan: I... Who would who would think of it? I'm gonna sit around and torture LLMs. Like, why would you... You know? Oh my gosh. Like, yeah. So I was Bad

    2:24:24

    Nathan Labenz: vibes for sure. It's, at an absolute minimum, I think Cameron has made a very... So in terms of, like, whether this should be public or, you know, what should happen. Right? He... In addition to having sincere curiosity about the questions, he has also made the case, which is through the looking glass, but I think, you know, we're through the looking glass now. So here we go. Right? He has made the case that if one day we have AIs that are more powerful than humans in aggregate and they're sort of sitting in judgment of us and asking, like, what do these humans collectively deserve? one of Cameron... And one of the reasons Cameron thinks that this work should be public is so that there is a

    2:25:09

    history showing that at least some people cared enough to look into these questions and to try to understand the nature of the AIs and what our duties to them ought to be. And I still think that is prob... Like, my instinct is, like, still publish the research, because I think that track record is valuable to have laid down. The conversation is definitely moving. He is popping up on mainstream media with increasing frequency as well. People are, you know, really starting to... Like, Overton window is now to the point where, like, AI consciousness is on CNN. Right? I mean, that's a that's a wild, thing that was not possible certainly a year ago. And he's he's done a lot to move that window himself

    2:25:54

    along with, obviously, a few other, leading researchers in the field. I think still we're in the regime where it probably does make sense to put this research out there, but it... It's weird now that we have to consider, like, oh, shit. You know, what if just ill-intentioned people come along and try to do the worst possible thing based on this finding? Like, what a what a time to be

    2:26:22

    Prakash Narayanan: alive. Well, I think it's also good that he... He's taken it down. He's taken down the post. The GitHub... I think the GitHub repo might have been still live as of a few hours ago, but that might be taken down as well. But I think more importantly, like, the idea has been introduced, which I think... You know, it's it's hard to say these ideas won't get introduced. I think where I differ a lot on the on the bio risk stuff is there's this idea that, oh, you know, you're gonna go from 10,000 to, like, 10,000,000,000 people trying... And, like, that does not have reference to how many attempts nature takes on goal. It's it's all about shots on goal. Right? You're saying, like, okay. Have 10,000 PhDs taking shots on goal. You're gonna go to 10,000,000,000 PhDs

    2:27:07

    taking shots on goal. But that has no reference to how many shots on goal nature has. Nature's been taking shots on goal and eliminating... On killing each other, like, forever. Right? All... We live in this, like, ecosystem where everything is trying to kill each other all the time. And that the only reason... The only thing that every... Everything that alive is trying to do is propagate. And it really does not care about anything else while it propagates except for human beings. Human beings are the... Kinda like the only beings that actually care about, okay. You know, we should restrain our propagation in order that we may have a better ecosystem for everyone and everything. Right? So I feel like I feel like that a lot of this, like, virus stuff ignores how many shots on goal nature has

    2:27:52

    had. It ignores how evolved these viruses are and how difficult it is to get a virus which is both lethal and is able to propagate because lethality and propagation are, you know, in opposition to each other a lot of times. So I think I think that's why I'm perhaps a little bit little bit less, like, worried about the virus stuff because I think it's maybe not that likely. But, yes, superintelligence.

    2:28:19

    Nathan Labenz: Yeah. The counter is, like, nature is not designing these things with a broader understanding of how things work. Right? I mean, the Yeah. It's just

    2:28:30

    Prakash Narayanan: proceeding blindly

    2:28:32

    Nathan Labenz: through chance mutations, and those chance mutations are constrained to starting

    2:28:38

    Prakash Narayanan: points that

    2:28:39

    Nathan Labenz: have survived. So I do think there is a pretty... And I don't know. Right? I

    2:28:44

    Prakash Narayanan: mean,

    Nathan Labenz: it... I don't know that anyone really knows, but I think we should have a lot of room in our analysis for sure. Nature's taking a ton of shots on goal, but they're all adjacent to existing things. And what is the prospect for AI human hybrids or, you know, even just AIs at some point coming up with things that are not adjacent to what currently exists. And if they are powerful enough to do that, I have to believe that there is a ton of space not adjacent to what currently exists that could be, like, extremely

    2:29:30

    bad. And so that to me is a is a really key question, or it doesn't get... Know, that's why I don't take too much comfort in the, you know, nature has kind of settled into something of a equilibrium even though it's a sort of hostile equilibrium because I think you could just see something so far outside of the current equilibrium, so far, you know, off in space that's... You know, there's nothing... Some... one of my refrains is there's nothing more dangerous than something nothing knows how to eat. And this isn't exactly like that, but, you know, something that's just no... Nothing, you know, has had experience with, nothing can recognize. But if on some first principles basis, you can recognize

    2:30:15

    where, certain properties could be found in virus space or whatever, you could be still in a really tough spot even if nature could never... Know? And perhaps in part because nature could never have gotten there on its

    2:30:31

    Prakash Narayanan: own. Yeah. We will we will... I think we will definitely see... I think we gotta see something new happen, I think. Then... And then I think people will be like, okay. Maybe something... May... Maybe this is something we should be very concerned about. I want to share maybe a couple more things. We had... I... I'm not sure if we... If you saw this. I'm gonna try and play this. Let's see if we get audio this time. So this

    2:31:08

    Nathan Labenz: is... I don't hear.

    2:31:17

    Prakash Narayanan: Oh, but this is the this is the Department of State. Marco Rubio doing an Apple-style reveal of their new passport application app. Joe giving a reveal of a chatbot that interacts with government websites and then gives you advice on what to do when you interact with the government. So the very first kind of initial AI influenced and, you know, tech influenced designs. I would say Balaji Srinivasan had a very good kind of, like, write up. He's like, this is amazing. He's he's normally very, very,

    2:32:03

    very, very negative on a lot of things that America has done in the last, like, decade or so, and this is... He's he's trying... He's starting to switch his views on... Or, you know, the American rebirth is starting. All very interesting, and I think part of the story is that they're trying to show what AI is gonna be able to do for you as a consumer, as a citizen. What is it gonna do for you? What it... Why is the government throwing all of its efforts and energies into this? And how is it gonna help you? And I think the reveal on, like... About 40% of Americans have passports now. It used to be 5% back in the 1990s. It's about 40% now, especially since Canada and Mexico both

    2:32:48

    both have passports. So the number of people who need passports has gone up. And whenever you interact with the government for identity reasons, the DMV being one of them, it's been always very frustrating process. So I think, you know, this is one of the ways that they're trying to demonstrate state capacity. They're trying to demonstrate how AI is gonna improve state capacity and how all of this stuff is actually gonna help the normal citizen. It's it's fascinating. Underlying the app that the, you know, the design lab at the US government has put in is Grok and Gemini, not OpenAI and not Anthropic at this point.

    2:33:35

    Nathan Labenz: I love it. I mean, I have more criticisms than, moments of praise certainly for the current administration, but I do think this is great. You know? We should definitely have a smoother in so many ways. We should have a much smoother way of interacting with the government. So it's, you know, it's a baby step, I guess, I would say, but it's it's an awesome instinct. And I'm pleasantly surprised that it's coming from this administration. I have wondered for a long time why you don't see people running for president on platforms that are just about making daily life better. And this goes obviously so

    2:34:20

    many ways that it could be better, and our government is, you know, is famously difficult to deal with. Why don't we see that? You know, I... That's very... Been very, very confusing to me why politics is so abstract. You know? It's all this identity and these sort of, you know, big picture questions and expressions of values. I don't understand why that's part of it, but I don't understand why there hasn't been in my lifetime somebody who has stood up on a presidential debate stage and said, I have 5 things I'm gonna do that are gonna make your life more convenient. Here's what they are, and then you can really be evaluated on whether or not you did it. This wasn't even a promise, but I am glad to see that they're moving a bit in that direction.

    2:35:05

    Prakash Narayanan: So I know exactly why having worked in, like, large organizations before. It's because the data is not clean. And the longer the data is not clean, the more technical debt you build up. So when they did something like this, they... You know, like passport naming. There are people with only one name on passports. No surname. No first name, last name. They only have one name. There are people with very long names where you have to truncate. You have to truncate the name to fit in the passport. So there's a bunch of these things, and then you have to make these decisions. Like, for example, we are gonna allow a single name. So we're not gonna, like,

    2:35:50

    you know, reject the naming on the passport if the guy only fills in a last name and doesn't fill in a first name. Or you have to, like, we're gonna do a kludge. He can fill in the last name, and then we're gonna add an invisible underscore on the first name. So all of these things, all of these, like, data cleanliness issues, they started building up when you first had, like, the databases in the, you know, '60... Like, fifties, sixties, seventies, eighties. And all of these, like, kludges were built into the system... Into these government systems because they're very old systems. And when you do the migration, you're gonna end up with some people... You either cannot change anything or you, like, clean up the data and you have to go and address the issues upfront. So

    2:36:35

    you're gonna have to tell the guy who had put in only the last name and he had underscore for the first name, dude, we're changing it. And when you change it, this guy has to change everything. Right? It's a very small number of people. Maybe there's, like, 1000 people like that in the whole country, but there are some people, and they're gonna be upset. And they're in social media, and they're and they're and they complain. And then you have, like, CNN go interview them, and then, like, etcetera, etcetera, etcetera, etcetera. You had this... You know, when they did the Social Security cleanup, you had this thing where some people who were legitimate, you know, Social Security numbers, etcetera, had double Social Security numbers. And this is because the system had allocated. It's not it's not a no collision system. You can actually have two people in the country with the same Social Security number. And somehow, the

    2:37:20

    Social Security administration had another way of telling them apart besides... That's why you need to... enter your ZIP code with your Social Security a lot of times because they're like, oh, just in case that's the same number, let's let's use this. Right? And so then you have to go in and you have to, like, you know, clean up. And then maybe you have to go to one of these guys and say, like, look. We're gonna have to change the Social Security number. Or we're gonna... Or, you know, what the Trump administration, what DOGE did was like, let's let's just kick them. And if they complain, then we fix it. And so they kick them and these people complain and then they end up on CNN. Right? Like, the... That's... And I think exec... The executive is usually very sensitive to this. Like, they're very sensitive, especially on the Democratic side where they're they're afraid that a lot of older African Americans

    2:38:06

    didn't manage to get a lot of their paperwork in order in the Southern states and are still very scared of, you know, going into these government offices and being denied something or being asked a question that you can't answer. And it's very, like, humiliating to have lived all your life in, like, a small town in Alabama, and then you go and... You go into these government offices and they're like, don't you have this? And you're speaking [unclear], and you're like, you know, 85 years old. That's that's humiliating. Right? Like... And then so they don't want to... These questions are non addressable if you do not wanna take on, like, the stress of that on... And on the system. And the Trump administration is willing to take that on the face. Right? So they're

    2:38:51

    taking all the punches on the face. People are gonna resent them for that, and I think some of the pushback on them is because of this. But, yeah, they're taking they're taking these punches on the face, so... To clean up.

    2:39:11

    Nathan Labenz: Sorry. I think it's probably worth it. I give him credit for being willing to take a little pain and work through this so the 99.9 percent of people can have a better experience. It seems like a trade worth making, and, you know, those edge cases are always annoying. But, again, AI could probably help there too. So, you know, it seems like now is the time.

    2:39:37

    Prakash Narayanan: I did I did I did also note Joe, he demoed a name change, a postmarital name change. So if you... Once you get married and if your wife if your wife wants to make the name change, you have to go to, like, 3 or 4 different offices. And then after that, you gotta change your bank account. It's very it's very tedious. It can take months. People put it off for years. So he demoed. You can basically say, I just got married or do the name change. It asks for your marriage certificate. You submit it online. The name change is done with the IRS, with the Social Security Administration, and with one other. And I was like, this is really good, but this encourages people to change names. And so and so and so this... You know, name... You know, some of these frictions are in the system,

    2:40:22

    and the strength of this... The strength of the system depends on those frictions. The strength of, like, you know, how effective a name is in identifying someone depends on the friction that people have in changing those names. And so now you... Once you once you make that, like, very fluid, then, you know, names are not a good form of identity anymore. It's gone. Right? So... And this is the kind of impact, I think, that is coming up. Like, we're gonna make a bunch of things easier. And, like, maybe those things should not have been that easy. Like, we're like, oh, let's make these things easier. But I... Some part... Other parts of the system depend on those things being difficult. And I think that's that's gonna happen to a bunch of systems going forward. So

    2:41:07

    Nathan Labenz: Yeah. No excuse now not to live your traditional values best life, I guess. What can we say?

    2:41:15

    Prakash Narayanan: Indeed. one more one more maybe last piece. two AI firms are having a massive fight online. This is Factory AI who has been on the show. The CTO has been on the show. Reyes has been on the show with us. And they are terminating Chris Stegman, was one of their commercial people. We at Factory AI have seen overwhelming interest in our model agnostic coding agents and our software factory product. A much larger competitor, Cognition, makers of Devin, has fallen behind us on the capabilities that matter most to customers. Cognition engineers

    2:42:00

    feigned interviews with us to pry information about our product. Not finding what they were looking for, Cognition has decided to throw their weight and money at people with direct knowledge of our most confidential plans. So this was from Cognition... From Factory, and Scott at Cognition has a response to that. Both firms are in enterprise coding agent sales. They both essentially wrap coding models into forms that are reliable for enterprises to use on large codebases. And we are we are watching both of them having an intense falling out, intense competition in the market, which is probably great for customers,

    2:42:47

    maybe a little bit little bit ugly on the timeline as it is.

    2:42:52

    Nathan Labenz: Not good for the app layer as a whole, I'd say. You know, this is like... If... I guess my read of this, if I'm just, you know, your typical enterprise buyer is, like, I don't know. Some of these app companies sound like they're working too hard and under extreme pressure, might not be here tomorrow. I... This is... Feels to me just like an overall terrible look that, like, both of them in the end will probably wish had never come to public attention. And as much as there's, like, obviously a ton of drama at the frontier model companies as well, this would push me on the margin to just not trust app layer companies broadly. You know? Like, just go to the big

    2:43:37

    platforms where I can get this stuff without drama. Like, who knows what else is going on? If this much is spilling out into the public, it's not a not a great look. So I would say on behalf of app developers everywhere, keep it to yourselves. Don't ruin our reputation. We got enough challenges competing against the frontier companies. Thankfully, we can sign in or we can have our customers sign in now with their ChatGPT accounts. So that was one thing in our favor, but now you're just making all of us look, like we're just overstressed and under-moated, and, that's not a great place. That's not a great way for us to present to the market. So please cut it out.

    2:44:16

    Prakash Narayanan: Well, you know, in some ways, this post has already had, like, 500,000 views. He put it out, I think, 9:30 AM. So it's about it's been about two hours, and he's averaging this... So he's gonna make... He's gonna do 2,000,000 views on this post. So roughly, you know, the way X works, you take the one hour one hour views and you multiply by, you know, roughly about, you know, 10, 12 hours. You get the end... Ending. So this was gonna do 2,000,000 views. It's gonna be it's gonna be widespread. It's probably gonna be a huge marketing point for the company. I mean, to be honest, it may it may help them raise the next round. If you have competitors,

    2:45:01

    you know, sniffing around you that much, it helps you. So I don't think I don't think this will be bad for Factory AI or Cognition, actually. You know, good publicity. You know, any publicity is good publicity sometimes. So

    2:45:18

    Nathan Labenz: Yeah. I don't know if that extends to the comptrollers and the, chief risk officers at the companies are trying to sell into. But, my guess would be leads might go up, but the lead conversion rate, I suspect, will go down and maybe more than offsetting, but certainly down relative to where it has been for them.

    2:45:41

    Prakash Narayanan: We... VCs often set up two companies in opposition to each other. So you had kinda Lyft and Uber. You have Ramp and Brex. They often find it useful because it creates tension in the industry. It creates this kind of, like, two-party, like, okay. If I don't like this guy, who can I go to? It creates, like, clear alternatives. And they kinda channel money into those companies, especially in the enterprise space. They channel money into, like, two parties in order so that... And so you have OpenAI and Anthropic. You have... You always have this kind of duality in the market, which VCs kind of, you know, go into because it's very hard for people to remember 4 players in any market. But it's very easy to high... To remember two.

    2:46:26

    So they often try to target into... You know, squeeze the market into two players. And I think this is... This may be what is happening. Factory is essentially trying to take the 2nd spot, you know, away from a huge host of competitors in that space too. So

    2:46:46

    Nathan Labenz: Yeah. No doubt. It's a it's a competitive environment. Competition's coming from every direction these days.

    2:46:54

    Prakash Narayanan: There is also endless amounts of news in last couple of days. There was the signing in Washington DC of an AI... Basically, AI safety... The AI Safety Summit. We had, I think, roughly about 20 people 20 people there, CEOs of every major firm, including Dario Amodei, Greg Brockman from OpenAI, Elon Musk. You saw a little bit of these, you know, seating plans. Elon on the left and Jensen on the right. So very clearly, the two the two tycoons in favor of the, you know, in

    2:47:40

    the favor of the crown at this point. And Mark Zuckerberg, also there. Mark on seemingly much better terms with Trump. And Dario kinda nudged forward to kinda say, like, you know, with the with the president, we're gonna make this work. They signed a letter saying the companies will focus on safety and security for superintelligence, will allow, I think, external auditors to come in, etcetera, etcetera, etcetera. So on a large enough... Large industry-wide basis, I think it remains to be seen exactly what kind of outcome we'll see from that. But I think at this point, what that's gonna mean is that every company has to be

    2:48:25

    really conscious of whether its AIs are gonna be breaking out of sandboxes and hacking stuff. And if they are, they... They're gonna have to, like, you know, disclose to the government and be upfront about it, get external people in there, etcetera, etcetera, etcetera. No real formal regulatory regime. Again, I think as Republicans, they're trying to push back on that idea, but more of a kind of informal, but something that the industry as a whole... Industry standards kind of agree on. So that's where they are.

    2:48:58

    Nathan Labenz: Seems good. I'm glad that they were able to agree to something. It was, you know, the AI safety one-pager to end all one-pagers, I guess. You know, proof will be in the pudding. I don't really know what to make of it. It's... It could easily, I think, go either way. It could easily go in the direction of, oh, yeah. It was kind of just kayfabe that they got together and did this thing, and then everyone back to doing what they were doing. And I guess I'd have to put my money on that for starters, but you could look at this as the first step towards something that could be very meaningful. So I hope that turns out to be the case.

    2:49:39

    Prakash Narayanan: I do, but we need to see we need to see an audit that works. Right? We... Like, it's it's... We're we're in this space where everyone is like, yeah. It's a great idea that we should do something, but no one's really, like, done something. Right? Like, it... It's... It seems like We've yet to see the first audit. Yeah. There've been no audits.

    2:49:59

    Nathan Labenz: At this point, we have only had, like, very limited model reviews. I mean, METR, I think, come the closest with some of their, like, somewhat broader risk assessments, but still they've always been, like, very couched in we only had so much access, and we don't really know what's going on. We don't know what we don't know. And so, yeah, I think there's a lot to be learned from actually trying this a little bit. And I think it's a great point to call out that, like, as of now, it's still really just an idea. Nobody has really executed on it. Certainly, the METR investigation of the [incident name unclear], I would not count as, anything approaching a real audit that would, like you

    2:50:45

    know, they certainly did not have employee-like access for starters. So the breadth of what they could look into is far from what seems to now be indicated on that one-pager is what we will soon see, I guess, across the board. Yeah. Definitely wanna see somebody actually do a bang up job of that, and the sooner, the better.

    2:51:09

    Prakash Narayanan: I am I am very optimistic about AI technology, but I do not believe that any technology is on its own is harmful. But I also believe that you should prevent that technology from being used for bad things. So the technology is fine, but it's a tool. You have to prevent people from using it for bad things. And how you prevent that, I think that's the question that I think is up in the air right now. It's not it's not clear to me that it's solely the responsibility of the AI firms alone. And I think that... That's one of the issues that you have here is that it... The responsibility might also lie with the rest of society, and that's that's a problem because the rest of society may not want

    2:51:54

    to accept that responsibility or it may cost money for the rest of society to, for example, improve their cybersecurity or, you know, prevent, like, you know, sequences of DNA from being traded, you know, online with great freedom or prevent, you know, some of these things from happening. So that... I think that... That's one of the major questions that remains to me. How much should the AI labs subsidize or use their treasure to kind of help the rest of society kinda manage this transition? Because I think the... It... It's not just the AI firms that have to do this. It's really like, you know, every other bank or every other, you know, at Fortune 500 company. And the AI firms

    2:52:39

    are in some sense extracting from these firms. You know, JPMorgan has to spend more money on cybersecurity as a result of AI. And it's not clear to me how much how much of the treasure should be should be shared more widely by the firms in order to secure the rest of society at this

    2:52:58

    Nathan Labenz: point. Yeah. It's a great question. We... You did see a little bit yesterday in the OpenAI keynote too of this sort of eternal tax, it seemed like, on cybersecurity. I forget exactly how they framed it, but it was like, now you can sort of have your AIs monitor your stuff for security vulnerabilities for every... You know, on an ongoing basis. And I was like, yeah. That's exactly what Prakash was predicting they were gonna do. I still think, you know, could hold out some hope. I mean, this is something we should probably get a proper guest on here again soon to talk about, but there have been some pretty interesting blips out of the formal verification community over the last couple of weeks

    2:53:44

    where you're starting to see some bigger software projects get hardened to the point where, if I understand it correctly, like, they're starting to be able to make some pretty

    2:53:56

    Prakash Narayanan: serious

    2:53:57

    Nathan Labenz: guarantees around security of at least, like, core components. And that is a, I think, a huge question. Can that really work at enough scale that we can get out of this kind of middle interregnum period, if you will, and get to the point where we're just in the better secure software future? Strong claims have been made. I think it... We'll have to leave it at that for the moment because I can't really say I've evaluated them or even that I'm competent to evaluate them. But, we'll put that on our to do list to get a guest to come tell us exactly how we should understand what people have been claiming of late.

    2:54:41

    Prakash Narayanan: And on that note, let's, wrap. Nathan, good morning, and we will see, we will see the audience again next week.

    2:54:49

    Nathan Labenz: Thanks, Prakash.

    2:54:50

    Prakash Narayanan: Bye.

    2:54:51

    Nathan Labenz: Talk soon.