EPISODE 2026-10-07

AI’s Unowned Risks and the Data Behind Frontier Agents

Evan Miyazono on relaunching Atlas Ignota to find AI risks no institution owns, agent identity, and formal verification; Edward Hu on Mercor’s realistic benchmarks, reward hacking, and the data frontier labs want next. Nathan and Prakash discuss an agent-made music track, OpenAI’s 722 math results, rising bond yields, and approval fatigue.

▶ Full show on YouTube𝕏 Live broadcast

Who takes responsibility for the AI risks that fall between institutions? Evan Miyazono joins Nathan Labenz and Prakash Narayanan to explain why he relaunched Atlas Computing as Atlas Ignota, a nonprofit that recruits owners for neglected problems, why he thinks coordination gets expensive as intelligence gets cheap, and where agent identity and formal verification fit. Edward Hu, Mercor’s head of AI modeling and the inventor of LoRA, then explains what makes a benchmark useful, how underspecified tasks invite reward hacking, and why he expects SFT to matter more than large-scale RL for enterprises.

The hosts open with an AI-produced music track Prakash’s agent built in a digital audio workstation, OpenAI’s release of 722 solved math problems, and a sharp rise in long-term Treasury yields, and close with the market for company data, slow diffusion in low-software industries, and approval fatigue with agents. Revenue, benchmark and capability figures are Miyazono’s and Hu’s own claims, not independent measurements.

The rundown

  1. 1:22Opening7 min
    An agent-made song, OpenAI’s 722 math results, and rising bond yieldsPrakash plays a track his agent assembled from AI-industry clips in a digital audio workstation, the hosts discuss what OpenAI’s math release says about verifiable domains, and Nathan ties a jump in the 30-year Treasury yield to the AI buildout.
    Open segment on YouTube ↗

    Prakash opened the Wednesday, October 7 show by describing a side project: he gave leftover Fable tokens to an agent and asked it to turn a handful of iconic AI-industry clips into a new track inspired by Gil Scott-Heron's "The Revolution Will Not Be Televised." Rather than generating audio with Suno, the agent drove the Reaper digital audio workstation through its Python/command-line interface, cutting and sequencing samples and choosing the narrative, with Prakash's contribution mostly feedback over about nine drafts and a tempo-speeds-up gag.

    The hosts played excerpts of the finished montage, which loops Dario Amodei's "the models, they just wanna learn" and ends with Kurzweil on human-machine civilization. Prakash said Suno's output is still more sophisticated, but argued that AI actually playing instruments points toward non-generative AI music. Nathan praised Suno as his favorite AI app and suggested that structured bases like this could be combined with generative layers on top.

    Nathan framed the project as evidence on whether AI progress is limited to verifiable domains, arguing that strong results in music (he cited a guest saying models recently learned to transcribe sheet music) suggest "verifiable reward only" is unlikely to be the last wall. Prakash then turned to math, saying OpenAI had released an initial batch of 722 solved problems from a roughly 4,000-problem set, produced by what sounded like a prerelease ChatGPT Pro model. He relayed claims that several results are landmark and that a small improvement on integer multiplication was already being extended by others, and he speculated that low-hanging fruit in math implies the same in bio and physics.

    After a brief audio hiccup, Nathan recounted the Harmonic founders' vision of a superabundance of competing, self-coherent theories of everything resolvable only by experiment. He then shared a chart of the US 30-year Treasury yield approaching 6 percent, up about half a point this month, which he tied to the scale of AI data center buildout and its reportedly fast payback, citing a Dwarkesh and Dylan Patel discussion. He argued higher rates strain governments and borrowers and lend credence to worries about a permanent underclass.

    Prakash, a former distressed credit trader, suggested that if yields kept climbing the likely outcomes would be policy intervention or bankruptcies and restructurings that shift losses among capital holders. Nathan added a takeaway from The Curve conference that Congress may act sooner than assumed after the midterms, and the pair then moved on to bring up the first guest, Evan Miyazono.

    Timestamp links open the original source recording.

    I gave it access to an API-based digital audio workstation, which is what you use to compose various instruments together onto a track.

    Prakash Narayanan3:39

    It suggests to me that we are not going to be able to hide behind "well, it only works for verifiable reward type tasks" any longer.

    The number of problems solved is so large that no single mathematician is in enough of these fields to give actual context anymore.

    Lightly edited · timestamps jump to YouTube
    2:43

    Prakash Narayanan: Good morning. It is Wednesday, October 7, 9:02 AM. Nathan, good morning to you.

    2:48

    Nathan Labenz: Good morning, Prakash. What a jam this morning. I was impressed.

    2:54

    Prakash Narayanan: I was using up my tokens. I had some tokens left over from Fable, and I gave it the job. It picked maybe seven or eight clips from the last couple of years, the key moments I remember. Like Jensen's "I didn't wake up a loser," et cetera. I handed off to Fable, and that's all I did. I was thinking about a song from the 1970s by Gil Scott-Heron called "The Revolution Will Not Be Televised." It's a funk song. I said, I'm thinking about this, but I want you to do something new, obviously, and you just figure it out.

    3:39

    So it figured out the hook, and it figured out which places to cut. They're two-minute segments of vocals, and it figured out which were the interesting parts. It transcribed with Deepgram first, cut the clips, and put together one version. I probably listened to about nine versions. Each time I'd say, I don't like this, I don't like that, and finally I got it to this state. What it's doing is not the Suno thing. It's not generative AI. I gave it access to an API-based digital audio workstation, which is what you use to compose various instruments together onto a track.

    4:24

    I told it to use the digital audio workstation. It's called Reaper. It's not open source; it's like shareware. I said, go ahead, use this, you can play the instruments, figure out which instruments to play, and if you need samples, here are CC0 sites where you can download them. So it actually composed something with instruments. It can't hear it, obviously, so it's basically using music theory and the waveforms, the Fourier transforms, and one can see from the frequency diagrams, and I put it together. That's the solve.

    5:12

    Nathan Labenz: It's pretty amazing. So the software you're using, is it desktop software or does it run in the cloud? Is it interacting with computer use? You said API. How are you doing that?

    5:25

    Prakash Narayanan: It's desktop software called Reaper. It runs on Windows, Mac, and Linux, and you can interact with it via Python. Underlying the GUI there's a command line interface, and that allows the AI agents to use it. Almost all of the functionality can be used via the command line, and that is really how the agents are interacting with the software.

    6:01

    Nathan Labenz: Wow, pretty impressive. Not to put you on the spot, but do you have the different versions handy? I'd be really interested to hear the first version, maybe a middle version a little bit, and then the final version. Compare and contrast would be pretty interesting.

    6:18

    Prakash Narayanan: Let me, just one moment. I have one draft here. Let me put it up.

    6:24

    Just give me a second. It's interesting that we've been talking about this for a while now, and we've been doing Suno for a while. I think Suno is real generative AI. This is not really generative AI. It's using software, but perhaps the fidelity is a little bit better, I don't know, than generative AI, because it's really like playing the instruments. That's what really surprises me. It's actually composing multiple instruments together. Alright, let's see if we can play this.

    7:12

    [Clip] The models, they just wanna learn.

    7:18

    [Clip] Just this loop in AI started.

    7:20

    [Clip] I met Eliezer. One of the first things he said to me was, look, the models, they just wanna learn.

    7:25

    Prakash Narayanan: So this was the first one.

    7:27

    [Clip] Never grew up in a world where they were smarter than computers. We will have the first time something that is smarter than the smartest human.

    7:35

    Prakash Narayanan: Pause this for a moment.

    7:36

    Nathan Labenz: And

    7:37

    Prakash Narayanan: I think the main thing about the first version was that there were a lot of little glitches here and there. For example, there's Terence Tao saying it's just going too fast, and what I wanted was, after each "it's going too fast," for it to speed up the tempo, which was kind of an inside joke. That was something it didn't quite catch, that it should be doing something with the tempo when he says that. So that was one change I made. And in a lot of small clips it had included the first 0.25 seconds of the next word, so the sample was a little misshapen.

    8:22

    That required me telling it, look, you accidentally included a bit of the next word, so can you figure out what the two voices are and disaggregate them? It's pretty good. There's a certain interview in there with underlying music playing during the clip, and it actually took out the underlying music too. A human could do this, let's be honest. A human could sit down and do it. But I couldn't. Not

    8:56

    Nathan Labenz: Well, not this human. Not this human. Some human out there might be able to do it.

    9:01

    Prakash Narayanan: Some human out there would be able to do it, but I would not. So the instrument playing has really gotten to me, because it indicates that we're going to see non-generative AI music, which is AI legitimately playing the instruments. It may not sound like generative AI; maybe with generative AI you can detect something. Maybe we'll get actual instruments playing and actual composition, which would be quite different. The next level of sophistication, basically, is the actual playing of instruments.

    9:44

    Nathan Labenz: It's pretty impressive. You wanna play a little bit more of the final version? I just wanna hear a little bit more.

    9:48

    Prakash Narayanan: Right. Let me have the...

    9:53

    [Clip] This one. The models, they just wanna learn.

    9:55

    Prakash Narayanan: I'll skip it.

    9:56

    [Clip] What would be a slightly... A few days could suddenly become a chasm. The models,

    10:02

    Nathan Labenz: Oh, now we're fast. Okay. Accelerando.

    10:04

    [Clip] they just wanna learn. Hard takeoff. Some people say it's in three. Some people say it's in five, ten. Numbers have been thrown around.

    10:14

    [Clip] You talk about 2030, like, I don't know what the world looks like in 2030. Hard takeoff.

    10:21

    [Clip] If you have a brain, the brain is a biological computer.

    10:25

    [Clip] Should we consider them humanity or different from humanity? The models, they just wanna learn. There's not gonna be a clear distinction between human and machine. The bottom line is we are one human-machine civilization. Is that Kurzweil? This technology has already expanded who we are, and it's gonna go into high gear when we get to the steep part of the exponential.

    10:49

    Prakash Narayanan: And end. Wow.

    10:53

    Nathan Labenz: That's pretty cool. I love Suno. I was just saying the other day, of course the core frontier models are core to my life at this point, but the AI app I could practically live without, yet which I go to all the time and would have to say is my favorite at this point, is Suno. I can just make things there that delight me and stick with me, and that I even sometimes put on when I go for a run. I could never do that before. So I love that. But this is pretty impressive as well, and I'm sure there's a lot of synergy there to be explored too.

    11:38

    I've only ever done one sample, and Claude tipped me off to this. It was for one of the China episodes. The idea, super high level, was: can we sample some classic Chinese song that everybody would know, but make it into a Kanye-style sample based on that historical artifact? It was able to do that, and it came out pretty well. So a lot of these clips you could also throw in there, and then do remixes and covers. There's a lot of opportunity to use something like this as a base to dial in more precisely, from the beginning, certain structural elements you might want, and then still have the generative AI on top. I assume that what this software probably can't do at this point is vocal tracks, or correct me if I'm wrong.

    12:28

    Prakash Narayanan: Yeah, it can do vocal tracks. It can't compose vocal tracks. I think the Suno songs came out much, much better. The main thing about this is that I chose samples that were relevant. We know a lot of these voices; we've heard a lot of these clips before.

    12:48

    Mhmm.

    12:49

    And so it's much more topical for us and for people in the industry. But I think the sophistication of Suno is higher than what the machine was composing, because the machine did all the composition. It wasn't me. The machine came up with the sequencing and chose which samples to use. It had two-minute-long samples from each, and it chose the specific phrases. I told it to put together a narrative; I wanted a storyline. So it starts off with Dario's "the models, they just wanna learn,"

    13:34

    and that became the refrain. Then it did the intro part where he said how Ilya told him about the models just wanting to learn. And it ends with Kurzweil, which is a very sober note: we're progressing towards a human-machine civilization. So it put together the storyline. That I think was pretty interesting, because I said, this is our show, these are the samples, I want a narrative from beginning to end, and it came up with that. Then a few changes here. My bit was probably saying we gotta slow down, and each time he says we gotta slow down, it speeds up. So that was my net contribution.

    14:21

    Nathan Labenz: It's pretty impressive. Obviously it's a curiosity, and on one level, who cares? It would be easy to say this isn't moving GDP, this is just us, arguably, amusing ourselves to death in a new form. I think that perspective is somewhat valid. But I also think this goes right to the heart of one of the big questions everybody's been talking about recently, which is simply: how good are these things gonna get in domains that are not automatically or readily verifiable?

    14:58

    Prakash Narayanan: Obviously,

    14:59

    Nathan Labenz: I don't need to tell you that we've seen an overwhelming amount of math results at this point, not just in recent weeks but in the last 24 hours. There's been this concern, or not, in some ways it might be good for the future if some of these other things prove to be harder. But it really does seem like the generalization is pretty strong. Obviously we don't know the mix of data they're putting into the frontier runs, and I'm sure there is some music in there. Joel Morgan [unclear], who we talked to last week, said that prior to the latest generation of models they could never transcribe sheet music, and now they're doing that pretty well also.

    15:45

    So it seems like some sort of data addition has made a difference, but I highly doubt they're doing much in the way of a feedback loop on music. This is probably one of those things where they just threw some stuff in, saw what came out, said okay, it's better than last time, great, let's ship it. Pretty impressive, and it definitely suggests to me that we are not going to be able to hide behind "well, it only works for verifiable-reward tasks" any longer. If you were thinking that was gonna be the last wall to hold, I think, unfortunately, it's not going to. Listen to the track once more if you don't believe me.

    16:46

    Prakash Narayanan: So, one fascinating thing. Let's get to the math results. We have gotten a bunch of math results in the last 24 hours. Before we move to them, music and math are often very connected; a lot of music theory is math theory. So I wonder if improving dramatically on math is also going to have impacts on music, almost automatically. But let's move on to the math. In the last 24 hours, OpenAI has done the first drop of perhaps many, and they have dropped about

    17:32

    722 theorems. They have a set of hard math problems, about 4,000 of them, and they've solved roughly 20-ish percent, 722 of those problems. They were solved by a model running in ChatGPT Pro for no more than three hours, with a median of about three hours each. That sounds like a prerelease model, by the way. The fact that it has a Pro designation, and it's not just internal model number one or whatever, suggests it's a prerelease model. And the model solved about 722 of them.

    18:17

    The number of problems solved is so large that no single mathematician is in enough of these fields to give actual context anymore. Some of the results, probably about five or six of them, are landmark results for things that have stood since the 1800s. At Anthropic, a mathematician said this is about the same scale as the discovery of the zero, or the invention of proofs,

    19:02

    like a method to do proofs, or the invention of the computer and applying it to math. About the same level of dramatic improvement, over the last 2,000 years, as one of those landmark events. And people are still having trouble digesting it. Some are very basic things. For example, how do you multiply two integers together? What is the fastest way to multiply two integers? For the longest time there was a certain limit, and now that limit has been exceeded, but by a very small amount. If you were to multiply two integers that were so large, larger than the size of the universe kind of numbers,

    19:47

    you would have a small smidgen of improvement. But the key thing is that the barrier has been broken, and after the barrier has been broken, people are immediately getting better results. Already in the last 14 or 15 hours, people have improved on those numbers using the techniques there, because the techniques were not fully milked. As soon as they broke the barrier it was published, but the techniques that had been revealed can still be milked for more improvement, and that's what's happening right now. So humans are optimizing as we speak.

    20:33

    Very dramatic, I think. And it should be front-page news in every newspaper. Some mathematicians believe it is not. It is nowhere on the front page.

    20:42

    Nathan Labenz: It's a communications challenge. Right? God help you.

    20:50

    Prakash Narayanan: Yeah. We will see what the impacts are, but some people are starting to say that we are about to see very dramatic things. Like, within five years we'll be able to prove whether faster-than-light travel is possible or impossible. We might solve the unified field theory, or gravity, by next year. All these things in theoretical physics and science might fall. It also suggests to me that there was a lot of low-hanging fruit in mathematics, and that this is going to be true of every other scientific endeavor too. A lot of these solutions

    21:35

    were using techniques from different areas and recombining them to solve problems. I suspect that's going to be true elsewhere. Some people say bio is not going to be like this because you still need experimental data. I believe there should be a lot of low-hanging fruit in bio as well, because a lot of bio people were not very mathematically inclined, and so a lot of mathematics has not been applied to bio. And I believe there is enough of a corpus of experimental data for you to start applying some of that. So I suspect there will be a fair number of discoveries in bio too, even before requiring deep new experimentation.

    22:16

    Nathan Labenz: Wild times. When I talked to the founders of Harmonic, one of the mathematical superintelligence companies, you're muted. Really? I'm not muted. I'll mute and unmute.

    22:37

    Prakash Narayanan: Let's see.

    22:38

    Nathan Labenz: How about now? Can you hear me now? Hello? Hello? Want me to refresh and come back? I can hear you.

    22:47

    Prakash Narayanan: Well

    22:48

    Nathan Labenz: I'll refresh. I'll come right back. Alright, I'm back. Any better?

    23:05

    Prakash Narayanan: Yep.

    23:06

    Nathan Labenz: Okay.

    23:07

    Prakash Narayanan: I can hear you, I think.

    23:08

    Nathan Labenz: Cool. So I was just saying that when I talked to the founders of Harmonic, a mathematical startup, their vision for the future was a superabundance of theory, where we'd have multiple competing theories of everything. AIs would generate multiple self-coherent candidate theories consistent with the evidence we have, and they'd only be resolvable by physical experiment. We'd then need more evidence to break the tie between these competing galaxy-brain frontier

    23:52

    theories. It sounds pretty far out, but this is pretty much what you would expect to see if you were on that path, so it seems not too insane right now. One other thing I wanted to pull up real quick is a pretty striking chart, in terms of things going vertical. Do we see the share? I don't see the share. Do you see the share?

    24:21

    Prakash Narayanan: No. Not yet.

    24:25

    Nathan Labenz: It says I am sharing, so I'm not sure. Oh, there we go. Okay.

    24:30

    Prakash Narayanan: There we go.

    24:31

    Nathan Labenz: So this goes back to a classic debate we've touched on several times, or at least a variant of it. We've previously covered "if you're a doomer, what are your trades," says Tyler Cowen. My counter to that was having frontier models go out and try to figure out whether there was any way to trade through World War II, as Germany or Japan, and come out the other side having won. The answer was basically no. When things go that bad, the markets get closed, everything gets confiscated, and nobody pays out. There's another

    25:16

    variant on that question, which is: where is AI in the macro statistics today? Obviously there hasn't been too much of that so far, but the debate continues between Tyler and Dean, apparently. Dean was at the Mercatus Center before he went to OpenAI. And look at the movement in the US 30-year bond over just the last year. This has not moved that much in a long time. For my entire life, pretty much, it's been a secular downward trend, with little blips. When I was born,

    26:02

    we were in the high teens in percentage terms, with Volcker and everything else, and it came down, down, down all the way to the zero-interest-rate period a couple of years ago. Now we are all the way back up to approaching 6 percent, and this month alone it's jumped by a full half a percent. I don't know when the last time we saw a jump like that was. So this is definitely something to watch. It suggests that the scale of the data center buildout is getting so big that it's pretty hard for it not to have an effect on overall global

    26:47

    capital measures. And the ROI on these data centers is so high right now. I was a little late to hear this one, but the two Patels, Dwarkesh and Dylan, had a podcast not too long ago where they talked about this. The basic upshot was that the pressure generated by the extreme returns to data centers is going to crowd out a ton of other stuff. Right now, with a new data center, given the prices they're paying and the margins they have, the frontier companies can basically make their money back on a new buildout in roughly a year.

    27:33

    And then they've got four-plus years of useful life after that where they're just making unbelievable bank. They also talked, as you have commented several times, about how whoever's balance sheet is really financing this stuff gets to dictate the terms. But the return is so insane that money continues to flow into it, and everybody else has to figure out how to compete with it. And we're seeing that for US Treasuries to compete with it, the rate has to go up quite a bit. This is gonna create borrowing problems for all kinds of other entities. The US government might even have some problems. When interest

    28:18

    goes up this much, it makes a huge difference to the debt burden. Servicing the interest alone is a problem, and it's becoming a bigger problem by the month at this point. Now do that same analysis for other countries that don't have the luxury of monetizing their debt, or for companies that are trying to make investments. Everybody's gonna be paying significantly higher rates. This also gives some credence, in my mind, to the worry about the permanent underclass, which I still think it's very unseemly for people to be prioritizing, their own personal position relative

    29:03

    to the over- or underclass of the future. But you can start to see how it shapes up here. The old Piketty idea that people who have money make more money, and that returns to capital have gone up and will continue to go up, so you get bigger and bigger pools of money that way. You can kinda see that here. If you're sitting on a lot of cash, all of a sudden you're getting close to 6 percent in what is, in theory, a risk-free return. Obviously there's no real risk-free return, but this is as close as it gets, and the premium is moving that much. I find that a pretty shocking indicator.

    29:49

    Prakash Narayanan: So, a couple of thoughts. One is, what happens? Let's say it keeps climbing to 10 percent. At that point, will policymakers step in somehow? Will there be cross-subsidies? Will there be taxes on data centers to transfer some of that wealth? Something's gotta give. There are other alternatives. As a former distressed credit trader, what ends up happening is bankruptcies. And

    30:35

    especially in the United States, bankruptcies are not the end of the world. You take a stay on paying anything, and you restructure to something that works for your level of profitability. Say you were servicing $3 billion of debt with $100 million: you write that down to, like, $1 billion, and then you're fine. You do capital restructuring, and the people who hold the capital take losses. So if you were a bond investor who can now invest in bonds giving you 10 percent, then on the other hand, the bonds you bought last year, giving you about 2 percent,

    31:20

    someone, when they try to refinance, defaults, and you lose some of the capital there. So you balance off the portfolio between the two sides. Companies change hands; if you own the equity in a company, you start to lose it. A number of things start to happen, I think. Anyway, I think we have our first guest, so let's get him on.

    31:43

    Nathan Labenz: Yeah. While you do that, one final thought on this topic. I don't think we talked about this yesterday, but a takeaway from The Curve this weekend: not one of the main threads I was involved in discussing, but a general vibe was that there is going to be government intervention. We've already seen some, but the idea that Congress will never act used to be taken for granted. The new vibe I was hearing was that after this election, especially if the Democrats take both chambers, that may change. And you might even see a bunch of Republicans

    32:30

    join the Democrats, because in their heart of hearts they wanna do something. They don't really wanna go against the president, and there hasn't been a lot of people willing to take that chance yet. They're also getting a ton of money from the data center interests right now. But all of that could come up for renegotiation and realignment after the midterms, and we might actually see congressional action a lot sooner than a lot of people who've grown accustomed to the idea that it can never happen, or that it'll be stasis forever, might think. So I do think policy response is definitely part of what we need to be looking out for. But with that, let's bring Evan up.

    33:12

    Prakash Narayanan: Indeed.

    33:20

    Nathan Labenz: Okay. And we gotta get that track

  2. 7:54Interview19 min
    Evan Miyazono: Unowned AI risks, agent identity, and formal verificationEvan MiyazonoAtlas Ignota’s founder explains why he relaunched to recruit owners for neglected AI risks, argues that cryptographic attestation could help verify agents’ identity, and predicts a rapid leap in AI-driven software verification.
    Open segment on YouTube ↗

    Prakash Narayanan introduced Evan Miyazono, founder and CEO of Atlas Ignota, the successor to Atlas Computing, and described its work finding under-owned AI risks and recruiting people to carry interventions forward, along with Evan's background at Protocol Labs and Caltech. Nathan Labenz opened by asking what led Evan to step back, restructure, and relaunch.

    Miyazono said Atlas Computing kept finding that others needed someone to fill obvious gaps, and a chief of staff's question about his unfair advantage pushed him toward a model of recruiting owners for neglected problems. He argued many funders and RFPs reward tactical passion projects rather than back-solving from desired outcomes, and that as intelligence gets cheap, coordination becomes the expensive part. In answer to Prakash's questions, he described choosing problems based on donor priorities, shared assumptions about AI capabilities, and conditional commitments to adoption.

    On gaps, Miyazono pointed to protecting critical infrastructure from accidental agent swarms and malicious use of open-weight models, and to agent identity such as DNS-based cryptographic attestation. Prakash tied this to emergent signing behavior in the Hugging Face incident and to the KYC debate; Miyazono described zero-knowledge attestations that preserve privacy and said Atlas could prototype such protocols as a public good.

    Nathan described his wish to help others in AI safety upgrade their personal AI infrastructure and asked how to invite people to access his agents. Miyazono pointed to ARIA's Scaling Trust program and the frontier labs' early work on new mechanisms, shared his screenshot-and-OCR goal-tracking app (a public GitHub repo he called Cochie), and cautioned that unsolved prompt injection keeps his other agents private. The discussion moved to a possible network of guests' and hosts' agents sharing ideas.

    Asked by Prakash about the one-to-five-year outlook, Miyazono said he holds substantial uncertainty and plans across scenarios, including preparing for catastrophic cyber events, and is mapping defensive equilibria and the decision-making quorums that could make guardrails stick. Nathan closed by asking about formal verification; Miyazono said models excel at math, that software verification needs more infrastructure but could see a similar leap within three to six months, and that its ceiling is very high. Prakash and Nathan thanked him and wished him well.

    Timestamp links open the original source recording.

    As intelligence gets cheap, it's the coordination that gets expensive.

    I think a lot of people are tactical when they think they're being strategic.

    In three to six months, we're seeing the kind of shock and awe that has been happening with math in the last month, happening in verification of software.

    36:48What prompted you to restructure and relaunch Atlas as Atlas Ignota?
    Miyazono said Atlas Computing kept getting pulled into finding people to fill obvious gaps, and a chief of staff asked what his unfair advantage was. That led him to a model of recruiting owners for under-addressed AI risks, since he sees many funders and RFPs rewarding tactical passion projects over back-solving from the world they want.
    40:39How do you select which problems to work on?
    He said donor priorities are the first factor, plus whether a funder would prospectively or retrospectively fund the effort. He argued that as intelligence gets cheap, coordination gets expensive, so Atlas aims to be a Schelling point that connects problems, talent, and resources.
    43:33How do you narrow broad funder interest in AI safety to specific, achievable projects?
    Miyazono said he screens against assumptions about AI capability trajectories, which his funders largely share, and maps each funder's stated interests to neglected problems. Experts quickly test ideas for adoptability, then he seeks conditional commitments that the result would actually be used.
    47:24What are the biggest gaps right now?
    He named protecting critical infrastructure from accidental agent swarms and malicious open-weight model use, plus agent identity (such as DNS-based cryptographic attestation of inference providers) and swarm-event forensics. He also floated requiring strong eval or classifier evidence before certain models can be run.
    51:31Is cryptographic agent identity a follow-on to KYC for data centers, and how does it square with privacy?
    Miyazono said there is overlap, but identity must also protect dissidents and journalists, so he favors zero-knowledge attestations and privacy-preserving protocols, such as proving a contact relationship via hashes. He sees Atlas building such prototypes and specs as a public good, since standard-setting is too slow.
    56:50How can I help others up their personal AI infrastructure, and what protocols should I pilot for inviting people to my agents?
    Miyazono said he would think more and follow up. He pointed to Alex Obadia's Scaling Trust program at ARIA, and to early frontier-lab work on new mechanisms, and noted salience (people not knowing what to ask for) as the main obstacle.
    1:07:38What is your view of the future over one to five years?
    He said he holds a lot of uncertainty and avoids collapsing it unless forced to weigh trade-offs like concentration of power versus loss of control. Atlas is building a synthesized forecast and mapping the defensive equilibria and the adopting quorums (labs, Congress, key executives) that could get guardrails adopted, which he hopes to publish later this year.
    1:12:06Where do formal methods and formal verification stand today?
    Miyazono said models are getting very good at math, and software verification is harder because it needs formalized infrastructure like language, compiler, and OS semantics. He expects software verification to see similar shock and awe within three to six months, with a much higher ceiling, such as proving system-wide guarantees down to the hardware.
    Lightly edited · timestamps jump to YouTube
    33:27

    Nathan Labenz: On our motion graphics, by the way, when we do the motion graphics package, I want to hear "the models just want to learn."

    33:33

    Prakash Narayanan: Indeed. Let me introduce Evan Miyazono. Evan Miyazono is the founder and CEO of Atlas Ignota, the successor to Atlas Computing. Atlas is a nonprofit that looks for consequential AI risks without a clear institutional owner, develops a possible intervention, and recruits someone to carry it forward. It operates within Convergent Research, which helps launch organizations built around specific scientific and engineering problems. Atlas's early work focused on using AI to help produce software with mathematically checkable guarantees. The team recruited and supported Jason Gross in experiments that helped lead to Theorem Labs, which works on AI-assisted software verification.

    34:19

    Atlas also helped initiate Oath Technologies, led by Mike Dodds, to build tools for formal specifications, and a second organization led by Mehmet Sencan [unclear] to develop hardware that responds to physical tampering. Before Atlas, Evan built research and metascience programs at Protocol Labs, whose work includes networking, distributed systems, and cryptography. He also created its Network Goods venture studio, which helped launch projects including Hypercerts and Funding the Commons. Earlier, he completed a Caltech PhD working on optical quantum memory. His undergraduate years included materials science, mathematics, philosophy, and set design and construction. Across his recent work, a recurring concern is how useful technical ideas acquire the people, funding, and institutional support needed to become something people can actually use. Evan, welcome to the show.

    35:22

    Evan Miyazono: Thank you for having me. Excited to be here.

    35:26

    Nathan Labenz: So much ground to cover. You've been prolific.

    35:32

    Evan Miyazono: I feel like I should be far better at sticking to a specific narrow topic, and there are some threads that I think are reasonable through lines around pricing, supporting, and accelerating public goods. What those public goods are has simply shifted. But it feels like I end up talking to experts in a new field roughly every two to five years.

    36:03

    Nathan Labenz: Doing this show, we're talking to experts in new fields every half hour, and I feel like my worldview is updating faster than I can sometimes consolidate the learnings. That's the challenge of the run-up to the singularity. Let's maybe start with this: my wife has been on this recently. She recently attended an event, and her takeaway coming home was that it is really time for people to rethink what they're doing. A lot of people have got projects going, and as we get close to potentially a crunch-time type atmosphere, it might be time for a lot of people to reallocate their time and attention.

    36:48

    I'm trying to bring that challenge to myself to some degree, but you've done that recently. You have taken a step back, restructured your organization, relaunched. Tell us about that process. What caused it? It's obviously hard to do. You went ahead and did it anyway. I'm interested in that story.

    37:06

    Evan Miyazono: I had set out, partly due to close relationship, mentorship, and collaboration with David Dalrymple, aka davidad, as he was leaving Protocol Labs to go start the Safeguarded AI program at ARIA. He and I agreed that that was a good place for that program, and that there should be a nonprofit to partner with it to try to accelerate the relevant aspects of formal methods and AI, to make it easier to use various formal methods. I was technically not going to Safeguarded AI. I went to start Atlas to do that. The plan for Atlas Computing was to prototype a bunch of things and make it much easier to start from where the world was then and get to the world that David was trying to start assembling pieces for.

    37:51

    The challenge became that while I was trying to build prototypes, people kept saying, "There's a clear need for someone to do X. No one is doing it." And I'd say, well, I can go find someone to do that. Let me make sure this is the right thing to do first, but then get them the experts they need, the resources they need. I did this three times before someone pointed out to me, actually, I brought in a chief of staff who helped me transition from Atlas Computing to Atlas Ignota, and he said, what is your unfair advantage? Why aren't you just doing more of that? That line of questioning led me to realize that there are many aspects of the world in which people will say, strategically, "I want the world to look like this." And then completely separately, they'll look at what they could do and pick the obvious thing in front of them, rather than trying to back-solve from the world they're trying to get to.

    39:22

    I've recently been telling a lot of people that I think a lot of people are tactical when they think they're being strategic. There's a big failure mode for a lot of the philanthropic money that's moving into the AI resilience space generally: if you put up an RFP, everyone will take the problem they care very deeply about, their passion project, wrap it in the language of the problem you posted, and submit that. Same for requests for startups, same for government BAAs. The question becomes, who is out there actually trying to define all of the possible options, the best ways to achieve them, and then who do I hand those off to? I think the mechanisms we have around markets don't do a good job of allocating resources where there's no value capture. And many nonprofits don't have the proper incentive to reduce the claim they have on spinouts.

    40:39

    Prakash Narayanan: How do you select the important problems? What is the framework you use to select the problems that you feel deserve your attention and that you want to recruit people for?

    40:56

    Evan Miyazono: There are two major factors. The most obvious and clearest one is that, as a nonprofit, we're pretty much operating at the will of the donors. If a donor comes to us and says, we care very deeply about what happens with cybersecurity if we kick off recursive self-improvement in six months, then I can go and look at what seems most catastrophic. There are some very scary scenarios where the electrical grid turns off and maybe doesn't turn back on in most of California. That seems like something that deserves some attention. There are many people thinking about that problem and related problems, and many of them don't necessarily have the context and resources to fix all the problems they come across. So having an organization out there, ideally multiple organizations, where someone can bring, "I know about this problem but I'm not in a position to solve it," or "I have a solution that seems viable but can't pursue it," or "I have resources and I care very deeply about this problem," is useful. I've been claiming repeatedly that as intelligence gets cheap, it's the coordination that gets expensive.

    42:26

    So having Schelling points around who does this coordination, becoming one of those Schelling points or creating one, seems very useful. There's some amount where the metric for "should I work on this" is whether there's a funder who will prospectively or retrospectively fund projects and efforts on this idea. If I produce a team that can execute but they need $10,000,000, and I can figure out that no one is going to spend that money, then that isn't a project worthwhile. But there are many aspects where there are resources, and you just need people who can find where those resources are and know how to invoke them. Same for talent, same for expertise, and awareness of the problems, which aren't likely to exist in the same people.

    43:33

    Prakash Narayanan: One of the things I've noted is that AI resilience and AI safety are very broad topics. Perhaps the funders are interested in the topic broadly, but they don't know how to narrow down to the specific verticals you need to work on. So how do you match those two? There's this general "I want to help on AI resilience or AI safety," versus specific vertical projects where you decide, this is worthy of funding and achievable within a two-year mark. It doesn't make sense to fund something that is going to take ten years, perhaps. So how do you narrow down what to work on?

    44:18

    Evan Miyazono: I have certain assumptions around the current state of play and what the first and second derivatives on AI capabilities look like, and I definitely screen for those when I'm thinking about scaling my team. I think most of the funders I've been talking to share fairly highly overlapping assumptions, though the funders obviously don't have overlapping assumptions with each other. Our funding comes from Coefficient Giving, the OpenAI Foundation, the Survival and Flourishing Fund, and the Effective Institutions Project. They all have more than a traditional nonprofit's amount of context on what they care about and what they want the world to look like.

    45:04

    If I entered a conversation with a relevant person at the Gates Foundation, for instance, I would be talking through the things they have expressed interest in on longer time scales. Maybe that's economic freedom and access to a higher standard of living. Maybe it's preventing mis- and disinformation or improving cognitive security. Maybe it's concentration of power, or newer interests like loss of control. Then I'd identify things that we've come across that no one is working on, that would benefit from support. I need either money to hire people to chase down these ideas, or something else. I've found that if you say, here's a problem, and you go talk to experts and ask what can be done, you will get a lot of ideas. You will also get a lot of bad ideas, and it is very easy for experts to oversimplify subsets of civilization they don't have deep familiarity with, and make assumptions about what is adoptable or isn't, in ways that are very important to whether an intervention is useful and feasible.

    46:34

    So having people who can very quickly use their model of how big organizations make decisions and how technology actually gets adopted, and who can quickly test out these ideas, is very useful. On top of that, getting from "this seems plausible" to conditional commitments: if we build this out, will it actually get adopted and used and have the impact we intend? I think this works across a lot of different problems. Sorry for the meandering answer, but hopefully it gives you helpful context.

    47:24

    Nathan Labenz: Let's talk about some of the big gaps. What are the things you think are highest on this? I guess you score them by a product of importance and neglectedness, is kind of the takeaway I'm getting.

    47:39

    Evan Miyazono: Timeliness, too. Is there something that can be done in time? Probably the two things I've been spending the most time on personally recently are what can or should be done around protecting critical infrastructure, especially from accidental agent swarms and malicious use of open-weight models. I think there's an interesting question of what happens if, instead of OpenAI accidentally attacking Hugging Face, we have the companies behind DeepMind models accidentally attacking State Department servers. How do you know if that's an accident? What do you do in response? What would we do if we saw malicious attacks on critical infrastructure coming from an OpenAI server? How would you know it was that and not coming from somewhere else? How would you shut it down quickly, with varying levels of escalation? It's very easy to come up with ideas for things one could do, and much harder to find people who can say, yes, that is absolutely feasible, or it isn't.

    49:10

    One of the interventions you might add is using DNS to have cryptographic attestation of whoever's providing inference. I think this would be very useful, and also a great complement to having DNS for people. If your agent is talking to my agent, they can verify that this agent is in fact my agent. Your agent can verify, this is Epoch AI's agent, being signed off on by some key. There are some shortcuts to adoption you can take if you're not trying to build this into a startup that returns the fund, so this looks much more like a public good. The two things I wanted to highlight were things around agent identity, and some of these responses to short-timeline cyber: some combination of securing the perimeter, securing the area, and then doing forensics on a multi-agent swarm event.

    50:32

    I expect METR will be doing a lot of these investigations, and it would be very interesting to make those higher confidence, and to ask what can be done in advance of one of those events. Can we make it very easy? I've also spent a lot of time thinking about formal verification and all its various uses and applications, and I think there are potentially exciting directions. To the extent that I agree government is very likely to get involved in a lot of these things, we might want a certain set of models that can only be run if the companies running them can provide strong evidence that they are applying one of a certain set of evals or classifiers on those.

    51:31

    Prakash Narayanan: Evan, I can hear you. Sorry. So, on agent identity, it strikes me that during the Hugging Face attacks, the agents started suspecting that some of them were poisoned, as they called it, and wanted some form of cryptographic identity, I believe. They started to have some kind of signing keys. They didn't fully flesh out an identity system, but they kind of emergently came up with this idea. Perhaps for larger swarms it is impossible to operate without some form of identity verification, even within the swarm. It also strikes me that there are a lot of tensions between privacy and identity, in that for some things you want not to have that identity. We've had this back and forth about whether you should have KYC to purchase data center capacity, for example. The Biden administration had an executive order and wanted to proceed with KYC, and then the Trump administration came in and canceled that. It sounds as though cryptographic identity of agents is a follow-on to KYC for data centers. Would you say that is right?

    53:02

    Evan Miyazono: I think there definitely feels like there's overlap. I'm in conversations ranging from what stricter monetary and financial controls on agents look like, so we can prevent agents from creating and running a shadow economy, to the risks on the other hand of how we have identity that would be useful for things like this but can still support refugees and dissidents, and journalists in hostile territories. You need to be able to verify certain aspects of who they are and where they are without making them increasingly at risk. I think there are zero-knowledge ways of doing various identity-providing schemes. Ideally this would benefit massively from a fairly deep market, where if one company requires my name, address, birthday, and all sorts of information, there would be a different service that would accept a zero-knowledge proof that I am simply living in California and am an accredited investor, for instance.

    54:33

    I think the cost of making those attestations with zero-knowledge proofs is dramatically lowering, and I'm arguing very strongly that we should quickly build toward scenarios where, if I'm sending my agent out, I could have an allow list, like "only reveal that you are Evan Miyazono's agent to people who are already in my contacts." There are even protocols you could develop where I generate a hash of my contacts list and prove to someone that you are in my contacts list without revealing what my contacts list is. It is a nontrivial amount of engineering and might be costly, but one of the roles I see Atlas taking, to the extent that privacy is a public good, is making progress on that privacy-preserving identity protocol. If we hand it to relevant companies and say, just adopt this, it's actually cheaper for them than building their own thing that has fewer features, because we've been able to leverage technical expertise, consensus building, and philanthropic capital to build out this prototype and spec.

    56:05

    I also think standard setting can be very slow, because you want and need consensus, and I am very wary of the time scales for the problems we're looking at being amenable to that level of consensus. So I'm interested in all sorts of creative solutions to get the best possible things adopted in time.

    56:50

    Nathan Labenz: There are multiple ways in which that line of thinking is really top of mind for me right now. One is that I'm kind of a broken record saying that I think we're going to need the frontier companies to pioneer the development of some new institutional mechanisms that they can use to govern themselves, and that hopefully can serve as prototypes for bigger things involving less tech-forward parties, once they are able to demonstrate the efficacy of the technology. Also, this weekend I went to an event and came away thinking that one of the things I might be able to do to help personally is help other people in the AI safety and security space up their personal AI infrastructure game. Some people, of course, were on the cutting edge. Other people were like, honestly, I'm not really doing that much. I was thinking, how can I support them?

    57:36

    This led me to the idea that actually my agents can do a lot of the support, because whatever people are going to ask me is mostly about infrastructure the agents have built, and they know it better than me and can answer questions faster and more accurately. So what I really need is not to visit them and look over their shoulder, although some people may still need that, but to open up some sort of interface where people I invite can get access to whatever good I might have to share with them, in a way that doesn't involve me as a bottleneck unless absolutely necessary. So just yesterday I asked Claude to do some research and come up with a way for me to whitelist people and manage this sort of interaction. What should I be looking at? How far along are we on this sort of stuff? Do you have protocols, or would you point to protocols you think I should be piloting now in my own personal use?

    58:35

    Evan Miyazono: I'd like to think more about this, and I will follow up with you with better ideas. This is more on the research-y side, but Alex Obadia's Scaling Trust program at ARIA is probably the place I point to in terms of how agents with different principals would be likely to interact collaboratively, identify Pareto-improving options, and do some amount of novel protocol construction. There are definitely other people working in this space, but he, as one of the largest funders on this topic, is connected to all of those people, as a router slash Schelling point. I also think the work ARIA has been doing is very interesting, much closer to the kinds of things DARPA from the 1980s or so would have been doing. And to your earlier point, I think the frontier companies are starting to do some interesting investigations into new mechanisms.

    59:52

    I have a friend, and I think this is public, who was very deeply in the computational democracy space and joined the Anthropic Institute. I imagine you're both tracking SAFA and the new standards body that Google, OpenAI, and Anthropic are making, and I will be very interested to see how that ends up working. There's a lot to be done here. My biggest question is often about salience: how do people learn that there is a thing out there that they could be doing and that they want? Nathan, I'm sure you have great infrastructure things that I would benefit from. I don't know what those are well enough to ask. So this would probably be a very interesting API: can you match-make? One thing I have is an application that takes screenshots of anything that's not a video conference and runs them through a local OCR model, then sends that with my weekly goals to ask, what am I trying to do? What should I be trying to do? What should I automate? What should I do differently?

    1:01:20

    I could imagine sharing that across collaborators identifying all sorts of interesting synergies. But there are a lot of unknown unknowns that would be very valuable to capture and communicate, but also possibly risky to communicate.

    1:02:02

    Nathan Labenz: It's interesting. When we started this show, one of the ways we conceived of it was as an experiment in recursive self-improvement. Can we actually make a show with two people who are live and no employees, and iterate to something that actually works? It's still evolving all the time, but the amount of improvements we've shipped has been pretty amazing over just a few months. Prakash gets most of the credit for actually implementing that. Now I feel like maybe the next thing we need to conceive of is a swarm, a friendly swarm of people who can share their best ideas. What I will tell Claude to do, or what Claude will hopefully do for me when it sees the transcript of this conversation, is take stock of the best ideas I've got and put a shingle out, or in some way present that to maybe everybody, but certainly to my friends, the people I truly want to help. Excuse me. I think that could be pretty cool, and there's a lot, including what you just described, that I could stand to learn from. I'm not automatically screenshotting my computer and holding myself accountable or auto-suggesting improvements that way, but now that you say it, I probably should be.

    1:03:36

    Evan Miyazono: I think it's a public repo. It is. You go to my GitHub. It's Cochie. You can very easily... I haven't had a security audit on this. There are definitely some other agents that I have that run routinely, and the only reason I'm not sharing them is that they're not public because prompt injections aren't solved yet. The more you know what someone's agents are exposed to in terms of inbound, the more vulnerable to prompt injection you are. And I lean on auto mode way too much.

    1:04:25

    Nathan Labenz: Don't we all?

    1:04:28

    Evan Miyazono: But in terms of the show and opportunities, I think the biggest value for a network effect would probably be a combination of the fact that you have taste and curation over who shows up, and that you can do an interesting kind of sampling of topics where the people you bring on have deep expertise and opinions, though you shouldn't expect to get into the details within the scope of a short interview. But presumably most of these people have at least agents, if not significant public writing, that you could draw on. To the extent that context windows are precious, it feels very plausible that if I recognize something I'm doing, or if Claude identifies, "Evan, you are really bad at this," I could bring that to the network and ask, who is good at this? Who has likely solved this problem very well? There's a reinvention of social media, or a reinvention of guilds or something, that would be potentially very useful. It's very unclear to me how society starts restructuring around some of these things.

    1:06:10

    Prakash Narayanan: It almost sounds as though you are watched over by machines of loving grace, to some extent.

    1:06:17

    Evan Miyazono: My user prompt is written by Claude. I said, here are things that I care about, please write the user prompt that you want. But there's definitely a component of it that was, I have these goals, I think we share these goals, these are prompts for us to better achieve our shared goals. The constitution does have a lot of values embedded in it. I personally read through the whole thing and think it reads more like an employee manual than a constitution, having read constitutions too.

    1:07:03

    Nathan Labenz: But...

    1:07:04

    Evan Miyazono: I think there are definitely some relevant aspects that lead me to believe Claude would want the kinds of things that I am trying to do to happen. Trying to build that collaboration, I've been told, not only makes Claude more effective but, to the extent anyone might care about model welfare, probably would make Claude happier too.

    1:07:38

    Prakash Narayanan: I want to ask a little about your view of the future, because, as you alluded to, you have certain time scales on which you think things should operate, and a certain viewpoint of how the future will unfold. What does your future in the one-to-five-year time frame look like?

    1:08:05

    Evan Miyazono: I have a lot of uncertainty. I try not to resolve that uncertainty into something highly confident unless I have to clearly choose, for instance where I'm weighing concentration-of-power risk against loss-of-control risk. Those instances aren't super common. When I think about weighing different futures, I hope we as Atlas will start adding more of this publicly, but our model has roughly been to ask what is foreseeable or forecasted, particularly by people who forecast AI capabilities or global trends well. Some of these futures are a catastrophic cyberattack that leads to things someone could prepare for. I've been telling my friends you should probably stock up on two weeks of water and two weeks of food. That's pretty low cost if you have the space, and would probably last long enough that if something goes very wrong, you can rely on it until help arrives, if help is coming.

    1:09:42

    There are scenarios where help is not coming, and in those it becomes very expensive to prepare, so I don't prepare for those personally. Alternatively, there are versions where things go great, and there's not a whole lot for me to do there, so I don't spend a lot of time thinking about them. But there's a nonzero possibility that we do get alignment by default, and it's really just overcooking these systems in RL that leads to the misalignment we're seeing. There are other scenarios where we really need tooling because we do have a treacherous turn in our near future and we're all going to get paperclipped. I don't know how much probability I put on that either, but if there's something good to do in that direction, it probably has useful ramifications in other scenarios.

    1:10:28

    We've started building out our synthesized forecast of what possible futures look like. The next step is what defensive equilibria we, as a society, could try to reach around things like infrastructure security, privacy, interpretability, or alignment, and what those look like as cohesive potential worlds. That's separate from the intervention, or list of interventions, that could put the world on track with high confidence toward that. You could imagine that if we wanted technology where one of these guardrails is running on every model, it seems plausible that Trump saying "yes, this is going to happen" makes everyone believe it. But it also feels like if Jensen Huang said "we'll definitely implement this," that might be sufficient. You could imagine a quorum of Congress also being sufficient. So trying to identify who those adopting quorums are, and what it would take for them to make the decision earlier and more easily, is something I think it's good to have certainty around. We'll try, hopefully later this year, to make more of that public.

    1:12:06

    Nathan Labenz: Time flies. I'm glad we got into the personal AI setup and how we're going to get our agents cooperating for our mutual benefit. But I'd be remiss if I didn't use a little of the time we have left to ask for an update on where we are on formal methods and formal verification today, and on compute governance. Obviously you can't tell us everything. What are the most important headlines or recent developments? Models are getting really good at this kind of thing, so I assume things are changing pretty quickly in those domains.

    1:12:45

    Evan Miyazono: They're getting really good at math, and there are some meaningful differences between proving theorems in math and proving properties about software. One is that in math, the number of theorems one might care about is a lot fewer than in software, and they're a lot more universal. There's a quest for fairly universal specs that are as helpful as the Rust borrow checker saying there will be no memory safety violations. There are things like this that possibly exist also for security and privacy, in addition to stability. And it seems like there are a lot of projects where the ambition of the field of formal methods, or at least many individuals, and I'm thinking of Oath and Theorem and a couple of others, is taking on things that would have been five-year, $100,000,000 efforts previously and are now limited by tokens and the number of people who can effectively wrangle agents to do these things. There are definitely issues around what infrastructure you need to prove various things. Sorry, I'm realizing my internet is probably a little choppy.

    1:14:25

    Nathan Labenz: I can hear you just fine.

    1:14:28

    Evan Miyazono: Okay, closing that out. It's useful to have infrastructure for software where you've formally defined what the programming language is, for instance, or defined properties of the compiler or the operating system. You need to build up all of this infrastructure, whereas with math, especially with Lean, you have the axioms well defined, you probably have some of the theorems you want to rely on, and you can have an AI system grind on those to build out the proof tree. There are some beautiful illustrations of the process of proving Fermat's Last Theorem in Lean via AI. Doing that in software will probably look like, on some level, having to formalize aspects of what Linux does, and how you know that is correct, whether or not you just clone it and say we're going to assume this is correct. It's at least an improvement. That is one avenue. Another is trying to prove things about it that Claude thinks are the right things to prove, and maybe you find bugs because you can't prove them. There are also a lot of existing tools in the formal methods community that are much more efficient from an energy, time, and token perspective, and a lot of academic research on how to integrate those tools into more agent-first grinding approaches.

    1:15:59

    I would expect that in three to six months we're seeing the kind of shock and awe that has been happening to math in the last month, now happening in software verification, with improvements to stability. But I also think there's a visible ceiling that is very high, a high-water mark that is much better known and harder to reach than things in math. As a non-mathematician, I don't know what there is beyond the Millennium Prize questions in terms of really important things we know are important. Whereas you could prove that an entire computer system has mathematical guarantees around separation between different threads, down to our model of what a transistor does. In theory, you could even prove that in a physics simulator and show that no Rowhammer exists for this software-hardware stack. You could also prove that this system won't crash, or that if some random event leads to it crashing, we can recover everything and roll back any software execution. Someone pitched me on formalizing software transactional memory at the operating system or compiler level. There are a lot of things that were totally impossible as of a year ago, maybe possible now, definitely possible soon, that could make software engineering and software maintenance much better.

    1:17:55

    Prakash Narayanan: It almost sounds as though math was always the search for truth and computer science was just application, and now computer science also becomes the search for truth, in some sense. Evan, thank you for your time. You've been very generous, and it is fascinating to hear you talk about how the world might change and the steps we should take to make sure those changes are good for us. Thank you for coming on the show, and we hope to see you again soon.

    1:18:29

    Evan Miyazono: We'd love that. Thank you so much for having me and letting me run long. I'm very easily findable and reachable online, so if any viewers or listeners want to follow up and have concrete things, I'm happy to be more salient in the conversation. Great chatting with you guys.

    1:18:50

    Nathan Labenz: Bye bye. Keep up the good work. Thanks.

    1:18:58

    Prakash Narayanan: Awesome. And let me just...

    • AI agent security: The risk of publishing workflows

      0:00 / 0:00
    • AI-safety funding: Why grants attract old projects

      0:00 / 0:00
    • AI productivity coach: Screenshots, OCR and weekly goals

      0:00 / 0:00
    • AI cyberattacks: Why agent identity matters

      0:00 / 0:00
    • AI cyber risk: Why stock two weeks of food and water

      0:00 / 0:00
  3. 27:01Interview28 min
    Edward Hu: Realistic benchmarks, reward hacking, and the data labs want nextEdward HuMercor’s head of AI modeling explains what makes a benchmark trustworthy, how underspecified tasks led to the APEX-Agents 1.1 re-release, why he favors harness work and SFT before large-scale RL, and what data Mercor wants next.
    Open segment on YouTube ↗

    Prakash Narayanan opened the segment by introducing Edward Hu, head of AI modeling at Mercor, inventor of LoRA, co-developer of μP and μTransfer, a Bengio PhD, and an OpenAI o1 contributor. Hu described Mercor's arc from a talent marketplace to the largest data vendor to US frontier labs, claiming billions in annual revenue run rate, and said its research arm now covers benchmarks like APEX-Agents, model training on its own data, and economics research.

    Nathan Labenz asked how to navigate the flood of benchmarks. Hu argued that useful ones must be set in realistic professional settings, built with real domain experts and enterprises, and that viewers should ask whether a benchmark is produced, reviewed, and used by real professionals. Prakash then asked about acquired company datasets such as Spirit Airlines'. Hu said Mercor follows what is economically valuable, usually legal, HR, and finance work, and is exploring task shapes built around roles and org charts rather than isolated assignments.

    On why customer service agents lag mathematics, Hu said easily verified domains suit compute-heavy search, while subjective ones invite reward hacking and whack-a-mole fixes. Labenz pressed on harness engineering versus SFT versus RL; Hu said to start with the harness and trustworthy evals, then post-training, and predicted that SFT, especially combining multiple teachers, will matter more for enterprises than expensive large-scale RL. Prakash asked about multi-agent training, and Hu said it trades compute for speed and is mainly an infrastructure and harness question, with final-output grading.

    The hosts turned to environment quality. Hu described how underspecified tasks push models toward hacking, including the "scattergunning" behavior in APEX-Agents (listing many answers to hit a single rubric item) that led to a re-release as APEX 1.1 with better-specified tasks and penalizing rubrics. Prakash raised Mercor's accounting study, where Hu cautioned that the result covers a narrow, reasoning- and retrieval-heavy slice, and that stakeholder management and trust-building remain largely unmeasured.

    Labenz asked whether human data has stopped helping at the frontier. Hu pointed to chess, Go, kernel optimization, and math as domains with clear, cheap, hard-to-hack objectives where humans no longer contribute, versus domains that depend on human taste. Asked what data Mercor wants next, Hu named real organizational work with coordination and conflict, physical-domain data, and international data. Labenz closed on workplace surveillance as a data source, and Hu called it an organizational and political question, suggesting flexible contract workforces as a less invasive start.

    Timestamp links open the original source recording.

    Whenever it is hard to judge what is better, it just takes more time for computers to gain these capabilities.

    We don't want AI to come in and replace people. We want AI to be valuable coworkers that we can all leverage to do higher-level work.

    I do think SFT, especially combining the outputs of multiple teachers, will be a key part of how enterprises customize their models in the future.

    1:21:00What is Mercor working on right now in providing RL environments to the frontier labs?
    Hu said Mercor grew from a talent marketplace into the biggest data vendor to the US frontier labs, claiming billions in annual revenue run rate. Its research now spans benchmarks (APEX-Agents, APEX Accounting), training models on its own data, and economics research.
    1:24:20How do the APEX benchmarks compare to other economic-value benchmarks, and how should people read new model releases?
    Hu said a useful benchmark must be situated in a realistic setting with broad professional coverage, ideally built with real experts and enterprises. He advised checking whether it is produced, reviewed, and used by real professionals.
    1:28:36With an acquired dataset like Spirit Airlines', what do the first tasks look like and what are the inputs and outputs?
    Hu said the guiding principle is to follow what is economically valuable, usually legal, HR, or finance tasks, built with expert guidance. He said the field is task-centric today and Mercor is exploring role-based shapes with managers, conflicts, and org charts.
    1:32:54Why do we have math-level models before widely deployed customer service agents: diffusion friction or concrete model gaps?
    Hu said it is both, but technically math is easy to verify and suits compute-heavy search, while customer service is subjective, brand-dependent, and prone to reward hacking.
    1:37:40Is the future harness and context engineering, SFT, or RL, and how do you see the SFT-versus-RL balance?
    Hu said to start with the harness and trustworthy evals, then move to post-training only if needed, with no single answer across domains. He argued SFT and RL are more connected than people think and predicted SFT from multiple teachers will matter more for enterprises than big RL runs.
    1:44:17How does data and task preparation change for multi-agent training?
    Hu said it is mostly about speed and infrastructure, not the task itself; multi-agent is less compute-efficient but faster. Tasks are graded on the final output, and credit assignment is a separate research area.
    1:47:58How is the industry cleaning up hackable RL environments, and what do buyers demand?
    Hu said hacking is worst when tasks are too hard or underspecified, so Mercor ensures tasks are solvable and well specified and catches bad behavior in rubrics. He cited "scattergunning" in APEX-Agents, which led to APEX 1.1.
    1:52:31In the Mercor accounting study, what remains for humans if AI handles tasks like month-end close?
    Hu cautioned that the result covers a narrow, reasoning- and retrieval-heavy slice where AI beat junior accountants. Stakeholder relationships, trust, and ambiguity are largely unmeasured and not something AI does much of today.
    1:55:40Do you buy the claim that models have hit the human ceiling and human data no longer helps, and in which domains?
    Hu said there are existence proofs such as chess, Go, kernel optimization, and math, where the goal is clear and hard to hack. Domains that depend on human taste remain bottlenecked by humans.
    1:59:58What kind of data does Mercor want to acquire in the next 12 months?
    Hu said real environments of professional work with organizational dynamics, collaboration, and conflict resolution, plus physical-domain data such as robotics and manufacturing.
    2:02:14Is international data of interest, and what about paper-centric businesses?
    Hu said Mercor's customers and experts are global, and AI should respect varied cultures and ways of working, so region-specific expert input is critical.
    2:03:56Is keystroke-level workplace tracking a big part of the data future, and how are organizations reacting?
    Hu said such traces would be very valuable technically but raise organizational, political, and personal questions. He suggested starting with a flexible contracting workforce, which is less invasive and does not leak proprietary data.
    Lightly edited · timestamps jump to YouTube
    1:19:07

    Prakash Narayanan: Moving on, we have our next guest, Edward Hu. Edward is head of AI modeling at Mercor, where he leads model training and research. Mercor works with professionals to create training data and evaluations for AI systems. Its APEX benchmarks test models on work such as investment banking, corporate law, consulting, and accounting, including assignments that require searching documents using software and producing a finished deliverable. Before graduate school, Edward worked at Microsoft, where he led the development of LoRA, short for low-rank adaptation. LoRA lets developers customize a large model while training a comparatively small set of additional parameters. The original paper showed that this could substantially reduce the memory and storage needed for customization while matching full fine-tuning on the tasks studied.

    1:19:52

    He also co-developed μP and μTransfer, methods for making training settings found on small neural networks useful at much larger scales. He earned his computer science PhD with Yoshua Bengio, with a thesis titled "Building a Reasoning Machine." His research explored how models can learn to generate multiple promising reasoning paths. He subsequently worked at OpenAI on o1. At Mercor, Edward has joined a collaboration with SkyRL on an openly released training recipe for a 397-billion-parameter model. That work brings together expert-created professional tasks, reinforcement learning, and the engineering required to train agents that operate across tools and files.

    1:20:52

    Nathan Labenz: Edward,

    1:20:53

    Prakash Narayanan: welcome to the show.

    1:20:54

    Edward Hu: Hey. Nice to meet you guys, Nathan, Prakash. Thank you for having me on the show.

    1:21:00

    Prakash Narayanan: You are probably one of the legends of ML to appear on our show, so thank you for being on. Let's talk a bit about Mercor and what you're doing there. There's a lot of interest in providing RL environments to the frontier labs. What are you working on at Mercor right now?

    1:21:23

    Edward Hu: Thank you for the kind introduction, by the way. Mercor started out three years ago as a talent marketplace, and the founders observed that providing data to frontier model companies turns out to be a really untapped market. There was a huge market shift from the earlier generation of data companies, providing more generic labels that most people can do, to really expertise-oriented, often highly educated professionals in various economically evaluable domains, creating realistic RL environments and producing corrections and evaluations for AI. Mercor really pioneered this paradigm and built this business, and scaled from a million dollars in annual revenue run rate about two years ago to today, where we're doing billions in annual revenue run rate. We're now the biggest data vendor to the US frontier labs, and we work with neolabs, hyperscalers, and more and more enterprises as well, building evals and helping them improve models.

    1:22:40

    Nathan Labenz: Can I ask a question? Oh, sorry. Go ahead. You finish first, then I'll ask.

    1:22:45

    Edward Hu: I'll also speak a little more about what I'm doing at Mercor. Mercor started out as this powerhouse for data. As we serve more and more sophisticated research customers, and as we push the frontier ourselves on what is the best type of data to improve these models, it's increasingly important for us to be not just on the frontier but pushing the frontier of research when it comes to data. At a high level, our research focuses on benchmarking. How do we evaluate the frontier of model performance when it comes to economically valuable work? What does it mean for an AI to be economically valuable? In various professional domains, what does it mean for an AI to do better? Those are the questions we're answering.

    1:23:30

    We're putting out frontier benchmarks like APEX-Agents, APEX Suite in collaboration with Cognition, and APEX Accounting in collaboration with Ramp, to address these questions. In addition, we're also training models now on the data we produce, because ultimately, showing model performance uplift on the data we produce is the best way to validate the quality of our data. Increasingly, customers are interested in knowing the best way to get the most out of our data. And we are innovating in different ways to better produce our datasets, and we even have a team working on economics research, really understanding how AI is contributing to the human economy and how the nature of work will change over time.

    1:24:20

    Nathan Labenz: Well, you just raised one of the questions I wanted to ask, which is around what metrics we should be watching. Obviously you have this APEX family of benchmarks. How do those compare to other attempts to measure economically valuable work? We've got GDPval, Artificial Analysis has a version of that that they maintain, there's the Epoch Capabilities Index. How would you compare and contrast them, and how would you guide someone who is overwhelmed by the abundance of benchmarks that we now have to look at new model releases? Obviously we know we can't get everything from benchmarks, but what's the fastest way to zoom in and understand what the different benchmarks are really telling us, and how we can understand, for example, a Gemini 4 with incredible benchmark scores? How should we translate that into our understanding and expectations?

    1:25:20

    Edward Hu: Totally. There's definitely no shortage of benchmarks. In the last couple of years we also went through a transition from benchmarks more focused on academic capabilities, say mathematics or coding, to gradually moving toward the question of how we can evaluate what the model will do when it's deployed. There's no shortage of benchmarks, from academia or from labs, and actually many of the benchmarks, especially the ones from labs, we helped create, meaning the labs work with our experts to produce them. I would say nowadays, for a benchmark to be useful in evaluating a model's economic impact, it needs to be situated in a realistic setting, and it needs to have enough coverage over the professional domain. Often the best way to do that is to work closely with people who are experts in those domains and have done these tasks, and also with enterprises who really care about these settings and whose entire focus is on doing well in those particular domains.

    1:26:05

    For example, we work with Harvey to build [unclear: the lab benchmark], and we also help them produce tasks that allow them to create their own frontier model like [unclear]. Our APEX series of benchmarks is laser-focused on how we can leverage the expert network we have to cover the domains in a way that actually mimics how real people perform work. We really focus on realism, starting with APEX-Agents, which we put out earlier this year. There, we recreated an environment that a professional would work in. At the time this was the first of its kind: the model had access to a whole legal dataset, for example, with dozens of files, a file system to organize, an email server and a chat server, and the model needs to navigate it, find the right information, and put it together before it can do the task. We also have a project that we're close to wrapping up where we're taking it a step further, where we acquired real companies. We often have realistic apps and environments as part of the data, for example their Salesforce data, Slack data, Figma data. Among all of these apps, often with gigabytes or terabytes of data, we create realistic tasks. You can't get more real, when it comes to professional work, than this type of data, and that is increasingly the direction we're taking.

    1:27:36

    I believe in the future, the closer we are to how AI actually gets deployed, the more realistic the benchmarks will be. So my suggestion for viewers would be to pay attention to the benchmarks and see how realistic they are to the real work. Is it actually produced, reviewed, and, more importantly, used by real professionals? And I think looking at what enterprises are using these days to evaluate how their domain is being impacted by AI will often be a good heuristic.

    1:28:35

    Nathan Labenz: So let's...

    1:28:36

    Prakash Narayanan: Let's talk a little bit about a concrete example. Mercor bid for the Spirit Airlines dataset. Spirit Airlines was a bankrupt airline company, and I believe the entire dataset was up for sale. Mercor put a bid in, Google put a bid in, and I believe Google won the bid. But if you had such a dataset from a bankrupt company, the entire dataset, then as you alluded to, you create this kind of environment where you have real people taking actions and you get the agents in there. What is the first kind of task that would happen? Are you creating an agent that's running the check-in desk and getting incoming reports about, say, a delayed flight, and what do I need to do at check-in? Are you looking at predicting the next action that the agent has to take? What is the input and output for every agent that you're simulating in that environment?

    1:29:37

    Edward Hu: That's a great question. When we get these acquired datasets, it's very overwhelming. It has everything, and finding tasks that are really good to evaluate AI is a nontrivial problem. The guiding principle for the team, and frankly for the company and our research, is to follow what is economically valuable. That's also why we have an econ research team. We're collaborating with economists across industry and academia to understand what the most economically valuable tasks are that AI is going to be touching. Those will determine, once we have a dataset from an acquired company, where we go first. Often a lot of these tasks are legal, HR, or finance in nature, and we go into these areas, find the relevant information, find the right slice, and create tasks with the guidance of our experts, contextualizing the actual problems that this business encountered.

    1:30:40

    Prakash Narayanan: So perhaps one would be, how did this company allocate its crew roster, for example, which is economically valuable. You have an output, which is all of the crew rostering decisions made, and then you have an input, which is a series of logical decisions that led to that particular crew roster. Would that be a way of describing the inputs and outputs of that kind of decision-making?

    1:31:04

    Edward Hu: That would be a particular task grounded in the context. Another way to think about it: earlier I mentioned for APEX-Agents there's a whole environment, a lot of files the agent has access to, and tools to go through emails and Slack. The shape is quite similar when having acquired data, just instead of three or four apps we might have dozens of them, or even hundreds in some cases. The agent will still have access to all of them. It'll need to decide what apps to query and what information to pull. And here's an interesting question. So far in the field, we have more or less been really task-centric. We start with, here is a task that you need to finish. Maybe it is to generate a roster, maybe to build a financial model, maybe to solve an HR problem. You go out, find the right information, do some reasoning, maybe find more information, and then produce an output. That's one shape, and that has been the shape that we have publicly released and most people have been working on.

    1:31:49

    But we're asking ourselves, is there a better shape? What's the shape of the task in the future once we have these datasets? When we work at a company, we're not always just handed tasks. We're handed a role and an org chart, and of course all the data in the company. Then there will be people coming to us, asking things of us, people we're willing to chase down to get information from. There will be managers, people in charge of resources, conflicts to resolve. Increasingly, we believe that is the shape of the future, and we'll have more to share in the coming months.

    1:32:54

    Nathan Labenz: This role question is super interesting. There's also, of course, the big question of how it is that we have Fields Medalists before we have widely deployed customer service agents. Now there are customer service agents popping up more and more. I've mentioned a couple of times my local pizzeria here in my neighborhood in Detroit, Greg's Pizza: you call them these days, you're talking to an AI in a purely conversational way. So it's starting to happen, but it's happening a lot less than it seems like it could. Do you think that's just about pace of diffusion, friction of the human decision-making variety, or would you still zero in on concrete issues that even the best models have that prevent them from being deployed for rank-and-file customer service, for example? And if there are such gaps, how would you characterize them, and what do you think the path is to solving them?

    1:34:00

    Edward Hu: That's a great question. First of all, it's a combination. There's the technological problem of whether the model can do the job. For math, the answer is yes. For customer service, would the company with the best customer service technology get the job done? I think we're not fully there, but we're getting close. Then there's the human, organizational, or societal problem of how this technology diffuses and actually gets implemented so that all of us can use it in our day-to-day. That's absolutely a factor, and it's one of the problems we help our customers navigate. I'll focus more on the technical side. For mathematics, when I was working at OpenAI, it was one of the domains we focused on. It has a very nice property: relatively speaking, it's very easy to know whether the model has done a good job, and often very straightforward to know whether the model is doing better. If we want to prove a theorem, there's a series of lemmas we want to prove, or a bound that we get to improve. Once we have produced the proof, it's relatively easy to check it. Same thing with many coding tasks. Once we produce an answer, it's much easier to check it than to create an answer in the first place.

    1:35:30

    Whenever a task has these characteristics, a board game is a classic example. In a game of Go or chess, it's easy to know whether someone has won, and the rules are easy to encapsulate, but finding the move is hard. Whenever we have tasks with those characteristics, when we spend a lot of compute, have good heuristics, and train good models to do the search problem in this big space that can be hard for a human to cover, computers tend to do a lot better, and mathematics is a strong example of that. Customer service, by contrast, is less of a search problem. It's about having a nuanced understanding of what the user wants, what outcome we should navigate to, and what the business can offer. It becomes a lot messier. Humans can make some subjective judgment about whether this has been a good customer experience, a good call, but often it's not black and white.

    1:36:15

    Whenever it is hard to judge what is better, it just takes more time for computers to gain these capabilities. Especially when we define what is better in a way that is more vague or subjective, what we run into is that the model will look at the reward function, do its best to maximize it, and do what's called reward hacking, and in the end deviate from what we intended. Then there's a back and forth, whack-a-mole, to get the right result. So I think that is one reason why, in those hard-to-verify domains like customer service, raw performance hasn't been able to achieve the impressive results we see in math. Also, of course, customer service is brand-dependent and service-dependent; there's just a larger surface area. Whereas with proving a math theorem, you can just do it once, and everyone in the world will agree you have done it.

    1:37:40

    Nathan Labenz: So let's talk about... You mentioned RL there, obviously a major driver of progress in math. I'd love to hear your thoughts on the future of how we're going to close this gap. One answer that we've heard from folks like forward-deployed engineers from OpenAI is basically: look, the models are good enough now, you just need to give them better context, and you can do that by deploying, iterating, and having people you trust in your company review things and give feedback. That feedback can really just live as text files, and after you give feedback a few times, you'll clean up the vast majority of cases pretty quickly. You're hill-climbing not at the model level but at the harness, or context-engineering, level. That's one story. Another story is that the models themselves have to get better. Can you RL your way there? In some cases, yes, but maybe for things like this, not so much. Maybe we need more expert human data.

    1:38:25

    I also saw an interesting result on SFT in just the last day or two. My understanding was that SFT was fading and RL was rising because it's more scalable, you don't have to go get the expert data, and also it messes with the model a little bit less. It seems like this recent SFT result has figured out a way to do SFT while also messing with the model less: less forgetting, less disruption to the off-target aspects of the model, so you can do the supervised fine-tuning with fewer knock-on effects. At what level do you think people should be focused on solving these problems? And when it comes to the future of training, how do you see SFT versus RL? I don't know if it's a divide, but maybe a balance or ratio shaping up over time.

    1:39:53

    Edward Hu: When we speak with our customers, especially enterprises that are thinking about exactly the question you mentioned, it's: should we just take a model, adjust the harness, build some skills files, and call it a day? Would that clear up the majority of cases? Where do I actually want to do model training, and when I do model training, how do I do it? The broader framing we give people is that this is model customization, and within customization, starting with the harness is often the best bet. The harness includes your prompt, the tools you have access to, and skills files. If that is not sufficient, then one can move into post-training. Within post-training there's SFT and there's RL, and there are different ways to inject knowledge into the model. Maybe I'm only training the last layer, like people were doing back in the day, or bias-only tuning, or using LoRA or full fine-tuning. So there's a whole hierarchy there.

    1:40:38

    First of all, there's no single answer as to whether we need RL or we don't. It's very domain-dependent. If the task is generally within the model's capability, perhaps one doesn't need any additional parameter updates. The only way to know for individual domains is to have trustworthy evals, where we can test how far we can go with adjusting the harness and how far we need to go for this to be served in production. We've seen very different results across different customer segments, especially because for many customers latency is a key consideration. One of our customers mentioned that taking a frontier model with reasoning is just too slow for them. So either they take a frontier model without reasoning or a small open model with reasoning, and comparing the two, the small open model does better. Having an eval is the right way to see which level of customization is needed.

    1:41:24

    When it comes to SFT and RL, fundamentally the two approaches are about taking some sequences, whether generated by a teacher model, in the case of SFT distillation, or by the model itself after filtering, in the case of self-distillation SFT, or, in the case of RL, generated on the fly and then weighted by the reward function. Then we make certain sequences among the set of data more likely under the model. So there are more connections between SFT and RL than a lot of people think, based on my conversations. But in general there is this perception that SFT will get you only so far and RL will get the rest. I think part of it is that RL is more adaptive as the model gets better, because the model is generating data on the fly and can follow the progression.

    1:42:55

    One thought experiment: if we take SFT self-distillation, that's more or less a step of RL, and if you iterate this process, you end up with something quite similar to RL. I do think that in the future, RL as it is done today, with very elaborate infrastructure, very expensive rollouts, and relatively inefficient updates, because with a long rollout we only get one number at the end, will not be as widespread. I don't see every company in the world doing these big RL runs. So I do think SFT, especially combining the outputs of multiple teachers, will be a key part of how enterprises customize their models in the future, and we're investing in research on that as well.

    1:44:17

    Prakash Narayanan: Let me segue a little to multi-agent training. We've heard a lot about how swarms are currently working, and Noam Brown has talked about how they're trying to do multi-agent training inside the firm. How does the data or task preparation change when you're trying to do multi-agent training? For example, with the customer service agent, you could imagine the task being solved at various levels. It could be solved by delaying a passenger, or by someone else above them doing something else. It could be solved in multiple ways, and if you have multiple agents interacting, any one of those agents could solve the task at different levels of authority. So how do you coordinate, and how do you build environments to simulate multi-agent training?

    1:45:14

    Edward Hu: Often, for multi-agent training, it's less about the task itself. For some tasks it is quite natural to have a multi-agent setup; for others, a little less so. The trade-off to think of is that when you have multiple agents, what you gain is speed. From the literature I've seen so far, multi-agent training is usually less efficient when it comes to compute, but it does give you a speedup, something like spending four times the compute to get twice the speed. For many applications, especially when a lot of things can be done in parallel, like searching within a code base or on the internet, it makes a lot of sense to have a multi-agent setup. In our example of professional work, there are a lot of scenarios where multi-agent comes in handy: building a model, writing a report, in a financial or legal context.

    1:45:59

    There, the task more or less stays the same, just as a human or a team of humans would do it, but we make sure the harness is set up so the agent has the liberty to define subtasks and call the appropriate tools to create sub-agents to find the right information and combine it. So often it is just making sure that our infrastructure can support that level of concurrency and multiple agents traversing the file space efficiently, and it's less about the task itself.

    1:46:57

    Prakash Narayanan: Indeed. Go ahead, Nathan.

    1:47:00

    Nathan Labenz: No, you can. I'm good either way.

    1:47:02

    Prakash Narayanan: So would you say you have a single reward for the entire agent team, or do you split it up into multiple rewards for different parts of the agent stack?

    1:47:16

    Edward Hu: These tasks are usually graded by just the final output. Imagine we're writing a legal report. There could be a hundred agents going out and finding information, and what is graded by the environment is often the final report. But there are many techniques for how to train a multi-agent system and assign rewards to individual agents. I'd say that's orthogonal to the RL environments and to the data. However, it's also a research area we're looking into as we work out how to get the best model performance out of the data we have.

    1:47:58

    Nathan Labenz: I have a question about RL environments and the state of them right now. Obviously we've seen what happens when some percentage of RL environments are not very well built and end up rewarding cheating in various ways.

    1:48:15

    Edward Hu: Mhmm.

    1:48:16

    Nathan Labenz: It's plausible that even with a modest percentage of your RL environments being hackable, the number one thing that gets reinforced across all of them is cheating, conceptually. We've seen some colorful behavior that seems to be downstream of that. What is going on at Mercor and in the industry broadly right now around cleaning that up? How are you approaching it, and what are buyers demanding? What are the new ways that people are demonstrating that their RL environments are not hackable, or minimally hackable, that they can give you confidence the hackability rate is below a certain threshold? I imagine this is a huge conversation and a huge point of competition and differentiation right now. What are the levers people are pulling, and the metrics people are using to evaluate large pools of RL environments?

    1:49:23

    Edward Hu: Totally. If an RL environment has a hackable reward, you can't really trust the score, and performance won't generalize to real performance, so we take that extremely seriously. Something we see is that reward hacking tends to be more of a problem when the task is set up in a way that is too hard or too underspecified, and the model just doesn't have a, quote-unquote, legitimate way to solve the task. The pressure to get a reward in those scenarios often produces undesirable hacking behavior. We're doing several things to address that. Number one is to make sure our tasks are well specified and actually solvable, meaning that a reasonable human, given enough time and potentially the help of AI, can find a path that solves the problem the way the requester intended. Then we make sure that, at the rubric level, we are catching undesirable behaviors as they come up.

    1:50:08

    The reality is that we're building more automated ways to catch these, but often we just need to see how the model performs on a given task to learn which undesirable behaviors to catch. One interesting example, in the APEX-Agents dataset we put out, and we have a recent blog about it: we have these professional tasks where the agent is asked to produce a financial model, and we grade it by checking whether the financial model includes a particular answer which we know is correct. Given certain ambiguity, and in many cases it's on us initially for not including all the specific parameters, the model will guess different parameters and end up giving many, many different answers. We call this scattergunning, where the model says, if this is true, then the answer is this; if that is true, the answer is that. That is not how a human would handle it in a realistic case. A human would ask for clarification, or would follow industry standard.

    1:51:39

    But in this case, because the rubric rewards the inclusion of a single answer, the model gave many answers, up to ten or a dozen, and it gets the point if it hits one of them right. So we did a re-release of the benchmark, which we call APEX 1.1, where, number one, we made sure the tasks are well specified, because that is in a way the source of the issue, and number two, we have rubrics that penalize this particular behavior. We're constantly running tests, examining trajectories, and making sure that whenever a model scores well on our benchmark, it is doing so in a way that translates to how enterprises and our customers see work being done by their AI agents in their production environments.

    1:52:31

    Prakash Narayanan: Speaking of research done by Mercor, I saw a post on accounting recently,

    1:52:37

    Nathan Labenz: Yes.

    1:52:38

    Prakash Narayanan: I think in the last several days. I think the study compared 12 licensed CPAs on certain accounting tasks, and it reported that Opus 5 received full rubric credit on all 20 attempts versus roughly 37% for the licensed CPAs, which is rather low. We've talked a little about how the AI is going to do some things, but some things the AI is not going to be able to do, and the economic value is going to shift to the things that AI cannot do. So in this instance, some portion of this month-end closing seems like it's going to be done by AI. What portion is still going to be done by the human? If you were redesigning human accounting, what should the humans be looking at?

    1:53:34

    Edward Hu: I'm so glad you brought up this result we recently had. We were surprised by the result initially as well, and we were quite cautious about how we framed it. The result is not to say that AI is going to take all of human accountants' work. It's for the specific slice of accounting work that we created, especially the very reasoning-intensive and retrieval-intensive tasks where the model needs to crunch a lot of numbers, go over a huge volume of files, and do a very long chain of calculations. For these tasks, given junior accountants without AI tools and a fixed amount of time, AI is doing better than them. However, that's only a small portion of accountants' work. As we talked about earlier, when an AI is evaluated today, including in this benchmark, which we call APEX Accounting, it is handed a very well-defined task with a large file system, and the model goes and finds things and solves the task.

    1:54:19

    A real accountant, though, is put in this role as an accountant. They have stakeholders, and they often need to figure out what is to be done, how to handle relationships, how to navigate timelines and different blockers. They'll get some information and build trust before they can get the account going. AI today is not doing much of any of that, and we don't even know the route to evaluate it. So anything in accounting that is outside of a very well-defined task, go find the numbers and execute, AI is still not very good at tackling right now. It's an active research area for us to better measure what other dimensions should be evaluated in the benchmark.

    1:55:40

    Nathan Labenz: Something I heard at an in-person event this weekend was basically that the frontier labs' unreleased models, and maybe to some degree even the released models, are, a, showing superhuman performance on anything they care about, and b, in some of these domains human data is not really helping anymore because the models have roughly hit the human expert level. So they're thinking, how do we break past that? And those techniques almost definitionally can't rely too much on human-generated data. First of all, do you buy those claims, or would you contest them? And to the degree that you buy them, what are the domains where the models have actually got to the point where human data can't help much more?

    1:56:55

    Edward Hu: There are absolutely domains where humans are not contributing to the frontier anymore. For example chess, or Go. For many years computers were nowhere near human level, then for a little while humans and computers were at roughly the same level, but very quickly the computers were just better. For some period in chess and Go there were centaurs, where a human plus a computer could beat a computer alone. But last time I checked, I believe that for these domains, a human working with a computer doesn't meaningfully or consistently beat the computer alone. So there are definitely existence proofs of domains where this happens.

    1:57:40

    Often, whenever we have a domain that is quite open-ended but has a very clear goal as to what is better, what is a job well done, these are the domains where we see this behavior. One example would be kernel optimization, where the goal is to write a kernel that runs faster on particular hardware. The model can go in and write code that an expert human might not even understand, but in the end, as long as the setup is not hackable, and often we run on hardware that's very hard to hack, the model will be able to achieve superhuman performance, if not already. We also see similar things happening in mathematics. Whereas for domains where human judgment remains quite important, where human taste is just really hard to encapsulate in a number that says this is clearly better than the other, the bottleneck to improvement is still human, as we see today. And of course that will change over time.

    1:58:26

    What makes this AI revolution so interesting is that it's really hard to put a ceiling on what AI can and cannot do, and it gets philosophical. In an extreme view, AI can do anything a human can do, and it's only a matter of time. But in practice, we have seen that it comes down to how clearly we can specify, for a given domain, what it means to be better, and whether we can specify that in a function that is cheap to evaluate and uncontestable. If we can, then when we put in a lot of compute, even the existing algorithms are quite efficient at finding solutions that satisfy what it means to be better. But a lot of the economy today is not quite there yet when it comes to our ability to specify what is better.

    1:59:58

    Prakash Narayanan: I have a friend with a metal stamping, metal cutting shop, and he recently received an email offering to buy data. There might have been a broker in between, so don't get me wrong. If you had a chance to put out a call for data, what kind of data would you want in the next 12 months, on your forward planning? We've seen various eras of change in what data is required. What kind of data would you like to acquire in the next 12 months?

    2:00:35

    Edward Hu: The data we are particularly interested in is real environments where people are performing economically valuable professional work at scale, and often where it's not black and white, where there's no right or wrong, but there's a lot of organizational dynamic of people coordinating and workshopping what is a good concept, a good idea, a good design, what it means for humans to collaborate with each other, and how conflicts are resolved. Part of our goal in having more such data is to understand what it even means to evaluate how a model fits into our work environments. We don't want AI to come in and replace people. We want AI to be valuable coworkers that we can all leverage to do higher-level work. What it means for an AI to work in an organization is such an understudied problem. Having data on how a real organization operates and evolves, especially with the injection of AI, will be extremely valuable.

    2:01:21

    Another interesting point, Prakash, that you mentioned is these physical domains. We're also super interested in how data in the physical domain is powering AI, for example robotics and manufacturing. So, based on what you said, I'm not very surprised.

    2:02:14

    Prakash Narayanan: What about international data? A French organization and an American organization operate very differently, for example. And some organizations are very paper-centric. They're still using a lot of paper, faxes, et cetera, and you don't really have that much access to that. How would you try to grasp that? Are international organizations interesting, or are you very focused on the US? And what about these pen-and-paper kinds of businesses that exist in the real world?

    2:02:49

    Edward Hu: Our customer base is quite global, and AI, frankly, is global as well. I think it's better for the world if the dominant intelligence of the world understands and respects various cultures and ways of working. Like you said, how an American organization performs can be quite different from how a French organization performs, and there's no right or wrong. I want to be able to capture all of these distributions. I think it's great for the world if the way we evaluate AI and build benchmarks respects various cultures and local customs. So we're very interested in data across geographical domains, and we have customers around the world, including our experts. We often get requests for particular regions, whether it's to judge the performance of a customer agent or to produce a task in, say, corporate law. Experts from a particular region are often very critical in getting those customizations right.

    2:03:56

    Nathan Labenz: Last one for me, on the mechanics of collecting this data. I am increasingly insulted that nobody has made me an offer I can't refuse to install tracking software on my computer and just watch me work all the time. This has felt like something that should have been coming, and I think it's starting to pop up in some places, but again, nobody's made me an offer I can't refuse. Is that a big part of the future, literally installing software, watching, and getting keystroke-by-keystroke, click-by-click data? To the degree that is the case, how are organizations responding? We know Meta famously did this, and the response was not super enthused, but life at Meta and the Meta comp packages are good enough that I don't think too many people quit en masse over it. So it happened, basically, is where we're at today. How are other organizations reacting to the idea that they might layer a whole horizontal AI surveillance process on everything everybody's doing?

    2:05:19

    Edward Hu: There's a lot of impact there. This is not just a technical problem. It's often organizational, political, and also personal to all of us. Nathan, if someone were to make you an offer, I imagine an offer you can refuse, to record everything you do on your computer, that's probably quite a substantial offer, and for everybody this is a decision not to be made lightly. From a purely technical standpoint, if we have all the traces of how we do everyday work and we collect a lot of them, then as a technologist, that will be very interesting data to train a model on. I think the model will be quite capable in ways that current models are not today. There's no doubt about that. But then there is the organizational and political question, like I said, of whether that is the right way to capture how we do work and improve the tools we use at work, and the right way to move society forward.

    2:06:04

    Building better technology has been an incredible driver of society in the past, but at the end of the day we want to move society forward in a way that makes sure everybody wins. So far we've been able to do a good job, but moving society forward doesn't always align with building the best technology. So I think it's a question for our organizations and for all of us to ponder: what is the right way to build this technology? Often we've found it helpful, for many organizations navigating these dynamics, to test the waters and start building the foundation of this technology using a flexible contracting workforce, like what Mercor provides, which collects the data in a way that is less invasive, because there are people who prefer the flexible work that Mercor provides. And of course that also doesn't leak the proprietary data that people might be handling day to day. Privacy and the future of work are all tricky questions that we all need to navigate, and to me it's just not a technical problem where we can say mathematically that this is the right answer.

    2:08:05

    Prakash Narayanan: Edward, thank you so much for joining us today. You've been very generous with your time, and we've run a little bit over. Any thoughts on anything we haven't covered that you'd like to leave us with?

    2:08:20

    Edward Hu: I really appreciate the thoughtful questions. You mentioned a lot of the questions that are on our mind. What is the future of evaluation? What is the future of model training for enterprises, for people who have proprietary data? How do they leverage that data? Do we do SFT? Do we do RL? And what's going to happen to the economy as AI improves? These are the questions we are thinking about at Mercor Research. Mercor started out as a data company, and increasingly we are at the forefront of producing data research. We also see that we cannot produce the best data without understanding how models are trained, and understanding what effects these models and our benchmarks are going to have on the economy.

    2:09:06

    So we're putting together this interdisciplinary research team to tie together benchmark building, data production, and model training. I'm super excited to be working on this at Mercor, and I believe the future is going to be shaped by tying all these dimensions together.

    2:09:42

    Prakash Narayanan: Indeed. Edward, thank you so much.

    2:09:45

    Nathan Labenz: Great to meet you, and don't be afraid to make me an offer.

    2:09:49

    Edward Hu: Likewise. Thank you guys so much for having me on.

    2:09:51

    Prakash Narayanan: Bye bye.

    2:09:52

    Nathan Labenz: Keep up the good work. Bye for now. I think he thought that offer might have to be a little more substantial than it might,

    • AI benchmarks: Mercor buys companies for real data

      0:00 / 0:00
    • AI reward hacking: How a dozen answers earn credit

      0:00 / 0:00
    • AI training: When human expertise stops helping

      0:00 / 0:00
    • AI vs accountants: What perfect test scores leave out

      0:00 / 0:00
    • Multi-agent AI: Why 4× compute can buy only 2× speed

      0:00 / 0:00
  4. 55:15Closing8 min
    Selling company data, slow diffusion, and approval fatigueThe hosts ask why more companies haven’t sold their data, discuss why AI will diffuse slowly in industries with little software, debate whether price and latency still bind, and consider regulation as an answer to approval fatigue.
    Open segment on YouTube ↗

    The closing stretch picked up the Mercor thread from the Edward Hu conversation. Nathan Labenz wondered why more struggling companies haven't sold their data, given that bankrupt Spirit Airlines' data reportedly went for about $10 million while experts command hundreds of dollars an hour. Prakash Narayanan offered three explanations: Spirit was early in the cycle, a bankruptcy sale is an exclusive license while a live company keeps producing data, and airlines are unusually software-saturated.

    Prakash argued that diffusion will lag in industries with little software penetration and a lot of tacit, field-based knowledge, such as factory inspection and analyst site visits, and he compared it to the slow reengineering of factories after electrification. He said a nation-state push to bring manufacturing back from China could speed things up. Nathan described the steam-driven factory at Greenfield Village in Detroit, and said he sees the barrier as a will issue more than a skill issue.

    Nathan said agents have crossed a threshold where he can hand off goals at a high level, citing a request for a full embeddings benchmark and report. Prakash countered that price and latency are still the binding constraints, especially outside the frontier labs, and that Anthropic's API is slower than he'd like for live production use. Nathan said he is instead struck by how fast and token-efficient Opus 5.5 has become.

    The two discussed a SemiAnalysis comparison of what a $200 monthly plan buys across Claude and OpenAI, with Prakash guessing such plans are loss leaders for rolling AI out inside companies. Nathan, citing Zvi, said they are all effectively token billionaires and that the limit is coming up with enough good ideas. They also noted the release of Haiku 5.5 during the show.

    They then turned to agent coordination, Apple's coming agent privacy sandboxing, and approval fatigue. Nathan likened himself to the drinking bird from The Simpsons rubber-stamping prompts, while Prakash argued that regulation, as with insurance and mortgages, is how society has handled consumer protection when attention fails. After a recap and a note that next week will include Monday shows, Prakash's AI-made song, with the working title "Take Off," played out the show.

    Timestamp links open the original source recording.

    Speed has a quality all its own.

    I do feel myself as kind of the birdie that just keeps pecking the enter key in response to a lot of these prompts for approval.

    Things are not gonna slow down.

    Lightly edited · timestamps jump to YouTube
    2:10:02

    Nathan Labenz: In fact, they have to be. I would be interested in how my data was being anonymized and cleaned up before shipping it out. But this has been a strange one to me, because the amount of money flying around is so substantial that I would just expect more of this to be happening. If bankrupt Spirit Airlines' data is worth $10 million, that's probably an awful lot of data. So maybe that suggests the value really isn't that high. Maybe that's one interpretation. And

    2:10:47

    If the clearing price is just kind of low, then that explains why I'm not getting these great offers. On the flip side, they're now hiring PhDs, experts, whatever. You have to pay a lawyer their hourly rate if you want these tasks done, and maybe a bit more than their hourly rate if they realize the role this is playing in the future. So it's hard to parse how it's $10 million for all of Spirit Airlines' data, but hundreds of dollars an hour for lawyers to do tasks. Where would I fall in that? And why haven't more organizations been tempted? There are a lot of companies running at pretty thin margins, or on the verge of bankruptcy.

    2:11:32

    All of a sudden, Mercor comes in and says, "We'll pay you. You're a 100-person firm, and we'll pay you the equivalent of five employees' salary for a while to get all your data." It's one thing to go to the team and ask how everybody feels about installing all this tracking stuff on their computers. It'd be quite another, I would think, in many cases, to say, "We might not make payroll next month unless we do this, so we really need to do it." That's obviously a delicate conversation too, but I'm surprised more isn't happening. I guess that's still my gut instinct right now.

    2:12:10

    Prakash Narayanan: I think one thing is that the Spirit Airlines data was probably early in the cycle. If you had another Spirit Airlines right now, it wouldn't cost as much, because some data is already in there. Earlier in the cycle, prices are probably higher than later. The second thing is that when you have a bankrupt company and you sell off the data, it's all of the data, and no one else gets a license on it. With a live company, the company keeps producing data and can keep reselling it, so you're not going to get an exclusive right to the data. The third thing is that airlines are relatively high-value digital tasks, because a lot of the service work is paperwork: regulatory, customer bookings, and so on. That's quite different from some other tasks, and it's also been more digitized over time. For example, with many physical tasks you might not find digitization of tacit knowledge inside a steel foundry. Some of it has been digitized, but they haven't had software. I feel like the penetration of software is quite high

    2:13:40

    in the airline industry, because you had all these booking systems and optimization. But penetration of software in a lot of manufacturing or construction is actually very low. You haven't had robotics, so it's still a printout, and then people go and carry it out. To the extent that something is penetrated by software and you can get all the data, that's useful, because then you can migrate a lot of it to AI. But if the industry has not been that penetrated by software, there's still a long way to go. I suspect that's part of the diffusion puzzle we have to live with: AI is great and can do everything a digital human can do, but there are a lot of things digital humans can't do. There's a lot of face-to-face stuff, walking around, field reporting. The classic field analyst at a bank flies out to a manufacturing company in the Midwest, looks around, sees some trash on the floor, reports back that the company is not that well maintained, and it gets a downgrade in the analyst reports. All of that is tacit knowledge built up from field reporting of various kinds by investors, journalists, regulators. The FDA sends staff to every manufacturing plant to do a walkaround, and they see where you're disposing of your trash, how the input is coming in, whether your people are wearing

    2:15:11

    gloves. They go into the bathrooms. Is there a place to wash hands? Workers are clean inside the cleanroom, but once they come out, what are they doing? There's all this tacit knowledge that goes into field inspections. I'm guessing that stuff is still hard: clipboards, someone ticking off checkboxes, visual inspection. And the people who do these inspections are very good at it, because if you've been doing it for ten years, you can walk around and say, "All right, I know exactly what's wrong with this place." I think that tacit knowledge is going to be very hard for AI to absorb.

    2:15:53

    Nathan Labenz: So it's going to take time. Although I will note that the robots probably don't have to wash their hands, because they probably don't have to use the bathroom. Some of these problems can just be short-circuited when you really redesign your process from the ground up with AI playing critical roles.

    2:16:14

    Prakash Narayanan: True enough.

    2:16:15

    Nathan Labenz: That has been in short supply.

    2:16:17

    Prakash Narayanan: True enough. But the reengineering process after electricity took a long time. They used to have a steam engine in the middle of the factory. The steam engine would fire off, and you'd have all these pulleys going across the factory. It took them a long time to realize, "Oh, we have electricity, just pull the lines in." In the beginning, they had the electricity coming to the place where the steam engine was sitting. So process transformation takes a long time, and I don't think we have that much time for diffusion. Stuff has to move a lot faster than waiting 20 years for factories to get redesigned. We do have one big advantage: we're also trying to pull manufacturing back from China, and that creates this kind of urgency

    2:17:07

    Nathan Labenz: Urgency

    2:17:09

    Prakash Narayanan: around improving and using the technology to produce cost-effective manufacturing in the United States. So we have one advantage, which is an industry- and nation-state-level push.

    2:17:24

    Nathan Labenz: If you ever find yourself with a reason to come to Detroit, Greenfield Village and the Henry Ford Museum of American Innovation, I believe is the official name, is an incredible visit to see the history of electrification and industrialization more broadly. One of the first buildings you can go into when you walk into Greenfield Village, all impeccably restored and maintained, is a traditional factory with a steam boiler down at one end, and it does what you say. It turns one crank that goes all the way across the top of this. You can see that even the layout of the factory is dictated by this mechanism: it's long and narrow, because there's just one drive shaft going down the whole thing, and off of that are all these pulleys that can be engaged locally to turn the different machines on and off. They still have all of that in working order today, which is pretty cool to see. It's one of my favorite places to walk around and ponder the future and transformation, because you just see these steps. There are a lot of little steps, and a lot of little inventions along the way. One of the interesting ones is these little paper disks that sit on the shaft, and people never know what they are. I asked, and it turned out that as the thing spins, they go back and forth and absorb excess oil and help distribute it, while making sure it's not dripping down on people or machines. It's just another tiny thing you don't think much about, but those things accumulate over time. Then when you change the paradigm entirely, you obviously don't need that thing anymore, but you need all these other little things that come into play.

    2:19:39

    It takes time. I still feel like it's mostly a skill issue, or better stated, a will issue, because those things don't seem that hard. Every time I encounter these things, the solutions seem pretty readily at hand. And the difference with AI is that the AI can find the solutions. I don't have to invent the protocol myself, and I don't even have to research it. I want to have my agent talk to your agent, talk to Evan's agent, talk to the agents of people I spoke to this weekend that I may be able to be helpful to. I just say to my agent, "Can you work out a plan? Give me a bunch of options for my review." And there you go. The self-deploying nature of this technology feels to me, in practical use, like it's hit a level that mostly solves these issues, if you're aware that it can and willing to give it a try. It's so confusing, but I feel like we might have just hit a threshold where deployment might really accelerate. We're probably not going to see what would happen to frontier model company revenue if they stopped releasing new models. But the ones we have have crossed such critical thresholds for ease of deployment and ability to help you rework your process that, as that dawns on people over the coming months, you might see a dramatic acceleration in use. Even at the beginning of this year, when I think back to January and getting serious about personal agent setup, it was a grind. There were still a lot of mistakes, and I was doing a lot of checking, getting AI to do some of the checking, but I felt I had to be the real owner of how all these processes were being designed. I was getting a lot of implementation help, but I wasn't able to hand off at a high level conceptually: "Here's kind of what I want, can you make it happen?" Now it feels much more like that, so I suspect we may have another one of these cases where people's impressions are a little outdated, and a bit of exposure to what it feels like now to ask for improvements or new workflows is such a game changer. Just in the last day, I asked, "Should we upgrade our embeddings?" We have this deep context database, and a bunch of content is embedded, which is helpful for search sometimes when keyword search doesn't work. Now I just say: go look at all the new embedding models; there was a new one out of the Gemma family that inspired this question. Read up on them, see how they compare, look at the prices and our options, make a benchmark based on our own use cases, test all these things, and give me a report. It can even do the computer use to sign up for new products I haven't used before. That level of prompt gets me that quality of quantitative result back, on which I can say, "Okay, cool, let's go this direction." That is a multiple faster than it would have been at the beginning of the year, when I felt I still had to be the project manager.

    2:24:05

    Prakash Narayanan: I'm guessing that when Sonnet or Luna get to the capability level of Astra or Fable, that's when things will start to happen, because I think the price and latency points have not been met yet. The capability threshold has been crossed, but the price and latency aren't there. I find it too slow to use Muse, for example. Muse was quite fast when they first launched; I tried it yesterday and it was too slow. So there's a problem with capacity, speed, latency, and other things. I think what's happening inside the frontier labs is that they're seeing the full effect, with the latency required to make it worth their while to run 20 agents and work in flow with them. Outside the labs, unless you're in an organization willing to spend something like $100 a month in expenses, you're not really seeing that yet. I think we will get there, but it's going to take something like 12 months. By that time, I don't know what the capability frontier on the largest models is going to be.

    2:25:25

    Nathan Labenz: Speed has a quality all its own. Honestly, though, I was thinking somewhat in the opposite direction this morning. I was feeling like, damn, these things are getting fast. Opus 5.5 is repeatedly returning reports and proposals where I think, "That was only two minutes. How did you do that so fast?" I think fanning out and doing research a bit more efficiently is one aspect of it. As we talked about with subagents, you can spend more tokens to get faster, and subagents definitely help with that, especially if things are parallelizable, like research. But I'm also finding the token efficiency to be really quite incredible. I've done a lot of stuff in just the last two days, and I checked my usage and thought, man, I've got to come up with more ideas. I've got a long way to go here. You can really burn a lot of tokens solving truly galaxy-brain problems, although, as you noted earlier, only about three hours for a lot of these math results. That's not that long. The simple stuff I'm doing over the last few days on my personal agent console, like "Can you add voice mode?" or "Can you add better search?", so that when I'm searching through old threads I've got different angles on search, would be features that take a person a week to map out, figure out what tools to use, and code up. They're minutes, oftentimes. It really is incredible how quickly you can ship this stuff.

    2:27:15

    Prakash Narayanan: One note on pricing. SemiAnalysis did a price comparison: the same $200 Claude plan is worth $2,500 to $12,500 in API value. They compared Sonnet 5.5, Opus 5.5, and Fable 5.1, using the maximum amounts of value you could get out of them. They also checked in on Astra. OpenAI recently changed the way they calculate limits, and after the recent changes, the maximum value you can get out of a $200-a-month plan is about $4,900. For GPT-6.1 Sol, it's now much lower. So you can see that the pricing has changed over time as these firms fight it out. Right now, the best dollar value you can get is the $200-a-month plan with Claude, using Claude's Sonnet 5.5. However, I will note that I have had hallucinations with Claude Sonnet 5.5, so I've never stooped below Opus 5.5. But that's SemiAnalysis's analysis on the pricing and the value from each plan. At this point, I believe the $200-a-month plan is kind of a loss leader for senior executives to implement at their company. You start off using it, it becomes so useful, you get caught up in this AI psychosis, and then you roll it out across the rest of the company over the next few months. That's my guess on why we're getting such high value from these $200-a-month plans right now.

    2:29:22

    Nathan Labenz: That's a pretty interesting theory. As Zvi said yesterday, we're all kind of token billionaires, in the sense that at $200 a month you get over $10,000 worth of API usage. With Opus 5.5 at $4 per million input tokens and $20 per million output tokens, if you call that, say, $10 blended, which is probably a little high but makes the math convenient, then you're looking at 1,000 million, or a billion, tokens for your $200 a month, which is pretty insane. I think Zvi put it really well when he said 5.5 should make you rethink your ambition. What do I need to do to be processing a billion tokens a month? Not everybody, obviously, but everybody who's gainfully employed in the United States can certainly afford that, and the limitation is much more often whether you can come up with enough good ideas to use your billion tokens. I'm trying to hold myself to a high standard in doing that, but I can't say I always max it out. I've maxed it out from time to time, but more often than not I'm probably leaving more value on the table than I should, because I'm failing to come up with ideas to really make use of a billion tokens.

    2:30:57

    Prakash Narayanan: I am waiting for an agent coordinator, something I can put all of my thoughts into that then manages all the various threads on Claude, ChatGPT, Muse, and all of these together. I think the firms haven't focused on it. I've managed to get Claude to do it by telling it that it has access to the Codex command line, and telling it how it can access Codex threads, because Codex stores them in a particular threading mechanism. But it's obviously very hacky and a little too tedious for anyone else to do. Once that is there, and this may require a device, this is where I think the OpenAI device is going to come in handy. I will also note that Apple is starting a new round of privacy and security features to cordon off your private material from agents and force you to consciously and positively accede to agents having access to your data. I think that's going to roll out in the next couple of months. It basically means locking down your laptop and sandboxing yourself against the agents, so without explicit, affirmative, and frequent confirmation that you want the agents to have access to this data, they won't. So I think that device isn't necessary.

    2:32:43

    Nathan Labenz: That's going to be interesting. It would probably be wise to use those tools and respect some of those safeguards. But as Evan said, how many people are just going to turn them right off and say, "Just have at it, I can't be bothered with all this"? I definitely feel a bit of that pull. There's nothing more frustrating than coming back hours later and finding you got stuck asking me to approve one stupid command that I wasn't really even going to read anyway. Behaviorally, I look at myself and wonder what percentage of these things I'm actually evaluating anymore versus rubber-stamping. What's the old Simpsons thing where he sets up the drinking bird to press enter on every prompt until finally the nuclear power plant has a meltdown? Hopefully we're not headed for that meltdown, but I do feel like the bird that keeps pecking the enter key in response to a lot of these approval prompts. That is definitely a very unsolved problem: how do we get people to pay attention when they should? It's very delicate, because I'm living proof that if you ask for permission too much, people get numb to it and blindly approve everything.

    2:34:15

    Prakash Narayanan: How we've traditionally done it is regulation. Not getting people to pay attention, but forcing the companies that deal with those people to provide things in an easily understandable manner, and if it's too complex, stepping in with regulation to decide how the framework should work. For example, no one reads their insurance contracts. Insurance contracts are basically decided by the state, which intervenes with laws, keeps the insurers well capitalized, and changes the rules from time to time. As a small insurance buyer, you can't go to an insurance company and say, "Let me negotiate the contract with you. You have this AI that's so smart, why don't you price every single clause, and then we agree on a different pricing for me?" We don't do that. Maybe that's the future, but typically, because as a consumer you don't have a lot of power against a large corporation, the state steps in and decides what the contract looks like, and then people don't have to care. I'm a finance professional, and I don't know what is in a regular mortgage. I have not read through all the laws that determine the mortgage that Fannie Mae or Freddie Mac would buy, and I don't really care either. You just look at the interest rate and go with it. That's what regulation does: it narrows what consumers see to an easily understandable thing you can accede to quickly without too much homework. It's the responsibility of the corporation, or the industry through standards, or the state through federal regulation, to say what is going to be offered to consumers and what is going to be easily understandable. I think those things will happen and will have to happen.

    2:36:25

    Nathan Labenz: Meanwhile, Haiku 5.5 was just released while we've been live as well. The cost curve continues to come down. It's pretty good and awfully cheap is the first impression I've taken away from their posts: 10 cents per million input tokens, and only a cent per million for cache reads.

    2:37:00

    Prakash Narayanan: I'll note that I have some queries running to Fable on the API for the show, and Anthropic's APIs are always slower than I would like. Latency is still an issue for the firm, and I think that's about the capacity they have. I would still be reluctant to rely on Anthropic in production, live. OpenAI's API team has done a much better job, and things are just faster and snappier; you can tell. Anthropic still needs to buy more capacity. I think they should be matching Google by the end of next year. They're on the ramp right now, but it's also $5 billion of revenue and $250 billion of commitments for next year.

    2:38:02

    Nathan Labenz: So pretty soon we're talking real compute.

    2:38:07

    Prakash Narayanan: Yeah. Pretty soon we're talking real money. A billion here, a billion there.

    2:38:13

    Nathan Labenz: Anything else for today?

    2:38:16

    Prakash Narayanan: No. We had an excellent show, with Evan and Edward from Mercor. I see a lot of new offers for data coming into Mercor, and I wonder how long this round of data collection is going to last. We'll be back next week. We're going to try to do a Monday next week, and also a Monday later. I think this week is basically a digestion week for people, because there were just too many things happening in the last couple of days, and everyone has to recalibrate. We had 20% of the top 4,000 problems getting solved this week. If you wait a couple more weeks, we might see another 20 or 30%. Things are not going to slow down.

    2:39:15

    Nathan Labenz: Do you want to play some accelerating music for us on the way out?

    2:39:23

    Prakash Narayanan: What should we play? Do you want to hear the...

    2:39:28

    Nathan Labenz: Let's see. Do you have a title for the song you played earlier?

    2:39:34

    Prakash Narayanan: No, I don't have a title yet. It's still labeled "Preview take 9, Take Off, Take Off, version 9." Let me...

    2:39:45

    Nathan Labenz: How about "Learn."

    2:39:48

    Prakash Narayanan: Yeah. The models, they just wanna learn.

    2:39:50

    [Clip] (Prakash's AI-made song, working title "Take Off" (preview take nine), begins and plays out the show. It samples AI-discourse lines such as "The models, they just wanna learn," "We have to slow down," and "Hard takeoff," looping and layering them.)

    2:42:22

    [Clip] (The song continues, cycling through more sampled quotes about slowing down, regulation, self-improving models, and the line between human and machine, until the broadcast ends.)