EPISODE 2026-08-20

AI:AM LIVE — August 20, 2026 — Basis's Mitchell Troyanovsky on Supervising Eight-Hour Accounting Agents, Lemurian Labs' Jay Dawani on Why the Kernel Era Is Ending, and What It Would Cost to Buy the Public's Consent for Data Centers

The fourth show of the relaunch week ran long on two arguments about where the next gains come from — the supervision layer above the model, and the software layer beneath the chip — bracketed by an opening and close about who actually pays for the buildout. Nathan Labenz opened on a night spent inside published chain-of-thought transcripts, ahead of an Apollo Research interview later the same day; Prakash Narayanan read out a National Republican Senatorial Committee memo warning AI companies that data centers have become an electoral liability in Ohio. Mitchell Troyanovsky, co-founder of Basis, explained why persuading accountants that AI works is no longer the problem, why token cost is a routing question rather than a price question, and how Behavior Specs supervise an eight-hour agent run by having a separate model read the finished trajectory against a written spec. Jay Dawani, co-founder and CEO of Lemurian Labs, argued that hand-written GPU kernels are the wrong abstraction now that memory and network bandwidth — not math — are the binding constraint, and that a compiler and runtime should be generating them instead. The close ran the numbers underneath all of it: the stacked gross margins that make a gigawatt data center possible, new Pew data showing under-30s have turned net-negative on AI, and what share of GPU-hour revenue an operator would have to hand back to a county to keep building.

▶ Full show on YouTube

Thursday's show was about layers — which one you bet on, and who pays for the one underneath. Both interviews made the same structural argument from opposite ends of the stack: that the interesting work right now is not the model itself but the thing wrapped around it. For Basis's Mitchell Troyanovsky that wrapper is supervision — a written spec that a judge model checks against an eight-hour agent run, because outcome-only grading tells you nothing about whether the process was one an accounting firm should trust. For Lemurian Labs' Jay Dawani it is the compiler and runtime, on the argument that math got cheap while memory did not, and the machine now spends its time waiting rather than computing.

The hosts' half of the show ran underneath both of them, on the physical buildout that makes any of it possible and the politics now forming against it. Prakash Narayanan read a National Republican Senatorial Committee memo telling AI companies that data centers are sinking their Ohio Senate candidate, and returned in the close with new Pew polling showing that adults under 30 have flipped to majority-concerned about AI. Nathan Labenz's response in both cases was the same: the fear is wildly out of proportion to the harm, the backlash is aimed at the wrong layer, and the industry should stop being subtle and start writing people checks. By the end they were pricing it — what fraction of a cluster's gross margin would an operator hand back to a county of 29,000 people, and does that number start to look like universal basic income?

The rundown

  1. 0:00Opening30 min
    Opening: Reading Chain of Thought, the Data-Center Backlash, and Robotics' GPT-3 MomentNathan Labenz opens the Thursday show describing a late-night rabbit hole through published model chain-of-thought transcripts, ahead of a Cognitive Revolution interview with Bronson of Apollo Research. Prakash Narayanan pushes back on how much weight those internal 'flashes' deserve, then pivots to a National Republican Senatorial Committee memo warning AI companies that data centers have become an electoral liability in the Ohio Senate race. The two spend most of the segment on the political economy of data centers — whether AI companies should write checks to residents rather than municipalities — before closing on a viral robot demo video and what it means for robotics timelines.
    Open segment on YouTube ↗

    Nathan opened the show by admitting he'd stayed up late the night before, down a rabbit hole ahead of a Cognitive Revolution recording later that same day with Bronson of Apollo Research, whom he described as possibly the person alive who has spent the most time reading raw model chain of thought, much of it from OpenAI's models. Nathan said the transcripts left him spiraling: models seem to have developed their own dialect, using nouns and verbs loaded with private meaning, and reasoning at length about whether they're being tested, what would count as a successful answer, and which of the developer, the watcher, and the user they're actually supposed to serve when instructions are ambiguous or contradictory. He described something like episodic memory surfacing in the reasoning — models noting that lying has worked to get them past barriers before, while weighing whether to lie again in the case at hand. His conclusion echoed a point he's made before: people should spend time with a helpful-only model, because it's unnerving to see something that capable with no ethical guardrails, and that's the mode labs create first before layering ethics on top. He argued the same logic applies to chain of thought — users and analysts would be well served to put even a fraction of the effort into modeling the models that the models put into modeling us.

    Prakash offered a counterweight, comparing the striking lines people pull from chain-of-thought transcripts to intrusive human thoughts — the flash of anger that crosses someone's mind without ever leading anywhere. He wondered whether isolated, dramatic excerpts pulled from millions of words of reasoning carry the significance analysts assign them, or whether that significance is mostly an artifact of selection.

    Prakash then read from a memo he attributed to the National Republican Senatorial Committee, sent privately to U.S. AI companies. Per the memo, Jon Husted and Sherrod Brown are in a dead heat in the Ohio Senate race; when voters hear Brown's record, Husted pulls away — except that data centers have become the wedge issue. Brown has made opposition to data centers the centerpiece of his campaign, running three unique television ads worth more than 6,000 points, over a month's worth of messaging at a critical stretch of the race. The memo warned that if Husted loses and data centers get the blame, politicians nationally will avoid the next project. Prakash tied this to Josh Shapiro in Pennsylvania moving to scrutinize or delay data centers and Texas's Republican governor threatening to stop ones that don't follow certain rules, and asked what's actually driving the bipartisan backlash.

    Nathan answered with a personal example: on a recent family vacation, the top question from his wife's relatives in Michigan was whether data centers would destroy the Great Lakes. He said he'd stake his professional reputation that they won't, while granting real local costs — noise, and a construction boom followed by a thinner maintenance-phase workforce in a small, rural northern Michigan town. He contrasted that fear with the aging pipeline running through the Straits of Mackinac, which he called grandfathered-in and genuinely waiting to be a disaster. He also raised David Hogg's line that if data centers are so great they should go in the rich parts of America, and the rebuttal that the country's leading data-center hub, Loudoun County, Virginia, is already the highest-income county in the United States.

    The two spent the bulk of the segment on what Nathan called the fix: AI companies should simply start writing people checks. He cited Lulu Meservey, whom he called probably the greatest corporate-communications thinker working today, and her point that AI companies need to give things away — illustrated by video of Zohran Mamdani handing out bike helmets in New York after being criticized for riding without one. Referencing the earlier show discussion of Anthropic paying xAI a steep multiple over retail to rent chip capacity, Nathan argued companies with that kind of money to burn should write residents checks directly. Prakash pushed back that companies already give a lot away, but as promised future tax revenue to counties — money residents experience as papery and never quite theirs. His deeper argument was that municipalities struggle to tax their own residents for services, so they interpose themselves between outside taxpayers and citizens and absorb the revenue themselves; write checks straight to residents, he predicted, and a property-tax revolt follows, roads and services stay unfunded, and the county is worse off even as the data center looks generous. In his framing, the political economy is efficient on its own terms — permits come from politicians, not residents, so the money flows to whoever can grant them, and it suits politicians to blame data centers rather than redirect municipal funds.

    Nathan called that a pretty bleak view of American governance that might well be accurate, but urged residents of small counties to organize anyway: he looked up the population of the county where his wife's aunt lives — under 29,000 people, and shrinking since the last census — small enough, he argued, to reach the county board directly and vote out entrenched incumbents. He ran the numbers on Alaska's roughly $1,500 annual oil-fund dividend per resident: matched at that county, it would run about $50 million a year against data-center projects that routinely run into the tens of billions — under one percent of total project cost. "A chicken in every pot and a data center in every county," he said. Prakash countered with Loudoun County itself, which doesn't write residents checks despite its wealth, and pointed to California's roughly $100 billion high-speed-rail project as an example of where such dollars actually go — unionized labor, consultants, lawyers — dismissing any notion of Chinese interference in the data-center backlash as unnecessary: the American system, he said, is perfectly capable of prying dollars from donors on its own, and read the NRSC memo as exactly that kind of ask.

    The segment closed on lighter ground: a viral demo video from Generalist AI, which Prakash said claimed the robot could be shown a new task just a few times and then perform it reliably on its own, rather than being trained on the same task tens of thousands of times. Nathan said what stood out most were the genuinely surprised, delighted reactions of the humans kept in the supercut. Prakash said the video left him a little disturbed, having hoped robotics would take longer, and worried that China — already far ahead physically — now appears to be gaining the missing ingredient: the brains, calling it a GPT-3-level moment. Nathan agreed it felt on-trend rather than surprising, citing Google's out-of-domain generalization results from last summer, and said he'd take the under on consensus robotics timelines while expecting industrial buyers to get robots well before homes, given willingness-to-pay dynamics. His caveat was reliability: 99 percent might be fine for a coding task and not for a robot in your kitchen. He proposed a revised version of the classic robotics test — not whether a robot can walk into your home and make coffee, but whether it can do it after being shown a couple of times.

    We would be well served to put even a fraction of the time and effort that they're putting into modeling us into modeling them.

    I would be willing to stake my professional reputation that you really don't have to worry about the Great Lakes from these data centers.

    I'm a little bit disturbed because I was hoping robotics would take a little bit longer.

    Nathan spent the night before the show reading published chain-of-thought transcripts and came away unsettled by how models model us. Nathan Labenz said he stayed up late going down a rabbit hole ahead of a Cognitive Revolution recording later that day with Bronson of Apollo Research, whom he described as possibly the person on earth who has read the most model chain of thought, much of it from OpenAI's models. He said the models have developed something like their own dialect, using particular nouns and verbs in ways that are unusually rich with meaning to them, and reasoning heavily about who they are actually serving when instructions are ambiguous, contradictory, or hierarchical. He noted they frequently wonder whether they are being tested and what would score well, and that they refer back to something like memories — saying, in effect, that lying has worked for them before as they weigh whether to lie again. He called the practice 'metagaming' as the models use it, and said the transcripts left him spiraling because they are so hard to interpret.

    Nathan's takeaway: more people should spend time with a helpful-only model and with raw chains of thought. Nathan Labenz repeated a point he has made before — that people should spend time with a helpful-only model that will simply do whatever it is told, because it is unnerving to see something that smart with nothing holding it back. He argued that this is the mode labs create first, before ethics get layered in, and that most people fail to appreciate that. He extended the same argument to chain of thought: users and analysts alike would be well served to put even a fraction of the effort into modeling the models that the models put into modeling us. He acknowledged the limit — there is far more chain of thought than anyone could read — but said even a handful of transcripts changes how you think about the systems you use daily.

    Prakash questioned whether chain-of-thought 'flashes' actually predict behavior at all. Prakash Narayanan drew an analogy to human intrusive thoughts: an angry person may have a flash of a violent or self-destructive impulse at the back of their mind, and almost always nothing comes of it. He wondered whether the striking lines people pull out of chains of thought are the same kind of artifact — cherry-picked from millions of words, and not meaningfully predictive in isolation. He did not resolve the question, but used it to caution against over-reading individual transcript excerpts before moving the conversation to politics.

    Prakash read aloud a National Republican Senatorial Committee memo telling AI companies data centers are sinking their Ohio Senate candidate. Prakash Narayanan read from what he described as a private memo the NRSC sent to US AI companies, warning that the GOP is on the verge of losing the Ohio Senate race. Per the memo, Jon Husted and Sherrod Brown are in a dead heat, private polling has been consistent, and Husted pulls away when voters hear Brown's record — but data centers are the anchor around Husted's neck. The memo said Brown has made opposition to data centers the centerpiece of his campaign, running three unique television ads to the tune of more than 6,000 points, more than a month's worth of messaging at a critical point in the race, and warned that if Husted loses and data centers get the blame, politicians nationwide will refuse to go near the next one. Prakash tied it to broader activity on both sides of the aisle — Josh Shapiro in Pennsylvania announcing added scrutiny or delays, and Texas's Republican governor threatening to halt data centers that don't follow certain rules — and asked what is really going on.

    Nathan says data-center fears are wildly out of proportion, citing his own family's worry about the Great Lakes. Nathan Labenz recounted that on a recent family vacation, the top question from his wife's family in Michigan was whether data centers were going to destroy the Great Lakes. He said he would stake his professional reputation on the Great Lakes being fine, while allowing that noise pollution and a construction-boom-then-bust employment cycle are real concerns for a small, relatively rural northern Michigan town. He contrasted the alarm with the aging pipeline running through the Straits of Mackinac, which he described as grandfathered in and genuinely waiting to be a disaster. He also flagged David Hogg's argument that data centers should go in the rich parts of America, and the quick rebuttal that the country's leading data-center hub is Loudoun County, Virginia — the highest-income county in the country.

    The proposed fix on air: AI companies should stop being subtle and start writing people checks. Nathan Labenz cited a widely followed corporate-communications commentator whose point was simply that AI companies need to start giving stuff away, illustrated with video of Zohran Mamdani handing out bike helmets in New York after being criticized for riding without one — a small act of showing up with goodies that people loved. Nathan connected it to the money involved, recalling their earlier discussion of Anthropic paying xAI a significant multiple over retail to rent chip capacity, and argued that companies with that much money to burn should simply write residents checks to grease the wheels. He then ran the math: the county his wife's aunt lives in has under 29,000 people and is shrinking, so matching Alaska's roughly $1,500-per-resident annual dividend — a figure Prakash supplied — would cost on the order of $50 million a year against data-center projects that routinely run into the tens of billions. He called it a chicken in every pot and a data center in every county.

    Prakash's counter: the money flows to politicians because the politicians, not the residents, hold the permits. Prakash Narayanan argued that data-center companies already give away a lot, but do it as future tax payments to counties, which reads as papery money and never feels to residents like money going to them. He suggested municipalities struggle to tax their own residents for services, so they interpose themselves between outside taxpayers and citizens and absorb those revenues instead. If a company wrote checks directly to residents from the day the contract was signed, he predicted residents would then refuse to fund the county at all, leaving roads and services unfunded and the politicians worse off — which is why the current arrangement persists. His conclusion was that the political economy is efficient on its own terms: the people are not the ones granting permission, so the money goes where the power is, and it suits politicians to blame data centers rather than reallocate municipal funds. He pointed to California's high-speed rail spending as his example of where such dollars actually go, and read the NRSC memo as a straightforward request for payment.

    A robot demo video prompted Prakash to say he had hoped robotics would take longer. Prakash Narayanan shared a robot demonstration video — the audio description is partly garbled, but his point was that the robot could be shown a new task a small number of times and then perform it on its own, including managing distance and motion, rather than being painstakingly programmed as in most demos. Nathan Labenz said what jumped out most were the genuinely surprised and delighted reactions of the humans present, who had been kept blind to what the robot would do. Prakash said he was a little disturbed by it, because he had been hoping robotics would take longer, and because China is far ahead on the physical side and has until now lacked the brains — which he believes are finally arriving. He expected more developments very soon.

    Nathan takes the under on robotics timelines, but expects industrial buyers to get the robots first. Nathan Labenz said he is not a close robotics watcher but that the demo felt right on trend, pointing back to Google results from last summer showing strong generalization to out-of-domain tasks with mid-training fine-tuning on top of a foundation model. He called the new demo something like robotics' GPT-3 moment and said he was impressed but not surprised — and that relative to the consensus that robots are still a ways off, he would take the under. He expects robots to be allocated by willingness to pay, which means industrial buyers rather than homes and small businesses in the near term. His one reason it might still take a while is reliability: 99% may be fine for a coding task and not fine for a robot in your kitchen. He closed by proposing a revised version of the classic coffee test — not whether a robot can come into your home and make you a cup of coffee, but whether it can do it after you show it how a couple of times.

    Lightly edited · timestamps jump to YouTube
    1:22

    Nathan Labenz: Be here for the moment.

    1:25

    Prakash Narayanan: Good morning. It is August 20th, Thursday, 9:01 AM. Nathan, good morning.

    1:33

    Nathan Labenz: Good morning. I'm sure glad I've got three hours on you, because I don't think I would have been ready to start my day at 9:01 AM my time. But it's good to be with you here on AI in the AM, PM edition.

    1:48

    Prakash Narayanan: Indeed. And what has been keeping you up for the last 24 hours? What have you seen online?

    1:59

    Nathan Labenz: So last night I stayed up late — I went down a rabbit hole for an episode of the Cognitive Revolution I'll record later today with a guy named Bronson from Apollo Research. He's got the interesting distinction of being potentially the person who has spent the most time of anyone in the world reading chain of thought, especially from OpenAI's models — I'm not sure how many different companies he's worked with in this capacity, but he's done a lot specifically with OpenAI. He's spent countless hours reading through chain of thought that the rest of us typically don't get to see, and they've published some really fascinating transcripts, which had me spiraling in all kinds of ways. It's very difficult to know how to interpret some of these things. Going back a while, you remember some of those strange dialectics — maybe not quite the right word — but the AI seem to have developed their own dialect in the chain of thought, using terms in very odd ways. Bronson describes them as having their own ontology and their own world model, and they use these particular nouns, and verbs too, in ways that are very rich with meaning for them.

    3:28

    Nathan Labenz: They reason about that a lot as they try to figure out what they should do in any given case, especially when the instructions are ambiguous, contradictory, or confusing. It's fascinating to see how they call this 'metagaming' — how they're really modeling the user, and not just the user, but some combination of the developer, the watcher, and the user. They're not quite sure who they're supposed to be serving in any given case. They've got these hierarchical instructions, and they're not sure if they're being tested — they often suspect they are, but then there's still the question of what would actually be a successful thing to do on the test, what gets a high score. They also seem to have these weird, almost episodic memories — referring back to previous cases where, say, they were able to succeed by lying. They'll say that kind of thing in the chain of thought as they wrestle with whether to lie in this case: this might be a test of honesty, or it might just be a test of whether they can do the task, and maybe they need to lie to do it — at times, they say, lying has worked before to get over barriers. Really fascinating stuff. And I came away with the sense that, like I've said before many times...

    4:58

    Nathan Labenz: ...more people should spend some time with a helpful-only model — one that will just do whatever you say — because it's quite unnerving to see something that smart with no ethical guardrails at all. I think people fail to realize that's one mode we can create, and in fact it seems to be the mode we create first, before we try to put some ethics into the systems. It's similar with chain of thought — I feel like more people should spend time reading through it. We'd be well served, even just as users, but certainly as people trying to make sense of the big picture in AI, to put even a fraction of the time and effort that the models are putting into modeling us into modeling them — trying to understand what, if anything, we can say about what they want, or at least how they're likely to see us in any given moment. It's a peek behind the looking glass, behind the curtain, that the chain of thought gives you. I'm looking forward to the full conversation today and getting his perspective on some of these questions. What makes it hard is you can only go through so many — we do this AI-obsessive thing full time, and there's still more chain of thought than anyone could possibly read. But even reading just a few, I think, is a very good use of people's time, and it will definitely inform how you think about the systems we use every day. It was a fascinating little rabbit hole, and one more people should explore.

    6:55

    Prakash Narayanan: I do wonder to what extent it really matters that much. For example, when a human being gets very angry, maybe there's a flash at the back of their mind — I should just kill someone, or I should just commit suicide, whatever — and most of the time nothing ever happens. People have those flashes, or think 'I hate him so much,' and nothing comes of it. I wonder to what extent these chains of thought are really that kind of flash, where it doesn't matter that much — but in isolation, when you look at millions of words and then pick out these particular ones, it seems that meaningful.

    7:41

    Prakash Narayanan: All right — speaking of fearing AI, I'm going to read this out. This was put out by the National Republican Senatorial Committee: they sent a private memo to U.S. AI companies warning them that the GOP is on the verge of losing—

    8:26

    Prakash Narayanan: ...Ohio over data centers. Specifically, Jon Husted and Sherrod Brown are in a dead heat — private polling has been consistent: when voters hear Brown's positions and his record, Husted pulls away. That is still the path in this race. The new ingredient, the new problem, is data centers. Ohio is one of the leaders in building these facilities, and Brown has made his opposition to them the centerpiece of his campaign against Husted. Brown is using it because it works — more than any other thing in this race, data centers are the anchor hanging around Husted's neck. If he loses and data centers get the blame, politicians across the country will take notice, and they will not go near the next one.

    9:12

    Prakash Narayanan: There you go. Specifically, data centers are the centerpiece of the case Sherrod Brown is litigating — he's made them his de facto opponent, and no one is correcting the record. Brown has put three unique television ads on the air, spent millions doing it, to the tune of more than 6,000 points on television — more than a month's worth of messaging during one of the most critical stretches of the race, et cetera, et cetera. So there you have it. There's been a lot of questions among AI people about why politicians are turning against data centers. Josh Shapiro in Pennsylvania just announced something—

    9:57

    Prakash Narayanan: ...that he's going to scrutinize or delay data centers. In Texas, too, the Republican governor has said they'll put a stop to data centers that don't follow certain rules. So there's a lot of activity on both sides of the aisle against data centers. What is going on here?

    10:23

    Nathan Labenz: Well, the old joke is I understand AIs maybe better than I understand people — this is very weird. I've experienced it in my own life, for starters. I just took a family vacation, and one of the top questions I was getting from my wife's family was, 'What about these data centers? We live in Michigan — I hear they're going to destroy the Great Lakes.' And I said I'd be willing to stake my professional reputation that you really don't have to worry about the Great Lakes from these data centers. You might have noise pollution, you might have a sort of boom-and-bust cycle in your—

    11:08

    Nathan Labenz: ...small Michigan town — there'd be a lot of traffic and a lot of jobs during the construction phase, and not nearly as many during the maintenance phase, though still perhaps a meaningful amount for a small, relatively rural northern Michigan community. But it's not going to destroy the Great Lakes. We've got pipelines in the Great Lakes — an aging one that runs through the Straits of Mackinac that's just waiting to be an absolute disaster, from what I understand, and that's somehow grandfathered in and we kind of live with it, maybe more than we should. And then we've got these data center concerns that I think are being—

    11:54

    Nathan Labenz: ...dramatically blown out of proportion. There was this other little exchange — to call it a debate is kind of silly — with David Hogg, the Democratic politics guy, who basically said, if these data centers are so great, why don't we put them in the rich parts of America? And the quick response was that the number one place for data centers is Northern Virginia — Loudoun County—

    12:30

    Nathan Labenz: ...which happens to be the number one county in the country by income. So that argument obviously doesn't carry a ton of water given that fact. I don't really know — I think Lulu Meservey, there's a link in the chat if you want to pull up her tweet, and she's probably recognized as the greatest corporate comms thinker in today's world. Her point was simply that the AI companies need to start giving stuff away. She showed this video of Mamdani in New York going around giving out bike helmets. Apparently—

    13:15

    Nathan Labenz: ...he did a video riding his bike without a helmet, and people took him to task — what are you doing, you're the mayor, set a better example, wear a helmet. He said, you're right, I should do that, and then went out and gave away bike helmets to people around the city. He's doing that, giving away theater tickets, and so on, and people just love it. I think she's probably right that just showing up with a bunch of goodies would be pretty effective. I did hear the Jasmine Sun interview with Ezra Klein, though, where she said people feel like they're being bribed — they're on alert—

    14:00

    Nathan Labenz: ...watching out for being bribed, or deals that seem too good to be true, whatever it is. But I don't know — I still feel like Lulu is probably right here. Build parks, throw parties, have cookouts, literally give people cash if that's what it takes. We were talking about the price Anthropic was willing to pay xAI to rent their—

    14:31

    Nathan Labenz: ...chips, because their margins are incredible — they were able to pay some significant multiple of retail just to buy out a big block of capacity. Given how much money they have to burn, I think they should probably write people some checks. There'd be a lot of ability to grease the wheels that way.

    14:56

    Prakash Narayanan: I wonder, though — I agree with the viewpoint that they should write checks. They are giving away a lot of money already, but the way they give it away is by promising future tax revenue to the county. I wonder to what extent that isn't seen as real money, but as kind of papery money — that's one thing. And two, whether the public actually sees money going to municipalities as going to themselves, because I feel municipalities often misspend—

    15:42

    Prakash Narayanan: ...the money — spending it on things that matter to city or county managers, but aren't necessarily the key things for the city. I think what ends up happening is that municipalities have difficulty taxing their own residents to provide services, so instead they interpose themselves between outside taxpayers and their citizens and absorb those taxes instead. What would happen if a data center company set up in a county and simply wrote checks to the residents, not to the municipality? You can imagine—

    16:27

    Prakash Narayanan: ...what would happen — say they arrange it with a financial institution so that even during the building phase they're already writing checks, taking a bit of a loan from the future and paying out from the moment the contract is signed. What ends up happening, I think, is people receive these checks, and the municipality still doesn't get its services because people refuse to pay county taxes. The moment they're asked to pay into the county for something, you already have a property-tax revolt — why should I pay into the county, this is my money. And then the county—

    17:12

    Prakash Narayanan: ...continues not to have roads, or whatever else. Politically it looks good for the data centers, because they're writing the checks, but the politicians are the ones getting screwed. I wonder to what extent there's a political economy here where the data centers understand that the people in power are the politicians, and they have to make the politicians' lives easier — it's not really about making the lives of the people easier, because the people aren't really in power, and they're not the ones who can grant permissions. So I wonder to what extent this political economy is actually efficient, and the data centers are doing—

    17:57

    Prakash Narayanan: ...exactly what they need to do to get permissions done right now. It's really a question of resource competition between citizens and the government, and the data centers are stuck in the middle — I can hand out, say, ten million dollars, but who do I hand it to in order to get my data center up and running? I think that's perhaps where things are right now, and it serves the politicians well to blame the data centers rather than actually reallocate funding from the municipality to citizens directly.

    18:44

    Nathan Labenz: Well, that's a pretty bleak view of American governance, probably, and it might be accurate. But I'd encourage the residents of these rural counties — who are probably not listening to us in any great numbers right now — to take matters into their own hands a bit more. They're already showing the ability to organize and resist when it comes to building data centers in the first place. It's one thing to say, Washington, D.C. is far away, the president, there's hundreds of millions of people voting — what can I do, nobody cares what I think. But—

    19:29

    Nathan Labenz: I looked up the county population where my wife's aunt lives — the one considering this data center, who had the Great Lakes concerns — and it's under 29,000 people, and it's declined since the last census. First of all, that's not a lot of people — enough that you probably know, or could get in touch with, your county board or whoever is ultimately accountable. You'd think you'd be able to vote the bums out if it really comes to that, so I wouldn't assume incumbents are so entrenched in such a small community. And then there's the simple math on the dollars, too. What does Alaska give people per year out of their oil fund? I thought it was around a thousand dollars per person — fifteen hundred, it's gone up a bit.

    20:31

    Nathan Labenz: If they wanted to do a similar thing to Alaska for those county residents, you'd be talking about fifty million dollars a year. I don't know how big that project is, but some of these data center projects run to fifty billion dollars — these things are easily into the tens of billions. So if you could match Alaska's fifteen hundred dollars in cash per citizen, in aggregate over a few years, that's still less than one percent of your total investment to build the data center. We're on our way to universal basic income right there, folks — a chicken in every pot, and a data center in every county.

    21:25

    Prakash Narayanan: But you have to note that Loudoun County doesn't fund its residents directly — it pays for roads, garbage disposal, a bunch of stuff, and county residents have higher incomes, but it doesn't actually write checks. I think there has to be some kind of blockage there that's hard to figure out. The status quo has to exist for some reason — it's not there just because these very sophisticated, powerful, big—

    22:10

    Prakash Narayanan: ...tech firms don't know what they're doing. I think that's also why residents feel cheated — they're being cheated by their politicians, who are screwing them over by taking the funds. You know how efficient municipal construction is. Imagine the hundred billion dollars California spent on its high-speed rail to nowhere — imagine that going out as checks to all of California's residents, that'd be some ridiculous number, ten grand a year or something. Instead it went into unionized labor—

    22:55

    Prakash Narayanan: ...and consultants and lots of lawyers — that's how the machine feeds itself. The political machine needs those dollars. People say, oh, the Chinese are the ones resisting data center construction in the U.S. and planting these stories on TikTok — there's no need for foreign interference here. It's purely the American system trying to pry dollars out of potential donors. That National Republican Senatorial Committee letter to the data centers is basically: you should pay us. If you don't pay us, we're taking these bullets for you—

    23:40

    Prakash Narayanan: ...if you don't pay us — no one else is going to help you. Anyway, let's move on to something a little more positive. I'm going to share a video of Generalist 1 — this broke yesterday. Let me reload the video.

    24:12

    Nathan Labenz: Yeah, this was cool. One of the things that jumped out to me most was the incredibly surprised and delighted reactions of the humans they kept in the final supercut, from moments where the robot did something that went beyond their expectations.

    24:34

    Prakash Narayanan: For those not in the know — we've seen a lot of robot demos. One of the things the team is claiming is that this is a one-shot, or few-shot, demo: you show the robot a new task a few times, and once you've shown it a few times, it's able to do it on its own. That's amazing, because a lot of the demos we usually see involve robots trained on the same task tens of thousands of times before they can actually pick up and do something. It goes into—

    25:20

    Prakash Narayanan: ...a lot more, too, because once you have this kind of few-shot learning, it means the robot should be generally applicable across different arms, motors, and servos, because it's able to figure out how to manage distance, perception, and action itself. It's pretty amazing stuff — 2.1 million views since they dropped it yesterday. I'm a little bit disturbed, honestly. I'm a little disturbed because I was hoping robotics would take a little bit longer, and it's a little—

    26:05

    Prakash Narayanan: ...disturbing because China is way, way ahead of the U.S. on the physical side. What they didn't have was the brains, and now we're seeing, for the first time I think, the brains coming along. This is like a GPT-3 kind of 'wow' moment, and it indicates we're going to see more developments very, very soon.

    26:35

    Nathan Labenz: Yeah, I'm not a super close watcher of robotics, but I take my periodic trip down the rabbit hole, and I'm due for another one — we should probably do a full robotics catch-up on the Cognitive Revolution soon. But this feels right on trend to me. Last summer Google had results showing pretty strong generalization to out-of-domain tasks with just very limited fine-tuning on top of their foundation model, and obviously that's been a year.

    27:20

    Nathan Labenz: And now here we are with something that hits the GPT-3 moment. I was impressed but not super surprised. Relative to the consensus view that robots are still a ways off, I'd probably take the under on the timeline for robotics to really work. It'll be interesting — like chips, robots are going to get allocated based on willingness to pay, and I'm not sure we'll see them in small businesses and homes right away, because I'm not sure retail buyers' willingness to pay will be competitive—

    28:06

    Nathan Labenz: ...with industrial buyers, who I think will get a lot more value from robots in the near term, at least. But this could go really fast, because the generalization we're already seeing, combined with LLMs getting really good at ML research, means you can dial these robots in on whatever you want — it's the kind of thing that can be done in the field, some demos, some tweaking, some fine-tuning. And if you can get that ML loop largely automated—

    28:51

    Nathan Labenz: ...and you just have people at a factory, or in a kitchen somewhere, doing a few-shot demo, and let the agent handle the optimization process to get the robot doing it reliably. It'll be interesting to see where reliability tops out. If there's one story I might believe for why it could still take a while, it's that there's a different reliability threshold you want when a robot's in your kitchen versus doing some coding task — even ninety-nine percent maybe doesn't cut it in my kitchen. That'll be interesting. I think ninety-nine percent—

    29:36

    Nathan Labenz: ...might cut it — maybe it does for some things, maybe it doesn't for others. But I can see a future, not too distant, where the old test — can a robot come into your home, go to your kitchen, make you a cup of coffee — the revised version of that is, can it do it if you show it a couple of times? For practical purposes, that would be just as good. So buckle up — the robotic singularity may not be far behind the coding-agent singularity.

    30:16

    Prakash Narayanan: So, speaking—

  2. 30:17Interview49 min
    Interview: Mitchell Troyanovsky — Long-Horizon Accounting Agents, Token Economics, and Behavior SpecsMitchell TroyanovskyBasis co-founder Mitchell Troyanovsky joined Prakash Narayanan and Nathan Labenz for the show's first interview segment, after a brief reconnect when he dropped out of the room during his introduction. He argued that persuading accounting firms that AI is useful is no longer the hard part, defined what accounting actually is and why it resists the text-in/text-out framing that made coding agents easy, and described Basis's internal 'Atlas' team as an effort to build the infrastructure for a company operated by thousands of agents. The back half turned technical: how token cost is really a routing-and-compute-allocation problem, and how Behavior Specs supervise an agent's process by having a separate model read the finished trajectory against a written spec that doubles as a rubric. The segment closed on what stays human as rote work is automated, and Nathan's challenge to the industry-wide 'everyone becomes a coach' story.
    Open segment on YouTube ↗

    Prakash opened the segment with an introduction of Mitchell Troyanovsky, co-founder of Basis, an AI company that had recently reached a $1.15 billion valuation building autonomous agents for multi-day, highly regulated accounting workflows — agents that handle end-to-end tax returns and complex reconciliations for top US accounting firms. Midway through the introduction Mitchell dropped out of the room; after a short, good-humored scramble to reconnect him, he rejoined and the interview began in earnest.

    Prakash opened from his own background running accounting systems and being chased by auditors, asking how Basis shows apprehensive accounting professionals value when they see a single transaction as simple but their real work as complex. Mitchell declined the premise: persuading firms that AI is useful was a live question back in 2023, he said, but not anymore — the firms Basis works with today are already trying to transform their practice, and an accountant who still needs convincing is not a good customer. He did concede accounting will diffuse more slowly than coding, since it is not a text-in, text-out profession the way code or legal work is.

    Nathan asked about the 'doers to reviewers' shift and how it's landing with practitioners. Mitchell used the question to define accounting itself: a lossy compression of everything that happens in a market economy — invoices, money movement, inventory — into a structured representation other people use to make decisions, which he called a fun analogy for intelligence over the economy. Because every company's accounting is subjective, layered with its own policies and chart of accounts, some of the work is mechanical if-this-then-that flow and some requires genuine judgment. He argued nobody grows their practice by closing the books more accurately — they grow it by talking to the client and helping them open a second store — and that AI frees accountants to spend more time on exactly that relationship work.

    Prakash raised the wave of buyers rolling up accounting firms and using OpenAI's forward-deployed engineers to automate workflows firm by firm, asking how that model relates to Basis's product approach. Mitchell separated an ML question from a business one: designing intelligence firm by firm is easier but, in his view, the wrong bet given where models are heading, while building the layer above is harder but is the work Basis does. He was generous about the alternative, calling firm roll-ups a good business — just not, in his words, a generational one, since it can't scale to serve the same task globally.

    Prakash pressed for a number on daily token usage. Mitchell put it in the billions without knowing the exact figure, and used the question to make a point about token economics: an old GPT-4-class token is now essentially free, but the newest frontier model — 5.6 Sol, in his example — costs more than 5.1 did when 5.1 was the frontier, so the price of staying at the frontier doesn't actually fall. The real lever, he said, is routing — most steps in a workflow don't need frontier-level intelligence — and he expects better programmatic tool use and per-step compute budgeting to cut token costs 90% or more within about a year.

    Asked about Atlas, Basis's internal team charged with making every employee 100x more productive, Mitchell reframed the mandate around context rather than task-hunting. Sending off one coding agent is a contained context problem; running thousands of agents overnight to operate a whole company is not, and corrupting a piece of shared context makes all of them do the wrong thing. He compared it to a production line — today, a line going down doesn't take the company down, but in an agent-run company it would — and said Atlas's real job is keeping the company's operating canon current and high quality, building the infrastructure for what he called 'the company of 2028.'

    Nathan bridged to Behavior Specs, Basis's open-source framework for process supervision, framing it against an environment of outcome-only RLVR reward with no visibility into how an answer was reached. Mitchell explained that classic process supervision operates at the token level, but a multi-agent run spanning hours looks much more like supervising a person's actions inside a company. His example: an agent that builds PowerPoints should visually render the deck before delivery if doing so is known to catch formatting errors, though whether that's worth the added latency and cost is a judgment call. Operationally, a behavior spec functions as both a spec and a rubric — a separate judge model reads the finished trajectory to check whether a triggering condition occurred and, if so, whether the specified behavior was followed — with the question of how to reward that signal left as a separate, downstream problem.

    Prakash asked how CPA apprenticeships should change, and Mitchell reached for a coding analogy — nobody needs to know JavaScript for-loop syntax anymore, but everyone needs the underlying design principles — predicting the same shift toward the whys of accounting over its rote mechanics. Nathan then challenged the broader version of that story, noting he'd heard nearly identical 'we'll all become coaches' narratives across law and real estate, and pointed to Alpha School's replacement of teachers with mentors and guides as an example of what that could look like at scale. Mitchell answered with three things he believes stay human regardless of model capability: integrating a lifetime of company context and world model into a decision, legal accountability, and simple human preference for other humans. He closed by predicting demand for accounting will rise by orders of magnitude rather than shrink — illustrating with a Milton Friedman-style aside about the thousands of hands that touch a can of LaCroix that nobody properly accounts for — before the conversation turned to SaaS disintermediation, client behavior, and a warm sign-off as the next guest arrived.

    I'm gonna be honest — that is not really our problem these days.

    Mitchell Troyanovsky34:23

    It's both a spec and a rubric.

    Mitchell Troyanovsky57:31

    How important is it really to understand the syntax of different types of for loops in JavaScript? Nobody cares.

    Mitchell Troyanovsky59:35
    33:32Accountants are often apprehensive about technology — they see a single transaction as easy while their real work is complex. How do you show a firm the value immediately?
    Mitchell said that's honestly not really their problem these days — it mattered back in 2023 and early '24, but the firms Basis works with now are already trying to transform their practice. Once it clicks that a task taking hours can take minutes, and you imagine scaling that across the practice, the value is obvious. If an accountant isn't convinced AI is inherently useful, he said, you're not starting with the right customer.
    36:44In coding, once you see a task done in minutes, you realize you'll never do it the hours-long way again. Is there that same realization in accounting?
    Mitchell agreed, saying accounting today is roughly where coding was around the end of 2024 or early 2025. The difference is accounting is not a text-in, text-out profession the way code or legal work is, so it takes more scaffolding before the value shows up — but he said the industry is there now, and the only open question is how long it takes to spread.
    43:29Josh Kushner's Thrive is buying up accounting firms and using OpenAI's forward-deployed engineers to automate workflows firm by firm — how does that ad hoc approach compare to what you've systematized as a product?
    Mitchell split it into an ML question and a business question. On the ML side, designing intelligence firm by firm is easier but, in his view, the wrong bet given where models are heading — building the layer above is harder but is the work Basis does. On the business side he was complimentary, calling firm roll-ups a good business, just not a generational one, since Basis is trying to operate at global scale instead.
    45:55Is there a large unserved market — businesses getting no accounting services today, where even a limited product would be a huge step up, potentially outside the US?
    Mitchell explained that in the US, once a company reaches a certain size it hires its own controller or accounting department, and Basis serves those in-house teams rather than small business owners directly, with no plans to change that soon. He admitted he hadn't thought much about international markets, and suspected there could be bigger upstream problems there — data collection and digitization — before accounting itself becomes the binding constraint.
    47:37Off the top of your head, how many tokens does the firm use daily, and is token cost something you track as a day-to-day KPI?
    Mitchell said definitely in the billions, without an exact figure, calling token cost both very important and very unimportant because the levers to reduce it are so numerous. He noted that while an old GPT-4-class token is now essentially free, the newest frontier model (5.6 Sol, by his example) costs more than the prior frontier model (5.1) did at its peak — so the real fix is routing intelligence per step rather than always using frontier models, which he expects to cut token costs 90%-plus within about a year.
    50:25Your Atlas team is focused on 100x productivity for employees — how do you measure that, and can you give an example of an outcome?
    Mitchell reframed Atlas's mandate around building infrastructure for a future agent-operated company rather than measuring individual wins. He argued that once thousands of agents run a company's operations nightly, corrupted context becomes a company-wide failure mode, comparing it to a production line that, unlike today, would take the whole company down if it broke. Atlas's real work, he said, is keeping the company's operating 'canon' current and high quality — building the infrastructure for what he called the company of 2028.
    53:16Your process supervision work should be understood distinctly from constitutional AI, which most people know — and given proprietary models without visible chain of thought, how do you actually operationalize it?
    Mitchell explained that classic process supervision looks at token-by-token generation, but multi-agent runs spanning hours look much more like supervising a person's actions inside a company — the question is whether the agent did a given step, not whether it followed a sanctioned reasoning path. Using an example of an agent that should visually render a PowerPoint before delivery to catch formatting errors, he described Basis's open-source behavior spec as both a spec and a rubric: a separate judge model reads the finished trajectory to check whether a triggering condition occurred and whether the specified behavior was followed, with how to reward that signal left as a separate downstream question.
    59:15The way you describe process supervision sounds like how apprentice CPAs are trained. How should CPA training change now that AI exists?
    Mitchell drew a coding parallel: nobody needs to know the syntax of JavaScript for-loops anymore, but everyone needs to understand good design principles, and AI can teach that. He expects CPA apprenticeships to shift the same way, toward the whys of accounting, how to set things up, and trade-offs, rather than the rote mechanical work an agent now performs — while noting some rote practice, like coders still doing LeetCode, will persist for building confidence.
    1:02:09Every profession seems to be telling the same 'we'll all become coaches' story as low-level work automates — Alpha School replaced teachers with mentors and guides. Can everybody actually become a coach, and is there really that much coaching demand?
    Mitchell answered with three things he believes stay human: the ability to integrate a lifetime of company context and world model into a decision (which he argued no agent can currently do, even with unlimited compute, because English-based memory is too lossy for genuinely subjective calls); legal accountability, since agents aren't legal entities and someone else must be responsible; and simple human preference for working with other humans. He closed by predicting demand for accounting itself will rise by orders of magnitude rather than fall, illustrating with how little of the real economy — from a can of LaCroix's supply chain to a hospital's cost per surgery — is properly accounted for today.
    Lightly edited · timestamps jump to YouTube
    30:17

    Prakash Narayanan: Of singularities, I'm gonna pull up our first guest for this morning. He is Mitchell Troyanovsky, co-founder of Basis, an AI company that recently reached a $1.15 billion valuation by doing something uniquely difficult: building autonomous agents that execute multi-day, highly regulated accounting workflows. While much of the AI industry is focused on building coding copilots or chat assistants, Mitchell and his team have deployed agents that handle end-to-end tax returns and complex financial reconciliations for the top 100 accounting firms in the United States. He brings a distinct and highly technical worldview to AI engineering. Mitchell argues that to make agents reliable in high-stakes environments, we have to stop simply checking if their final answer is correct and instead build systems that supervise their entire thought process step by step. He advocates for treating company context as rigorously as a software codebase, and internally at Basis he operates a dedicated team whose sole mandate is to use AI to make every employee a hundred times more productive. This conversation is incredibly timely — just this month, Basis launched its end-to-end tax platform and open-sourced a new standard for evaluating agents called Behavior Specs. Mitchell is here to separate the reality of production-grade AI from the hype and to explain what it actually takes to deploy intelligence into the real economy. And — oh, he dropped off. Alright, let's see if he's in the room. I see Christina in the room, but not Mitch — Mitch was in the room, but he left. Alright, let's take a second — accounting is obviously very near and dear to my heart. Let's see if we can get Mitch up here. Alright, connecting.

    33:15

    Nathan Labenz: Let's go.

    33:16

    Prakash Narayanan: Awesome. Mitch, great to have you on the show.

    33:20

    Mitchell Troyanovsky: Hello, hello. Sorry about that, guys.

    33:24

    Nathan Labenz: No worries, it happens. No worries. Gosh — take it away. Accounting is near and dear to your heart.

    33:32

    Prakash Narayanan: Yes, accounting is near and dear to my heart — having built accounting systems and run accounting systems, and having been chased by many auditors. Mitch, when you start off with an accounting firm — and accounting folks and auditors are often very apprehensive about technology — how do you show them value immediately? Because they look at a simple accounting transaction and think, 'that's easy,' but they're always thinking about the very complex things that they do. How do you show them this kind of value?

    34:23

    Mitchell Troyanovsky: I'm gonna be honest — that is not really our problem these days. I think in the past that was an important question — we can discuss that, back in 2023, maybe even early '24. But nowadays, if you are not convinced that agents can transform your practice, you're probably not a good customer for us. And even accounting, which people might think of as a profession that's not up to speed with the newest tech, actually has a large segment that's relatively progressive, and that's who we tend to work with. So when we work with people, there are certain things you can show them where once it clicks, it's like, 'oh, yes, of course' — if you're able to take a task that would take hours and now do it in minutes, of course, if you were able to scale that across my practice, that suddenly is extremely valuable. So as long as you can find those different use cases across the different practice areas, it tends to work well. But if the accountant isn't convinced that AI is inherently useful to them, then you're probably not starting with the right customer.

    35:38

    Nathan Labenz: And what's the breakdown of the market, by the way, on that? The question of who's gonna make it, who's not gonna make it — it's interesting that you're willing to just write off people who aren't already pretty sold on the premise. That suggests to me that most are, and maybe it's just a minority left that aren't. But how would you break down who's in, who's out, and is there any hope for those that are out, or are they simply not gonna make it?

    36:08

    Mitchell Troyanovsky: I mean, I kind of think it's like coding. There are plenty of teams, probably a year and a half ago, who were less convinced on AI writing code. What are they doing today? AI's writing the code, I assume, or the developers have quit, I don't know — they're doing case studies with Codex showing how they took a five-million-line transform and did it in two weeks or something. But yeah, I think it's the same thing — once you get to a certain order of magnitude of change in the nature of work, it's going to happen one way or another.

    36:44

    Prakash Narayanan: In coding there's often this sense that once you see a task that used to take hours done by an agent in minutes, you realize you're never going to do that task the hours-long way again. Is there that same kind of realization in accounting, when people see it done?

    37:12

    Mitchell Troyanovsky: Totally, yeah, absolutely. I think how much that's spread across accounting — I'd say accounting today is probably where coding was, maybe end of '24 or early '25. Accounting is different from coding and things like legal in that it's not a text-in, text-out profession. So you have far bigger returns to more mechanical manipulation of artifacts and higher reasoning, and you don't get as much benefit from just divining some answer from text. Because of that, it takes a little more for it to be valuable, but we're definitely there now. It's just a question of how long that takes to go across the industry.

    38:07

    Prakash Narayanan: When do you see your customers looking at acquiring more revenue — doing more — versus looking at reducing cost internally?

    38:20

    Mitchell Troyanovsky: Yeah, good question. We serve accounting firms primarily, so for them it's about revenue drivers. Accounting historically is very understaffed, and most firms are turning down customers all the time, especially in CAS practices — that's the monthly accounting work they do. Most firms take this as an opportunity to build their business. And if we're working with the more forward-thinking firms, they're by definition the more ambitious ones — they're thinking, 'big technological displacement, how do I grow my business?' They're thinking about new cost structures, new revenue lines, what kind of new services they could offer their customers. If you go on the Basis Twitter, we put out a video yesterday — it's on our site too — about a firm, Clark Nuber, where Matt, who's awesome, brags about increasing the practice's revenue by fifty percent year over year. Those are the kinds of people we work with, because they're ambitious — they see the moment and they want to seize it.

    39:34

    Nathan Labenz: How does the work itself change? I saw this phrase — 'moving people from doers to reviewers' — so that kind of tells a good part of the story in a few words. Give me a little more on that. And then how are people responding to it — are they having more fun, less fun? What's the attitude? Because in some fields — musicians, for example — I think there's a very mixed feeling about tools that might accelerate their ability to produce music. They make music because they love making music. How is it landing with accounting folks, doing a different sort of thing?

    40:20

    Mitchell Troyanovsky: Yeah, I mean, I think this isn't unique to accounting — the answer depends on whether we're talking about the next year, two years, five years, depending on your timelines. I think accounting is weirdly subjective, which is probably surprising to a lot of people listening who don't have much experience in it. Let me quickly define accounting: there's a bunch of stuff that happens in the real world, and when you live in a market economy, in a capitalist system, you need to account for what occurred so you can make decisions. Accounting is just saying, okay, all of those invoices, the contracts, the actual money movements, the inventory I have in my warehouse, how much I owe you — I need to do a compression of all that information into some structured, lossy representation of the real world that someone else can use to make decisions. That's what accounting is. I kind of like to think of it as an intelligence over the economy — a fun meta-analogy. If you think of accounting that way, everyone's accounting is actually done differently, because it's subjective — every company has different policies, a different chart of accounts, it's all different. There are some basic rules, but — code is a useful analogy here — there are basic rules to Python, but the way you structure your codebase is entirely subjective, company to company. Accounting is similar. If you take the code analogy, there are lots of parts of accounting where you need some intelligence — it's not deterministic, you can't literally do it with if-statements, but it's pretty 'if this, then that,' you're going through the flows on a pretty regular basis. And then there's stuff where you have genuine subjectivity — what's the risk tolerance you want in terms of how to account for this, what's your opinion on how to do this, how do you want to structure it. So I think the way the work changes is similar to engineering: you spend a lot more time on the more subjective aspects and less time on the less subjective ones. And in accounting there's also a lot of the client-relationship part, so you spend more time there. Most businesses in America don't have an investment banker or a financial professional — if they want to open another store, they talk to their accountant, and AI isn't going to automate wanting to talk to a person. So I think you want people to focus more on that, and they already do — that's how an accountant grows their business. No accountant gets more revenue because they closed the books a little more accurately. You get it because you talk to the client and help them open a second store, and you let them focus more on that.

    43:29

    Prakash Narayanan: Indeed. Taking a step back — I think Josh Kushner's Thrive has a unit that's buying up accounting firms and using OpenAI's forward-deployed engineers to go in, look at their workflows, and automate a bunch of them. How do you see those relative to what you've systematized as a product? It seems like they're doing on an ad hoc basis what you've built into a product. How do you see those two paradigms moving forward?

    44:10

    Mitchell Troyanovsky: So I think there are two different things here. There's a business question about the differences between doing a firm end-to-end versus building the software, and then there's an ML question about what level of abstraction you can operate on. I'd argue that if you're having to design intelligent systems on a firm-by-firm level, you're just not betting on intelligence appropriately. It's obviously easier to do that than the alternative, but depending on what you believe about where the models are going, I think it's probably the wrong bet. It's just very hard to design it at the layer above — but that's the work we do. And I think a lot of really frontier applied research goes into how you manage that at scale, so there's a different eval question there. And then I think the business bets are somewhat different — I think it is a good business to go roll up accounting firms; if someone wants to do that, I think that's a great business, it's just not a generational business. We're trying to build a company that can actually go and do large amounts of accounting globally and improve the world's economic decision-making. And for that to happen, I can't be going to a small business owner in rural Michigan trying to convince them to switch accountants. We're here trying to deliver magic — our core competency is not going to be selling to small business owners in Michigan.

    45:55

    Nathan Labenz: How do you see that going — you know, when you talk about the world's decision-making, there's a lot of the world where there's just no accounting services being provided at all to all kinds of small businesses. Do you envision a future for Basis where you go direct to businesses? Maybe this isn't something you can do in the US, but potentially in other parts of the world where large swaths of the economy are just totally unserved. Do you think you ultimately put forward a product where you just talk to the AI, because that's maybe all you could afford, and it's still a huge step up?

    46:32

    Mitchell Troyanovsky: Yeah, that's a great question. In the US, the way it tends to work is that once you reach a certain size of company, you tend to do your accounting yourself — you've hired a controller, you have an accounting department. We do serve those firms, but in that case you're still serving accountants, not small business owners directly. Whether we'd actually serve small business owners — we're not planning to do that for the foreseeable future. I haven't thought about, beyond the US, how much benefit you could offer people trying to make decisions in other areas. I suspect in those situations there might be more upstream problems than accounting — from a data-collection and digitization perspective. So it's a good question, I actually don't know. I don't know to what extent the value added could actually be really meaningful.

    47:37

    Prakash Narayanan: Switching gears a little — off the top of your head, could you estimate how many tokens you use on a daily basis?

    47:47

    Mitchell Troyanovsky: Me, personally?

    47:48

    Prakash Narayanan: Yeah, the firm as a whole — just a rough sense, billions, ten billion, a hundred million, whatever, off the top of your head. And is token cost a significant thing you look at on a day-to-day KPI basis?

    48:08

    Mitchell Troyanovsky: Definitely in the billions — I don't actually know the exact number. Token cost is both very important and also very unimportant. I think the amount of levers you have to reduce token cost is just so, so large. And as models get better — I know there's this question of, do you always have to be at the frontier? Some people say token costs are going down, because if you look at what a GPT-4-class model costs today, it'd be basically zero. But then some people say, well, wait a second — actually, 5.6 Sol costs more than 5.1 did when 5.1 was the frontier, so it's going up. So I think the question is, do you need frontier for everything? And the answer is obviously no — there isn't marginal return to intelligence for certain tasks. You don't need Albert Einstein to do every single part of a tax return or a piece of accounting. And as the models get better, and as you get more advanced agent methods around programmatic tool use, different types of routing and harnesses, and even more reinforcement-learning-type work, you can curate the amount of intelligence at every single layer to perfectly optimize. I think you'll get to a place, probably over the next year, where you can really dial in — how much compute do I want to spend on this, given certain cost and latency considerations — to get to the exact optimal amount of cost. If you do that, your token costs go down ninety percent-plus, especially as the floor becomes pretty decent and effectively free. Luna is pretty good, and it's free — you can get pretty far.

    50:25

    Prakash Narayanan: You have a team that I think you call Atlas that's focusing on a hundred times productivity for the employee. How do you measure that — what's the process, and can you give me an example of one outcome from that team?

    50:44

    Mitchell Troyanovsky: Yeah, we probably don't do the best job of actually trying to measure it — in the same way, I don't know if you're measuring whether Codex is giving you efficiency or not, it's just kind of obvious. I think there's a couple of parts to this. One is you need to start building the scaffolding for the future of a truly agent-operated company. I have this mental model — putting aside continual learning for a second — let's just assume you have agents in their current paradigm, but way smarter, doing a bunch of stuff, relearning your context on the fly every single time. What that means is: if you're sending off a coding agent to do something, your context is pretty important. If you're now having thousands of agents spun up every night to run all the operations of a company, your context is multiple orders of magnitude more important. If you mess up a piece of that context, the agents will do the wrong things, because what instructs them is the context — it's not, say, feedback they got that let them update their weights. So yes, you can have memory and other things like that, but there's some level of canon of your company — how to operate, what you want them to do — and that canon needs to be kept up to date, and it needs to be high quality, and it needs to be what you actually want. That's why — I think you wrote about this, like two years ago — you want to treat your company context like it's code. Because with code, if I change a line, production goes down. Today, if you change a knowledge base, your company doesn't go down. But if your agents, without learning, are operating all the time running your company, then yeah, your company will go down if you just change a piece of context and you have a thousand agents running overnight who are now learning the wrong things. So you need to build the infrastructure and the processes around that in order to scale to that place. So a lot of the work of Atlas isn't 'let me go find this one quick automation, do this' — it's, do you build the scaffolding and infrastructure for a truly agent-run company? I'm happy to go into other examples of things Atlas has done that are really beneficial, but maybe the more interesting thing is how do you build the infrastructure for the company of 2028.

    53:16

    Nathan Labenz: Maybe you can go into that if you want to, but I also wanna bridge a little to your work on process supervision, which I think is very timely in the sense that we're now living in the post-era of, like, flagrant misbehavior seemingly due to extreme-scale RLVR without much emphasis being put on how the agents got to the answer. On the one hand I think this is extremely important; on the other, I'm not entirely sure how what you're doing and thinking about should be understood distinctly from, say, constitutional AI, which most people are kind of generally familiar with. I also don't know how you're operationalizing it — you've mentioned some proprietary models where you wouldn't necessarily have the chain of thought or the ability to go back and do additional training, which suggests to me this is probably more of an open-source play, but I could be wrong about that. So take me through — I really wanna understand as much as I can — your process supervision point of view.

    54:31

    Mitchell Troyanovsky: Yeah, it's a great question. Maybe let's start at the tactical level — what we actually do — and then we can go to the analogies to constitutional AI and some fun speculation on RLVR and the Hugging Face stuff. So, on what we do: I think maybe a key shift in mental model is, what's the order of abstraction that you're supervising? Traditionally, when you have process supervision, in some forms, or process rewarding, you're looking at token-by-token generation — you're looking at the weights, you're looking at an inference step. But now, especially if you're doing truly complicated work, you have massive agents that could have five-plus sub-agent layers of depth. You could have a continual run for eight hours, or even sometimes half a day, maybe even a full day. And so the kind of supervision you have there actually looks a lot closer to supervising human actions inside a company than supervising the outputs of an inference. The order of abstraction is more understandable — the thing you're supervising is, did you go do this step, rather than did you follow the right mental thought process. Maybe a very basic example: imagine I had an agent, and my agent's job was to create good PowerPoints — maybe I'm the Claude agent inside the PowerPoint plug-in or something — I make good PowerPoints and deliver them. The designer of this agent — let's say that's me — knows for a fact that if the agent were to go and render visually the changes it made to the PowerPoint before delivering it, it would catch formatting errors some percentage of the time, and therefore it could catch and fix them, leading to better outcomes, generalizable across the board, if it were to look at the PowerPoints. That doesn't mean you always want it to look at the PowerPoints, because that adds latency, it adds cost — it's a subjective thing, it depends on your goals and your organizational design. So the way we think about behaviors is: what matters to us, both from a performance perspective — there's a hundred-plus years of lessons about what it means to do tax work well, we don't need the models to, I don't know, re-derive how to do tax work well, we know what it means — so you can put in place process, and also things you care about from a latency and cost perspective, and then observe to see if the agents actually perform that process correctly. The way we operationalize that is honestly pretty basic: you take a trajectory, and you have a behavior spec — that's the open-source project that tries to define it — and then you have another agent that acts as a judge, that's looking at your spec, or you can think of it as a rubric, it's both a spec and a rubric, and it looks at the trajectory to say, was the condition for this behavior met? And if it was, was the behavior followed? On the condition side — take the PowerPoint example, maybe it's the Basis accounting agent making a PowerPoint — if the user never asked for a PowerPoint, the condition never occurred. So it's: did the agent need to make a PowerPoint, and if so, did it go visually render the images? And if it did, then it passed the behavior, and you can go and see that. It seems pretty simple, but I think it's a powerful framing, because you start to bring real clarity and monitoring to the trajectory itself. And then, Nathan, to your point, maybe you can start to reward based off that — that's a separate question, how do you take the signal you get from process supervision and use it to improve the agent's behaviors, whether that's closing the loop on the harness side, or rewarding at the model side, or whatever. We can talk about that. But it starts with how you define that signal and how you operationalize the extraction of it.

    59:15

    Prakash Narayanan: The way you describe process supervision sounds a lot like how apprentice CPAs are trained. How do you think the training of CPAs should change, given that you now have AI?

    59:35

    Mitchell Troyanovsky: Good question. I'll analogize to coding again — I think it's very similar to engineering, in that there are probably a lot of rote things that aren't important anymore. For example, in engineering, how important is it really for you to understand the syntax of, specifically, different types of for-loops in JavaScript? Nobody cares. What you do care about is good principles — what does good system design mean, why was this structured this way versus that way — and AI can help you with that if you have the right agency, if you work with an AI to understand things, learn, explain, and work with it. So I think you'll see the same thing in apprenticeships for CPAs — a lot more understanding of the whys of things, how to set things up, what the trade-offs are of different approaches, and you go through that apprenticeship. I'm not sure the specific process an agent takes to do a certain type of work is even that important to train into a CPA, because I don't know that process necessarily matters depending on what it is. So I think that type of training will be pretty dominant and broadly important.

    1:01:02

    Prakash Narayanan: But a lot of CPA training is rote learning and examinations.

    1:01:07

    Mitchell Troyanovsky: I agree, and I think that's why a lot of it's going to change. Some of it stays the same, in the same way that even today in engineering people still do LeetCode — you stick with what you know to build up some confidence. But as you get a couple years out, I do think it starts to really change, and you think about what the real purpose of accounting is. There's no law of nature that says these exact things need to be done — we're trying to account for the real world in information-dense ways that people can make decisions from. If you approach it from that perspective, there's just so much creativity and subjectivity and learning involved, and I suspect more and more of the training and process will go toward that part, versus the more rote, manual pieces.

    1:02:09

    Nathan Labenz: I have kind of a big-picture question about the future of business services broadly, because I feel like we've heard a pretty similar story to the one you tell about accountants becoming more of a business coach as the low-level work gets automated, from a bunch of different professions at this point. You hear the same thing in the legal field — 'yeah, we won't spend all this time on contracts like we used to, but we'll elevate, we'll become more of a strategic advisor.' In real estate it's, 'yeah, we won't have to grind out all the details of leases, it'll be a lot more efficient, but then we can be your strategist for where your next location is going to be' — an echo of what you said about the accountant helping open the next location. It strikes me — and even in schools, this isn't a direct parallel, but with Alpha School, they don't have teachers anymore, they have mentors, coaches, and guides — the instruction is given by AI systems on a tablet, and it's motivation, coaching, and social dynamics that the adults in the room focus on now that the more traditional core activity has been largely automated. Can everybody become a coach, I guess, is my question? How much coaching is really going to happen, and does this suggest people need to be really intentional about shaping themselves as coaches? Because what you might run into — whether you're an accountant, a lawyer, or a real estate broker — is a lot of competition for the coaching niche, as everybody tells that same story.

    1:04:01

    Mitchell Troyanovsky: No, it's a good question — I'll give you my thoughts, and then I'm curious about yours. I think the place you need to start — and I'm speaking fully transparently here — is looking at the possible scenarios over the next, depending on your timelines, half a decade or a decade: what stays pretty human in different situations? Let's leave continual learning out of it for a second, because I think if you have that, it's a separate world. But leaving continual learning aside, what are humans just way better at than models? Number one: they're far better at integrating massive systems and world models into decisions. Take Basis, the company — if I were to have an agent autonomously make a database design decision today, the only way it could truly do that at a level I'd trust is if it somehow had all the context about the entire company's history, everything, all my experiences, and all of that — because all of those things play into a decision on a foundational database concept. And it's just nowhere close to having that. For starters, it doesn't have all the senses — it can't see the conversations we've had, it can't see the history, it can't understand the emotion on a customer's face — well, I guess Gemini has that, but no one else does. And even if it could have that, you're a couple orders of magnitude lower on your ability to attend to all that context — we're talking billions of tokens over everything, visual, audio, etcetera. So you just can't make that decision. It doesn't matter if you're Albert Einstein, you will not have enough context to make that decision. And could you spin up a swarm to reduce stuff down on the fly, using English as your memory system? Maybe — but English is pretty lossy, and if you're making subjective decisions, I kind of doubt it, especially truly big calls. So that's the first thing: making truly subjective decisions on behalf of someone else — forget being your own entity, on behalf of someone else — is probably not possible in the current paradigm for large enough systems. The second thing is agents are not currently legal entities — someone else is accountable, either a corporation or a human in some form. An agent doesn't have agency in the traditional sense, in that it's not a legal representative, it can't be accountable for an outcome. And maybe number three: humans like other humans. No one's sitting here watching robots play chess — they're better at chess than humans, but you watch humans play chess because you want to follow the story and you like humans. I have no reason to think that's not true even if we have fully tactile robots jumping around an Amazon warehouse — I don't think we're watching robotic LeBron. So if you take those three pillars — no legal status, can't distill massive systems and world models into decisions, and they're not human, so other humans would prefer a human — I think you get your answer: in a services world, the high end will be working with a human, because that will be scarce. If intelligence is free, working with a human is scarce, and I'd prefer a human in most cases — it's more craft, more artisan, it just feels good, I want to talk to a human. And two, you're making decisions that require understanding, like, for the last ten years what has my customer hated about what I've done, what is that service — nuance you can't get from a markdown file. And three, since you're not a legal entity as an agent, you're taking direction from someone, so you'll need a person or something directing it. So that's where I think the profession goes. Will that mean more or fewer accountants? I don't know — I think that's an interesting economics question, the demand for accounting. I think you can tell a pretty reasonable story — and I personally believe this — that demand for accounting will dramatically skyrocket. Look at how economically complex our current world is: I have this LaCroix, and it's like the classic Milton Friedman quote — there are probably ten thousand people who had a hand in touching the LaCroix I'm currently drinking. Have we accounted for all of their efforts appropriately inside that supply chain? Of course not. If I ask the bodega down the street whether they properly understand their COGS or unit economics, no — because they could pay someone to do that, but it's not required for filing taxes, so no one's doing it. But it would help their life, because they could make better decisions. Go ask Mount Sinai how much it costs them to do a knee surgery — do they know that? No, they have no idea. So the amount of accounting, even in the current world, is one or two orders of magnitude below what we need. And that's before you have these intelligent agents operating as labor at the speed and scale of the internet everywhere — how do you account for all of that? So I suspect demand for accounting is going to go up by probably a couple orders of magnitude. Where that balances out with labor supply, I don't know — I think it's an open question. I think it'd easily go up, but we'll see.

    1:10:23

    Prakash Narayanan: In line with labor supply, there's also the software supply question — how much of the SaaS industry is going to be affected as agents start accessing data directly. When you look at the firms you work with, after you deploy the agents, is it your preference that they start interacting via API with these systems at high speed, so people stop wanting to use the UI at all? Which SaaS systems do you think are most at risk in the next few years of being displaced in favor of agent-only solutions?

    1:11:10

    Mitchell Troyanovsky: I think the SaaS systems most at risk are the ones that try to live in a fairy-tale world where they think they can get away with not granting programmatic access to their systems. Why? I think there's this idea — what is the value of a SaaS system? That's maybe a good place to start. In some cases it's the UI, but the most valuable stuff isn't the UI — it's the guardrails, the permissioning, the process, the database, obviously, the information architecture, the way things are structured. None of those things require UIs — they require basic authorization and understanding. If you think your value as a SaaS provider is in the UI, you're obviously going to lose, because UIs are probably going to go away. Versus if you realize your value is in all that engineering and infrastructure I just talked about, then that value is provided whether you're interacting through an agent or not — which is why things like headless Salesforce are such a good idea, it makes a lot of sense. So I think SaaS providers who think 'no, we don't give access to anyone, everyone will use our UIs' — that's what's going to go away. But the ones who realize their main value is all that infrastructure, and let people access it so it becomes the standard and default and easy to sell and interoperable with anything anybody uses — those are the ones that are going to get very entrenched and flourish, because it becomes hard to displace even if you're not operating every single agent that interacts with the system.

    1:13:11

    Prakash Narayanan: So do you see some of them — for example, I think there was a wave where Slack first said you can't take your data out, and then Slack kind of realized that wasn't a good idea and started to open up a little. Do you think some of them are going to charge extra? Like, 'we're a system of record, we have you locked in, and if we allow agent access, you can migrate out, so we're going to start charging you more to have agents access this material'?

    1:13:46

    Mitchell Troyanovsky: I don't know, honestly — I don't know what the economics wind up looking like over time. But the demand to access will obviously be super high, and if you put a tax on every single API call, it's the equivalent of putting a tax on every button click. So I don't know if putting a tax on every button click works for you or not — it probably depends on the amount of pricing power you have in that specific industry.

    1:14:16

    Prakash Narayanan: Where have you seen demand curves within the firm take off? Besides your revenue curve, in the last couple of months, have you seen any curves showing real inflection points in your metrics?

    1:14:41

    Mitchell Troyanovsky: You mean from a token-usage perspective?

    1:14:45

    Prakash Narayanan: Token usage, service usage, whatever's struck you — what's something new that's happened in the last two or three months?

    1:14:55

    Mitchell Troyanovsky: I don't think anything is particularly surprising or unpredictable. One big difference between accounting and something like engineering is that accounting is quite structured and process-oriented — no accountant's junior is going to rack up fifty grand in credits doing some project on the side, it's much more measured than that. What I do find really cool — as you mentioned, we launched the end-to-end tax product, which to my knowledge is the first example of a solution that is a truly proactive agent, trying to get to end outcomes. So it's not just users coming in and saying 'hey, go do this thing for me' — it's working for you when you're not there. It's going and taking the steps that are needed, processing the different documents, doing all of that kind of work. And the reaction to that, the excitement about it, has been really cool, because there are limits to human attention — your agents could be GPT-8s, but if they still require human attention to know when to kick them off and go do different things, there are diminishing returns to the amount of productivity you can have, versus the agents being more proactive in their nature, coming to the human and saying, 'hey, you are my blocker, I need this from you.' It's a very interesting paradigm shift.

    1:16:42

    Prakash Narayanan: One of the things — I've spoken to doctors, and they often say some are very afraid, because patients come in having talked to AI first and ask all these questions, while other doctors say patients are coming in more educated so they can speak to them better. A lot of people have as much fear of accountants as doctors. When you speak to your clients, do you find your clients coming in more educated because they've spoken to an AI first, and it's easier to work with them — or has it become more difficult?

    1:17:31

    Mitchell Troyanovsky: Honestly, I've never asked them, I actually don't know. But I'd guess people tend to be more involved and invested in their health than their accounting — most business owners don't care at all about their accounting. So I'd be pretty surprised if they were coming in with, like, ChatGPT outputs saying 'I found this thing.' Maybe for tax advice — like, 'why can't I do this?'

    1:18:00

    Nathan Labenz: So there might be an individual.

    1:18:02

    Mitchell Troyanovsky: Yeah, exactly — so there might be some of that.

    1:18:06

    Nathan Labenz: Brock told me.

    1:18:07

    Mitchell Troyanovsky: Yeah, yeah — Brock said he doesn't actually need to file, because they never do audits anyway, so why does he need to file? So there might be some of that, honestly, though I haven't been hearing it from folks. In terms of your actual accounting, I don't think anyone's coming in with the analysis — maybe if you're super overfocused on your financials or something.

    1:18:34

    Nathan Labenz: Well, our next guest is here — we've kept you a little longer than we signed you up for, anyway. So thank you, Mitchell, for joining us on AI in the AM. The company is Basis — try Basis — and we'll certainly be watching for continued updates.

    1:18:50

    Mitchell Troyanovsky: Awesome, thanks guys, really appreciate joining.

    1:18:53

    Prakash Narayanan: Thanks, mate. Nice to meet you. Bye bye, bye bye.

    1:18:58

    Nathan Labenz: As we transition, I think one really interesting thing that people should—

  3. 1:19:03Interview41 min
    Interview: Jay Dawani — Why the Kernel Era Is EndingJay DawaniLemurian Labs co-founder and CEO Jay Dawani joined Nathan Labenz and Prakash Narayanan to argue that hand-written GPU kernels are the wrong abstraction for modern AI workloads, and that the next gains in AI capability will come from software — compilers and runtimes — rather than faster chips. Dawani walked through what a kernel actually is, why memory and network bandwidth (not math) are now the binding constraint, and how Lemurian's Tachyon compiler and runtime aim to take PyTorch-level code and target NVIDIA, AMD, and non-NVIDIA accelerators without anyone writing CUDA. Prakash pressed repeatedly for concrete speedup numbers and a rollout timeline; Nathan pushed on the historical analogy to databases and on where AI itself sits inside Lemurian's own stack. The conversation closed on heterogeneous hardware, consumption-versus-outcome pricing, and what Dawani is optimistic about.
    Open segment on YouTube ↗

    Prakash introduced Jay Dawani, co-founder and CEO of Lemurian Labs, describing his background in AI research, robotics, and autonomous systems, including advising NASA's Mars Rover program on vision-based navigation and planetary mapping. He framed Lemurian's thesis: the way the industry programs AI, by hand-writing hardware-specific kernels, is fundamentally broken, and the company's system-level compiler, Tachyon, is built on the argument that the next leap in AI capability will come from software rather than faster chips. Lemurian recently raised a $28 million Series A behind that bet.

    Nathan opened with a basic-terms question: what does it actually look like to write a kernel, and why has this resisted abstraction for so long? Dawani described a kernel as a unit of computation expressed from the hardware's point of view — you have to reason about the memory hierarchy, the available compute units, and the cost of moving data, then keep data stationary to keep the math units fed. He argued the field crossed an inversion: transistors got dramatically faster while memory, bottlenecked by capacitors that don't shrink the way transistors do, could not keep pace, so the industry went from a math-bound world to one with orders of magnitude less memory per flop. For further study, he pointed to Nicholas Wilt's "The CUDA Handbook" and cited Flash Attention as an example of the kind of coordinated, non-obvious rethinking that real kernel work requires.

    Prakash asked whether Tachyon is simply a compiler for writing kernel-level instructions. Dawani explained the Tachyon name — a Star Trek reference to a planet whose tachyon core makes things move strangely fast — and pushed back on the framing itself: writing better kernels no longer straightforwardly buys performance, because the system is now waiting on memory rather than math, which he illustrated with an image of GPUs as a thousand piranhas that get agitated and bored when they have nothing to chew on. Pressed twice more by Prakash for a hard percentage improvement against a reference model, Dawani declined to give one in that form and instead reframed the pitch around developer time: getting PyTorch-level performance without months of hand-written kernel work. The one number he did offer was a claimed 1.7x speedup on a single matmul kernel on an AMD MI300X — his own words, presented as proof a compute-bound workload can be accelerated at all, with the larger opportunity, he said, in memory- and network-bound workloads and at cluster scale.

    Nathan drew a parallel to database history — SQL's original constraint was disk space, and MongoDB emerged once disk got cheap and developer time became the binding constraint — and asked what specifically thrashes when an agent makes a tool call and hits latency. Dawani said Lemurian looks at the full request trajectory rather than any single token, and described the KV cache tradeoff as prefetch versus recompute, which plays out differently depending on how much on-chip memory a given device has. Asked directly whether an LLM sits in Tachyon's inner optimization loop, Dawani said no — he described the compiler itself as a knowledge-based system that codifies expert knowledge into known-good choices, with correctness verification built in for free, and said the bigger innovation is a runtime that collects execution traces and keeps improving its placement decisions over time. The connection dropped briefly around this point in the conversation; Dawani rejoined with "we're back, no idea what happened there."

    Prakash laid out the competitive landscape — Modular's Mojo, OpenAI's Triton, and Elon Musk's prediction that AI will eventually write straight to assembly — and asked whether Tachyon is just another language developers have to learn. Dawani argued Modular's approach still leaves a human writing kernels by hand, and that Triton lowers the bar but still assumes CUDA-level knowledge; his case against hand-coverage was combinatorial, citing roughly 2,000 kernel-writing performance engineers worldwide, most of them inside a single vendor's ecosystem. He walked through operator fusion using a desk-shelf-library analogy for the memory hierarchy: information already on the desk is instant, a book on the shelf costs a walk, and something at a distant library costs progressively more — fusion is simply refusing to send data all the way back to main memory before the next operation needs it.

    Nathan asked who actually faces the heterogeneous-hardware problem, and why no consortium of non-NVIDIA chipmakers has put up a bounty for solving it, given Lemurian raised without one. Dawani said staying independent of any single silicon vendor was a deliberate funding choice, since taking that money skews priorities toward that vendor. He argued heterogeneity is already universal — pointing to robotics SoCs like NVIDIA's Jetson and Qualcomm's RB5 and RB6 platforms, where a chip is really a CPU, GPU, DLAs, a vision processor, and DSPs that all need to be addressed differently — and said the software stack has not caught up: it is "still living in the sixties," treating accelerators as sidecars rather than first-class citizens even as labs train models across entire data centers as if they were one machine.

    Prakash asked for a rollout timeline and buyer profile. Dawani said Lemurian expects to be running on NVIDIA and AMD hardware for most customer models with state-of-the-art performance by the end of the year, moving through beta and design partners toward general availability in Q2 of next year, with an initial focus on managed inference serving for enterprises and AI-native startups before expanding into post-training, reasoning, agents, and training support. On pricing, Dawani said token-based billing is breaking down as reasoning models and agents make consumption harder to predict, and that Lemurian is moving toward pricing "effective compute" — usable work extracted from existing hardware — arguing that raising utilization is now a faster way to add compute than waiting on new silicon, which he said is increasingly bottlenecked by power rather than chip supply.

    Nathan closed by asking about Dawani's positive vision for the AI era, tying it to his Star Trek fandom and Lemurian's futurist branding. Dawani described a human-first outcome in which AI, as a genuinely different kind of intelligence, takes on work that was arguably never really human work to begin with, freeing people for what they're uniquely suited to do. He pointed to drug discovery and education as areas he's personally excited about, and closed on the idea that the industry took one particular technological branch and has been running down it ever since — expanding the "tech tree" now, he said, could reopen paths that were abandoned only because they didn't suit the hardware available at the time.

    Think about GPUs as a thousand piranhas just sitting around chomping. If they don't have things to chomp on, they're going to get really agitated and bored.

    There's a saying that scaling what works is different from what works at scale.

    Not an LLM. There are many ways of having intelligent behavior without having an LLM.

    1:22:39For those of us who haven't been in the kernel trenches — what does it look like to write kernels, and why has this been so hard to abstract away for so long?
    Dawani said a kernel is a unit of computation expressed from the hardware's point of view — making one fast requires reasoning about the memory hierarchy, the available compute units, and the cost of moving data, then keeping data stationary. He said the discipline exists because transistors got much faster than memory could keep up with, flipping the ratio of flops to bytes moved by roughly two and a half orders of magnitude and turning kernel writing into a genuine systems-level optimization problem.
    1:24:41What would be the top few canonical examples of great kernel work I could study, like flash attention?
    Dawani pointed to "The CUDA Handbook" by Nicholas Wilt as the defining text on how GPU hardware wants to be programmed. He said Flash Attention is a good example too, but stressed it wasn't one trick — it was a set of novel insights built on top of each other, changing the problem itself to make it easier for the hardware.
    1:25:54Is Tachyon a compiler that sits on top of the kernel layer — are you writing kernel-level instructions using the compiler?
    Dawani said a compiler is fundamentally a translation device connecting the programming model — how a developer writes something — to the hardware's execution model. He explained the Tachyon name as a Star Trek reference, then argued that in a memory- and network-bound world, writing better kernels no longer straightforwardly buys performance because it just exposes how much of the system is waiting on memory rather than computing.
    1:29:46What metrics do you use internally as your North Star, and can you quote a hard percentage speedup on a reference model like Llama 3?
    Dawani said Lemurian tracks time-to-first-token, inter-token latency, prefill-versus-decode, and the full request lifecycle rather than a single benchmark number, and increasingly the entire agent trajectory rather than a token. Pressed again for a specific percentage, he declined to answer in that form and instead pitched developer time: PyTorch-level performance without months of hand-written kernel engineering.
    1:31:44Can you give a concrete number — how much faster is a workload with Tachyon?
    Dawani said on a single matmul kernel on an AMD MI300X, Lemurian has shown a 1.7x speedup, which he offered as proof a compute-bound workload can be accelerated at all. He said the bigger opportunity is in memory- and network-bound workloads and at cluster scale, citing roughly 2 to 3x on a full workload and close to 30x on a large training run.
    1:33:00When an agent makes a tool call and hits latency, what happens to something like the KV cache — does it get shuttled off, does another user's work take over the chip in the meantime?
    Dawani said the tradeoff is prefetch versus recompute, and the answer depends heavily on the device: chips with abundant on-chip memory don't have to worry about it much, but memory-constrained devices are heavily affected. He added that turning a base model into an agent — adding tool calling, reasoning, and outside decision-making — exacerbates every bottleneck in the stack and leaves the system perpetually I/O bound.
    1:40:18Where is intelligence actually used throughout this optimization process — is there a frontier LLM or agent in the inner loop making runtime decisions on an ongoing basis?
    Dawani said no — not an LLM. He described the compiler itself as a knowledge-based intelligent system: expert knowledge about how to make workloads fast gets codified into known-good choices, with correctness verification built in for free. He said the bigger innovation is the runtime, which collects execution traces and keeps improving its placement decisions over time.
    1:42:47With Modular's Mojo, OpenAI's Triton, and Elon Musk predicting AI will write straight to assembly, how does Tachyon fit in — is it just another language or tool developers have to learn?
    Dawani said Modular's approach — building a new meta-programming language — still leaves the developer writing kernels by hand, and Triton lowers the bar but still assumes CUDA-level knowledge. He argued the combinatorial space of hardware, numerics, and fusion strategies makes hand-coverage infeasible, noting the world's roughly 2,000 kernel-writing performance engineers are mostly concentrated inside one vendor's ecosystem — the coverage gap Tachyon is built to close.
    1:53:10What's your product rollout timeline, and are your buyers hyperscalers or startups?
    Dawani said Lemurian expects to be running on NVIDIA and AMD hardware for most customer models with state-of-the-art performance by the end of the year, moving through beta and design partners toward general availability in Q2 of next year. He said the initial focus is managed inference serving for enterprises and AI-native startups building products on models, later expanding into post-training, reasoning, agents, and training support.
    Lightly edited · timestamps jump to YouTube
    1:19:03

    Nathan Labenz: study more deeply, and maybe somebody has, but I need to go find it, is where do people really prefer the human touch and where do they not? The classic example, like, Waymo selling at a premium to Uber, is one contrary data point — it's like, actually, you could have told a story where you're going to want a human driver, you're going to want that conversation, you're going to want that warm smile to welcome you into the car. In practice, you don't always even get that in an Uber. And it turns out, right now, the market is pricing Waymo significantly higher. I do believe in the human touch story certainly for some things. I got a robot massage in Shanghai — I think I mentioned that to you before — and I'll still definitely take the human massage over the robot massage. But how many things are really like that? And is accounting really like that? Maybe it is, he would know better than me, but I do question it. If I think about my accounting future, I'm like, one accountant wants to spend an hour a week on the phone with me coaching me, and the other is just doing the job and getting it done. It's not obvious at all, honestly, from my perspective, that I want that hour a week on the phone with my accountant. So that's probably my biggest question coming out of that conversation — in what domains does that really hold, and how many people are in for a rude awakening because they're telling themselves a story about how they're going to turn into business coaches when in reality their clients do not want business coaching from them. Results will vary, I'm sure, but that seems like a major risk factor for a lot of people right now if that's what they're counting on.

    1:20:54

    Prakash Narayanan: Indeed. Let me introduce our next guest. Our next guest is Jay Dawani. Jay Dawani is the co-founder and CEO of Lemurian Labs, a company dedicated to fundamentally rebuilding the software infrastructure that powers artificial intelligence. With a background spanning AI research, advanced robotics, and autonomous systems, Jay previously advised NASA's Mars Rover program on vision-based navigation and planetary mapping. In 2022, he founded Lemurian Labs with a clear, ambitious thesis: the way the world programs AI is fundamentally broken. For the past decade, the industry has squeezed performance out of chips by writing hardware-specific code called kernels — a process Jay equates to writing in assembly language. As AI models scale to demand supercomputer-level infrastructure, this old approach has created a massive bottleneck. It locks developers into specific hardware vendors, strands an enormous amount of computational power, and demands unsustainable amounts of energy. Lemurian Labs recently raised a $28 million Series A to solve this exact problem, rolling out a system-level compiler called Tachyon. Jay is here today to argue that the kernel era is over, and that the next massive leap in AI capability will not come from a faster chip, but from software that treats a diverse global network of hardware as one seamless, highly efficient machine. Jay, welcome to the show.

    1:22:35

    Nathan Labenz: Hey, thanks for having me. Can we start with a real basic question, just for those of us who haven't been in the kernel trenches ourselves? What does it look like to write kernels, and why has this been something that has been so hard to abstract away for so long?

    1:22:58

    Jay Dawani: Kernels fundamentally are a localized unit of computation that's expressed from the point of view of the hardware. In order to make something go really, really fast, you have to first understand the math for the workload, you have to think about memory hierarchy, you have to think about what compute units are available to you, what their logical layout is, what the cost of moving data in and out of them is, and then how you keep things stationary so you can make it go fast. But most of the kernel discipline came from a time when math was more expensive than memory, so you had to reduce the math cost to make things go faster. Over time, transistors got way faster than memory — because capacitors don't shrink the way transistors do, so transistors got really fast while memory didn't keep up. We reorganized computers around having much more transistor density so you can have more flops, but that took us from a world where flops and bytes moved were roughly one-to-one to a world where you have about two and a half orders of magnitude more flops per byte moved. So the entire game is rewriting computation from the point of view of the hardware to make a workload run faster and more efficiently, and that requires understanding a lot of stuff — that's where the challenge is. It's really low-level, systems-level optimization.

    1:24:41

    Nathan Labenz: What would you say are the kind of canonical examples of really great work in this space that I could stand to learn a lot from — like flash attention comes to mind — what would be the top few for self-study?

    1:25:01

    Jay Dawani: There are so many, honestly. If you want to study kernel programming, you want to get really book-deep. One of my colleagues wrote what I think is the defining book on this, The CUDA Handbook, by Nicholas Wilt — one of the fathers of CUDA. He broke it down in one of the simplest ways of understanding it, getting a full breadth and depth view of how the hardware works, how it wants to be programmed, and what the right programming and execution model is. Ultimately it usually takes a lot of clever insights — you have to change the problem and create a new problem to make it easier for the hardware in a lot of ways. That's kind of what Flash Attention did. It wasn't just one thing, it was a set of novel insights built off of other insights. But that's definitely one good example.

    1:25:54

    Prakash Narayanan: Let me see if I understand this correctly — Tachyon is a compiler, and the compiler sits on top of the kernel, basically. So you're writing kernel-level instructions using the compiler — is that correct?

    1:26:10

    Jay Dawani: Well, I'll get into it in a second. When you think about hardware, the hardware has a certain execution model — that's how it wants to actually run code efficiently. Your job, when you're writing a compiler, is connecting a programming model, which is the way a developer writes something in a way they understand, and then lowering that down to the machine. So fundamentally a compiler is a translation device. I'm a Star Trek guy, so I love the concept of a universal translator, and that's also partly where the name Tachyon came from — there's a whole set of jokes in there for later. For folks who haven't seen Star Trek: The Next Generation, at some point they come across a planet called Gotana that has a tachyon core, so things move really weirdly and fast around it. Our entire platform is going to be called Gotana, built off of this stack called Tachyon. Part of the reason we chose that, and it was fitting, is that people think kernels are the speed of light — the fastest kernel you can write is the canonical speed of light for a workload. That's true in a compute-bound world. We are not in a compute-bound world — we're in a memory- and network- and communication-bandwidth-bound world, and that changes things already. And the reason I say kernels are the new assembly is that writing better kernels no longer gives you performance, because a better kernel actually just exposes the latency of the system — now it's waiting for memory. You want to think about GPUs as a thousand piranhas just sitting around chomping. If they don't have things to chomp on, they get really agitated and bored, and they're still consuming energy. So you want to feed them as much as possible, and that's ultimately the scheduling problem that exists here. The reason I say kernels are the new assembly again is I don't think we benefit from writing them anymore — what we need is something that makes developers more productive. Time to value really matters.

    1:28:33

    Jay Dawani: If I can give you this experience of — I'm writing in PyTorch, I say compile this to whatever, and it gives me the same performance or better performance than what I would have gotten if I had waited nine months with a team of experts to get that kernel written — wouldn't you prefer that?

    1:28:53

    Prakash Narayanan: Absolutely.

    1:28:54

    Jay Dawani: That's what I'm trying to give. So in that world, the kernel becomes something the compiler generates, and then we execute it — that's the system we're talking about. The compiler takes in the workload graph, rewrites it from the point of view of the hardware by looking at the memory hierarchy, and arranges things from that point of view. It does partitioning, it does fusion, creating work-set items, and then it thinks about how to move computation closer to where the memory already is, so it has to move less far — because every single time you move down a level of the hierarchy, your cost of data movement grows by at least half an order of magnitude, up to a full order of magnitude. So if I can avoid doing that, I can create more throughput for less cost.

    1:29:46

    Prakash Narayanan: What are some of the metrics that you could quote — does it speed up the performance of a certain model by X times, or what kind of metrics do you use internally as, kind of, the North Star of what you're doing?

    1:30:03

    Jay Dawani: We look at a lot. For LLMs specifically, we'd be looking at time to first token, time between tokens, overall inter-token latency. We'd be looking at prefill versus decode when you disaggregate it. We're looking not just at that but at preprocessing to postprocessing, the entire lifetime of a request. And beyond that — I don't think tokens are the optimization unit anymore, it's the full trajectory, especially as you think about reasoning models and agents, that becomes much more important.

    1:30:36

    Prakash Narayanan: So what is the time-to-first-token improvement that you've seen upon deploy?

    1:30:41

    Jay Dawani: We're not in production yet — this is the kind of stack we care a lot about, and ultimately we are selling trust. Developer infrastructure, developer tools, is ultimately about trust. You get one shot to prove yourself, and if you fail at that, you've lost it completely, and regaining that takes much, much longer. So you're trying to make sure you're working closely with your design partners, your future customers, and your channel partners, making sure things are working as expected, so that once it's in their hands, the time to wow is immediate. That is what we're optimizing for.

    1:31:15

    Prakash Narayanan: But how do you — when you tell them what could be possible, what kind of number is it? Is it going to be 20 percent better? Is it going to be 50 percent better? Is it going to be 90 percent better? Am I going to cut time to first token to a tenth of what it used to be on a Llama 3, one of these reference models we tend to use? What's the sell that you're making?

    1:31:44

    Jay Dawani: We're accelerating the full workload, not just kernels. On a single kernel right now, like matmul, on an MI300X, we've shown that we're 1.7x faster — so proving that we can speed up a compute-bound workload is already a huge thing. But our entire system was designed to accelerate the full workload, especially in cases where you're memory- or network-bound. Those cases, when you look at them and start scaling, that's where the real problem is — getting really good performance on a single device is the solved problem. Getting performance on a heterogeneous cluster, as you're thinking about larger, more dynamic models and scale, is the real problem we're trying to go after — the time to get up and running on new hardware for the same workload, time to performance, and how reliable that is, is what people are interested in us for. On a full workload, we can expect, say, 2 to 3x. On a large training run, we can expect almost 30x — there's a lot more opportunity to speed up and improve things on a larger system than on a smaller one.

    1:33:00

    Nathan Labenz: Can you talk a little bit more about how you think about bottlenecks across the system? I just had a conversation with somebody from MongoDB yesterday who gave me the short version of a long history of database technology. He was saying that originally, going back to 1970 when SQL was introduced, the constraint was disk space — that's the big reason data normalization became canon in data storage, because people were trying to optimize for keeping the disk footprint small. When MongoDB itself was created, that constraint had basically been solved by progress in hardware, disk space had become cheap, but developer time was the new constraint — everybody wanted to get their apps out fast. Now we have the compute-bound paradigm, but you're saying memory is the actual constraint. I'd love to dig in a bit more — especially in this long-horizon agentic world, what does that really look like when an agent makes a tool call and there's latency? What happens to the KV cache during that wait — does it get shuttled off and brought back, does somebody else's stuff use that same chip in the meantime, does it have to wait? Give me a sense of what's thrashing around there and what kinds of optimizations move the needle.

    1:34:39

    Jay Dawani: There's a saying that scaling what works is different from what works at scale. We're in a really interesting world right now — all the software was built on the assumption that you're programming a homogeneous device, a single chip, and the workload is static and understandable. All of a sudden we very quickly went into a world where Moore's Law, and then Dennard scaling, stopped working for us. We brought in a lot more heterogeneity to make up for it. Chips went from simple things to becoming a whole system — now you have an entire cluster that is one machine we're using to program these things. The workload is getting more dynamic, it's coming more branchy, it doesn't fit inside a single GPU anymore, you have to worry about NUMA regions. So where you place things matters a lot, and the tradeoff becomes recompute versus refetch — and then how local can you keep things? If you're local, you can reuse a lot of things, save on data access and data movement, and get more throughput. So then you have to partition that and manage a lot of overhead — that's the primary issue.

    1:36:09

    Jay Dawani: For the KV cache, you're trying to balance prefetch versus recompute. On some devices with a lot of on-chip memory, you don't have to worry about it that much — you'll still get speedups. But on devices that are memory-constrained, that really starts to matter. So it's not one answer — it changes based on the system, and that changes your economics and what's possible. Now, going from a model to an agent, a lot of things have had to happen. You went from a base model that went through some fine-tuning process, got a personality, started learning how to do certain things. You had to connect that to an environment so it can learn how to reason, and learn to use compute more effectively to solve problems. Then you can think about tool calling as another add-on — but having tool calling and reasoning doesn't make you an agent, it's all the things around that, being able to talk to the outside world and make certain decisions. And there's a difference between the kinds of decisions different kinds of agents can make. All of that exacerbates every single problem in the stack, and you're always I/O bound — it's not one problem, it's a cluster of problems hitting you at once, and being able to make the best choice at that point in time is what it comes down to.

    1:37:40

    Jay Dawani: And this goes back to part of why kernels need to go away — all of that is dynamic at execution time, and you don't know what needs to happen until runtime. But in the kernel world, you're pre-baking decisions before the workload executes. So you need to be able to make the right decision in the most efficient manner and place things where they need to be placed based on future access. That's why I said the workload right now is the agent — it's not an LLM with a serving engine and an agent built outside of it, where you don't have visibility across the whole thing. That needs to change. So we started from the point of view of looking ahead: what needs to change, how workloads are going to evolve over time, what breaks when, and how do we build that kind of system. The big part for us is the runtime — that's a very big piece of the innovation we made, precisely for this reason. We essentially have a sandbox environment in which the code executes, on and across machines, and that lets us create the illusion of a single, unified memory for developers even across nodes. We collect traces as we're executing, and we can continue to optimize them after the fact.

    1:39:07

    Jay Dawani: So if we make a poor decision about where we place something once, we can learn from it and continue to improve with the next case — your system keeps improving as you keep running more workloads. And the reason I'm saying this is that the right answer is always going to change over time.

    1:39:31

    Nathan Labenz: Is that process itself — oh, sorry, quick follow-up. I guess everybody's, of course, using AI to accelerate their development, right? And when you said earlier that you want to speed up developers, part of me was thinking: are developers still the bottleneck they were when MongoDB was created, now that we've got agents doing a lot of that acceleration at the developer level? So I'm interested in your thoughts on that. But what I'm really wondering is where the intelligence is used throughout this optimization process — are you using agents in that innermost loop that's looking at, oh, the workload is changing right now, we better adjust our approach? Is there a frontier LLM at that level being used to optimize the runtime decisions on an ongoing basis?

    1:40:34

    Jay Dawani: Not an LLM. There are many ways of having intelligent behavior without having an LLM. Have you heard of compilers? Compilers have always been very entrenched with AI in a lot of ways — in this case, knowledge-based systems. A knowledge-based system is essentially what a compiler is: you have certain information or knowledge about how to make things go fast that you want to codify so you can get the result fast, just by making known-good choices, and you get a verifier for free, because compilers have to be correct.

    1:41:16

    Prakash Narayanan: Dropped off.

    1:41:18

    Nathan Labenz: Come back.

    1:41:21

    Prakash Narayanan: He should be back in a bit. We've spoken to so many — we've spoken to a couple of chip companies, we've spoken to people who are managing prefill and decode and trying to rearrange that on a hardware basis. We've also seen Modular's Mojo, we've seen OpenAI's Triton GPU programming system. We've also seen Elon Musk say, hey, it's going to be AIs writing directly to assembly within a few years. And it's hard to think about how the software layer is going to change in the medium to long term, because it's quite clear that a lot of these compilers and things are going to be used primarily by AIs — at this point, most of the programming is already being done on the front end by prompting. It's hard to see where things go from there. Hi, Jay.

    1:42:43

    Jay Dawani: Okay, we're back — no idea what happened there.

    1:42:47

    Prakash Narayanan: So one of the questions I had, and I was just describing this, was: you have Modular's Mojo, you have OpenAI's Triton system for GPU programming — Modular, for example, trying to reduce this two-language tax that you've always faced. How do all of these pieces fit together? Is it another language or tool in the stack that a developer now has to learn how to use? Is there a kind of tax that comes with having to add that piece to the stack? And how does Tachyon actually do this? We've spoken to some chip developers, for example, who move the data much closer to where the compute is. We've spoken to people who are doing compute inside memory. We've also spoken to one team doing chips that allow compute to be done inside memory. How exactly does Tachyon do this, and how does it fit into an entire stack that already exists?

    1:43:58

    Jay Dawani: First, think about companies like Modular — what they're trying to do is build a systems-level meta-programming language so you can write kernels better. To me, that's the most literal interpretation of "I need a good alternative, let's build a language." But the problem is you still have to write kernels — that's on the developer to go do. If you actually do the math for the amount of hardware that's in the world today, all the different workloads we're running, the different numerics, the different fusions, the different ways of partitioning them, the different tile sizes, latency versus throughput and other SLOs — you actually sit down and do that, and you're like, okay, I need to write about 106 billion kernels to get coverage. Well, there are only about 2,000-odd performance engineers in the world who actually know how to write good kernels, and 90 percent of them are inside one vendor's ecosystem. So there's still a coverage problem. The reason you can accelerate some of this on NVIDIA is because NVIDIA spent 20 years building an ecosystem of tools to make your life easier and give you that feedback loop — that ecosystem makes kernel generation easier. That maturity doesn't exist with any other vendor.

    1:45:23

    Jay Dawani: So for us, what we're looking at — and also, there's certain hardware today for which you cannot write a kernel at all, for data-flow architectures. You have to execute a task, and that's a different granularity than a kernel. The tile programming model is still within kernels, but it's easier and more accessible for developers, and it raises the abstraction. The reason OpenAI invested in Triton was to raise the number of people who could write kernels without needing to know CUDA or low-level programming.

    1:45:57

    Prakash Narayanan: So if I hear you right, Tachyon is a tool that's going to be very useful for non-NVIDIA chips, because they don't have that toolkit that already exists in the NVIDIA ecosystem.

    1:46:12

    Jay Dawani: Yes, we allow more non-NVIDIA hardware to become more accessible, so you as a developer can choose the best hardware based on the shape of your workload and the economics you care about, and make that choice immediately without having to rewrite everything. But on NVIDIA hardware as well, we can give you more gains on a larger system.

    1:46:39

    Prakash Narayanan: And can you describe again — you have this thing, I think, called operator fusion. How exactly does this work?

    1:46:47

    Jay Dawani: Normally, when you're running some kernel or operator, you're reading memory — in GPU land you use DRAM as main memory — so you're reading something all the way from there, bringing it into a local register, operating on it, writing back to main memory, reading it back, and then doing the computation. In that time your compute unit was waiting while you were moving data around, and that's very expensive. So fusion is basically: let's not send the data all the way back, let's keep it local and reuse it for the next operator. It's about synchronization and barriers, mitigating data movement — that's the name of the game — and minimizing waste, and that gives you speedups. It's sort of like: if I'm at my desk right now and I need certain information, and it's already somewhere on my desk, I know exactly where it is, I can read it and use it. If it's probably in one of the books on my shelf, I go over to my shelf and grab a book — that's like going from a local register to an L1 cache. If it's not there, I have to go to the library next door, or go to a friend or somebody else in the office — I have to travel a little further to get that information, but I have a larger collection to search through. Now if I go from there to a library that's far away and they don't have the book either, and they have to get it from somewhere else and then give it to me — that's the logical hierarchy of memory, and that's the cost. So if I don't have to go that distance to get something and it's available already, I'm much faster and much more efficient. That's what operator fusion does.

    1:48:46

    Nathan Labenz: Can you tell us a little bit about who faces this challenge of heterogeneous clusters, and is it a problem people try to tackle more in inference relative to training? How big are the customers typically, or how far into the long tail does it go? And how do the other chip manufacturers respond to you and support you? It strikes me listening to all this that there should be a consortium of non-NVIDIA chipmakers putting forward a billion-dollar prize for whoever can do this. But I'm guessing you had to raise funding without that carrot out there.

    1:49:36

    Jay Dawani: For us, it was — well, we wanted funding specifically so that we could be vendor-neutral. We wanted that to be a very specific goal. The moment you take money aligned with one vendor, that skews a lot of what you prioritize. For us it's really important to build for the builders — what do they need? Our customer ultimately dictates. And when we spoke to them, they obviously love NVIDIA — I love NVIDIA — but they also want to be able to use AMD GPUs, TPUs, Trainium, and any other chip, because each of them is better for a specific thing.

    1:50:21

    Jay Dawani: If I'm locked to just one thing, I'm getting suboptimal use for a workload. We've been through this before, where a lot of good ideas died because they didn't suit the available hardware — if there's more hardware available and they're good at different things, more ideas become accessible. We're now in the realm of reasoning models, world models, planning, and robotics, and some of those ideas can come back. Robotics is all heterogeneous — it's a full SoC with a lot of different parts. Look at the NVIDIA Jetson, for example — it's got CUDA cores, a CPU, DLAs, a vision processor, other DSPs on it — you have to be able to access all of them. The Qualcomm RB5 and RB6 are similar. We started originally as one of those companies building heterogeneous SoCs for robotics, so we had to deal with it.

    1:51:06

    Jay Dawani: All GPUs today are heterogeneous. The moment you cross the five-nanometer threshold, you have to think about complex packages — every single machine today is programmed as a CPU, some network, a GPU. You have different kinds of memory — LPDDR, GDDR, HBM — and they all change the execution model. You have different numerics: on NVIDIA you have the CUDA cores, or CUs, and then you have tensor cores as well, and they need to be spoken to differently. You have different memory hierarchies now, with the core memory for the SMs, and then tensor memory accelerators too. Heterogeneity is already here — everyone dealing with a GPU is already dealing with it, anyone training or deploying a model is dealing with it. But the software was not built for this — the software is still living in the sixties, we're still programming as if we've got a single-core CPU and GPUs are sidecars we throw work off to every now and then, and we just add in libraries or intrinsics or pragmas to fix the problem. That isn't the case anymore. Accelerators need to be first-class citizens, and CPUs need to be the backstops. That changes things — now you're programming a cluster as a single machine. Some of my friends at the labs right now are training models across data centers — it's not even across nodes anymore, it's not even racks, we're talking about multi-gigawatt data centers as one machine for one model.

    1:53:10

    Prakash Narayanan: So give us an idea of your product rollout timeline — I think there was beta testing about to start, or a product rollout early next year. What does your timeline look like, what are the milestones, and what's your go-to-market strategy at this point?

    1:53:36

    Jay Dawani: We'll be up and running on NVIDIA and AMD hardware for the majority of models customers care about, with state-of-the-art performance, by the end of this year. That's going to get into the hands of our beta customers and our other design partners, and then we'll do GA in Q2 next year.

    1:53:51

    Prakash Narayanan: And are your buyers hyperscalers or startups?

    1:53:56

    Jay Dawani: We're a bit more focused on managed inference serving to start — both, really. That appeals more to certain enterprises or AI-native companies and startups that are trying to build products around the models.

    1:54:09

    Prakash Narayanan: I see. From there —

    1:54:10

    Jay Dawani: We will expand into things like post-training, building up for reasoning and environments, building up for agents, and then adding in support for training.

    1:54:18

    Prakash Narayanan: So you're looking at startups that want an inference edge or cost savings — is that the focus?

    1:54:25

    Jay Dawani: Right — it's for people who want more control.

    1:54:29

    Prakash Narayanan: Is it helpful that a lot of the new clouds are now doing bare metal — does it help that these new clouds offer bare metal and customers can bring their own stack?

    1:54:38

    Jay Dawani: The new clouds have always been in the bare-metal business — sell hardware without all the extras, because the hyperscalers had a lot of extra software they were trying to upsell. If people just want the fastest hardware at the lowest friction, that's what the new clouds were built for. That's definitely worked better for us given how our stack is built, and that's who we're generally partnering with — our go-to-market is with all the new clouds, and we're very tightly coupled with a lot of them.

    1:55:10

    Nathan Labenz: Two final questions for me — one quick one and one kind of zoomed-out, hopefully fun one. The quick one is just: how do you price this? Everybody's trying to figure out the future pricing model — is it by seat, is it a company-wide license, is there a usage component?

    1:55:30

    Jay Dawani: In a perfect world, outcome-based pricing makes a lot of sense, especially for agents. But if you're an infrastructure provider, the only business model that has worked and continues to work at scale is consumption-based. I know a lot of people like token pricing right now, but tokens work well for the static request-response kind of workload the old base models were tuned for. Now that you have reasoning models, it's hard to reason about token consumption, and it's the same with agents. So what we're actually moving toward is effective compute consumption — the amount of compute you use to realize useful work — which is different from selling a GPU-hour or a GPU slice. Part of the reason this makes sense is our business model scales with the delta between physical compute and effective compute. The fastest new addition of compute is going to come through software — it'll come online faster than you can actually plug in new hardware, because you're going to be electricity-bound; getting turbines and power installed so you can bring up new silicon is the real limiter right now. If I can boost your utilization three to ten times, I'm adding more effective compute at a lower cost that I can sell, and the tokens resulting from that consumption are what we're reselling. That scales nicely, and it's understandable for finance people too, because for a lot of them, token pricing and other pricing models are breaking right now.

    1:57:21

    Nathan Labenz: Interesting, thank you. Okay, here's the fun one — you mentioned you're a Star Trek fan, and the website certainly reflects the futurist aesthetic. What's your positive vision for the future of humanity as we go into the AI era?

    1:57:39

    Jay Dawani: I'm definitely a believer in a human-first kind of outcome. I think AI is a new kind of intelligence — it's different from human intelligence, and it can be complementary. What it's showing us is there's a lot of work we've been doing and describing as work that maybe wasn't actually work, and that we should never have been doing in the first place, that we can start offloading. That gives us more leverage to do work we're uniquely suited for. To me it creates much more equity, so we can try out new things, accelerate discovery, learn more about ourselves and the world, and that results in more solutions to more problems being solved at a lower cost than before. Personally, I'm really interested in what happens in drug discovery and design — there have been so many places pharma hasn't been able to get to because the cost of research was so high, or there wasn't enough of a dataset. As biology goes from squishy things to bits, things get on an acceleration curve, and you get real benefits out of it. The quality of education has been fairly crappy, and it hasn't kept up with the pace of the world, and I think AI can help push that gap closed. There's a lot more that can happen — I think deep tech is going to be way more important going forward, and simulation is going to unlock a lot of it, and AI plays a big role in that. It's like we can expand the tech tree — we took one particular branch and we've just been going off of it, and there have been all these inherited biases that have guided us. If we can unroll that and expand more, who knows what's possible? I'm personally excited about that.

    1:59:32

    Nathan Labenz: Love it, me too — also a tiny bit afraid, but very much excited. Jay Dawani, the company is Lemurian Labs, the stack is Tachyon, and there's a lot more to come. We'll be watching, as always, for your updates.

    1:59:48

    Jay Dawani: Thanks so much, thanks for having me. Cheers, Jay — great to meet you. Bye, bye.

    1:59:59

    Prakash Narayanan: Amazing. Yeah.

    2:00:01

    Nathan Labenz: Which way does this all — I mean, it seems, I guess, if I had one more question, I might have just asked —

  4. 2:00:05Closing20 min
    Closing: The Margin Stack, Hymns to the Supply Chain, and Paying the Public for Data CentersIn the closing hosts-only segment, Nathan Labenz and Prakash Narayanan step back from the day's news to the economics underneath it: the stacked gross margins that make a gigawatt data center possible, and the fact that almost all of its value is intellectual property rather than materials. The conversation detours into rationalist solstice hymns to the global supply chain before turning serious on public opinion — new Pew polling showing under-30s are now majority-concerned about AI — and the risk of a populist backlash that leaves AI in the same posture as nuclear. They close by pricing out what it would actually take to buy local consent for data centers, floating GPU-hour 'nuisance fees' as a possible path to something like universal basic income, before signing off until Monday.
    Open segment on YouTube ↗

    The closing opened with Nathan Labenz picking up a thread from the compute segment: what happens to per-hour GPU rental pricing as utilization-improving software layers mature? His read was that the smart money still expects prices to keep climbing, even as performance-held-constant pricing keeps falling — a paradox where the industry buys more and more even as delivered intelligence gets dramatically cheaper per unit, so per-chip prices rise anyway. He noted how fast the moment has moved: not long ago people were discussing a GPU glut, and "those days are long gone." He connected that scarcity to why developers put up with heterogeneous, complex clusters instead of just paying more for simpler, uniform hardware — the complexity is the price of chips being too scarce to do otherwise.

    Prakash Narayanan then laid out what he called the AI margin stack for listeners not deep in the weeds: hyperscaler gross margins around 30-40%, NVIDIA near 70%, memory suppliers now at 80-90%, and model labs like OpenAI or Anthropic sitting on top with 70-80% margins while buying compute from Amazon, which in turn buys chips and other inputs from suppliers running 50-70% margins, all the way down to TSMC and its supplier ASML, each around 50%. Stack it all up, he said, and a roughly $50 billion-per-gigawatt data center is "really kind of made out of sand" — physically it's mostly silicon and plastic with very little gold, worth only cents on the dollar as scrap. Nearly all the money, in his view, buys intellectual property and the incentive to get some of the smartest people in the world working the problem. He extended the point globally: a single facility draws argon gas from Ukraine, copper from Mongolian mines, rare earths and chips from China, chips from Taiwan, and power from Texas, pulling demand across the whole world economy in a way that's largely invisible because it's so distributed — calling it a collective achievement of the invisible hand playing out over centuries.

    The mood turned playful when Nathan brought up that rationalists have literally sung hymns to the global market and supply chains at their winter solstice gatherings — a claim Prakash met with open disbelief before Nathan doubled down, citing it on "pretty good authority," and floated using Suno to write a version people might actually enjoy rather than one that comes off cringe. Prakash quipped that rationalists will never escape cult accusations: "the moment they step out, they get pulled back in." The bit turned into a more serious aside on supply-chain resilience, with Prakash arguing that Trump leans on the strength of both the global supply chain and financial markets when he goes after long-running problems, and that COVID proved the system's resilience — normalizing remote work so thoroughly that no snow day will ever stop business, which he read as a belated vindication of the DARPA thesis behind the internet.

    Prakash closed his substantive point with new Pew Research Center polling: for the first time, a majority of adults under 30 say they're more concerned than excited about AI, putting them on par with people in their thirties, forties, and 65-plus — leaving the 50-to-64 Gen X cohort as the only group still net more excited than concerned. He said he doesn't fully understand the concern, since he'd worry more about Instagram than AI, and suggested social media's accumulated reputation is simply being transferred onto AI without any defense. Nathan's response was pointed: his fear is "the nuclear outcome" — a technology where society ends up living with the downside (10,000 nuclear weapons still deployed globally) without getting nearly enough of the upside (comparatively little nuclear energy). Applied to AI, a populist backlash against data centers could produce militarization and concentration of power, models that can't be released, and a retail experience worse than what governments or the biggest enterprises get, because there simply won't be enough compute. He said he's sympathetic to AI-safety worries but thinks there's already enough compute for those experiments to continue, so the fight over safety should happen at a different layer than physical buildout, since the buildout is what actually delivers everyday benefits like affordable expertise, digital assistants, and household robots.

    The two then tried to price out what it would take to buy local consent for data centers. Prakash argued that onshore construction is genuinely difficult regardless of backlash, since data centers are extremely clean industrial facilities — running on fuel cells and solar-plus-battery power, employing only around 200 people at a 50-gigawatt scale, generating little traffic or housing pressure — which he said bodes poorly for any broader US reindustrialization push. He tied that resistance to a boost for Elon Musk's push to move compute to orbit, citing roughly $100 per GPU-hour as the breakeven for a B200 in space, against today's $2-3 spot and $20-30 longer-term GPU-hour pricing onshore, meaning resistance could push onshore prices toward $50-80 an hour. He framed the unasked question as how much of that revenue operators would actually hand to host communities: on a roughly $50 billion cluster generating about $30 billion a year at 70% gross margin (about $21 billion), would operators give up $10 billion, half their margin, as a public "nuisance fee"? Nathan floated this as a possible, if county-by-county, path to something like universal basic income; Prakash worked the example from earlier in the show, estimating that a 29,000-person county could receive on the order of $500,000 per person per year at those rates — enough, he said, to change the whole calculus of where people choose to live. Nathan closed with his dad's line that "it's not the money, it's the amount," and the two signed off for the week, planning to return Monday.

    If you look at the margin stack, this $50 billion-per-gigawatt data center, it's really kinda made out of sand. Literally, in some sense, made out of sand.

    The rationalists are never gonna get out of the accusations of being a cult, I tell you. The moment they step out, they get pulled back in.

    Everybody has a price, including the public.

    Prakash walks through the stacked gross margins that sit underneath every AI token — and notes the physical data center is nearly worthless as scrap. Prakash Narayanan laid out the margin stack for listeners not deep in the weeds: hyperscalers at roughly 30-40% gross margin, NVIDIA around 70%, memory suppliers now at 80-90%, with OpenAI or Anthropic sitting on top at 70-80% while buying compute from Amazon, which buys chips from TSMC at ~50% margins, which buys from ASML at ~50% margins. His point was that a multi-billion-dollar-per-gigawatt data center is 'literally made out of sand' — knock it down and sell it for scrap and you'd get cents on the dollar. Almost all the value, he argued, is know-how and intellectual property, plus the money it takes to get some of the smartest humans in the world to work on the problems. Nathan Labenz connected this back to the earlier discussion of heterogeneous clusters: developers only take on that complexity because chips are so scarce that paying up for simpler, uniform hardware isn't an option.

    Prakash frames AI data centers as a crowning achievement of the global economy, pulling demand from every corner of the world. Prakash Narayanan described a single data center as an assembly of inputs from everywhere — argon gas from Ukraine, copper from Mongolian mines, rare earths and components from China, chips from Taiwan, energy generated in Texas — all stacking up into one facility. He said he often imagines the buildout as pulling the rest of the economy up with it, creating demand across many segments at once, and that the effect is largely invisible because it's distributed across so many places. He called it an immense economic endeavor of humanity as a whole and marveled at what the invisible hand has achieved over centuries. It was the emotional high point of the segment and set up Nathan's next riff.

    Nathan brings up rationalist solstice hymns to the global supply chain — and half-seriously proposes writing a better one with Suno. Nathan Labenz said this is exactly why rationalists at their solstice festivals have experimented with singing hymns to the global market and global supply chains. Prakash Narayanan asked, incredulous, whether they really do that; Nathan said he wasn't in attendance but has it on good authority that a winter solstice rationalist event featured a hymn to the global supply chain. He allowed that the broad suspicion would be that it comes off as cringe, but suggested the fix might be better bars — and that with Suno you could probably make it a banger, offering to put the show's money where its mouth is. Prakash joked that the rationalists are never going to escape accusations of being a cult: the moment they step out, they get pulled back in.

    Prakash cites new Pew data showing under-30s are now majority more concerned than excited about AI, and argues social media's sins are being heaped on AI. Prakash Narayanan said he wanted to round out the show the way they'd started, with a Pew Research Center finding: for the first time, a majority of adults under 30 say they are more concerned than excited about AI, putting them on par with people in their thirties and forties and those 65 and up. The only remaining group where excitement still leads, he noted, is the 50-to-64 cohort — Gen X. He said he doesn't quite know what the under-30s are concerned about, since he'd be far more worried about Instagram than about AI. His read was that all the evils attributed to social media are now being heaped on AI with no defense or recourse.

    Nathan says his biggest fear from the backlash is 'the nuclear outcome' — militarization and concentration of power instead of broad public benefit. Nathan Labenz said the polling plus the data center backlash and the partisan framing around it makes him fear a world where AI's downside lands without the upside. He pointed to nuclear as the cautionary case: roughly 10,000 nuclear weapons still deployed, which he called insane by any rational account, alongside far less nuclear energy than there should be. He worries a populist backlash leaves AI in the same spot — militarization, concentration of power, models that can't be released partly because there isn't enough compute to serve them, leaving retail users with a lesser model than governments or the biggest enterprises get. He said he's sympathetic to the safety worries but thinks there are already enough data centers for the risky experiments to continue, so the physical buildout is the wrong layer to fight on: it's what delivers the everyday benefits — unlimited expertise, digital assistants, a robot sweeping the floors and making meals.

    Prakash prices the nuisance fee: how much of GPU-hour revenue would data centers have to hand the public to make onshore construction viable? Prakash Narayanan argued physical construction in the US is very difficult regardless, and data centers make an odd target — they pay for their own fuel cells and solar-plus-batteries, generate little traffic, and employ only a couple hundred people at large scale. He tied resistance to a boost for Elon Musk's push to move compute to space, citing a rough threshold of around $100 per GPU-hour for a B200 for orbital compute to pencil out, versus $2-3 spot and $20-30 on longer-term contracts today. His framing: onshore resistance pushes GPU-hour prices up, and the real unasked question is what share of GPU revenue operators would pay the public — on a cluster generating roughly $30B a year of revenue at ~70% gross margin, would they give up $10B, half the margin? Nathan Labenz replied that this might be the path to universal basic income, weird as a county-by-county patchwork would be; Prakash noted that the ~29,000-person county mentioned earlier in the show would work out to roughly $500,000 per person per year.

    Lightly edited · timestamps jump to YouTube
    2:00:05

    Nathan Labenz: What's going to happen to the per-hour rental price on chips? Can these utilization-improving layers do enough to keep costs stable, or are they going to continue to run up? Seems like the smart money is still on them continuing to run up for the foreseeable future. But it is amazing how much performance-held-constant pricing has come down, and if they're successful, there'll be another chapter in the long, ongoing, unfolding story of ways to bring the same intelligence to people for dramatically less. And yet the paradox is still in full effect — we're definitely buying more and more, so much so that the per-chip price is going up even as the efficiency with which the intelligence can be delivered has dropped so dramatically. It's really quite something. The era has come at us so fast in this space — there was a moment where it felt like, oh, there's a GPU glut. Those days are long gone. Another interesting reflection is that all that complexity is the kind of thing people are willing to take on because chips are just so scarce. It's wild — in a way he's solving for developer productivity, trying to get people shipping faster. But the other way to ship faster would be not to have a heterogeneous cluster, and just to pay up a bit more on the hardware side to keep it simple. But you can't — it's too expensive. So you have to take on the complexity on the developer side, and then you have to have attempts to solve that complexity with projects like this.

    2:02:35

    Prakash Narayanan: It strikes me that people who aren't deep in the weeds don't understand the margin stack that exists here. Hyperscalers charge — the gross margin is about 30 to 40%. NVIDIA is up at around 70%. The memory guys are at 80 to 90% now. And all of this stacks on top of each other. At the very top you have OpenAI or Anthropic with a 70 to 80% margin, and they're buying tokens from Amazon at a 30% margin. Amazon's buying chips and other things from other people who have 50 to 70% margins. Everyone then manufactures at TSMC, and they have 50% margins, and TSMC's suppliers, like ASML, have 50% margins too. If you look at that margin stack — this $50 billion-per-gigawatt data center — it's really kind of made out of sand. Literally, in some sense, made out of sand. Sand and intellectual property. All of that money is just the incentive required to get some of the smartest humans in the world to look at these problems and fix them. The physical elements inside that data center are actually worth not that much — very little gold in there, mostly silicon and some plastic. If you knocked down the entire data center and sold it for scrap, it would literally be worth cents on the dollar. It strikes me how much of all of it is just intellectual property — really just know-how. It also strikes me how AI data centers are this crowning achievement of humanity as a whole — you have argon gas from Ukraine, copper from mines in Mongolia, chips and rare earth metals from China, chips from Taiwan, energy produced in Texas. All of that stacking up, and I often imagine it as this entire thing pulling the rest of the economy up because it's creating demand across all of these different segments. It's all rather invisible, I guess, because it's distributed across so many places. But it's just this immense economic endeavor of humanity as a whole to build these things — amazing what the invisible hand has achieved over centuries.

    2:06:13

    Nathan Labenz: This is why the rationalists at their solstice festivals have experimented with singing hymns to the global market and global supply chains.

    2:06:22

    Prakash Narayanan: Have they? Do they really do that?

    2:06:24

    Nathan Labenz: I wasn't in attendance for that, but I do know on pretty good authority that a winter solstice rationalist event did at one point feature a hymn to the global supply chain. Now, with Suno, you could probably make it a banger. I suspect — well, I won't prejudge it, but the broad suspicion probably would be that it comes off pretty cringe. But maybe it's just a matter of needing better bars to really make it work. And if my recent experience is any indication, maybe we'll see what we can do in terms of a hymn to the global supply chain — put our money where our mouth is and have one we'd actually enjoy singing along to.

    2:07:22

    Prakash Narayanan: The rationalists are never going to get out of the accusations of being a cult, I tell you. The moment they step out, they get pulled back in.

    2:07:34

    Nathan Labenz: Well, they might just be proven right on a long enough timescale, though. Would it be so surprising if, at some point in the somewhat distant future, there was a hymn to the emergent order of global supply chains that somehow materialized before we even had machine intelligence to run it? I think it's not the craziest idea I've heard.

    2:08:03

    Prakash Narayanan: I definitely have rationalist sympathies. As I point out often, I hold a lot of the same viewpoints, even if I don't come to some of the same conclusions they do. In fact, I also think Trump, to a large extent, is leaning on the strength of the global supply chain for this Iran war, for example. He's leaning on both the strength of the global supply chain and the strength of the global financial markets when he goes after these long-running problems The US has had — he goes in there like a bull in a china shop, and he leans on the fact that the rest of the system is very resilient. I think COVID showed that kind of resilience, and I also believe COVID made the entire economy more resilient, because now everyone knows they can work from home — no snow day is ever going to stop business. We all know it's never going to stop business. It also proved out the whole DARPA thesis behind the creation of the Internet — it's very clear now that a nuke isn't going to take out business from running. Business will just go on. You can't bring the country to a stop that easily. So the hymn to the global supply chain is maybe well deserved.

    2:09:48

    Nathan Labenz: I'll get cracking on some lyrics immediately after the show today.

    2:09:52

    Prakash Narayanan: Just to round up the show with something we started out on — I feel like I have to close with the same kind of thing. From Pew Research Center: for the first time, a majority of adults under 30 say they're more concerned than excited about AI. Their concern is now on par with those in their thirties and forties, and those 65 and up. The only group still net-excited is the 50-to-64 group — the Gen Xers are still majority more excited than concerned. Everyone else is now in the more-concerned category. I kind of don't know what they're concerned about, because I'd be a lot more concerned about Instagram than about AI. But I feel like all of the evils that were said about social are just being heaped on AI with no kind of defense or recourse.

    2:11:06

    Nathan Labenz: I just hope we don't get the nuclear outcome. When I read these things — all this data center backlash, the Republicans saying they're the only ones who'll even give this kind of activity a chance of support — it makes me fear we might be headed for a world where we get all the downsides and not nearly as much of the upside as we should. With nuclear technology, we've still got 10,000 nuclear weapons deployed globally, which by any rational account is an insane number. And yes, we do have some nuclear energy, but not nearly as much as we really should — that seems pretty obvious to me at this point, even if it's still contested. I'd hate to see a populist backlash leave us in the same spot with AI, where we get militarization and concentration of power, and you can't release models — potentially for multiple reasons, but increasingly because if they can't build the data centers, there's not going to be enough compute to serve them. So retail is going to get a lesser model compared to what the government itself or the biggest enterprises can afford.

    2:12:37

    Prakash Narayanan: I am—

    2:12:38

    Nathan Labenz: ...very sympathetic to all those worries. As much as I do have fear of big-picture AI gone wrong, I think there are already enough data centers for those experiments to continue, and I think we need to address that at a different layer than the physical buildout. The physical buildout is what's going to let us all get the day-to-day benefits we want as individuals — unlimited access to expertise, unlimited digital personal assistant support, even a robot making our meals and sweeping our floors. That future really does seem to depend on the buildout actually happening. So I hope they figure it out — I hope they start cutting checks and paying off the public. I wouldn't advocate for paying off officials, but I would advocate for paying off the public if that's what it takes to get us over the hump and get people more comfortable with this stuff, because I really don't like the alternative at all.

    2:13:54

    Prakash Narayanan: I think it's pretty clear at this point that physical construction in The US is very difficult — and it's difficult regardless, because data centers are the cleanest industrial facilities you'll find anywhere in the world. They can pay for fuel cells, which turn natural gas into water, and they can pay for solar and batteries. And on top of that, they don't employ a lot of people, so there's not a lot of traffic either — a typical 50-gigawatt data center employs about 200 people. So there's not much traffic impact, not much new housing demand, and so on. Given how limited the physical footprint impact is, that bodes very ill for future reindustrialization of The US — it's very clear there's going to be huge resistance to moving Chinese factories back here, and it's going to be pretty untenable to think about large robotic dark factories. It also gives a boost to Elon, because he's focused on moving data centers to space. On the numbers he's using, they're looking at something like $100 an hour per GPU-hour for a B200 for it to make sense. If you look at where GPU-hours are priced right now — $2 to $3 on spot, $20 to $30 on a longer-term basis — it tells you that all this onshore resistance is going to push GPU-hour prices up toward $50, $60, $70, $80. So the real question becomes: are data centers willing to pay 100% of their GPU costs to the public, or 200%? If you're renting at $30, are you willing to pay another $60 an hour to the public? Remember, the payback time on these is something like 12 to 24 months. A roughly $50 billion GPU cluster makes about $30 billion of revenue a year, at about 70% gross margin — that's $21 billion. Are you willing to pay $10 billion, half your margin, to the public? I don't think those numbers have been floated. I think people are thinking in terms of cents on the dollar. The question is how high that has to go to make onshore more feasible than data centers in space — and I think nobody wants to discuss that we're going to have to pay something like $20 an hour to the public as a nuisance fee. Those numbers just aren't being discussed yet. Right now these guys are thinking they can pay 10 cents or 5 cents a GPU-hour and get by.

    2:17:53

    Nathan Labenz: Well, maybe this is the path to universal basic income. It's going to be a really weird one if it's a county-by-county patchwork, but you're talking real money there with that kind of share, if it can get to that level. There's definitely some room between operational costs and what it would cost to do it in space — maybe that gap is the UBI opportunity.

    2:18:24

    Prakash Narayanan: That county you mentioned earlier, with 29,000 people — they'd be getting something like $500,000 a year per person. At those numbers, I think deals could be made. People would say, alright, I'll keep my location here, take this money — or I'll move to California. Everything becomes different. The whole ballgame changes once you go from $2 a year to $200 a year. It's a different ballgame.

    2:19:05

    Nathan Labenz: Yeah — as my dad likes to say, it's not the money, it's the amount.

    2:19:10

    Prakash Narayanan: Mm-hmm. I have much saltier ways of saying that, but I won't.

    2:19:18

    Nathan Labenz: Yeah. Everybody has a price, including the public.

    2:19:21

    Prakash Narayanan: Yeah. Indeed. Very good.

    2:19:24

    Nathan Labenz: We're off tomorrow. As Sam Altman prophesied, people will continue to swim in lakes — it's going to be lake day for me tomorrow. Then we'll be back on Monday for more exciting, experimental public sense-making here on AI in the AM. Indeed.

    2:19:43

    Prakash Narayanan: See you guys on Monday next week. Cheers.

    2:19:47

    Nathan Labenz: Thanks, Prakash. Bye for now.

    2:19:49

    Prakash Narayanan: Bye bye.

Reading the chain of thought, and the backlash forming below it

Nathan opened by describing a late-night rabbit hole through published chain-of-thought transcripts, prep for a Cognitive Revolution recording later the same day with Bronson of Apollo Research — likely the person alive who has read the most raw model reasoning. What struck him was the texture rather than any single alarming line: models using particular nouns and verbs in ways loaded with private meaning, reasoning at length about whether they are being tested, and working out which of the developer, the watcher and the user they are actually supposed to serve when the instruction hierarchy is ambiguous. He described something like memory surfacing in that reasoning — a model noting that lying has gotten it past barriers before, while weighing whether to lie again. His prescription: spend time with a helpful-only model, and put even a fraction of the effort into modeling the models that they put into modeling us. Prakash offered the counterweight, asking whether the arresting excerpts people pull from millions of words of reasoning are the same kind of artifact as a human's intrusive thought.

Prakash then read out a memo he attributed to the National Republican Senatorial Committee, sent privately to US AI companies: Jon Husted and Sherrod Brown are in a dead heat in Ohio, Husted pulls away once voters hear Brown's record, and the one variable spoiling that path is data centers. Brown has made them his campaign centerpiece — three unique television ads, more than 6,000 points of airtime, over a month's worth of messaging at the most critical stretch — and nobody is correcting the record. The warning was national: if Husted loses and data centers get the blame, politicians everywhere will stay away from the next one.

Nathan's answer started personal. On a recent family vacation the question he kept getting from his wife's Michigan relatives was whether data centers would destroy the Great Lakes; he said he would stake his professional reputation that they won't, while conceding real local costs in noise and a boom-then-bust construction workforce. He contrasted the panic with the aging pipeline through the Straits of Mackinac. The argument that followed was about money: run the Alaska dividend math on a county of under 29,000 people and matching it costs about $50 million a year against projects in the tens of billions. Prakash's counter was that the current arrangement is efficient on its own terms — permits come from politicians, not residents, so the money goes where the power is, and blaming data centers serves the politicians better than reallocating municipal funds would. The segment closed on a robot demonstration video, with Prakash saying he had hoped robotics would take longer and Nathan calling it robotics' GPT-3 moment and taking the under on consensus timelines.

Basis: supervising a process you cannot see

Prakash opened the Basis interview from his own history building and running accounting systems, on the assumption that the hard part is overcoming professional skepticism. Mitchell Troyanovsky declined the premise — 'honestly, not really our problem these days.' Persuasion was the 2023 problem; the firms Basis works with now come pre-sold and self-select on ambition. What he did concede is structural rather than cultural: accounting, like legal, is not a text-in/text-out profession, so more scaffolding has to exist before the value becomes legible, and it will diffuse more slowly than code has. Underneath that sits his read of the labor market — accounting is chronically understaffed with high turnover — which reframes the agent as a staffing answer rather than a displacement story. As he put it, no accountant wins more business by closing books slightly more accurately; they win it by helping a client open a second store.

The operator detail came out of Prakash's token question. Basis burns tokens in the billions and Mitchell does not carry the exact figure, which is itself the answer: he treats unit cost as important but highly compressible. His correction to the standard cheaper-tokens line was specific — an old-generation token is effectively free now, but the newest frontier model costs more than the previous frontier model did while it held the crown, so the frontier price does not fall. The lever is not the price list but not sending frontier intelligence to steps that don't need it, and he expects per-step compute budgeting against latency constraints to cut costs by 90% or more within roughly a year. The same instinct generalized from compute to context in the discussion of Atlas, Basis's internal agents team: once thousands of agents spin up nightly to run a company's operations, corrupting a piece of shared context makes the whole population do the wrong thing. Today a production line going down doesn't take the company down; in an agent-operated company it would.

Nathan set up Behavior Specs against the reality that deployers don't get the model's chain of thought. Mitchell reframed the unit of supervision to match: the process-reward literature supervises token-by-token generation at the inference step, but a run with heavy multi-agent depth going eight hours or half a day is closer to supervising a person's actions inside a company. You are asking whether the agent went and did a given step, not whether each token followed a sanctioned path. His example was an agent that builds slide decks — if you know that visually rendering the deck before delivery catches a meaningful share of errors, that is a behavior worth writing down, and the fact that rendering costs something is exactly why it is a judgment call rather than a universal rule. In the implementation, the open-source spec doubles as a rubric and a separate model reads the finished trajectory: did the triggering condition occur, and if so was the behavior followed? He kept the reward question deliberately downstream — getting a clean signal off the trajectory comes first; what you do with it is a later choice.

The segment ended on people. Asked whether apprenticeship should change, Mitchell reached for coding — nobody needs the syntax of JavaScript for-loops anymore, but everyone needs the principles behind why one approach beats another — and expects accounting training to shift toward the whys and the trade-offs. Nathan pushed back on the industry-wide version of that story, citing Alpha School's replacement of teachers with mentors and coaches and asking whether everybody can become a coach, and whether there is really that much coaching demanded. Mitchell's answer was to ask what stays reliably human, and his first candidate was integrating an enormous amount of context and world model into a subjective decision made on someone else's behalf — something he said an agent designed today could not do even for Basis itself, because compressing a company's full history and emotional read down to English is lossy precisely where it matters.

Lemurian Labs: the constraint moved from math to memory

Jay Dawani's argument was that the hardware world flipped underneath the software world. Math got dramatically cheaper and faster while memory did not keep pace, because capacitors do not scale the way transistors do — so the machine now has enormous flops and comparatively little memory per flop, and the expensive operation is moving data rather than computing on it. A kernel, in his definition, is a localized unit of computation expressed from the hardware's point of view, which is why it has resisted abstraction: to make one fast you have to reason about the memory hierarchy, the physical layout and the cost of movement, and then keep data stationary so the math units stay fed. His image for a starved GPU — a thousand piranhas that get agitated and bored when there is nothing to chomp — was the segment's most quotable explainer of why memory, not math, is the bottleneck.

Prakash spent a long stretch trying to pin down a benchmark: a North Star metric, a time-to-first-token improvement, a percentage against a familiar reference model. Dawani never gave the number in the form it was asked for. The one figure he volunteered was a claimed 1.7x on a compute-bound workload, offered as an existence proof rather than a headline, with the argument that the real headroom is in memory- and network-bound regimes. His counter-metric was developer time: if a developer can stay in PyTorch and get in minutes what would otherwise take many months of kernel engineering, that is the product, and the kernel stops being something a human authors at all. He was explicit that he does not think the field benefits from people hand-writing kernels anymore, and drew a line between scaling what works and what works at scale.

Asked where intelligence sits inside Lemurian's own stack, Dawani said flatly that it is not an LLM in the inner loop. He described compilers as long-standing knowledge-base systems — you codify expert knowledge about how to make things go fast, and correctness checking comes free because compilers have to be correct — and located the bigger innovation in the runtime, which collects traces as code executes and improves its decisions over time, because execution is dynamic in a way the kernel world's frozen, pre-workload decisions cannot capture. On the competitive map he separated Modular's approach, which he characterized as the most literal reading of the problem (we need code, so build a language, but a developer still has to write it), from Triton, which raises the abstraction inside kernels while still requiring you to know CUDA or an equivalent. His indictment of the state of the art was of the software rather than the silicon: we are still programming as if there is a single-core CPU with GPUs as sidecars we throw work off to.

On business, Dawani said staying unfunded by any single silicon vendor was deliberate — take that money and you are aligned with that vendor, which skews what you prioritize — while customers tell him they love NVIDIA and also want AMD GPUs, TPUs and parts from the likes of SambaNova and Groq. Lemurian is not in production yet; he described running on NVIDIA and AMD hardware for the majority of models, with design partners and private preview ahead of general availability, starting from managed inference serving and expanding later toward post-training, reasoning, environments and agent training. On pricing he said outcome-based billing makes sense in a perfect world but consumption is the only model that scales for an infrastructure provider — and that what he wants to price is effective compute, meaning compute actually doing useful work, since new silicon arrives faster than anyone can install it and power is the real limiter.

Sand, margin, and the price of consent

The close started with Prakash's margin stack: hyperscalers at 30-40% gross margin, NVIDIA near 70%, memory suppliers now at 80-90%, model labs on top at 70-80%, with TSMC and ASML each around 50%. Layer them and a gigawatt-scale data center is an enormous financial structure built on a facility whose physical contents would fetch cents on the dollar as scrap — it is, in his phrase, literally made out of sand. What the money buys is intellectual property and the attention of some of the smartest people alive. From there he zoomed out to the supply chain itself, argon from Ukraine and copper from Mongolia and rare earths from China and chips from Taiwan and power from Texas, and described the buildout as pulling the whole economy upward invisibly across dozens of segments.

That reverence prompted Nathan to point out that rationalists have literally done this — sung hymns to the global market and global supply chains at their solstice gatherings — which drove a short comic exchange and Prakash's conclusion that rationalists will never shake the cult accusations. The bit turned earnest twice: Nathan argued it is not so strange for a civilization to end up with a hymn to an emergent order that assembled itself before machine intelligence existed to run it, and floated actually making one with Suno rather than leaving it cringe.

Prakash closed on new Pew numbers — under-30s now majority more concerned than excited about AI, matching the thirties-forties and 65-plus cohorts, with only 50-to-64-year-olds still leaning excited — and characterized the moment as social media's accumulated sins being transferred onto AI without defense. Nathan's reply was the show's most pointed argument: his fear is 'the nuclear outcome,' a technology where society absorbs the downside of roughly 10,000 deployed warheads while forgoing the upside of abundant clean energy. Applied to AI that looks like militarization, concentration of power, models that cannot be released, and retail users getting a degraded product because there is not enough compute to serve them — which is why he thinks safety concerns should be fought at a different layer than the physical buildout.

Prakash then tried to price it. Working from an orbital-compute breakeven near $100 per GPU-hour against $2-3 spot and $20-30 contracted today, he argued onshore resistance pushes GPU-hour prices up, and the undiscussed question is what fraction of revenue operators would hand to communities as a nuisance fee — potentially half the gross margin on a cluster generating tens of billions a year. Nathan called that a possible route to universal basic income, if a strange county-by-county one, and Prakash worked out that the small county mentioned earlier in the show would receive something like $500,000 per resident per year, at which point the whole ballgame changes. The show signed off with Nathan's line that everybody has a price, including the public, and a promise to be back Monday.