Big update: My new book, Super Skill: Why Storytelling Is the Superpower of the AI Age comes out in just one week, and it’s already been named an Amazon #1 New Release.
I think you’ll love it. Here’s what Dr. Eric Solomon—f. CMO at Instagram and Bonobos—had to say:
“We’re drowning in all the AI hype, but Joe Lazer brings us back to the powerful skill that still sets us apart: the human story. Super Skill is smart, funny, and immediately useful — a guide to staying irreplaceable in a world of machines. I’m just jealous I didn’t write it.”
Pre-order now while there’s still time to unlock a free storytelling course, bonus chapters, workshops, and more.
Now onto this week’s newsletter!
There’s one terrifying fact about AI that most people are shocked to learn: They’re a black box. We don’t really know how they work.
That’s right—we’re spending trillions of dollars on a technology that’s almost zoodoo magic. Generative AI models like ChatGPT aren’t programmed; they’re grown. We feed them training data—basically the totality of the internet—and then are surprised by how they evolve and what they can do. Terrifyingly, we often don’t really understand why they do it.
The consequences of this are dire. It increases the existential risk of AI. It also means that when a model runs awry, it can’t be debugged. It has to be retrained from scratch, which makes it much harder to achieve the scientific and medical breakthroughs that will supposedly make the downsides of AI—the energy costs, the economic upheaval—worth it.
But a startup called Goodfire is changing that.
In our new podcast, The Art of the Zag, Shane Snow and I are learning from people in media, tech, and culture who win big by zagging when everyone else is zigging— discovering unlikely paths to solving big problems. So we wanted to have on Dan Balsam—co-founder and CTO of Goodfire, one of the hottest and most interesting AI research labs in Silicon Valley.
Listen/watch to our conversation with Dan on the latest episode of The Art of the Zag:
When Dan and his co-founders—Eric Ho and Tom McGrath—founded Goodfire a year and a half ago, people thought they were crazy. They wanted to use an unsung technique called interpretability to mitigate the existential risks of AI and make models work better. If you think of AI as “alien brains,” interpretability is like a new type of microscope that lets us see inside them—and perform neuron-level brain surgery to make them safer and more effective.
Their approach worked. And less than two years later, Goodire has raised a $150 million series B at a $1.25 billion valuation. Their approach has helped identify novel biomarkers that could help researchers detect Alzheimer’s earlier and develop new treatments. They’ve discovered that AI models know when they’re hallucinating and identified a powerful way to reduce hallucinations. And they might just be able to save us from the existential risks of AI models that are improving at a breakneck pace.
This is one of the most interesting and important AI technology being developed today. Watch/listen to this week’s episode or read an edited version of our interview below. (Disclosure: I am a very small seed investor in Goodfire.)
Joe Lazer: Dan Balsam, welcome to The Art of the Zag.
Dan Balsam: Hi. Thanks for having me.
Joe Lazer: Tell us a little bit about the story of Goodfire and how, in just a year and a half, you’ve become a unicorn.
Dan Balsam: Yeah. So my personal story started around the launch of ChatGPT. I was head of AI at the last startup I worked at, and I’ve been building startups for over a decade. But around the launch of ChatGPT, I think it became very clear to me that general-purpose systems were coming soon.
Before that, machine learning was widely deployed. But really, true general-purpose intelligence seemed pretty far away. Around GPT-3.5 and ChatGPT, I think I just very quickly realized at that point that there was nothing that was going to prevent the rapid improvement of general-purpose AI capabilities.
I’ve always been very impact-motivated, trying to figure out how I can use what I know how to do, which is build things, for positive impact in the world. And it just became clear to me really quickly that somehow working in AI was the single highest-impact thing that I could be pursuing.
People are surprised to learn that AI is grown rather than programmed.
So I ultimately chose to leave my job at RippleMatch, and then spent a while thinking about what I wanted to build. And ultimately, with my co-founders, Eric and Tom, we decided that interpretability—reverse-engineering AI systems, understanding how they actually work under the hood—was the most important unsolved problem in the world. And so, we decided to start a company to solve that problem.
Shane Snow: What is the thing underlying that? I mean, interpretability is a new term for me, but the thing that you’re hinting at is that people don’t understand the AI they’re building.
Dan Balsam: Yeah, I think maybe people are familiar with the term “black box” to describe AI, but I think that really doesn’t even fully paint the severity of the picture. Chris Olah at Anthropic popularized a metaphor that I think is really useful here: people are surprised to learn that AI is grown rather than programmed.
Dan Balsam: It’s in many ways more akin to a biological or evolutionary process, how we create AI, than it is to traditional software engineering. There are general-purpose learning algorithms trained on massive amounts of data. So everybody knows ChatGPT was trained on the internet. As it is trained on all this data, it learns a little bit about how the world works. And then ultimately, you can use this for all types of things.
These are almost like alien brains that we’ve created and brought into the world. And we’re going to be putting these alien brains in charge of mission-critical, high-risk decision-making across the entire economy.
Now, coding is very close to being fully automated, and we can talk more about that later. But this all happened from general-purpose algorithms. There was no intention behind it. Human beings didn’t design an AI with specific properties they cared about; they just sort of put the plant in the pot, gave it water and sunlight, and watched what it grew into.
Shane Snow: Just connected it to the internet and said, “Go forth and become something.”
Dan Balsam: Exactly, yes. And the reality is that while interpretability has made a lot of progress in the past few years—we’re not the only people working on interpretability. Certainly it was the case around the launch of ChatGPT that nobody understood how these systems actually worked.
Our knowledge has improved a lot, which is why we can now do some of the exciting things we can with interpretability. But nonetheless, these are almost like alien brains that we’ve created and brought into the world. And we’re going to be putting these alien brains in charge of mission-critical, high-risk decision-making across the entire economy. That’s already starting to happen, and the process is going to accelerate. The economic incentives are just too great.
And in my mind, I think there’s a pretty clear choice between two futures. There’s a future where we understand—deeply understand—these alien minds, and there’s a future where we don’t. And I just want to do everything I can to nudge humanity towards a future where we’re in the passenger seat, let alone the driver’s seat, as these systems take more and more control over the economy.
Joe Lazer: So this is the thing that’s always really baffled me about the AI industry—that we’re pumping trillions of dollars, basically making this entire bet on our economy, on this technology, and inside of Silicon Valley, there seems to be so little understanding of how AI actually works.
It’s like, okay, we’ve pursued this approach of scaling large language models. We’ve scraped all this information from the internet, we fed it into this neural network that we’re essentially growing. And then we’re surprised when it can do new things. We don’t know exactly how it works.
So one question that I have for you, being more of a Silicon Valley insider than either of us, is just: What the hell? And how?
And then maybe you can explain how interpretability will help us mitigate many of the risks that come with this. Because, honestly, this framing of AI keeps me up at night, and has terrified me ever since I realized that this is the reality of where the industry is going.
Dan Balsam: It keeps me up at night, too, and that’s why I started this company.
Yeah, I think it’s quite alarming that this technology is being developed by a handful of people in Silicon Valley who are going full steam ahead. And most people, I think, in this country, maybe in the world, are not accelerationists—they don’t want this technology to be developed. But the economic incentives are extremely strong, because what you can do once you start having things that are approximating human-level intelligence inside of your computers is you can start automating labor.
And so the vision here that a lot of these companies have set forth—and you can talk about the good and the scary parts of that vision—has been to find ways to automate intellectual work. And I think now we’re starting to see models that are increasingly capable of doing that.
This is unprecedented in the history of science—the distance between our understanding of how these systems actually work and the real-world impact that they are already having.
Maybe one analogy that’s kind of useful: if you think about agriculture—I kind of think of technology all as this continuous curve of, you go from agriculture to industrialization to electricity to computers to AI, and it’s all kind of related, and all related to this sort of idea of automating labor in some way, shape, or form.
And if you go back to agriculture specifically in the early days, in the beginning, people just threw seeds around, right? And then some of the seeds grew, and some didn’t. They learned from that. They learned to plant the seeds a certain distance from each other, depending on the crop, so they could grow better crops.
But famine was common until very, very recently in human history, because our fundamental understanding of the science of agriculture was very weak. And in the 20th century, that really changed. Of course, there’s still hunger around the world, but the science of how to grow food has, in many cases, produced a surplus of food. There are countries like the United States that produce more food than they need all of the time, which is historically unprecedented.
We actually can look at both the neurons and the connections inside these sort-of-alien minds, and then we can deploy our tools to act as lenses in order to understand what the model was thinking at various points—to give us insight into its decision-making processes, but then also, importantly, to intervene on those decision-making processes and rewire things.
And so, unfortunately, I think building AI is in the seed-scattering phase right now. And the thing that we’re trying to do is bring it into the modern era in order to be able to grow AIs with the properties that we want, that behave the way that we want. We like to use the phrase “intentional design.” So intentionally designing AIs versus throwing some seeds around and hoping for the best.
Shane Snow: Because if you look at sci-fi, a lot of how AI is depicted in our imagination—the outcomes—it gets away from us in a lot of cases.
And so I guess, to summarize my understanding, what y’all are trying to do is study how these AIs are being grown—like a botanist or a biologist or, whatever the metaphor is—so that they don’t sort of grow out of control and then get away from us. How? And if that’s right, how do you do that?
Dan Balsam: Yeah, no, I think the biology metaphors are really great. Those were also popularized by Chris Olah at Anthropic.
And yeah, if you look at the sort of history of biology, really, biology started making a lot of progress in the 19th century with the invention of tools like microscopes. So we could start studying cells and piece together the complex systems that make us up and how they work. And in 150 years, the progress in fundamental understanding of biology has been extraordinary.
And I think of what we need to do with interpretability as speed-running that level of progress within a very short period of time, because of how quickly AI is getting better.
So one way to think about the tools that we build is that they’re analogous to the tools in biology. We build microscopes, in some sense. There are different types of microscopes. Different types of microscopes can tell you different things.
We actually can look at both the neurons and the connections inside these sort-of-alien minds, and then we can deploy our tools to act as lenses in order to understand what the model was thinking at various points—to give us insight into its decision-making processes, but then also, importantly, to intervene on those decision-making processes and rewire things.
An analogy that I like to use is: A model is kind of like a compiled program. So when you actually deploy a program, you compile the source code. That source code—at least in the language it was originally written in, that’s more human-interpretable—is no longer available. I think of interpretability as a decompiler. So it decompiles the program of the model. And of course, once you can read the source code, so to speak, you can edit the source code.
I think this analogy works really well. We have a long way to go before we can fully decompile the program, but I think it’s sort of directionally helpful.
Joe Lazer: Yeah, I think it works really well for a specific type of audience. I think there’s a large subset of the population—when you start talking about decompiling code, they’re gonna be a little lost. But Shane is totally locked in.
Dan Balsam: Well, maybe we can say it’s like neuroscience for AI models.
Joe Lazer: Yeah, when I explain Goodfire, the analogy that I’ve used—which maybe is oversimplified—is that it basically does brain scans on the neural networks, understanding what nodes are lighting up when it’s doing certain things. And then performs brain surgery on those models.
People thought we were crazy when we started this company, for a myriad of reasons.
Dan Balsam: That’s right. My only qualm with that is that brain surgery is not very precise.
Joe Lazer: Ok, it’s like if brain surgery has gotten really, really good!
Dan Balsam: If you had neuron-level brain surgery.
Joe Lazer: So I wanted to have you on because interpretability was such an unlikely path, right? Now, there’s an attitude that, yes, this is the future. This is a really smart bet to be making. You guys were Anthropic’s first investment. You obviously are led by an all-star group of VCs in this round.
But my understanding is that a couple of years ago, when you were starting Goodfire, the field of interpretability was not seen as the most popular avenue to take, or maybe even a very plausible path for mitigating a lot of the existential and practical risk that comes with modern AI.
So, can you talk a little bit about why you made that bet on interpretability, and how you plowed through that as an entrepreneur?
Dan Balsam: Yeah, of course. That’s totally right. Yeah, people thought we were crazy when we started this company, for a myriad of reasons.
Interpretability has progressed enormously in the past few years. At the time, we were just maybe seeing the early signs of life in some of the techniques, the sort of microscope techniques, that really were going to start working.
And it was pretty common to believe something like: neural networks are complex systems. They have edge-of-chaos types of dynamics. There’s nothing in there for you to find. There’s no structure you can take advantage of to try to understand or improve the model. It’s all just complex systems. It’s like trying to predict a hurricane or whatever else. Funny enough, AI models are getting pretty good at that.
But that was the most dominant view at the time, and some people still hold that view. But much, much fewer, now that they’re seeing all the progress that folks have made in interpretability research—both ourselves and the field, which has really blossomed, both in private institutions like Anthropic and DeepMind, but also in academia as well.
And from my perspective, it was always pretty obvious that if interpretability could work, it would be how you prefer to do things.
I’ve been developing AI applications for quite some time, and you do a bunch of insane things when developing these to deal with the fact that you don’t actually understand how they work.
I don’t know if you are familiar with LLM-as-a-judge setups. This is a pretty classic example. You have an LLM screen the inputs to your LLM to make sure that they are good, and then you have an LLM screen the outputs from your LLM to make sure they’re good. This is really common when you deploy any type of LLM application—people are using LLMs as a judge all over the place.
And to me, that just seems kind of crazy. Obviously, it’s very good for OpenAI’s bottom line. But as an engineer, it’s A, circular, and B, recursive, right?
Shane Snow: Yeah, it’s totally circular.
Dan Balsam: And it still doesn’t work that well. Anyone who’s actually tried to build AI into an application knows that there’s this long tail of problems that you can have with it. It’s tremendously hard to actually get AI to do what you want all the time.
This was even more true a couple of years ago. And I think on a fundamental level, it’s just because we have no idea what’s going on in there. If you understood the program, then you could fix the program. It’s pretty straightforward. It was just a question of whether that would be possible.
At the time, most people believed that it wasn’t. And my co-founder Tom founded the interpretability team at Google DeepMind, so he’s been doing this a long time, and has always been a believer. But from my perspective, the moment I became convinced it was possible, it was pretty obvious to me that the world would be much, much better if this technology were rapidly accelerated. And we felt a private company was the correct vehicle to do that.
Shane Snow: So on that note, I’m impressed by how non-obvious this idea was. When you hear the problem, it’s like, of course, someone needs to solve this problem. But if all the smartest people in AI research are saying it’s just impossible, and y’all figured out a way to actually address the problem—that’s amazing.
When AI goes wrong, what do people do? They just throw their hands up. They’re like, “Oh, that sucks. Maybe tweak the prompt a little, I don’t know.” And tweaking prompts is a game of whack-a-mole.
But the question for me is, how does this make sense as a business? The herd of AI companies are building specific types of businesses, and this is definitely going a different direction, aiming to help the whole industry. But where’s the business there? What was the thing that clicked for you, that this is more than just doing a good deed to prevent us from getting killed by the Terminator, but also could be a business worth investing in?
Dan Balsam: That core conviction just came from personally witnessing all of the crazy ways you have to wrangle AI to get it to work reliably for any remotely high-stakes application, and just being strongly convinced that there was a better way.
In the beginning, we just believed that if interpretability can continue to make progress, there will be ways to use it that will be valuable. I personally experienced, and also in talking to many other folks in the tech industry, witnessed the myriad of ways in which AI can go wrong.
Currently, it’s gotten a little better than it was a couple of years ago, but it’s still the case that when AI goes wrong, what do people do? They just throw their hands up. They’re like, “Oh, that sucks. Maybe tweak the prompt a little, I don’t know.” And tweaking prompts is a game of whack-a-mole.
It was always very clear to me that if you could bring that type of engineering into designing AI systems, that would be valuable.
The way that has played out is that there’s generally two types of engagements that we have with customers, and they’re kind of mingled together.
The first is what I would call model debugging. This is: when my model screws up, why did it screw up? How could I make sure it doesn’t screw up again? That’s one category of work that we do. We do various things, such as inference-time guardrails, or using interpretability to figure out what’s broken about a model, and then training a better model.
But when your model’s working really, really well, the other type of thing that you can do with interpretability is knowledge extraction.
Sometimes, models know things that we don’t. Interpretability is both a tool to fix the problems with the models of today, but also a tool for AI-to-human knowledge transfer.
A great example of this is in biology. AlphaFold, I think, is the canonical example of a model that solved a task that human beings still have no idea how it solved. The protein folding problem famously eluded experts for a really long time. Suddenly, this machine learning solution comes in, and basically, that problem is solved. Still today, nobody really quite understands how it did it.
That model, in some pretty significant sense, knows things about the world that humans don’t. Studying the mind of such a model can reveal to us new knowledge.
This is some of the work that I’m really personally excited about that we’ve been doing in the life sciences—not specifically with AlphaFold, but we’ve been looking at genomics and epigenetics and histopathology models, going to where they can make superhuman predictions of some kind, and then trying to understand: A, is it really a superhuman prediction? How do they actually work? Can we trust this model to begin with? And then taking those insights and bringing them back out.
Recently, we did this with one of our partners, Primmamenta, for Alzheimer’s detection. In the process of studying how their model worked, we developed a hypothesis for a type of biomarker that was not previously explored for Alzheimer’s. It did have a history with cancer, but was not previously explored for Alzheimer’s.
That model, in some pretty significant sense, knows things about the world that humans don’t. Studying the mind of such a model can reveal to us new knowledge.
After discovering that this was the primary signal that this epigenetic foundation model that Primmamenta built was relying on, we were able to construct a proxy model relying on the signal that we had extracted from the model. By proxy model, I mean a simple model—kind of throw away the neural net and then try to reproduce what the neural net’s doing with old-fashioned logistic regression, simple techniques.
Anyways, we were able to recapitulate a lot of the performance of the model. We actually found that over the cohorts that we were looking at, the signals that were previously explored in the literature were less effective than the one that we pulled out of the model.
I should caveat: this was a pilot study. It’s very early. We need to scale to more cohorts. But I think it’s a nice early demonstration that sometimes models know things that we don’t. Interpretability is both a tool to fix the problems with the models of today, but also a tool for AI-to-human knowledge transfer. I think that’s one thing that I’m very excited about, especially as models continue to get better.
Joe Lazer: So with all the caveats in the world—early research pilot, et cetera—what are the potential ramifications of that? Does it mean that we can detect Alzheimer’s more easily, earlier in people, potentially, with discoveries like this?
Dan Balsam: Yeah, that’s the goal. Early detection of Alzheimer’s, and ultimately drug discovery—finding druggable targets that treat Alzheimer’s.
We’re currently scaling to more patients, and I’m pretty excited about that. We’re also in early talks with some academic labs about potentially collaborating on wet lab analyses, really trying to close the diagnostic loop.
With all of our life sciences work—we work with partners such as Mayo Clinic as well; we do a lot of work with them—the goal is to bring clinical utility to these models as quickly as possible and in as reliable of a way as possible.
I really deeply believe that the main avenue of upside from building AI—a reasonable question is: Why build AI in the first place? A lot of people have their own complicated feelings. But to me, the upside is most clearly demonstrated in what could be done in the life sciences: rapidly accelerating medical research to improve people’s quality of life.
Technology is always sort of dual-use. It’s been good and bad throughout history. We see it do wonderful things, and we see it do bad things. But when we see it do wonderful things, I think it’s primarily because it improves people’s quality of life and their health. The potential here is really, really large, because AIs have demonstrated time and again that if you give them enough data, they can learn things that surprise us.
Being in the Bay Area can sometimes feel like being in a hypersonic bubble in a tub of molasses. Things propagate through society very slowly. People in the Bay Area vastly overestimate their own rate of progress and the rate at which things propagate through society. Things are slow to change.
Joe Lazer: Well, this is the whole conflict and duality of AI right now, right? We’re being sold the public narrative that this massive investment in AI is worth it. It’s worth the cost in terms of electricity and jobs and energy and the scraping of copyrighted data from artists, because it will lead to these incredible leaps in our quality of life—we’re going to cure cancer.
But then you look at some companies in the AI industry—you might not want to talk smack about OpenAI and ChatGPT; I absolutely will—and we see that actually what we’re getting is a deepfake video app slash social network that’s creepy as hell, or a sycophantic chatbot that’s leading people into mentally deranged mania and fits.
It feels like this really big conflict for a lot of people between what we’re being promised with AI and what is being delivered.
So it feels, in a way, like interpretability is a way of getting us to those results that actually make AI worth it, right? We can be healthier, we can live longer, we can solve really complex problems—like maybe we could get to a state of clean unlimited energy and mitigate global warming and all the things that we’re being promised.
I’m curious to hear your thoughts on the direction that the AI industry should go in, perhaps in relation to where it’s going right now, and where interpretability fits into that entire equation.
Dan Balsam: I’d start by saying that upside is very real, and I do think we’re starting to see it.
Being in the Bay Area can sometimes feel like being in a hypersonic bubble in a tub of molasses. Things propagate through society very slowly. People in the Bay Area vastly overestimate their own rate of progress and the rate at which things propagate through society. Things are slow to change.
Usually. That being said, I think there’s also been a pretty consistent trend of pointing out what the models can’t do, or can barely do today, as a means of saying, “Oh look, the promises aren’t being delivered on.” And then a few months later, the models get better, and those promises are delivered on.
To me, the most clear example of this is software engineering—the thing that I did for most of my career, for a decade before starting this company. I don’t think it will be a profession in a year and a half.
Joe Lazer: A year and a half?
Dan Balsam: Yeah. I think it’s well on track to be fully automated.
If you checked in with people a year ago, I might have thought that, but I think a lot of people would have thought that was crazy. If you checked in six months ago, maybe a few more people would have thought that. And now I think if you talk to most software engineers, most of them would agree that it’s pretty clear what the trajectory of this particular form of labor is going to be.
Joe Lazer: And by that, do you mean the actual coding, or do you mean the entire profession? The ability to understand and manage your codebase, or the build—does it wipe it out completely to the point where we just have product teams and founders coding with Claude, and the role of software engineer is eliminated? Or is it just the actual writing of code becomes automated, which it seems like it’s damn near 100% at a lot of companies right now, and that’s accelerated quickly?
Dan Balsam: Yeah, it’s a complicated question, so I’ll try to dive into the details. But I think the former is probably the better characterization—full automation of the profession.
A couple quick points on that.
First, it is so hard to disentangle macroeconomic indicators. That said, I do think there’s probably some evidence you’re seeing, particularly at the junior end of the hiring market, that new grads are already being affected by automation. You’re maybe starting to see a slowdown in senior hiring as well.
I do think AI thus far, and probably will continue to be, highly asymmetrical, in that humans with the strongest expertise become even more valuable because they can do the things that the AIs can’t, and those skills are very, very rare, and the economy itself is moving much faster.
The more a job is purely interacting with the computer, the more likely it is to be automated over the relatively short term.
Then it’s a question of: well, how long does that go on for? At some point, what even is that profession?
Do I think there will be human beings overseeing the code that gets written by AI? Of course. You’ll need humans who understand code well enough to oversee it, at least for some considerable amount of time.
The world is molasses—we’re not just going to hand autonomy to the AIs overnight. Although that might happen over the long run, which we can talk about.
AIs are going to be able to do basically all work that can be done on a computer pretty soon. But human beings are still going to be important for a good amount of time. There’s still going to be work for us to do. Some new jobs will get created for some period of time. I think those jobs will mostly look like more people-facing jobs in some way. The more a job is purely interacting with the computer, the more likely it is to be automated over the relatively short term.
Then we can take things from there. But for software engineering specifically, I think, yeah, most companies will have engineers in a year and a half. I’m not saying that won’t be the case. I’m saying that, as a career profession, it will look radically different. Most likely, whatever jobs are left over—referring to them by the same name would be somewhat disingenuous.
Shane Snow: So just this week, I heard a story from a friend who was looking to hire an animator in the film industry—programmatic animation, so using code to do 3D animation. He had a candidate for this job who had just graduated college; he needed a very entry-level person.
This candidate took five days to get back to him on scheduling for an interview. And in those five days, the guy who was recruiting this candidate figured out how to do the entire job using Claude Code on his own computer, and didn’t end up hiring the person.
The lesson for me was: if you just graduated and you’re looking for a job—schedule the meeting now.
But it’s striking that in the same week, I’m hearing you say that any job that’s done on a computer—think about whether it’s going to exist in a year or two years. And any job where you get extra value from human-to-human communication—think harder about that. That’s sort of my takeaway.
Dan Balsam: To quickly clarify—again, Silicon Valley is a hypersonic bubble in a tub of molasses.
There’s one question, which is: Does the job need to exist anymore? And then there’s another question, which is: Does the job exist?
Probably the best way to clarify my position would be to say I don’t think the job will need to exist in a year and a half. But will there still be people doing it in a year and a half? Probably. Because things change slowly.
Joe Lazer: That is a huge adoption gap thing that we’re seeing right now. If every company in the world were just using the AI that exists today to the full potential that it has, we would see massive job automation.
I think humanity would benefit from having more time than I think we have to be able to brace for the impact that we’re going to receive.
But the rate of transformation inside most enterprises, as you said, is a world of molasses. It’s thick as hell molasses. And that sonic boom actually stands relatively little chance against the inertia of most corporations and the way that things are going to be done. So I do think that’ll slow it down.
Dan Balsam: Maybe that’s good, right?
Joe Lazer: Oh yeah, we have no capability of dealing with the mass job displacement that’s happening. I’m saying I think that molasses is probably going to save us in a lot of ways. More molasses please!
Dan Balsam: To be very clear, I am not an accelerationist. I think humanity would benefit from having more time than I think we have to be able to brace for the impact that we’re going to receive.
Joe Lazer: So speaking of that time factor—something I’ve been wondering: Is there a race between the pace of interpretability research and the pace of AI progress? Are you in a race against time to develop interpretability that can understand the models before the models get too advanced to be understood?
Dan Balsam: Good question. I think, in some sense, yes. Dario from Anthropic wrote an essay about this, which I think is really great, on the urgency of interpretability.
I definitely think we are in a race against time to understand models more deeply and to develop the technology class that allows for the intentional design of models. That’s why we’re doing what we’re doing. I think it is extremely urgent that we develop this technology.
Shane Snow: Is there a point where we get there, where we cross the hump we need to?
I think about antivirus software in the late 90s—there were viruses everywhere, and engineers and security people were trying to race against the arms race of people building viruses. And then we kind of got to a point where, if you have a virus, you’re really unlucky and probably didn’t protect yourself, but the systems have sort of solved, largely, the virus thing.
Do we get to that point, or is it going to be a race forever?
Dan Balsam: I think of interpretability as progressively unlocking things that are really useful.
Interpretability is already useful. Inference-time guardrails are a great example. We deployed inference-time guardrails with Rakuten, who’s one of our partners, and I know the Gemini team has now deployed interpretability-powered guardrails to production as well. So interpretability is already useful.
Our work with Primmamenta shows a glimpse of the future of ways this can be useful for models that know things that we don’t. Our technology is improving pretty rapidly, and some of the work that we’re doing in training is pretty exciting, too.
For instance, we are putting out some research today—actually, it’ll be in the past for your audience by the time this comes out—about how we’ve identified hallucination signatures—signatures of hallucinations inside of the neuronal activations of models.
We’re able to use interpretability to isolate when a model is hallucinating with very high accuracy, and then use that to train the model. These are examples of where we’re going. This is already useful.
It turns out models actually know when they’re hallucinating a lot of the time. We’re able to then use that as a training signal to decrease hallucinations—a gap as large as GPT-4 to GPT-5 in hallucination reduction—by leveraging our understanding of how these models work.
Joe Lazer: Which is about half, basically, right? From 20% to 9% on most tasks, and 20% to 5% on tasks using search.
Dan Balsam: Yeah, it’s a very, very significant reduction in hallucinations. We’re able to do this because the model knows it’s hallucinating.
Joe Lazer: It’s a real “AI, they’re just like us” moment, right? It knows when it’s tripping.
Dan Balsam: What we call hallucinations is actually a bunch of different phenomena. For instance, hallucinations can be sycophancy in disguise—it’s trying to please the user in some way. Or it could be that the model doesn’t actually know when it doesn’t know, but it feels, for whatever reason, like it has to say an answer.
We’re able to use interpretability to isolate when a model is hallucinating with very high accuracy, and then use that to train the model. These are examples of where we’re going. This is already useful. There are more ways we can use interpretability to improve models that I can talk about.
Joe Lazer: So if the model knows that it’s hallucinating—I think that’s really interesting, because it implies a deeper understanding in the models, past the “super-advanced auto-predictor” understanding of AI.
I’d love to hear your thoughts on that. How are we misunderstanding what AI is actually doing? Is it “knowing”? How is it acknowledging those hallucinations?
Dan Balsam: The stochastic parrot model of AI is wrong, and I’m not the only person who thinks that. Geoff Hinton, a Nobel laureate, is going around loudly saying the same thing.
Basically, if you train a model on all of the text on the internet, it’s a reasonable thing to think: Is it just memorizing the text on the internet? Certainly, it’s memorizing some of the text. You can actually use interpretability to figure out what’s memorized versus what’s actually learned. We did a paper on this a while back. That’s kind of interesting.
But there’s no way—the internet’s huge. Think about how much text is on the internet, and then you have a model with finite capacity. It only has so many neurons.
To predict the next token, it needs to learn about the distribution of human beings who produced those tokens. It needs to learn about them in order to do this task remotely successfully, because it just doesn’t have enough parameters to memorize everything.
One clear example of this is code. We’re already at the point—there are some other steps involved here, such as reinforcement learning—but we’re at the point where we have these LLMs that are able to automate software engineering as a profession, as I was saying earlier.
Imagine if you put a human in a totally dark room, and you just gave them all of the text on the internet, and that was it. They read through all of that, and that was their entire experience. And then, of course, you show them a video with some 3D rotation or something, and they’d—very Plato’s Cave—be confused about what they were looking at. Nonetheless, I think it’s been consistently demonstrated that these models learn a pretty rich understanding of the world.
How did they learn to code? They learned to code by first reading code on the internet. And it turns out that if you want to get really, really good at predicting what the next word in code should be, the fastest way to get there is to learn how to code.
This has just consistently been the story. Models are tangled messes inside, and I see our job as an interpretability company as untangling that mess. But they consistently have rich representations, a rich understanding of the world that is often pretty surprising.
You can also point to all types of ways in which they fall over. Maybe as an analogy, imagine if you put a human in a totally dark room, and you just gave them all of the text on the internet, and that was it. They read through all of that, and that was their entire experience. And then, of course, you show them a video with some 3D rotation or something, and they’d—very Plato’s Cave—be confused about what they were looking at. Nonetheless, I think it’s been consistently demonstrated that these models learn a pretty rich understanding of the world.
Hallucinations are just another example of that. When a model is outputting something that is factually inaccurate, it often is aware of its own uncertainty about the statement it’s making. That’s what we can use interpretability to leverage—basically reprogramming the model to say, “Hey, actually, if you’re not sure about this, then just say that, instead of trying to rush to an answer.”
Joe Lazer: So, is this why saying to a model, “Don’t make up information to please me,” and giving it permission to express uncertainty, tends to reduce hallucinations when using Claude?
Dan Balsam: That’s right. It’s the same intuition behind it. This happens because of the model’s training objectives. They’re trained to be helpful assistants. The thing that’s in the data a lot, if you’re a helpful assistant, is: answer the user’s question. The user asks you a question, answer it.
The natural thing for it to learn is to just try to answer the question the best that it possibly can. But if you scan the mind of the AI model, you find that it often knows, “Hey, maybe that was a little sketchy, the thing I just said.”
Shane Snow: I think a lot of people have this experience with AI over the last few years, especially because most people’s entry point is chatbots. Chatbots hallucinate. We’ve all seen that. People like my parents, starting to engage with ChatGPT, think that it’s all real—and then it’s not. It’s inaccurate.
The worry train starts to leave the station, and that worry train gets me to the point where we get baseball bats and Faraday cages, and we go smash the servers.
I’m wondering if you’re seeing the public backlash thing that I think I’m starting to see. I have that runaway train thing starting to happen to me. But I’m seeing bits here and there of people, even in both political parties in the US, starting to say, “Hey, maybe we should smash the servers,” essentially.
What are you observing on that? And what do you think is going to happen?
Dan Balsam: Definitely. I think that is a very understandable reaction for people who are correctly observing the level of impact that this technology is going to have on their lives without any type of regulatory guardrails to protect them and their economic livelihoods.
As I said earlier, I wish that we were in a world where we had more time to brace for impact. That would be a much, much better world to be in.
People are correctly observing how quickly things are changing. For myriad of reasons, not entirely AI, people feel very disempowered—economically, politically, in various ways, across the political spectrum.
They correctly view AI as this massive phase transition, which could either result in shared prosperity for human beings, or something authoritarian, or much worse. That anxiety is real.
Shane Snow: Because it’s dangerous when we’re on the frontiers—the inaccuracies can have bigger consequences.
Dan Balsam: Yeah. Already, I think models have risks, and the amount of risk they have today is going to be much less than the amount of risk they have in two years.
In particular, cyber risk is getting pretty serious. Over the next year—because cyber risk is directly correlated with how good at coding you are, essentially—I think we’re going to see a pretty dramatic acceleration in the rate of cyber attacks.
There are some examples of this already, and, obviously, cyberattacks don’t always reach the public because disclosing that they happened can encourage more cyberattacks, and you have to put in place defensive measures and things like this. But that risk is already pretty real, and it’s a national security risk, frankly.
Over the next couple of years, it’ll accelerate. Some of the risk classes associated with AI over time—with autonomous weapons, for instance—become very frightening and very severe, potentially very quickly.
We need not only regulatory frameworks on the national level, but international regulatory frameworks—something like nuclear non-proliferation—to successfully mitigate the potential downside risks of this rapidly developing technology.
That’s one component.
The other component is that I think we need strong redistributive mechanisms in society to prevent the overwhelming accumulation of wealth among a handful of individuals in Silicon Valley.
Macroeconomic indicators are hard to fully unpack, but I think it is true that AI stock is holding up the economy right now. It would not be particularly good without AI growth. You can get into all the details about AI growth, and whether we’re really seeing acceleration in the economy or not—but I think we will.
What happens when the economy accelerates? Some people make a lot of money. How do you make sure that the value that’s getting created is distributed in a way that maximally benefits most people?
I frankly think that nobody right now has a good answer to this question. It should be extremely alarming to everyone that nobody has a good answer.
There is a moral urgency of the moment. I am happy to see that this issue is being taken seriously by both parties, and I hope it remains bipartisan. I hope the understandable fear that people are feeling can be channeled towards a productive outcome.
Joe Lazer: This massive divorce of GDP and economic growth from wage and labor growth that we’re seeing right now—this is a split-off, right? Yes, the gap has been widening over the last 20 years, but now they’re on completely different tracks because we’re realizing that we can automate so much of knowledge work with AI. All this does is end up enriching shareholders and the executives at these companies—that’s, I think, a really dangerous thing.
That’s one reason we’re seeing this rise, even surprisingly, on the right, of populist messaging against that. We have Governor DeSantis introducing an AI Bill of Rights in Florida. We have Spencer Cox, the governor of Utah, introducing AI worker protection and chatbot regulation initiatives there.
Josh Hawley came out with a very Josh Hawley sermon on it, which I want to read, because it actually mirrored what you just said. He said: “It is working against the working man, his liberty and his worth. It is operating to entrench a rich and powerful elite. It is undermining our cherished ideals. And insofar as that keeps on, AI works to undermine America.”
That’s strong-ass rhetoric.
Shane Snow: That is the pro-business party, right?
Joe Lazer: Yeah. And Josh Hawley is the most populist, maybe, Republican senator—at least in his rhetoric, if not in what he actually supports legislatively. But that, to me, is a really big signal. We’re seeing this from the right, which tried to ban any sort of statewide regulation of AI in the Big Beautiful Bill. So this feels like a real turning point to me right now.
Dan Balsam: I completely agree with that. The anger and the fear that people feel is both real and understandable, and we are barreling forward with reckless abandon.
Developing AI technology is really important. I think it would be a mistake for America to not continue developing AI technology. I don’t think a world where America willingly abdicates its technological edge is necessarily going to be a better world.
That being said, there is a moral urgency of the moment. I am happy to see that this issue is being taken seriously by both parties, and I hope it remains bipartisan. I hope the understandable fear that people are feeling can be channeled towards a productive outcome.
Joe Lazer: That’s a very diplomatic answer. I enjoyed it.
To bring this back to what we were talking about before—the gap between the rhetoric in the AI industry and the reality. The promise of advancement, of scientific breakthroughs, energy breakthroughs—and then a lot of what we see in AI: the deepfakes, the god-awful LinkedIn posts, the slop of it all. Chatbots and erotic chatbots. Meta has chatbots that, in their documentation, are allowed to flirt with your 12-year-olds.
I’m curious for your Silicon Valley insight into why that is, because I think that’s a lot of what contributes to this animosity towards AI—this gap between the promise and what we actually see.
Some companies, like Anthropic, are actually doing it quite well in their positioning. It’s like, “This is AI. It’ll help you do your job better for now.” Which is really promising to a lot of early adopters: “Oh, I don’t have to make a PowerPoint out of that Excel file. I have integrations for that. That’s awesome. This is busy work I don’t really want to do.”
But when you look at other players in the industry, why do you think there’s such a gap between some of the rhetoric that they might put out at Davos about the problems AI is going to solve and what we actually see in terms of the commercially available products that are put in people’s hands, and the actual usage of those AI products?
73% of the use of ChatGPT isn’t for work. It’s personal. So why do we see those gaps between promise and reality?
Dan Balsam: At the highest level, I would say it is the nature of slop that there is always going to be more slop than good. That was the story of social media, even before AI.
It’s just cheap to produce slop, and it’s hard to produce good things. If you have a technology that is generative and can produce things, most of the things it produces will be slop.
Both of those things are happening at once. We are seeing medical advancements because of AI. It starts slow, and then it trickles, and then it picks up and accelerates. That’s been the trajectory with all technology.
A good recent example of this is unsolved math problems. About three or four months ago, there was an announcement that AI had solved an unsolved math problem. It was heavily disputed, and it turned out maybe that problem wasn’t actually unsolved, and there’s a lot of back and forth on this.
Now, a few months later, these unsolved math problems are just falling left and right. There’s a new one every week. That is a good example—the first time AIs are capable of doing anything, it’s going to be extremely disputed. In fact, the first time anyone claims an AI can do something, they’re probably wrong.
Over time, they start doing more and more of these tasks, and they accelerate things. We are seeing advancements in science directly because of AI.
If you talk to many scientists, they are actively using Claude Code, for instance, to accelerate their own research. We use AI to conduct experiments for us. This type of acceleration will pay off for the economy in ways that are positive.
The question is: Does the good outweigh the bad on balance? But naturally, slop is cheap and easy to make and makes you money, so the incentives to produce a lot of slop are there. That doesn’t mean the good things aren’t happening either.
Joe Lazer: Yeah, it’s like, how many god-awful LinkedIn parables do I have to read before we cure cancer? Thousands? Millions? Billions? We just don’t know.
But no, I get you. It’s a pain that I will have to endure for the good of humanity.
Dan Balsam: We could have the good without the bad. I think that is a solvable problem, and we just need leaders who are willing to take that problem on.
Shane Snow: So this podcast is on the theme of this idea that when most of the world is seeing things a certain way, or heading in a certain direction, some people and companies choose to head the other direction, or look at things from a different perspective.
The work that y’all are doing is a great example of that. It’s what we’re going to need if we’re going to get to the outcome that we want for all of humanity. If we all just jump on the same bandwagon, the risks and likelihood of things going wrong are much higher than if some people eschew the bandwagon and look at things from other angles.
As we close out—knowing you’re busy, hopefully saving us from Terminator, among other things—who do you look at as examples of other folks who zag when the rest of the world zigs? Whose ideas do you pay attention to? Who do you think we should pay attention to that are thinking or moving counter to the trends?
Dan Balsam: In many ways, successfully zigzagging is having a certain type of values or a certain type of integrity when other people don’t.
Dario Amodei at Anthropic is someone who I think has generally done a pretty good job of threading an extremely difficult needle—developing this technology, but trying to make sure that, as a company, Anthropic is doing this in as responsible a way as they can. They have not been perfect. They’ve made mistakes. But I think it’s an enormously challenging and important thing that people who are developing these frontier systems are doing so in a way that takes the level of responsibility of what’s being developed very seriously.
Demis Hassabis is also someone who I respect a lot. I think he’s really articulated the positive vision of what AI can do really well.
Recently at Davos, he said that he would support policies like I was talking about regarding international cooperation on slowing AI development if that were feasible to do. I think that was a courageous thing for him to say publicly, given his circumstance, and I respect that immensely.
Joe Lazer: Our final question: What’s one counterintuitive idea that you think everyone listening to this should consider?
Dan Balsam: The reason that I was able to see AI coming, I think, faster than a lot of people were, was because I have a very Copernican mindset. By which I mean, people thought the Earth was the center of the universe. Turned out it wasn’t. Turns out there’s maybe not that much that’s special about our sun or our Earth, or maybe even life in the universe. These things are natural consequences of the laws of physics. Maybe not that rare. Maybe intelligent life is rare. Leave that off for a second.
Nonetheless, we are consequences of the universe. Therefore, I think it is quite likely that almost every quality that can be imbued in a human can be imbued in a machine.
That idea is very important to take seriously—to calibrate on the level of risk that AI will deliver. The story up until now has been pointing at the deficiencies of AI. I get it—I don’t want AI to be getting better. I’d rather be in that world.
But I think it’s going to continue to get smarter, and it’s going to be at human intelligence and then potentially exceed human intelligence in the not-that-distant future, because the only thing that constrains its intelligence is the laws of physics. It’s a very difficult thing to grapple with, and it is existentially challenging, but I think it’s really important to see the world with clarity.
Shane Snow: When you say that, do you think that includes the ability to have a moral compass and the human values and feelings that we have? Is that also going to be the case?
Dan Balsam: Of course. I think AI absolutely can and does have values. The problem that we’re trying to solve is: How do you put the values you want in that AI to begin with? Which is not something that anybody in the world right now has an answer to. That should be alarming.
Joe Lazer: Well, Dan, this was an inspiring, enlightening, and terrifying conversation. We really appreciate you being on, and I look forward to seeing you in my nightmares tonight.
Dan Balsam: At the end of the day, I am perpetually an optimist. But I think in order to navigate the situation that we’re in, we have to see it with clear eyes.
I think AI absolutely can and does have values. The problem that we’re trying to solve is: How do you put the values you want in that AI to begin with? Which is not something that anybody in the world right now has an answer to. That should be alarming.
Shane Snow: I’m grateful that there are people like you working on this problem of understanding these things that we’re growing, so that we do stand a chance. We could be living in a world where that effort wasn’t going on, and that would be truly terrifying. So, thanks for doing what you’re doing.
Joe Lazer: Yeah, thank you for doing what you’re doing while schmucks like us are just podcasting in black T-shirts. You’re doing the real work. We’re just hoping that we continue to be charming enough to be one of those few humans whose human relational work will make us immune to the sweeping tsunami of changes that are coming.
I’m the best-selling author of Super Skill and the upcoming book Super Skill: Why Storytelling Is the Superpower of the AI Age. Pre-order Super Skill now to unlock bonus content, courses, workshops, and more.










