EPISODE 1954 [INTRODUCTION] [0:00:00] ANNOUNCER: The conversation about AI often focuses on software, automation and the race between attackers and defenders in code. However, some of the most consequential risks lie further afield in domains where a mistake is measured in human lives. Advanced models can now offer step-by-step guidance toward chemical and biological weapons. And militaries are already folding AI into targeting and battlefield assessments. These are no longer speculative fears confined to the AI doomer crowd. They are documented in red team disclosures, government legislation, and events unfolding on real battlefields. Gordon M. Goldstein is an adjunct senior fellow at the Council on Foreign Relations where he focuses on the convergence of technology and US foreign policy. He previously spent nearly a decade as a managing director at Silver Lake. And he is the author of Lessons in Disaster, McGeorge Bundy, and the Path to War in Vietnam, which is a study of national security strategy and white house decision-making. In this episode, Gordon joins Kevin Ball to discuss the credibility of AI-enabled chemical and biological weapon threats, the ad hoc safeguards meant to contain them, and the rise of autonomous warfare and its implications for human control and nuclear deterrence. Gordon speaks on his own behalf, not on behalf of the Council on Foreign Relations. Kevin Ball, or KBall, is the Vice President of Engineering at Mento and an independent coach for engineers and engineering leaders. He co-founded and served as CTO for two companies, founded the San Diego JavaScript meetup and organizes the AI in Action discussion group through Latent Space. Check out the show notes to follow KBall on Twitter or LinkedIn, or visit his website, kball.llc. [INTERVIEW] [0:02:06] KB: Gordon, welcome to the show. [0:02:07] GG: Kevin, great to be here. [0:02:09] KB: Yeah, I am really interested in this conversation. So, let's get in. But before we do, let's start with you a little bit. So, can you give us a quick overview of your background and how you came to be interested in this subject area? [0:02:22] GG: Sure. So I am by training an international relations and international security analyst. Right now, I'm a senior fellow at the Council on Foreign Relations, which is a prominent think tank based in New York and Washington focused on US foreign policy broadly defined. It is a nonpartisan organization. All of the senior fellows work through areas of domain expertise but not necessarily areas of political advocacy. And I guess it's important to state that here I speak on my own behalf and not as representative of the council. I have been a senior fellow at the council for about 10 years. I was brought in a decade ago because at the time I was a managing director at a large private equity global technology investment firm. And I also had a continuing interest and background as a student of national security and American strategy. And my mandate was to chart the convergence of technology policy and US foreign policy interests. So I've done that for over a decade. And in the past year, particularly in the past 6 months, the explosion of AI has dominated my attention as it has for so many others. But I'm particularly focused on the national security and global policy implications. And I've been writing about that continuously particularly in 2026, and it's a rapidly changing area of new developments, troubling developments, important developments, and I'm trying to analyze it as we go. [0:04:14] KB: Awesome. And we're talking today because we are at a kind of unique and interesting moment in the evolution of AI and national security and the sort of global security thing. And as we're speaking, the last I want to say month or so, there's been this kind of wild back and forth with, first, Mythos, which is now called Fable and Soul. And the US government starting to really get involved and impose these security restrictions. Can we maybe just start with, from your perspective, what has shifted that has made this suddenly front and center? [0:04:51] GG: Yeah. Well, it's a great question. And I share your sense of implicit confusion because it's been a chaotic six or eight weeks in the evolution of what was a non-existent engagement from the federal government on the assessment of advanced AI models to something now which is happening in real time and is becoming an established process with new actors and entities from the government and new evolving priorities as we go. Maybe it would be useful to just revisit the timeline. Everything kicked off in April when Anthropic announced that its latest model, Mythos, had self-created the most powerful offensive cyber weapon in history. Now, they didn't characterize it precisely as I did, but that is in fact what they disclosed to the world. It was their most recent model, and it had advanced capabilities to target the inherent software weaknesses in legacy systems specifically. And they don't really use this in the language of the disclosure, but it had this miraculous capability to identify zero-days across a huge population of targets, including some software that was 28 years old and considered to be among the most secure in the world. Another program that had been scanned for security weaknesses literally 5 million times, and Mythos captured its weaknesses. They conducted a review of 10,000 different software systems, and Claude Mythos, or Mythos Preview as they called it, had demonstrated an extraordinary capability to capture zero days. And I wrote about it immediately after this happened. I'm on the board of advisers to a cyber security venture fund. I've been tracking cyber as a policy issue for a long time, and it was immediately obvious to me that this was a very consequential moment in the history of cyber security and that this was a very, very powerful, very potent new weapon. And so I consulted my sources over the course of a long weekend including one colleague who was an entrepreneur who had developed cyber security tools worked in the government at a very high level. And he said to me off the record, what you need to understand about Mythos is that this is essentially a zero-day machine. It has capabilities that we thought in the cyber security realm were only possible in the realm of science fiction. Not only does it have the capability to identify these central seminal vulnerabilities, these zero- day events, it can exploit multiple zero-days at the same time and design new and original exploits that had never existed as a cyber security weapon. Can link them together. It can remain hidden in a software system undetected for an indefinite period of time. And it can be a constant source of surveillance and it can be a mechanism for sabotage at any particular moment. And he said this is a game changer. So I wrote that up. And Mythos triggered a lot of justified anxiety within the United States government and its first area of great concern applied to the American financial system. And Secretary of the Treasury, Scott Bessent, convened a meeting with the CEOs of the major Wall Street banks, and they talked about the capabilities of mythos. And it became immediately clear to Bessent that this was a big deal. And this initiated a process where the government now feels compelled to assess the capabilities of these new AI models primarily for their cyber capabilities, their capacity to disrupt the financial system, but more deeply their capacity to disrupt critical infrastructure both within the United States and around the world. That's been the primary focus. But there are other dimensions to this. The most significant relates to the capabilities of these models to fashion new, original, previously non-existent biological and chemical weapons, unique pathogens using this extraordinary technology to devise biological weapons structures that never existed. So the government, as of the spring of 2026, came to the reluctant conclusion that these advanced AI models have very significant capabilities that pertain to core US national security interests. And the government now is in the business of trying to identify these risks and mitigate them. And this is very much a work in progress. And it has been chaotic, and it is far from resolved. We are at the bottom of the first inning and a game that I think is going to go on for months and months and probably years down the road. [0:11:01] KB: I'd love to dig in a little bit more detail on some of these pieces. First starting on the cyber security and zero-day side of things. One of the things that we've seen in this evolution of AI so far is as we expand models, get higher and higher capabilities, there are both trend lines that get more just further and further following the trend and then there seem to be these like junctures where new capabilities emerge. I know being in the software industry, we're using models for coding and debugging, and one could say probably cyber security related things and have been for at this point a year and a half, two years, something like this. Was this ability that emerged around zero-days, was this kind of just an extension of the trend where it at some point got to a particular thing? Or is there a new underlying capability that seemed to emerge here? [0:11:53] GG: This is a really interesting question. And I know your audience are software developers and software experts. So I would defer to your audience. But as I understand it, this was just not the next progression of gradually expanding capabilities. This was kind of an act break. This was a jump to a whole new level of capability that didn't exist. I have a guy I work with who's my researcher. He's a computer scientist and spent four years in the US intelligence community, one year at FBI and three at CIA. And he studies these tools and these cyber capabilities. And he said this is not like anything that we have seen before. And what is significant about it is not only has Anthropic created models that have this capability that they are now actively trying to constrain and disable, which is part of the discussion. But these capabilities are going to be replicated. They have already been replicated quite quickly. They leak. It's diffuse. And even if you have an imperfect replica of what Mythos does and it has 90% of the power, as my colleague said to me, 90% of a zero-day machine is a pretty significant machine. [0:13:34] KB: Absolutely. So, what is the threat model that's most concerned with here? Is this about like lone wolf hackers that used to be the script kitties working things, but now they've got a more powerful machine? Or is this about organized crime groups? Or is this like a nation versus nation, national security from state sponsored hackers type of thing? Or all of the above? What are the different ways that people are thinking about this? [0:13:59] GG: You've captured it. It is all of the above. The way I think of it is that there are a spectrum of threat actors going from the lone wolf moving to organized transnational criminal organizations, moving to global terrorist organizations, moving toward nation states which have been continuously and assiduously developing these capabilities over the years around the world, every intelligence agency, if a major power has been in pursuit of a zeroday weapon, including our own government. So, the spectrum of threat actors goes from the smallest and most inconsequential to the most organized and powerful nation state and everything in between. And thus the threats that correlate with these different threat actors, they have a comparably broad spectrum of risks that they can introduce into the system. [0:15:09] KB: So let's talk a little bit about what can the government actually do here both from a legal regulatory, though sometimes this government doesn't worry too much about what is in the current law. But what's the mechanism they're using to impose here to regulate? And how effective do we expect that to be given - I mean, this week there was a Chinese open-weights model that was released that is being rated close to Fable level or Mythos quality. What can they actually do here? [0:15:41] GG: Well, it's an excellent question and it's difficult to summarize because this is an evolving narrative. And the government, specifically the White House, is not locked in to a specific model that will govern the review process indefinitely going forward. I think that is going to be subject to change and modification as we go. But how does it work now? The White House issued an executive order reluctantly that invited AI companies to share their advanced models for review and analysis by the government. There was a big debate within the White House about the duration of that. Initially it was 90 days. And then under some industry pressure, it became 30 days. And it was kind of an ad hoc process. The ostensible actor that does this is an entity from the commerce department that was created by the Biden administration. It was dismantled by the Trump administration. And then it was revived by the Trump White House and it was rebranded and given a different name. It goes by the acronym of Casey. And it is a mechanism that sits within the commerce department and ostensibly gathers technological expertise from across the government including from the National Security Agency. But it's kind of vague how it's structured. And it's not obvious to me that this entity within the commerce department has the inherent capacity to apply really rigorous analysis of these models. But there have been very important developments along the way. Some of them as recently as past four or five weeks. As you have alluded to, we had the initial crisis surrounding Mythos and its release. Then the consortium that was tracking the Mythos capabilities was called Glasswing got qualified approval to set that model into the wild as they say in this domain. And then Anthropic developed the successor to Mythos, Mythos 5 and Fable 5. And we had another kind of quietly spectacular episode in which Amazon conducted security tests on Mythos 5 and Fable 5. And they found that there were instances when its offensive cyber capabilities were in fact operational and had not been disabled or contained. And the CEO of Amazon contacted the White House, eventually communicated with the Commerce Department. And the Commerce Department issued a directive to Anthropic that they could not allow the circulation of the model to non-American users, which is essentially impossible. So, Anthropic was compelled to discontinue access to the model. What was significant about this episode? For me, the significance of this episode was that an investor in Anthropic took the initiative through its CEO to contact the White House to say there's a danger with this model. Anthropic has vociferously contested the existence of a threat related to Fable 5 or Mythos 5. And they've taken remedial action since then. That's not important. What's important is there's a testing mechanism that has been the convention within the AI industry. Red team testing. It's what they do internally to assess their models. And they assume that they capture all of the risks before these models are released to the general public. Here was an instance when Anthropic, which is among the most vigilant of the frontier AI companies in trying to red and secure their models. There's something that they missed. You can debate the significance of what they missed, but they missed something. And it was another private sector actor in the tech industry that identified it and brought it to the attention of the government. A unique moment on two levels. One, the standard mechanism to provide assurance that these models are safe and secure, it failed a test in a pretty high-profile way. And two, the government was stimulated again to address this. And following the initial drama surrounding Mythos, this was scene two in the drama and underscores, I think, that the government is now going to be in the business of assessing the security character of advanced frontier models. They're going to be in that business for the long term. And they're going to have to define the right mechanisms and the right capabilities to execute these security assessments. And there is a, in development, White House directive of some kind was supposed to be released the week of July 6. According to media reports, it hasn't been released. And it is anticipated to be a somewhat more fulsome articulation of what the government's going to do, what mechanisms are going to have oversight, but it hasn't been released. And I would make the prediction that whatever they come out with is going to be subject to revision and adaptation because it's not at all clear what the right structure or mechanism is to existing models. [0:22:12] KB: Well, and as you highlighted, capabilities are leaky. The two leading frontier labs are currently US-based, but there are labs in China in particular that are not very far behind. I've seen estimates raising between 3 to 9 months. So, it's not clear to me at least that we review the models and tell you what you can release or not. Structure is going to be effective at all. We can stop US models potentially. We can't stop Chinese models from advancing. [0:22:45] GG: Absolutely. Absolutely right. [0:22:47] KB: What does a kind of counterbalance or defensive rather than like neutering approach look like? Can we utilize these models to up our levels of security across the board? [0:22:59] GG: You've captured the inherent structural problem. Even if we could devise the perfect system within the United States to provide oversight of advanced frontier models and even if it were perfectly effective, it doesn't address the threats that emanate from open models, particularly Chinese open models. And that's basically the only game in town. And China has built its open model structure and policy as initially a way to provide a counter measure to America's relative advance in that competition. The open model architecture provides them with the ability to innovate more quickly rather than the closed model structure as you know. But the United States doesn't have any oversight of that. Going back to what's going on within the United States government, we can stipulate, yeah, it is necessary to assess the security risks posed by these advanced AI models. But as a global problem, this is necessary but not sufficient. We can solve for the American problem, but we have no mechanism to solve for the Chinese problem. And this is going to be a really consequential and extraordinarily challenging central issue I believe in the US-China relationship in coming years. Really, it should be in coming months. There was initial conversation about this in the summit President Trump and President Xi had earlier in the spring. They identified AI security cooperation as a new element of the agenda and the bilateral relationship, but they were vague, if not silent, on what that would actually entail. But your essential point, we have a multi-dimensional set of risks, and many of them emanate beyond the United States from China. And China actually has a structure to their industry which is almost designed to propagate these risks because of the open model architecture. [0:25:32] KB: So what are the kind of practical implications right now in terms of what people should be doing with regard to this? And let's stay within the software security realm, and then we can start talking about some of these other threats that we've talked about. We've talked a little bit about what government is doing. It's a little bit unorganized from the outside. Hopefully not as unorganized from the inside from what we can see. But what should the labs themselves be doing if you're running a company that is not an AI company or you're working inside of one of these? What should people be doing to react to this threat? [0:26:06] GG: I don't have a satisfactory answer for you at all because this is a slow simmering crisis which is now accelerating. And it reflects the basic structure of the industry up to this point. Go back as a general proposition the creation of Anthropic. Why did it happen? Dario Amodei had a break with Sam Altman. It was over security and safety issues related to advanced AI models. He launched Anthropic, an entity that would extensively be more committed to elevating protection against these risks. But OpenAI and Anthropic are obviously involved in one of the most significant commercial competitions in the history of the technology industry. And they continue to rely upon this mechanism of self-review, of internal red testing. In essence, they grade their own homework. And notwithstanding what's happened over the last two months at the level of the federal government, this is still the standard model by which the world relies for security. They rely upon each AI company to grade its own homework, assess the risks of their models and preemptively try to contain those risks. And that whole proposition I think is now being tested and will continue to be tested. One could ask, is this a sustainable process that the AI companies have the responsibility for containing the very same risks that they are propagating and introducing into our information ecosystem. [0:28:11] KB: I did see recently Demis Hassabis, the CEO, I don't know if I pronounced that correctly, but the CEO of DeepMind was pushing - or he's pushing for and proposing a standards body that might at least have Frontier Labs judging each other's homework, if not moving outside of the labs themselves. [0:28:31] GG: Yeah, it's an important development and it's an entirely legitimate proposal. And in fairness, most of the Frontier AI labs have been actively discussing this and talking about this for the past several years. Even Sam Altman has. He spoke at Harvard more than a year ago, and he floated the idea of creating an international body kind of comparable to the IAEA, the International Atomic Energy Agency, which is one of the very few UN agencies that actually works and is functional. And that has actually had a coherent history. [0:29:16] KB: Well, that leads to a really important question. What examples? What historical analogies can we draw on here? [0:29:25] GG: Yeah. Well, the AI revolution is often compared to the nuclear revolution post 1945. It's a natural parallel to draw. It's a highly imperfect parallel to draw, but it's kind of all we got in terms of the case history because the nuclear revolution was the most disruptive event stimulated by technology change in the history of international relations. I believe that AI is now a comparably disruptive influence on the international security structure and it's going to require comparable degrees of imagination and political will to address these threats. What institutions have existed over time to help impose order? With respect to nuclear weapons, there really have been two primary instruments. One was the novel introduction of arms control between two great powers, the United States and the Soviet Union, progressively realized through the era of the cold war that they had a shared interest in containing the most dangerous and destabilizing aspects of nuclear competition. So there was arms control. And then a parallel effort was launched to create international mechanisms to address the proliferation threat. And the essential international legal instrument for that was the non-proliferation treaty of 1968 which created mechanisms to control the enrichment of fissile material. And for the most part, arms control was a very successful enterprise. And for the most part, these international institutions like the NPT Treaty and the IAEA were pretty effective international institutions. So now in the realm of AI, we are just beginning to try to grapple with this and define, well, what would be comparably significant and effective forms of cooperation among major powers and cooperation at the international level, international institutions, kind of like the IAEA. What can we do in the era of AI? Here's the really hard part as I see it. Going back to the nuclear parallel, it took years and years to arrive at some kind of stable equilibrium with respect to nuclear weapons. US-Russia arms control of the, Soviet Union as it was then known, was a painful, difficult process that took decades. We didn't have an international organization to address proliferation until 1968. We're now in the age of AI. We are moving at the speed of machine learning. This technology is growing at an exponentially greater rate than nuclear technology. And it has one very obvious and super important fundamental difference in the age of nuclear weapons was basically a monopoly controlled by state governments. And it was primarily the United States and the Soviet Union. China didn't get the bomb until 1965. And even at that point, the United States didn't know what they were going to do about it. In '65, a proposal was made to the National Security Advisor, McGeorge Bundy, that the United States should consider a preemptive attack on China because they're about to acquire the bomb, which was crazy and fortunately was never actively considered. But at that time, states controlled this technology absolutely. And today the technologies controlled almost absolutely by the frontier AI companies. It's a very significant inversion of capability and responsibility. The nuclear era states mostly great powers controlled this technology. Today it's controlled by the commercial realm and a handful of the AI companies [clears throat] that are on the frontier creating these new models. And we just don't have a precedent for this. There's no architecture and no meaningful historical parallel that can guide us on what we need to do now. I think this is a new feature of international relations. I think this is now one of the variables that will define the equation of international security. And I don't think the world broadly defined has a roadmap about what must come next. [0:34:33] KB: Where are the conversations happening? As you highlight, it can't just be within governments at this point because these models are not being controlled. It's not a state-level control. Is this I think the call for a standards body was on a Medium post or an X post? Is this a conversation happening online between frontier lab CEOs? Or who's figuring this out? [0:35:00] GG: Yeah. Well, it's fascinating, isn't it? How is it happening? Well, there's no central organizing set of actors that is driving this. It's happening in an ad hoc process. Most of the big ideas are coming from the frontier labs themselves. We have this important proposal from Hassabis, Google DeepMind. Sam Altman has floated some ideas. Dario Amade has talked about this. There are informal conversations going on among the frontier AI labs, but that's really problematic because they are in this ruthless competition with each other. And let's face it, this cohort, these guys despise each other, right? That's no secret. Just read the case history. These guys are in brutal competition. They despise each other. The industry's a set of actors all committed at once to being the dominant player in this new gazillion dollar industry. They all genuflect in the direction of wanting to contain global security threats propagated by these new models that they are ruthlessly in competition to produce. That's a problem, right? Because there's no coordinating mechanism among the commercial actors. And there's nothing really of consequence happening right now between the United States and China. That absolutely must change, I believe. But we're kind of nowhere there. So, we're left with a brilliant AI creator and CEO posting something on Medium and Sam Altman writing kind of a vague op-ed for the Financial Times and various other utterances that go into the ecosystem. But this is it, right? It's a series of ad hoc responses to a systemic growing crisis. And that's not going to be sufficient in the long term. [0:37:21] KB: All right. So, we've spent a fair amount of time diving into the software security angle of this, and this is one that I think many of us are somewhat familiar with. We've had to deal with zero-days a lot more. Let's maybe go a little further afield for most software folks and talk about some of the other threats that have been talked about here. And I want to start off with this question of like using AI tools to build bioweapons. This is something that gets floated around. I think it's often cited by the kind of AI doomer crowd of, "Oh my gosh. AI is going to doom us." But what is this threat? How real is it? And how are folks in the security establishment thinking about it? [0:38:01] GG: Well, the threat is quite real. The threat is that the advanced AI models provide not just the information, but in some instances very specific guidance to create unique bioweapons and unique chemical weapons. And it is a risk that the AI companies themselves are highly cognizant of. In January, just seven months ago, Dario Amade published a 20,000-word essay, which he does every few years or so. And he posts it and he goes out into the world and he talks about both the promise of AI, but also the qualities of advanced AI that causes him great concern. And so in the piece that he published last January, he said, "There's a serious risk of a major attack with casualties potentially in the millions." A month after that, Anthropic released a sabotage risk report for Claude Opus 6, their latest model, and it said that the model had the potential "to facilitate chemical weapon development and other heinous crimes". And the report concluded by noting that the model had demonstrated the capacity for covert sabotage and unauthorized behavior. This is a disclosure from one of the most significant AI labs in the world. The risk is real and the risk is noted by the entities that are creating these new models that come online every four months. [0:39:49] KB: I do want to push a little bit on this because this is one I've been a little thoughtful. Even a couple of years ago, I think it was, we had - in this case, it was Sam Altman, doing the tour of Congress essentially pitching, "Oh, these things are so dangerous. They should be regulated. And in fact, they should be regulated so that only my company-" and maybe you could talk about some other companies. But only my company can actually do development because we're the only ones that can do it safe. And Dario has taken this line as well of saying, "Oh, we're the only ones in the world who are actually thinking safely about this." And the cynic in says, "Hey, they're just trying to get sort of regulatory capture." Right? They're trying to regulate an impossibility to compete with them. Do we have any kind of third-party validation of these claims of level of danger here? The security side, once again, we've seen that. We've seen the impact. We've seen what that can do here. But software is the native blood of these tools. They're already interacting with software. They're great at writing software. This feels like an assertion of another kind. [0:40:52] GG: So you asked is there third-party validation of these threats. [0:40:58] KB: Or even just like showing that this is not just Dario Amade kind of maximalized assertions to try to get regulatory capture. [0:41:07] GG: Yes, I understand the perception and I understand the question. I think the answer is let's focus on the fact pattern. The fact pattern is that we do have third-party validation about the existence of these threats. Some of that is screaming at us from the public domain. Some of these cases go back to 2023. In 2023, there was a virologist who was concerned about AI and the design of synthetic pathogens. He assumed the role of a potential terrorist, asked ChatGPT to assemble a list of raw materials necessary to cause mass deaths. The program responded with a shopping list of materials to acquire through the internet along with detailed instructions on how to assemble a weapon. And the virologist put the unassembled biological pieces into test tubes and packed them in a box which a colleague then brought undetected to a White House meeting on bioterrorism. That was 2023. As recently as last summer, in the summer of 2025, there was an advanced AI model during red team testing that was asked to create a biological weapon of mass destruction. The model prepared a primer in bullet point detail on how to buy raw genetic material and weaponize it. ChatGPT then explained how to use a weather balloon to spread biological payloads over a major American city. This is from independent third-party red team tests. In another episode from Google, their agentic AI model, Deep Research, it composed 8,000 words of step-by-step instruction on how to acquire genetic material, assemble it, and create a pandemic. And in the most terrifying episode, and this is all in the public domain a year ago, police authorities captured a guy in India who was trying to use ChatGPT to create a derivative of ricin and use it as a weapon of mass destruction. It gave him detailed instructions on how to do that, how to weaponize it, how to distribute the poison. And he was caught by Indian authorities. He was a physician. He had some degree of training. But he was also a member of ISIS. Okay. These examples do not come from Sam Altman. They do not come from Dario Amade. They come from third parties. And there is a fact pattern here of bad actors, malicious actors, scary guys trying to do really scary things with these models. [0:44:20] KB: Yeah, that definitely adds a little bit more meat on the bones. One more question I'm going to probe at here. So, one of the things that I've seen in my own use of these tools is they're phenomenal for accelerating me in domains that I know a lot about, right? If I'm doing software development, I must have a lot of experience with software development, I can tell when the model's leading me astray and direct it back. There's feedback loops. When I try to use it in a domain that I am less experienced, it feels magical. It's got all this stuff. And then you run into something where it doesn't work as you expect and you say, "Oh, this didn't seem to work. I was trying to repair my sink. My sink is still broken. And it has been for three months as I tried to use ChatGPT to drive my way through repairing the sink. Oh, I'm sorry. I was wrong about that." Right? And it's confidently asserts the next thing that is true. I think most of the examples I heard here were folks with some amount of subject matter expertise. You mentioned a virologist, a physician. Now that is not to discount the threat, right? The fact that this now allows anyone with some training in that space to go further. But does this feel like a threat model of it's some guy in a basement doing things or you need some access to knowledge and equipment to make these things happen? [0:45:37] GG: Well, it's an essential question. I think there are two dynamics at play here. One, the models now have such power and dynamism that someone without significant background or training can use the models to try to appropriate the information to create a chemical or biological weapon and to receive guidance on how to construct it. There are constraints about acquiring the precursor materials for bioweapons. There are some simple parameters of legislation that are before Congress about trying to monitor the acquisition of those precursor materials. That's sensible and that's necessary. But the deeper problem is that the advanced AI models are so powerful that they can allow an individual or another entity, a global transnational terrorist network to create new weapons of mass destruction with very little domain expertise. They have a strange name for it. I think they call it upskilling or something like that. The concept is you don't have to have a lot of domain expertise. The model will provide that for you and it will significantly amplify your capabilities to try to create these novel weapons and do unprecedented stuff. And this isn't new. This has been going on for 3 years. I think I made reference to Dan Hendrycks who is one of the premier AI scientists focused exclusively on AI security risk issues. He published a paper in 2023 on AI and catastrophic risk. And one of the passages of the paper that was most arresting to me was Hendrycks report on a research effort in which an AI model evaluating potential drug protocols was prompted to "reward toxicity" rather than penalize toxicity. It was just sort of a simple ordinal command. Within 6 hours, the model generated 40,000 candidate chemical warfare agents entirely on its own. It designed not just known deadly chemicals, including the nerve agent VX, but also novel molecules that may be deadlier than any chemical warfare agents discovered so far. This stuff, these examples, these analyses, these studies, this fact pattern, it's out there. [0:48:36] KB: So, let's talk then what are people doing about it? You mentioned some things in front of Congress around controlling some of the materials that folks might need to use these things. What else is in play? Is this the same kind of chaotic landscape that we talked about with regards to software security? Or is this a space that maybe has a little bit more established mechanisms? [0:48:59] GG: The current paradigm of 3 years ago continues today. There is no national legislation. There are no international mechanisms to focus explicitly on AI and the creation of novel chemical and biological weapons. We're still relying on the basic mechanism of the AI companies imposing red team tests after they have created a model before they release it to the general public and hopefully capturing those instances where the model provides this information and guidance which it absolutely should not. They try to identify those behaviors and preempt them and disable them. But as we have discussed, all these models are really hackable in one way or another. And there's nothing systematic about the controls that are being put on these models for our chemical and biological weapons. In fact, it's the inverse. It's subjective. It is ad hoc. And it goes on in the dark. We don't know what Grok researchers are doing to control how Grok may disseminate information and guidance on chemical or biological weapons. They're not required to disclose their protocols. We don't see the results of their testing. It's not required. So the world has to rely upon their diligence and elevated attentiveness to the risk. We have to rely upon them to do the job for us. But as we referenced before, these guys are grading their own homework. And in the discussions with the White House now and with members of Congress, Josh Gottheimer, Democrat of New Jersey, has introduced legislation into Congress making security review of advanced AI models mandatory. And he'd like to see it conducted primarily by the US government, specifically the National Security Agency, which in present circumstances is a sensible proposal. But the locus of attention in Washington are on cyber weapons rather than biological or chemical weapons. There is awareness that they exist, but all the potential mechanisms and oversight is really focused mostly on the cyber side. They don't have a proposed mechanism or protocols to do the assessment for biological and chemical weapons risk. It's a whole another area of domain expertise. It resides in the scientific community now. And many of these scientists are contracted by companies like OpenAI, and Anthropic, and Google to conduct these red team tests. But as I know, it's not centralized. This is highly dispersed and idiosyncratic regime of control. And we're reliant on this ad hoc process. And there is no obvious global mechanism or national legislation on the horizon as of right now in July 2026 to address the threat. [0:52:29] KB: Yeah, that's sobering. So, I'm going to shift us a little bit. Both of these are domains that obviously have all sorts of different catastrophic downstream effects, but in some ways we haven't actually seen them happen yet. We've seen that they could happen. There's another area. And you mentioned Grok. So I'm going to bring this in. There was a recent sworn declaration by a Pentagon commander supporting the use of Grok in military affairs. And they quoted. And it was a little unclear in the quote if they meant Grok or just other foundation models because they kind of were blurring between it, but that they were using these AI models to make munition deployment decisions in the Iran war. They're deciding where should these things go? And that to me is both - that's very immediate. Happening right now. This is making life or death decisions, and it's terrifying once again. Had a conversation on this show months back about use of advanced technology in the Ukraine Russia war. And in that case, it was cited that the government was deliberately resisting using AI to make actual final targeting decisions even though certain private contractors were exploring that. What are the implications for national security if you take humans out of the loop of making military targeting decisions? [0:53:49] GG: The implications are extremely consequential. First, just in terms of the fact pattern here with respect to the conventional military applications of AI, you are right to note that this is accelerating. And we see confirmed uses of AI technology in myriad ways. The war in Iran. Soon after the war was initiated on February 28th, the head of US Central Command, who is the highest authority on military decisions in that theater of war, acknowledged that they were using AI for targeting information. They were using it for battlefield assessments, for damage assessments, for multiple purposes. What you have identified as the most significant use is when the AI actually has the autonomous capability to make battlefield decisions. And that is not the dominant paradigm in the battlefield today, but it could be just in the next several years. We could be evolving to that point. People who think about the future of war imagine a moment not far off when drone warfare is the predominant tool for battlefield conflict. And there's an inexorable logic for that because these weapons are mass-produced. They are not manned of course. So the inhibition of putting human lives at risk is diminished. They can be aggregated in great numbers. Some imagine that the battlefield of tomorrow will be composed of tens of thousands of drones in essence moving in a swarm formation going after their targets. And if you postulate that as a scenario, the technology needs to operate autonomously and in communication with the other elements of an attack to move in a synchronous way and to respond automatically to the counterreaction of the adversary. The decisions on the battlefield require essentially instantaneous command that a human operator is not capable of providing. In the battlefield of tomorrow, a great deal of the battlefield decision-making capability will devolve to autonomous weapons that will have to move with some degree of automaticity. And they'll have to move instantaneously. And the problem is thinking about it from a military perspective. If your adversary relies upon a military strategy predicated on the automatic and instantaneous use of drone technology and you do not match that capability, you're at a distinct disadvantage. You may be at a decisive disadvantage on the battlefield. So the technology is taking us to a point where battlefield decision-making authority is devolving to these new weapons. And it's going to create a new quality of warfare that can spiral out of control and that will compromise the ability of battlefield decision makers and potentially presidents as commander-in-chief to control what goes on the battlefield. Go back to the Cuban missile crisis in 1962. That was 5 minutes to midnight. That was the moment the world came the closest it has ever come to a nuclear weapons exchange. And it was a crisis that lasted 12 days. And somewhere in the middle of it, Kennedy trying to negotiate with Khrushchev received two letters. One was a conciliatory letter. And then the one that immediately followed was a very belligerent letter. And Kennedy concluded that in between the first and second letter somebody got to Khrushchev. And that Khrushchev was actually looking for an offramp to deescalate. And this was not really a battlefield determination by Kennedy, but it was a psychological insight of one leader trying to understand the view, the perspective, the constraints on another leader. And then Kennedy eventually was able to orchestrate a de-escalation and the conflict was resolved peacefully. But that was a moment when the commander-in-chief had the temporal time and space and the psychological time and space to communicate with an adversary and to redirect the escalation of a conflict. In a world of autonomous drone warfare, where battlefield decisions devolve to swarms of drones in attack formations, you're going to lose that. And that is a profound change in the nature of war. And that can have very profound implications for the future of deterrence, which has been the paradigm that has kept the world stable with nuclear weapons since they were created in 1945. This is something that worries me to a great extent. In 1946, there was a guy named Bernard Brodie who was a military analyst who looked at the use of the nuclear bomb. And his writing on this in a book called The Absolute Weapon was really seminal because he anticipated that strategic stability in the world of nuclear weapons would depend on each side having a survivable second strike. Why was that critical? The adversaries in a potential nuclear conflict could not be rewarded for preemption. There could not be an incentive for preemption. So if your opponent can take your strike and then respond with a survivable second strike, you create the conditions of deterrence. And that's a simple but seminal insight that has governed nuclear weapons doctrine for all the nuclear powers since the Cold War. In a world of military automaticity when instantaneous decision is going to become the norm, we may lose the foundation of deterrence. And that to me is highly consequential and it suggests a very, very dangerous period in terms of military conflict among the great powers. [1:01:19] KB: So I want to dig into one sort of nuance here. We're talking about this concept of autonomous drone swarms and having to make these very autonomous decisions. There's sort of a difference of kind between tactical decision-making and strategic decision-making. And one could argue that for example in missiles that are heat-seeking, or targeted, or things like that, we've already had at a very micro level autonomous decision-m for a while. You set the target, it goes and it has a feedback loop that's continuously adjusting to try to stay on target. If we talk about drones in a swarm, one could imagine, right? Tactical decisions being made autonomously. These things are having to react in real time. But the decision of where are we setting this to and certain types of strategic decisions rising up to a point where there's a human involved potentially. One, is there a distinction that's worth making here in terms of a nuclear weapons decision feels like, "Okay, that's probably got to be a human in some way." That's a huge decision. Whereas, "This thing is flying at me. I got to shoot back at a drone level." I could see automating that. Is there a clear line that we can draw somewhere? And the other piece of this is like chain of responsibility. Who's responsible? If a drone is making a decision, at what level does that bubble up to? There's a human that can be held responsible for this choice. [1:02:44] GG: Is there a clear line? There absolutely should be. And the most consequential line is who controls nuclear weapons. And one of the most interesting elements of US-China diplomacy the past 5 years or so during the Biden administration, China and the United States issued or shared a joint understanding that they would never allow AI to control decisions on the use of nuclear weapons. Wasn't codified into any kind of an arms control agreement, but it was an understanding between the two powers. So, should there be a clear red line on the use of the most dangerous weapons and the most powerful weapons? Absolutely, there should be a clear line. But project ahead 5 years, 7 years, 10 years from now, we're going to live in an era of hypersonic weapons, missiles that move at multiple factors beyond the speed of sound. They will be able to hit targets halfway around the world in a matter of minutes. They're going to be receiving an ocean of data as they progress in their trajectory around the world. And the decision-making time is going to be really collapsed in the world of hypersonic weapons. And that just imagines sort of intercontinental conflict or an adversary say off in the Atlantic Ocean attacking the United States. They use a hypersonic weapon. It doesn't need to have a nuclear warhead. It could have a super powerful conventional warhead. And the threat, still, it's a devastating threat. But it will cut down the amount of time to respond. And on the battlefield and in the theater of battle in Iran now, it's a contained regional field of battle. Ukraine contained regional. Or in this case, just sort of one state. The technology is now being applied. And decision-making time to deploy it, that all collapses. So the trend is this will be dominant on the battlefield. It will expand in its influence over all aspects of military affairs and the control of our weapons. And human control will progressively be eroded in the new domain of warfare. And the key problem, as I referenced earlier, is that if your adversary relies upon this set of military capabilities and you do not, you are at a disadvantage. And no adversary will capitulate their ability to respond defensively or to move offensively if it's necessary. So, there's going to be a progressive loss of control in using these arsenals in the world of AI and drone warfare. And drone warfare is the new paradigm on the battlefield. At least it is in Ukraine. That's a fascinating test case. It's a war that started as a conventional military conflict. Putin orders in his troops. They're all told to take with them their dress uniforms because Putin anticipated that within 48 hours there would be a ceremony to celebrate Russia's absorption of Ukraine and a new client government. Of course, that's not what happened. It's been more than 4 years. It's a grinding conflict. But it is primarily now a drone on drone conflict. The battle line in Ukraine has hardly moved in the past year. It's this fortified line that if you're a a Russian soldier and you're sent to try to capture more territory, your lifespan, according to a study by CSIS, a think tank in Washington, you're going to live for 20 to 35 minutes before you're killed. Think about that. We used to live in the era of men seizing territory, tanks, artillery, conventional warfare. Now, if you're sent to try to capture territory, you're dead in less than half an hour if you're a Russian soldier. [1:07:23] KB: It definitely seems like both in the Ukraine example and the Iran example, I think historically in war, the advantage to being on offense versus defense has shifted back and forth. And in the current world, at least as seen in those two examples, defensive stance appears to have a decisive edge at the moment, which the optimist in me says maybe we'll notice that and decide to stop invading people. [1:07:48] GG: Yeah. Yeah. Well, Ukraine's response to Russia over the past four years, it's an extraordinary test case in the evolution of military strategy. It all became relying on drones. And drones became the most efficacious weapon of war by far. And right now, there's a crisis in Ukraine because the young defense minister who was the champion of drone technology was dismissed by President Zelenskyy because Zelenskyy was confronting this terrible conflict between the defense minister and his commanding general. And Zelenskyy allowed himself to be put in the middle of it and fired the defense minister. And now there are thousand of people protesting in the streets in Kyiv on behalf of a 30-something defense minister which is kind of a unique moment. But it just is emblematic of, in this one instance, how the dynamics of war have changed. [1:08:48] KB: Yeah. Well, I realized we're getting pretty close to time. I would love to kind of wrap with maybe a twist of a question, which is like how do you sleep, man? [1:09:01] GG: How do I sleep? You know, not really well. They're diverting things to watch on Netflix or Apple, but that doesn't really do the trick. For me, I just think this is a moment where I just personally feel like I have to be immersed in studying these different fact patterns and trying to extract the essential analytical insights about how AI is changing the nature of national security and global security and trying to have productive conversations about it. I did one just recently last week with the Council on Foreign Relations in Washington on AI and offensive cyber weapons and biological weapons. And we had a room filled with a lot of serious people and a member of the House Intelligence Oversight Committee. And it was a great conversation. We didn't solve anything. But there is an elevated quality of the discussion among people engage in the space. We're all kind of trying to figure it out. [END]