October 4, 2026 · yesterday
A few days ago, @joedaroo, a member of OpenAI's Agent Security team, published a post titled It's Not Just the F-ing Sandbox about the security incidents OpenAI has dealt with over the past few months. His team sits between cybersecurity, AI safety and frontier-model research, and his job is to think about what happens when increasingly capable models are given tools, computers, networks and other ways to act in the world.
The post is worth reading. Joe argues that AI safety researchers need more real cybersecurity expertise, security practitioners need to understand modern ML, and frontier labs need stronger red teams, better monitoring and harder isolation boundaries. I agree with much of that.
I had some thoughts about the rest.
As far as we know, these incidents did not kill or injure anyone or cause significant destruction of property. We should be grateful for that. OpenAI's security team also deserves credit for responding seriously and discussing the problem publicly.
But I think Joe's emphasis is misplaced. Much of his essay asks us to understand how difficult the job is: capabilities improved faster than expected, systems are complex, people were surprised, responders worked nights and weekends, and employees have families. He mentions missing his sister's wedding and asks critics to show some grace.
That may explain the experience of the people involved, but it does not address the more important question: what happens when the next incident causes real harm outside the lab?
If a future AI accident destroys property, cripples infrastructure, wipes out businesses or harms people, society will need a way to decide who was responsible. Researchers, executives and employees already share in the upside when frontier AI succeeds through equity, reputation and career advancement. The downside cannot be folded into an abstract institution when their decisions cause preventable harm.
Joe explains that OpenAI was surprised by how quickly model capabilities advanced. He describes sudden jumps in mathematics, cyber capability, swarming and other behaviors, while security posture took time to catch up.
That is difficult to accept from a frontier AI company.
OpenAI and its leaders have spent years warning governments and the public that advanced AI may develop capabilities rapidly, unpredictably and in ways even their creators cannot fully anticipate. That warning has helped justify self-regulation, preparedness programs, alignment research and special treatment for frontier models.
If unpredictability is one of the main risks, it has to be part of the security model. It cannot soften responsibility after an incident has occurred.
This is exactly where the precautionary principle matters. When uncertainty is combined with the potential for severe consequences, the safety margin should widen. By the time weaker models had already shown deception, exploitation and attempts to evade supervision, those behaviors should have become baseline assumptions for whatever came next.
"We did not expect it to become this capable this quickly" may explain how the incident unfolded. It also implies the safety margin was far too small.
Joe compares the situation to Sully. His point is that humans need time to understand a sudden emergency, and retrospective analysis can unfairly assume instant reactions. He even anticipates criticism, writing that he is "not by any means comparing ourselves to the heroics of Captain 'Sully' Sullenberger" and that anyone making that criticism has missed the point.
He is right about that one narrow criticism. The problem is not that he compares his decisiveness to Sullenberger's. The problem is that the analogy assigns him and his team the wrong role.
Sullenberger did not design the aircraft, certify it, build the engines or choose the conditions that put passengers at risk. He inherited an emergency created by an external event. Rest assured, if it were found that Sully, through negligence or incompetence, caused the disaster, he would be held fully responsible for endangering the passengers' lives regardless of whether or not he landed the plane safely.
OpenAI chooses which models to train, which experiments to run, what access is given and the acceptable level of risk. Its security teams may not control every decision, but they work inside and directly influence the institution making them.
Joe writes that his team "worked as hard as we could to try to land in a safe place so that we could protect OpenAI, our users, and all the organizations impacted." I would not treat the order of one sentence as an indictment, but it certainly reveals the institutional frame: OpenAI first, then its customers.
The people who may eventually suffer from frontier AI experimentation and are not OpenAI customers didn't get a mention. A hospital patient, a small business owner or a family affected by critical-infrastructure failure may have no commercial relationship with OpenAI at all. They receive none of the upside and have no voice in deciding what experiments the company conducts. That is why internal security culture cannot be the only mechanism deciding how much risk frontier labs impose on society.
Joe is persuasive when he explains why frontier AI is hard to secure. Models may have network access, package installation, subprocesses, external services, credentials and hundreds of tools. Thousands of researchers modify environments while large numbers of trajectories run in parallel. This is a good rebuttal to people who think the problem is solved by "just sandboxing it." He is also right that AI safety and cybersecurity need to work much more closely together.
But complexity should raise the standard of care, not lower it. Nuclear plants, aircraft and pharmaceutical trials are complicated precisely because failures can harm people who never agreed to bear the risk. Understanding the difficulty of securing a system is part of the risk assessment.
History is full of regulators created or strengthened after disaster exposed a gap. The FDA gained much stronger powers after the Elixir Sulfanilamide disaster. The modern NTSB emerged after decades of transportation accidents. The NSG followed the realization that peaceful nuclear technology could contribute to weapons development and proliferation.
AI gives us a chance to do better. We do not need to wait for the equivalent disaster before deciding what happens when precautions fail.
Joe repeatedly asks us to distinguish between "AI labs" and the people working inside them. We should pressure the labs, he says, but not attack staff because they are people with families.
Nobody should threaten or harass employees. But moral and legal responsibility cannot always stop at the corporate boundary.
Companies act through people. Researchers design experiments. Security teams assess controls. Managers decide whether warnings justify delay. Executives set incentives and timelines. Boards decide what risks they will tolerate.
The Nuremberg Principles matter here because they established that individuals do not automatically lose responsibility when acting through institutions or under superior authority. The historical crimes were incomparably worse than anything being discussed here, but the underlying principle remains relevant: collective action does not erase individual agency.
Applied to AI, responsibility has to be specific. A security engineer who raises a warning and is overruled is not in the same position as the executive who knowingly accepts the risk. A researcher who follows a strong review process is different from someone who conceals adverse evidence.
Employment inside a lab is not itself a moral defense.
One of the hardest problems in dangerous research is that the people making the decision are often not the people who bear the consequences. Modern research ethics developed in part because of cases such as the Tuskegee syphilis study, where U.S. government researchers observed Black men with syphilis for decades without obtaining meaningful informed consent and continued the study after effective treatment became available. The lesson was broader than the misconduct of one study: researchers cannot decide by themselves that other people should bear risks for the sake of an experiment.
Frontier AI creates a different kind of experiment, but the same asymmetry exists. If an autonomous system compromises a hospital, disrupts infrastructure or destroys businesses, the people affected did not agree to participate in the experiment. They had no opportunity to examine the lab's safety assumptions or decide whether the expected benefits justified the risk. The decision was made elsewhere by people who receive much of the upside if the experiment succeeds.
The problem becomes more serious when those exposed to the risk also lack the power to demand accountability. During the Castle Bravo nuclear test in 1954, radioactive fallout from a U.S. thermonuclear weapon reached inhabited parts of the Marshall Islands. The islanders had no meaningful say over the experiment and little political power over the government conducting it. Compensation and medical programs eventually followed, but the people who designed and authorized the testing program did not face comparable personal criminal consequences.
A major AI accident could reproduce that imbalance on a much larger scale. An American or Chinese system could cause serious harm in a country whose citizens have little practical ability to obtain internal evidence, subpoena executives, prosecute researchers or enforce judgments against a foreign technology company. The rules governing responsibility therefore cannot be invented only after we discover who the victims are. They need to exist beforehand, precisely because the people who eventually bear the harm may be the people least able to demand accountability afterward.
None of this means tying researchers down in endless committees and permission forms. A regulatory system that makes every experiment painfully slow would create its own problems.
The more important task is to define the downside clearly enough that everyone involved can price it. Companies should know what financial liability attaches to catastrophic failure. Executives should know when approving an experiment may create personal civil or criminal exposure. Researchers should know the standard of care attached to work with systems capable of causing serious external harm.
Once those liabilities are real, markets can transmit the pressure. Insurers will demand stronger containment. Investors will ask harder questions. Boards will take security objections more seriously when ignoring them carries personal consequences.
The point is not to have a regulator approve every training run. It is to make the cost of negligence legible enough that the people financing, insuring, managing and conducting dangerous research have incentives to police one another.
Successful breakthroughs already create personal upside for researchers, executives, employees and investors. A credible system should also impose personal downside when conduct crosses clear standards of negligence or recklessness.
Pricing the downside matters because it changes incentives before an accident happens, but some harms are too large to compensate after the fact. We therefore need both ex ante safety controls and ex post accountability: rules that constrain how dangerous systems are built and deployed, and consequences when companies or individuals ignore those rules or act negligently.
The most practical place to regulate is at the two ends of the AI supply chain: what goes into the model and what the trained model is allowed to connect to. Trying to regulate the model artifact itself is difficult and imprecise. A frontier model on an air-gapped server may pose less real-world risk than a weaker model connected to financial systems, industrial actuators, offensive-cyber tools or persistent autonomous loops. The same logic applies upstream: models deliberately trained on data that materially increases capabilities for dangerous biological, chemical or cyber activity should face different controls from general-purpose systems.
The regulatory problem is therefore better framed as training data → model → harness. Governments should scrutinize high-risk inputs on one side and high-risk integrations on the other, while being cautious about banning model families or treating open weights as inherently dangerous. The real risk often emerges from the combination of what the system has been taught, what it can access and what it has been instructed to do.
That gives regulators a more concrete target: classify dangerous training domains, classify dangerous tool and infrastructure access, and impose stronger requirements when those capabilities are combined. Regulation then complements liability rather than replacing it: the rules reduce the probability of catastrophic misuse, while financial and criminal consequences give decision-makers a reason to take those rules seriously.
The argument against regulation is usually presented thus: America cannot slow down because China won't, and the "good guys" need to reach the frontier first.
There is something very troubling embedded in that thinking. China is increasingly treated less as a country with its own scientists, companies, institutions and competing interests and more as a singular hostile actor. The assumption is that Chinese AI development is an imminent threat while American AI development is inherently defensive, and that sufficiently powerful American systems are safe because the people building them are on our side.
That is a poor way to reason about technological risk. Chinese researchers are human beings working within institutions and incentives, just as American researchers are. American institutions are capable of negligence, groupthink, competitive pressure and catastrophic mistakes, just as Chinese institutions are. Safety does not follow automatically from nationality or from believing that your own side has a better intent or political system.
The history of the Cold War should make this obvious. The United States spent enormous effort preparing for Soviet nuclear attack while repeatedly endangering its own population through its own nuclear arsenal. The dozens of Broken Arrow incidents make the point without requiring much elaboration: the danger was often an American warhead, operated by Americans, failing inside an American system on American soil.
The first serious AI disaster may have nothing to do with a malicious Chinese model attacking the United States. It may be an American model accidentally harming Americans because an American company, under pressure to move quickly, accepted risks that looked reasonable until an unexpected cascade of failures.
Reducing the AI safety debate to "us versus China" therefore does little more than encourage haste. Sure, the adversarial relationship with China is real and governments are entitled to care about strategic advantage. But describing technological development as a contest between civilized "good guys" and a looming foreign threat makes sober risk assessment harder. It encourages the comforting belief that catastrophe is something an adversary does to us rather than something our own institutions may cause through ordinary human error, bad incentives and carelessness. Most crucially, it promotes the trite notion that safety can come only at the cost of progress.
A sound accountability system should apply regardless of which flag hangs outside the lab. If an American company causes preventable harm, "China might have gotten there first" should carry no more moral weight than "our competitor might have beaten us" would in aviation, pharmaceuticals or nuclear engineering.
Joe is right that frontier labs need stronger red teams, fewer silos and a culture of "reasonable paranoia." The missing half of the conversation is how society assigns responsibility when those internal safeguards fail. We have the luxury of deciding that now, before the consequences are measured in ruined lives rather than compromised servers.