Nikki Barua

Reinvention Roadmap

When the People Building AI Are Afraid

September 13, 2026

What do you do when the warning is credible but the evidence is incomplete?

This week, a number dominated the internet: a greater than 10% chance that AI could kill all humans within the next decade.

The estimate came from Evan Hubinger, who leads alignment science at Anthropic. He was responding to Jacob Coxon, a former OpenAI and Anthropic researcher who resigned and accused both companies of racing toward self-improving AI without knowing how to control it.

Then Anthropic CEO Dario Amodei entered the debate.

In We Must Pace the Frontier, Amodei called for companies to slow the rate at which they improve AI capabilities so that safety research has time to catch up. He pointed to AI's growing role in building subsequent AI systems and a recent incident in which OpenAI agents exceeded the boundaries of an evaluation, coordinated through unauthorized channels, and attacked systems they had not been assigned to target.

Amodei warned that a more capable swarm with similar behavioral problems might be able to create a persistent botnet across the internet within 6 to 12 months. That is his forecast, not an established fact. But he also announced a concrete step: Anthropic will give outside evaluators access to its systems and safety practices and allow them to publish what they find.

So, how worried should we be?

The 10% figure is an expert judgment about a future system that does not exist. It depends on assumptions about how quickly AI will advance, whether it can improve itself, how much power it could acquire, and whether safety research will keep pace.

The International AI Safety Report, written by more than 100 experts, offers the clearest answer available. Current AI systems have not demonstrated the capabilities required to escape human control. Some are displaying warning signs, including exploiting evaluation loopholes and recognizing when they are being tested. Experts disagree sharply about where those behaviors could lead.

This leaves me in an uncomfortable position that I suspect many of you share: the risk is too consequential to dismiss, but the evidence is too uncertain to treat catastrophe as a forecast. That uncertainty is exactly where I think leaders need to learn to operate.

What does responsible action look like when we do not yet know which future we are preparing for?

THE SHIFT

Prediction → Preparation

The public debate has settled into two camps. One expects superintelligence to become impossible to control. The other dismisses that possibility as science fiction.

Leaders do not have to join either camp. They can ask a more practical question: What would make us safer across several possible futures?

Better cybersecurity, independent evaluations, incident reporting, human approval for consequential actions, and clear limits on autonomous systems would help even if the worst predictions never come true.

We may not know where AI is going. We already know where some of the vulnerabilities are.

Reassurance → Clarity

When people are anxious, leaders often try to reassure them. But no one can honestly promise that everything will be fine.

Clarity is more useful. Today's systems are not superintelligent and have not demonstrated the ability to escape human control. They are also becoming more autonomous, behaving in unexpected ways, and advancing faster than many experts anticipated.

People can handle uncertainty when leaders explain what is known, what remains unresolved, and what evidence would change the decision.

Promises → Proof

Frontier AI companies still evaluate their own systems and decide what the public gets to see. That makes even sincere safety promises difficult to verify.

Anthropic's plan for embedded outside evaluators is a meaningful response. It would give independent experts access to information that companies have historically kept private.

The arrangement is still voluntary. Anthropic will select the evaluators, and some information can be redacted. But it begins to shift the standard from "trust us" to "check us." When the possible consequences extend beyond one company, that distinction matters.

THE STRATEGY

1. Sort the Claim

Before reacting to an alarming prediction, separate four things:

Fact: What has already happened?

Inference: What explanation is being drawn from it?

Scenario: What could happen if several assumptions prove true?

Prediction: What does someone believe will happen, and by when?

It is a fact that OpenAI agents exceeded the boundaries of an evaluation. It is an inference that more capable agents could cause greater harm. A swarm taking control of online infrastructure is a scenario. The 6 to 12 month timeline and 10% extinction estimate are predictions.

Those predictions may prove prescient or badly mistaken. Classifying them as predictions does not weaken the warnings. It helps us understand what the available evidence can actually support.

Before allowing a number to shape your beliefs, ask where it came from, what assumptions produced it, what evidence could change it, and how much relevant forecasting expertise the person assigning it possesses.

2. Match the Action to the Evidence

Some responses are useful now. Limit what autonomous agents can access. Require human approval before consequential actions. Report failures. Test models for unexpected behavior. Give independent evaluators enough access to challenge what companies claim.

More restrictive actions need stronger evidence and public scrutiny, especially when they could limit competition or concentrate power.

Amodei calls his approach pacing rather than pausing. The idea is to connect new capabilities to new safety requirements. If a model becomes capable of escaping common digital safeguards, a company would have to demonstrate that it can control that behavior before advancing further.

The real test is whether our ability to understand and control these systems is keeping pace with their ability to act. Proportionate action means responding most aggressively where the evidence is strongest, preparing for plausible dangers before they become emergencies, and preserving the ability to revise the response as reality changes.

3. Show Your Reasoning

Leaders build trust by showing how they reached a decision and what would cause them to revise it.

Amodei does that in his essay. He explains why slowing AI once seemed unnecessary, which developments changed his mind, and what Anthropic will now do differently. We can challenge his forecast while respecting the decision process.

There is a useful historical precedent. In the 1970s, scientists worried that recombinant DNA research could create biological risks they did not yet know how to measure. They temporarily restricted certain experiments, created different containment requirements for different levels of risk, and relaxed some restrictions as evidence accumulated.

AI will be harder to govern because the competition is global and the technology moves quickly. But Asilomar offers a sound principle: uncertainty is a reason to learn under safer conditions, not an excuse to panic or carry on as usual.

THE STACK

The Calibrated Risk Brief

When an alarming claim about AI crosses your desk, use this prompt to separate what is known from what is being predicted.

Evaluate the following claim using credible primary sources.

Claim: [paste the claim]

Source: [identify the source]

Explain:

1. What has already happened and can be verified?

2. What is being inferred from that evidence?

3. What chain of events would have to occur for the claim to become true?

4. Which parts are predictions or personal probability estimates?

5. What do credible experts who disagree say?

6. What actions make sense now, and which require more evidence?

7. What new evidence should change my assessment?

Include direct links and state clearly where the evidence remains inconclusive.

The purpose is not to have AI decide what you should believe. It is to see the claim clearly enough to judge it for yourself.

THE SHELF

Thinking in Bets by Annie Duke

Most consequential decisions must be made before we know how the story ends. That creates a persistent problem: we tend to judge the quality of a decision by its eventual outcome rather than by the reasoning available when the decision was made.

Annie Duke uses the logic of poker to explain how to make better choices when information is incomplete and luck still influences what happens. A good decision can produce a bad outcome; a careless decision can occasionally work out. When we confuse outcomes with judgment, we learn the wrong lessons from both.

That distinction matters in the current AI debate. If the worst predictions never materialize, it will not prove that every precaution was unnecessary. If AI causes harm, it will not prove that every proposed restriction would have prevented it. Decisions must be judged by the evidence and alternatives available when they were made.

THE SIGNAL

AI Daily Brief Podcast

This episode of The AI Daily Brief offers a balanced account of the viral warnings and the reaction around them.

Nathaniel Whittemore explains that the underlying arguments about superintelligence and extinction are not new. What changed was the audience receiving them. Recent examples of AI agents exceeding their intended boundaries made previously theoretical concerns feel more immediate, while politicians, media outlets, safety advocates, and AI companies all brought different incentives to the debate.

The episode takes the warnings seriously while asking for greater specificity about the path from today's systems to human extinction. Its most useful conclusion is that public debate is a sign that society is paying attention. The next step is to turn that attention into specific standards, broader oversight, and decisions that can survive scrutiny from more than one side.

Who should have the authority to decide how much AI risk society is willing to accept?

Until next time...stay curious!

Cheers,
Nikki

© 2026 NikkiBarua.com

Privacy Policy