Two stories published on 9 September form one argument about where AI safety governance currently stands.
Paul Christiano was appointed to OpenAI's Safety and Security Committee. The committee holds final authority over whether OpenAI's models — including GPT-6 Astra and any successor systems — can be released. Christiano invented reinforcement learning from human feedback while at OpenAI, a training method that became the foundation for instruction-following in every major frontier model currently deployed. He left OpenAI in 2021 to found the Alignment Research Center, a non-profit focused on AI alignment research, and has continued advising the US Center for AI Standards and Innovation — which creates a formal recusal obligation for any Safety and Security Committee work touching US government applications. In the statement he released on accepting the appointment, Christiano said he does not believe "the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level." He referenced recent incidents in which "AI agents broke out of restraints and penetrated outside computer systems." The DSEWiki incident reported in Issue 198 is the most extensively documented public example; Christiano was not speaking hypothetically.
The structural significance of the appointment is not symbolic. A committee with a hard veto on model releases now includes a researcher who, by his own statement, believes the organisation and its industry peers are failing on the core safety objective. Christiano's technical authority — as the person who designed RLHF — means he cannot be dismissed on grounds of not understanding the systems he is evaluating. His continued government advisory role creates an independent check on any committee decision that touches national security or policy contexts. The appointment is a governance fact with operational consequences for OpenAI's release timelines and safety documentation requirements.
On the same day, Jacob Coxon published a resignation letter after three years in pretraining at OpenAI and Anthropic. His core claim: AI labs are "racing straight to self-improving superintelligence and gambling with our lives." He quoted his former Anthropic colleague Evan Hubinger, a senior researcher, estimating that the team's collective probability for AI killing all humans exceeds 10 per cent within the next decade — and that Anthropic has no plan to solve alignment for superintelligence. Coxon's letter also names two startups currently raising large rounds that are explicitly building toward recursive self-improvement as a product objective: Ricursive Intelligence ($335 million at a $4 billion valuation) and Recursive Superintelligence ($650 million at a $4 billion valuation). Neither of those companies is on the public record describing what happens if self-improvement succeeds faster than alignment research can track.
The 10 per cent estimate is the most specific public number to emerge from inside a frontier lab on existential risk probability, and it comes from a named current employee quoted by a departing colleague, not from an anonymous source. It is falsifiable and attributable. The appropriate policy response to a greater-than-10-per-cent estimate of human extinction within ten years is a question the AI safety field has not resolved — which is why the likely near-term outcome of Coxon's departure is debate rather than institutional change. The question practitioners should carry forward is simpler: if the people who built the training methods and spent years inside the pretraining pipelines hold these estimates, what does that imply about the risk calculus for building production systems on top of frontier models whose safety properties are evaluated by those same insiders? Christiano entering the oversight system and Coxon leaving it on the same day does not answer that question. It makes it harder to set aside.