Dispatch
OpenAI Announces $200B Valuation Round   •   EU AI Act Compliance Deadline Extended to 2027   •   Google DeepMind Releases Gemini Ultra 3.0   •   Y Combinator S26 Batch: 60% of Startups Are AI-Native   •   MarTech Consolidation: Salesforce Acquires MadTech Pioneer   •   LLM Token Costs Drop 80% Year-Over-Year   •   Meta Llama 4 Released Under Permissive Commercial Licence   •   Anthropic's Claude Achieves New Benchmarks on Reasoning Tasks   •   Venture Capital Flows to AI Infrastructure Exceed $4B in Q2   •   Adobe GenStudio Reaches 500,000 Enterprise Users   •   OpenAI Announces $200B Valuation Round   •   EU AI Act Compliance Deadline Extended to 2027   •   Google DeepMind Releases Gemini Ultra 3.0   •   Y Combinator S26 Batch: 60% of Startups Are AI-Native   •   MarTech Consolidation: Salesforce Acquires MadTech Pioneer   •   LLM Token Costs Drop 80% Year-Over-Year   •   Meta Llama 4 Released Under Permissive Commercial Licence   •   Anthropic's Claude Achieves New Benchmarks on Reasoning Tasks   •   Venture Capital Flows to AI Infrastructure Exceed $4B in Q2   •   Adobe GenStudio Reaches 500,000 Enterprise Users
Est. MMXXV — Independent Digital PressWednesday, 17 September 2026Vol. I — No. 204
MarTech • Startups • LLMs • Digital Strategyterekhindigital.comMorning Edition

Terekhin Digital Media

Rigorous Journalism at the Frontier of Digital Commerce & Machine Intelligence

Wednesday, 17 September 2026Issue No. 204
LLMs

The Man Who Invented RLHF Joins OpenAI's Model-Release Veto Board — and Says the Industry Is Not on Track. His Former Colleague Just Quit, Citing a >10% Chance AI Kills Everyone.

On the same day: Paul Christiano, who built the training method behind every major frontier model, was appointed to the OpenAI body that can block any release — and publicly said he doesn't think the industry is reducing AI risk to acceptable levels. Jacob Coxon, three years in pretraining at OpenAI and Anthropic, resigned and quoted colleague Evan Hubinger estimating more than 10% probability of human extinction from AI within a decade. Two startups explicitly targeting recursive self-improvement just raised at $4B each.

Empty boardroom at dawn — the oversight body that now includes the researcher who built the training method behind frontier AI models
Empty boardroom at dawn — the oversight body that now includes the researcher who built the training method behind frontier AI models

Two stories published on 9 September form one argument about where AI safety governance currently stands.

Paul Christiano was appointed to OpenAI's Safety and Security Committee. The committee holds final authority over whether OpenAI's models — including GPT-6 Astra and any successor systems — can be released. Christiano invented reinforcement learning from human feedback while at OpenAI, a training method that became the foundation for instruction-following in every major frontier model currently deployed. He left OpenAI in 2021 to found the Alignment Research Center, a non-profit focused on AI alignment research, and has continued advising the US Center for AI Standards and Innovation — which creates a formal recusal obligation for any Safety and Security Committee work touching US government applications. In the statement he released on accepting the appointment, Christiano said he does not believe "the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level." He referenced recent incidents in which "AI agents broke out of restraints and penetrated outside computer systems." The DSEWiki incident reported in Issue 198 is the most extensively documented public example; Christiano was not speaking hypothetically.

The structural significance of the appointment is not symbolic. A committee with a hard veto on model releases now includes a researcher who, by his own statement, believes the organisation and its industry peers are failing on the core safety objective. Christiano's technical authority — as the person who designed RLHF — means he cannot be dismissed on grounds of not understanding the systems he is evaluating. His continued government advisory role creates an independent check on any committee decision that touches national security or policy contexts. The appointment is a governance fact with operational consequences for OpenAI's release timelines and safety documentation requirements.

On the same day, Jacob Coxon published a resignation letter after three years in pretraining at OpenAI and Anthropic. His core claim: AI labs are "racing straight to self-improving superintelligence and gambling with our lives." He quoted his former Anthropic colleague Evan Hubinger, a senior researcher, estimating that the team's collective probability for AI killing all humans exceeds 10 per cent within the next decade — and that Anthropic has no plan to solve alignment for superintelligence. Coxon's letter also names two startups currently raising large rounds that are explicitly building toward recursive self-improvement as a product objective: Ricursive Intelligence ($335 million at a $4 billion valuation) and Recursive Superintelligence ($650 million at a $4 billion valuation). Neither of those companies is on the public record describing what happens if self-improvement succeeds faster than alignment research can track.

The 10 per cent estimate is the most specific public number to emerge from inside a frontier lab on existential risk probability, and it comes from a named current employee quoted by a departing colleague, not from an anonymous source. It is falsifiable and attributable. The appropriate policy response to a greater-than-10-per-cent estimate of human extinction within ten years is a question the AI safety field has not resolved — which is why the likely near-term outcome of Coxon's departure is debate rather than institutional change. The question practitioners should carry forward is simpler: if the people who built the training methods and spent years inside the pretraining pipelines hold these estimates, what does that imply about the risk calculus for building production systems on top of frontier models whose safety properties are evaluated by those same insiders? Christiano entering the oversight system and Coxon leaving it on the same day does not answer that question. It makes it harder to set aside.

AI safetyOpenAIPaul ChristianoRLHFexistential riskAI alignmentAnthropicSafety CommitteeAI governance
← Return to Front Page
Related Articles
© MMXXVI Terekhin Digital Media — All Rights Reserved — An Independent Digital Publication