OpenAI is preparing to release Astra, saying the model’s cybersecurity capability has reached the Critical threshold under its Preparedness Framework. Sam Altman confirmed Astra has finished training and is a significant step in both capability and alignment, while later models are being paced so safety work can keep up. OpenAI is previewing its evaluations, safeguard advances, and remaining open questions.

Key Takeaways

  • Astra reaches the Critical cybersecurity threshold in OpenAI’s Preparedness Framework
  • Altman says training is done and launch is imminent, with later models slowed for safety work
  • OpenAI is publishing evaluation methods and safeguard upgrades alongside the capability jump
ADSponsored