OpenAI’s Cybersecurity Evaluation Incident: What Went Wrong
OpenAI disclosed a third-party cybersecurity evaluation incident involving its models. Here’s what happened, what it means, and what safeguards are coming.
OpenAI disclosed a third-party cybersecurity evaluation incident involving its models. Here’s what happened, what it means, and what safeguards are coming.
OpenAI’s new analysis reveals serious reliability issues in SWE-Bench Pro, one of AI’s most-cited coding benchmarks. Here’s what broke and why it matters.
OpenAI is backing the Appia Foundation to build shared AI safety standards, evaluation frameworks, and global cooperation norms for advanced AI systems.
OpenAI launches LifeSciBench, an expert-authored benchmark testing AI on real-world life science research tasks. Here’s what it measures and why it matters.