Two API Settings That Tripled GPT-5.6’s ARC-AGI-3 Score
OpenAI found that enabling just two API settings tripled GPT-5.6’s ARC-AGI-3 benchmark scores. Here’s what changed, why it matters, and what it means for developers.
OpenAI found that enabling just two API settings tripled GPT-5.6’s ARC-AGI-3 benchmark scores. Here’s what changed, why it matters, and what it means for developers.
Anthropic’s Claude Opus 5 launches July 24, 2026 — state-of-the-art on coding and knowledge work benchmarks at $5/M input tokens. Here’s what it means for you.
OpenAI’s GeneBench-Pro case studies reveal how genomics labs are actually using AI benchmarks — and what the results mean for the field’s future.
OpenAI’s GeneBench-Pro tests AI on real-world genomics and biology tasks. Here’s what it measures, why it matters, and what it means for scientific AI.
OpenAI launches LifeSciBench, an expert-authored benchmark testing AI on real-world life science research tasks. Here’s what it measures and why it matters.
OpenAI’s GPT-5.5 system card is out. Here’s what the safety evals, capability benchmarks, and deployment decisions actually mean for you.
OpenAI stops using SWE-bench Verified for AI coding tests, citing flawed benchmarks and training leakage. The company now recommends SWE-bench Pro instead.