Research & Benchmarks
Papers, findings, evals, and leaderboard results — the science behind the models.
Papers, findings, evals, and leaderboard results — the science behind the models.
OpenAI published results from an internal version of Astra, described as the company's next major model, solving ten open problems in mathematics and theoretical computer science that had seen no…
OpenAI said enabling two settings already available in its API — retained reasoning and compaction — nearly tripled its GPT-5.6 Sol model's score on the ARC-AGI-3 benchmark, taking it from 13.3% to…
Vercel published a blog post on July 27, 2026 introducing DeepsecBench, a new benchmark that measures how well AI models find cybersecurity vulnerabilities in real application code. In its own words,…
OpenAI published an exploratory field report on July 28, 2026 examining how AI coding agents perform on real scientific-computing work, drawing on eight case studies contributed by research teams who…
Anthropic published research on July 28, 2026 showing that its gated frontier model, Claude Mythos Preview, discovered improved cryptanalytic attacks against two cryptographic algorithms — the…
Together AI ran Moonshot's newly released Kimi K3 head-to-head against OpenAI's GPT-5.6 Sol on 904 real software-engineering rollouts, finding the two models trade wins depending on whether you're…
OpenAI published new research on July 27, 2026 showing that a large share of the work-related tasks people bring to ChatGPT don't match their job title — evidence, the company argues, that AI is…
Anthropic opened applications for a new sub-grant competition under its AI for Science program, offering researchers and early-stage biotech companies up to $50,000 in Claude API credits to work on…
Anthropic is committing $200 million to a new Economic Futures Research Fund that will pay outside researchers to study how society can prepare for AI-driven labor market disruption, spanning worker…
Google published the first edition of its AI & Economy ATLAS on July 23, 2026, a study built from 15 million de-identified interactions across the Gemini app, AI Mode, and the Gemini API that aims to…
A systematic review of 35 empirical studies conducted between 2022 and 2025 finds that children routinely attribute human-like qualities to LLM-powered chatbots, and that this anthropomorphism shapes…
Five U.S. national laboratories are running Meta's open-source Segment Anything Model 3 and DINOv3 in production to automate scientific image analysis under the Department of Energy's Genesis…
AWS researchers published a technical study on July 21, 2026 describing Self-Distilled Reasoning (SDR), a method for adding chain-of-thought training signal to fine-tuning datasets that were never…
OpenAI published a strategic framework this week for judging AI progress by economic value rather than raw capability, arguing that the right yardstick for enterprises is what it calls "Useful…
Anthropic published research on July 13, 2026 showing that the values Claude expresses in conversation are not fixed — they shift measurably depending on which model is answering and which language…
Kwaipilot, the AI coding team at Chinese short-video giant Kuaishou, published a technical report on July 6 for KAT-Coder-V2.5, an agentic coding model the team says ranks second only to Anthropic's…
Google DeepMind has digitally reconstructed a goal scored by Pelé on August 2, 1959 — the "Gol da Rua Javari" — that was never captured on film, using a combination of practical filmmaking and…
Google published a report on July 14 detailing Gemini adoption across Southeast Asia, saying the app is being adopted faster there than any Google product the company has previously launched in the…
Anthropic's latest Economic Index research pinpoints exactly where Australia sits in global Claude usage — and it's more nuanced than the "Australia is 1" framing some outlets ran with: high…
OpenAI has published a short paper claiming a full proof of the Cycle Double Cover Conjecture, a graph-theory problem open since the 1970s, crediting the work "entirely" to its GPT-5.6 Sol Ultra…
OpenAI has published the results of a quality audit into SWE-Bench Pro, the coding benchmark it previously recommended as a replacement for the widely used SWE-bench Verified, and concluded the eval…
Databricks has published results from an internal benchmark that tests coding agents against real engineering work pulled from its own multi-million-line codebase, rather than synthetic or curated…
NVIDIA and LangChain say tuning the agent harness around a model — rather than retraining the model itself — pushed Nemotron 3 Ultra to benchmark-leading results among open models on LangChain's…
Cursor has launched a CFO Council, a quarterly working group of enterprise finance leaders, alongside a set of internal usage data showing a widening gap between organizations that are capturing…
Anthropic researchers say they have identified a small set of internal neural patterns in Claude models that function like a global workspace for conscious access in the human brain — a distinct,…
A new academic study finds that U.S. export controls on advanced semiconductors had an unintended side effect: they pushed China to invest more heavily in open-source AI, strengthening an ecosystem…
OpenAI's economic research group has published an extension of its AI Jobs Transition Framework applied specifically to the European Union labor market, estimating how many EU jobs sit in occupations…
Cursor has published results from CursorBench 3.1, the latest version of its internal coding-agent benchmark, with Anthropic's Fable 5 Max topping the leaderboard and a cheaper Cursor-built model…
Anthropic has published the sixth installment of its Economic Index, subtitled "Cadences," shifting the research series from broad occupational trends to the hour-by-hour rhythm of how people…
Together AI published a preview of its ICML 2026 research contributions ahead of the International Conference on Machine Learning, running July 6-11 in Seoul. The company is presenting nine papers…
OpenAI on June 30 released GeneBench-Pro, a benchmark for evaluating AI agents on realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine. The…
Meta's Fundamental AI Research (FAIR) team published Brain2Qwerty v2 on June 29, a non-invasive brain-computer interface system that converts brain activity directly into typed text, reaching 61%…
OpenAI researchers published a data-driven paper on June 25, 2026 analyzing usage patterns across Codex, their AI coding and agentic task product. The findings show a dramatic expansion beyond the…
Researchers at Together AI have released ParallelKernelBench (PKB), a benchmark designed to test whether frontier large language models can generate optimized multi-GPU CUDA kernels. The results from…
Google on June 10, 2026 released DiffusionGemma, an experimental open-weights language model that generates text through a fundamentally different mechanism than standard autoregressive models. While…
Researchers at Boston Children's Hospital have used OpenAI's o3 reasoning model to identify new diagnoses for 18 children whose rare genetic diseases had gone unresolved by standard clinical workup.…
OpenAI's alignment research team published findings today showing that reinforcement learning applied to a targeted set of "beneficial trait" training scenarios produces safety improvements that…
OpenAI and Molecule.one published research on June 17, 2026 showing a near-autonomous AI system successfully improved a challenging reaction in medicinal chemistry — the first time an AI has tackled…
ChatGPT has lost its majority position in the AI assistant market for the first time since its 2022 launch, dropping to 46.4% market share as of May 2026, according to Sensor Tower's State of AI…
Google's AMIE (Articulate Medical Intelligence Explorer) has matched the performance of primary care physicians in managing chronic health conditions and significantly outperformed them in plan…
Anthropic published research on June 16 showing that the single strongest predictor of success with Claude Code is domain expertise in the task at hand — not whether a user has a background in…
Alibaba's Qwen team released a technical report on June 15, 2026 introducing Qwen-RobotWorld, a foundation model for embodied AI that predicts future visual trajectories across robotics, autonomous…
NVIDIA's Blackwell GPU platform claimed the top spot across all seven benchmarks in MLPerf Training 6.0, the latest edition of the peer-reviewed AI training benchmark suite. The results, published…
Anthropic published research on June 8, 2026, examining why AI agents struggle to reliably retrieve biological data — and how a single deterministic tool layer nearly eliminates the performance gap…
NVIDIA's Blackwell architecture has taken the top spot in AgentPerf, the first benchmark designed specifically to measure agentic AI infrastructure performance, according to results published June 12…
Anthropic on June 12 released the inaugural results from its Anthropic Public Record, a new survey series the company says it will repeat to track how U.S. public attitudes toward AI shift over time.…
Alibaba's Qwen research team has published Qwen-Image-Flash: Beyond Objective Design on arXiv, presenting a few-step distillation framework that accelerates the Qwen-Image-2.0 model family for both…
DeepSeek researchers published a paper on June 9, 2026 introducing Lookahead Sparse Attention (LSA), an inference-time technique that compresses the GPU memory footprint of long-context LLM serving…
Vercel published its AI gateway production index for May 2026, drawing on routing traffic across its infrastructure to document a widening divergence in the AI provider market: DeepSeek's share of…
Anthropic disclosed on June 3, 2026 that it detected and disrupted what it describes as the first large-scale cyberattack executed largely without human intervention — a campaign in which a Chinese…
Google DeepMind published details on May 20, 2026 of an accessibility agent that guides blind and low-vision athletes while running outdoors in real time, using a Pixel 10 Pro smartphone worn on the…
Anthropic on June 5 published benchmark data showing that Claude Opus 4.7 — a general-purpose frontier model with no chemistry-specific fine-tuning — performs comparably to specialized NMR software…
Google DeepMind's AI research assistant Co-Scientist helped biologists at MIT identify more than 20 genetic factors that may reverse cellular aging — and lab experiments confirmed that several of…
A team of researchers including prominent Google DeepMind scientists published a paper at the 43rd International Conference on Machine Learning (ICML 2026) arguing that the dominant paradigm in AI…
Anthropic has released a public best-practices guide for using Claude to find and fix vulnerabilities in source code, paired with an open-source reference harness on GitHub. The guide is grounded in…
Perplexity on June 1 published research introducing Search as Code (SaC), an architecture in which AI agents compose search pipelines programmatically from atomic building blocks rather than calling…
Anthropic on June 3 published a year-long study of how malicious actors have used Claude, examining 832 accounts the company banned for cyber abuse between March 2025 and March 2026. The headline…
On June 2, 2026, Anthropic disclosed that it is significantly expanding Project Glasswing, its program to harden critical-infrastructure software with frontier-model-assisted security analysis.…
Hugging Face users running the Transformers library below version 5.3.0 are exposed to a critical remote-code-execution vulnerability that turns a routine model load into attacker-controlled Python…
On June 2, 2026 a team at the University of Toronto and the Vector Institute released a paper on arXiv showing that an AI agent can power a computer worm capable of writing fresh exploits as it…
Anthropic's engineering team published a retrospective on 2026-05-25 describing how it contains Claude agents across three products — claude.ai, Claude Code, and Claude Cowork — with a focus on…
At CVPR 2026, NVIDIA Research presented three new physical-AI foundation models built around a common thesis: train at sufficient scale, and systems generalize across embodiments, scenarios and…
In the GPT-5.5 system card (April 23, 2026; updated April 24 for API safeguards), OpenAI classifies the model's biological/chemical and cybersecurity capabilities as "High" under its Preparedness…