Research & Benchmarks
Papers, findings, evals, and leaderboard results — the science behind the models.
Papers, findings, evals, and leaderboard results — the science behind the models.
Anthropic has published new research showing Claude can reproduce results from its earlier protein-binder-design work at roughly one-hundredth the computational cost, by having the model optimize the…
NVIDIA published its first MLPerf Inference results for Vera Rubin NVL72 on September 16, 2026, reporting up to 3.7x higher throughput than the current-generation GB300 NVL72 system on the v6.1…
Google published new findings from its AI & Economy ATLAS research initiative on September 15, 2026, alongside a new interactive, open-access data tool, showing sharp regional and occupational…
Anthropic published an engineering report explaining how it redesigned its continuous-integration (CI) test infrastructure after AI-assisted development caused CI job volume to explode across the…
Perplexity introduced Q2D-Web on September 9, 2026, a large-scale benchmark and public leaderboard for evaluating how well embedding models retrieve documents inside agentic retrieval-augmented…
OpenAI has published a case study on how a University of Pennsylvania bioengineering lab uses Codex and ChatGPT alongside its own deep-learning models to search genomes for candidate antimicrobial…
OpenAI says a graduate researcher at MIT's Engineering Quantum Systems Group (EQuS) connected GPT-5.6 Sol, routed through Codex, directly to the lab software controlling a superconducting quantum…
OpenAI says an internal AI system — described as significantly more capable than its flagship GPT-6 Astra — produced an analytical proof and a machine-verified Lean formalization showing that a…
Google DeepMind has released AlphaGenome Atlas, a public database that pre-computes the predicted effect of every possible single-letter genetic change in the human genome, giving researchers a free,…
OpenAI said on September 6 that it has reached a goal announced last fall by Sam Altman — fielding an "automated research intern" by September 2026 — and published internal metrics on how coding…
OpenAI has repeatedly altered the evaluation numbers published alongside its September 3 announcement of GPT-6 Astra, quietly revising several benchmark figures in the hours and days after the blog…
Anthropic said on September 4, 2026 that its Claude model produced the first complete, computer-checked formal proof of Fermat's Last Theorem, working largely autonomously over 11 days to encode the…
Alibaba's Qwen team has released Qwen-Drive-1.0, described as an initial step toward a single vision-language foundation model that can handle 3D perception, visual question answering, and motion…
A randomized experiment with more than 1,000 Bocconi University students, run in collaboration with OpenAI Economic Research, finds that ChatGPT access and a short critical-thinking exercise improve…
Together AI ran a head-to-head coding benchmark pitting Zhipu's GLM-5.3 against OpenAI's GPT-5.6 Sol, finding that while Sol edges out GLM-5.3 on raw accuracy, GLM-5.3 solves far more tasks per…
NVIDIA published results on August 21, 2026 showing its AVO agent architecture completed the entire ARC-AGI-3 public benchmark set, arguing the achievement demonstrates that agent system design — not…
Databricks has published results from the inaugural Grounded Reasoning Cup, a live competition that tested whether AI agents trained on one enterprise-document benchmark can generalize to an…
IBM Research published a study showing that giving an AI agent more memory of its own past runs does not help uniformly — the gain depends heavily on how strong the underlying model already is, with…
Cohere Labs published research showing that the cultural breadth present in large language model pretraining data is substantially stripped out by the time it reaches post-training, a pattern the…
Vercel's AI Gateway — which routes tokens between production applications and AI labs at massive scale — shows DeepSeek has overtaken Google to become the second-largest lab by token volume, while…
Anthropic has published research showing Claude can autonomously run substantial parts of a protein-design campaign, reporting that its models "designed protein binders against 15 targets, and…
Together AI published a benchmark comparing DeepSeek V4 Pro (the August 13 update) against OpenAI's GPT-5.6 Sol on DeepSWE, a software-engineering task suite, and found that routing between the two…
OpenAI published two studies this week documenting a widening gap between enterprises that have embraced agentic AI and those still using it as a lightweight assistant, alongside evidence that Codex…
A small research group called Pathway has published a paper claiming a new state of the art in cost-efficiency on ARC-AGI-1, one of the benchmarks used to measure abstract reasoning in AI systems,…
Anthropic has published new research examining how groups of AI agents behave when they interact with each other at scale, finding failure modes that look nothing like the ones researchers study in…
Google Research has extended its AMIE (Articulate Medical Intelligence Explorer) diagnostic AI system to conduct real-time audio-visual consultations, and a new study finds clinical evaluators rated…
An unreleased research version of Claude has extended a long-standing partial result tied to the Riemann hypothesis, one of mathematics' most famous unsolved problems, according to a paper Anthropic…
Together AI published a head-to-head benchmark of DeepSeek-V4 Flash 0731 against OpenAI's GPT-5.6 Luna on DeepSWE, a coding-agent benchmark built from 113 real open-source tasks run across 900 total…
Databricks has released OfficeQA Pro V2, a benchmark built to test whether AI agents can generalize to unfamiliar, multi-document enterprise research tasks rather than the narrower single-document…
Microsoft Research, working with Paige (now part of Tempus), has released PRISM2, a pathology foundation model that pairs tissue-image analysis with the language of pathology reports themselves — a…
Google DeepMind says its WeatherNext 2 model has achieved a major jump in forecasting accuracy for hurricanes and cyclones, publishing results August 6 that the company frames as a decade's worth of…
OpenAI published results from an internal version of Astra, described as the company's next major model, solving ten open problems in mathematics and theoretical computer science that had seen no…
OpenAI said enabling two settings already available in its API — retained reasoning and compaction — nearly tripled its GPT-5.6 Sol model's score on the ARC-AGI-3 benchmark, taking it from 13.3% to…
Vercel published a blog post on July 27, 2026 introducing DeepsecBench, a new benchmark that measures how well AI models find cybersecurity vulnerabilities in real application code. In its own words,…
OpenAI published an exploratory field report on July 28, 2026 examining how AI coding agents perform on real scientific-computing work, drawing on eight case studies contributed by research teams who…
Anthropic published research on July 28, 2026 showing that its gated frontier model, Claude Mythos Preview, discovered improved cryptanalytic attacks against two cryptographic algorithms — the…
Together AI ran Moonshot's newly released Kimi K3 head-to-head against OpenAI's GPT-5.6 Sol on 904 real software-engineering rollouts, finding the two models trade wins depending on whether you're…
OpenAI published new research on July 27, 2026 showing that a large share of the work-related tasks people bring to ChatGPT don't match their job title — evidence, the company argues, that AI is…
Anthropic opened applications for a new sub-grant competition under its AI for Science program, offering researchers and early-stage biotech companies up to $50,000 in Claude API credits to work on…
Anthropic is committing $200 million to a new Economic Futures Research Fund that will pay outside researchers to study how society can prepare for AI-driven labor market disruption, spanning worker…
Google published the first edition of its AI & Economy ATLAS on July 23, 2026, a study built from 15 million de-identified interactions across the Gemini app, AI Mode, and the Gemini API that aims to…
A systematic review of 35 empirical studies conducted between 2022 and 2025 finds that children routinely attribute human-like qualities to LLM-powered chatbots, and that this anthropomorphism shapes…
Five U.S. national laboratories are running Meta's open-source Segment Anything Model 3 and DINOv3 in production to automate scientific image analysis under the Department of Energy's Genesis…
AWS researchers published a technical study on July 21, 2026 describing Self-Distilled Reasoning (SDR), a method for adding chain-of-thought training signal to fine-tuning datasets that were never…
OpenAI published a strategic framework this week for judging AI progress by economic value rather than raw capability, arguing that the right yardstick for enterprises is what it calls "Useful…
Anthropic published research on July 13, 2026 showing that the values Claude expresses in conversation are not fixed — they shift measurably depending on which model is answering and which language…
Kwaipilot, the AI coding team at Chinese short-video giant Kuaishou, published a technical report on July 6 for KAT-Coder-V2.5, an agentic coding model the team says ranks second only to Anthropic's…
Google DeepMind has digitally reconstructed a goal scored by Pelé on August 2, 1959 — the "Gol da Rua Javari" — that was never captured on film, using a combination of practical filmmaking and…
Google published a report on July 14 detailing Gemini adoption across Southeast Asia, saying the app is being adopted faster there than any Google product the company has previously launched in the…
Anthropic's latest Economic Index research pinpoints exactly where Australia sits in global Claude usage — and it's more nuanced than the "Australia is 1" framing some outlets ran with: high…
OpenAI has published a short paper claiming a full proof of the Cycle Double Cover Conjecture, a graph-theory problem open since the 1970s, crediting the work "entirely" to its GPT-5.6 Sol Ultra…
OpenAI has published the results of a quality audit into SWE-Bench Pro, the coding benchmark it previously recommended as a replacement for the widely used SWE-bench Verified, and concluded the eval…
Databricks has published results from an internal benchmark that tests coding agents against real engineering work pulled from its own multi-million-line codebase, rather than synthetic or curated…
NVIDIA and LangChain say tuning the agent harness around a model — rather than retraining the model itself — pushed Nemotron 3 Ultra to benchmark-leading results among open models on LangChain's…
Cursor has launched a CFO Council, a quarterly working group of enterprise finance leaders, alongside a set of internal usage data showing a widening gap between organizations that are capturing…
Anthropic researchers say they have identified a small set of internal neural patterns in Claude models that function like a global workspace for conscious access in the human brain — a distinct,…
A new academic study finds that U.S. export controls on advanced semiconductors had an unintended side effect: they pushed China to invest more heavily in open-source AI, strengthening an ecosystem…
OpenAI's economic research group has published an extension of its AI Jobs Transition Framework applied specifically to the European Union labor market, estimating how many EU jobs sit in occupations…
Cursor has published results from CursorBench 3.1, the latest version of its internal coding-agent benchmark, with Anthropic's Fable 5 Max topping the leaderboard and a cheaper Cursor-built model…
Anthropic has published the sixth installment of its Economic Index, subtitled "Cadences," shifting the research series from broad occupational trends to the hour-by-hour rhythm of how people…
Together AI published a preview of its ICML 2026 research contributions ahead of the International Conference on Machine Learning, running July 6-11 in Seoul. The company is presenting nine papers…
OpenAI on June 30 released GeneBench-Pro, a benchmark for evaluating AI agents on realistic multi-stage scientific analyses in genomics, quantitative biology, and translational biomedicine. The…
Meta's Fundamental AI Research (FAIR) team published Brain2Qwerty v2 on June 29, a non-invasive brain-computer interface system that converts brain activity directly into typed text, reaching 61%…
OpenAI researchers published a data-driven paper on June 25, 2026 analyzing usage patterns across Codex, their AI coding and agentic task product. The findings show a dramatic expansion beyond the…
Researchers at Together AI have released ParallelKernelBench (PKB), a benchmark designed to test whether frontier large language models can generate optimized multi-GPU CUDA kernels. The results from…
Google on June 10, 2026 released DiffusionGemma, an experimental open-weights language model that generates text through a fundamentally different mechanism than standard autoregressive models. While…
Researchers at Boston Children's Hospital have used OpenAI's o3 reasoning model to identify new diagnoses for 18 children whose rare genetic diseases had gone unresolved by standard clinical workup.…
OpenAI's alignment research team published findings today showing that reinforcement learning applied to a targeted set of "beneficial trait" training scenarios produces safety improvements that…
OpenAI and Molecule.one published research on June 17, 2026 showing a near-autonomous AI system successfully improved a challenging reaction in medicinal chemistry — the first time an AI has tackled…
ChatGPT has lost its majority position in the AI assistant market for the first time since its 2022 launch, dropping to 46.4% market share as of May 2026, according to Sensor Tower's State of AI…
Google's AMIE (Articulate Medical Intelligence Explorer) has matched the performance of primary care physicians in managing chronic health conditions and significantly outperformed them in plan…
Anthropic published research on June 16 showing that the single strongest predictor of success with Claude Code is domain expertise in the task at hand — not whether a user has a background in…
Alibaba's Qwen team released a technical report on June 15, 2026 introducing Qwen-RobotWorld, a foundation model for embodied AI that predicts future visual trajectories across robotics, autonomous…
NVIDIA's Blackwell GPU platform claimed the top spot across all seven benchmarks in MLPerf Training 6.0, the latest edition of the peer-reviewed AI training benchmark suite. The results, published…
Anthropic published research on June 8, 2026, examining why AI agents struggle to reliably retrieve biological data — and how a single deterministic tool layer nearly eliminates the performance gap…
NVIDIA's Blackwell architecture has taken the top spot in AgentPerf, the first benchmark designed specifically to measure agentic AI infrastructure performance, according to results published June 12…
Anthropic on June 12 released the inaugural results from its Anthropic Public Record, a new survey series the company says it will repeat to track how U.S. public attitudes toward AI shift over time.…
Alibaba's Qwen research team has published Qwen-Image-Flash: Beyond Objective Design on arXiv, presenting a few-step distillation framework that accelerates the Qwen-Image-2.0 model family for both…
DeepSeek researchers published a paper on June 9, 2026 introducing Lookahead Sparse Attention (LSA), an inference-time technique that compresses the GPU memory footprint of long-context LLM serving…
Vercel published its AI gateway production index for May 2026, drawing on routing traffic across its infrastructure to document a widening divergence in the AI provider market: DeepSeek's share of…
Anthropic disclosed on June 3, 2026 that it detected and disrupted what it describes as the first large-scale cyberattack executed largely without human intervention — a campaign in which a Chinese…
Google DeepMind published details on May 20, 2026 of an accessibility agent that guides blind and low-vision athletes while running outdoors in real time, using a Pixel 10 Pro smartphone worn on the…
Anthropic on June 5 published benchmark data showing that Claude Opus 4.7 — a general-purpose frontier model with no chemistry-specific fine-tuning — performs comparably to specialized NMR software…
Google DeepMind's AI research assistant Co-Scientist helped biologists at MIT identify more than 20 genetic factors that may reverse cellular aging — and lab experiments confirmed that several of…
A team of researchers including prominent Google DeepMind scientists published a paper at the 43rd International Conference on Machine Learning (ICML 2026) arguing that the dominant paradigm in AI…
Anthropic has released a public best-practices guide for using Claude to find and fix vulnerabilities in source code, paired with an open-source reference harness on GitHub. The guide is grounded in…
Perplexity on June 1 published research introducing Search as Code (SaC), an architecture in which AI agents compose search pipelines programmatically from atomic building blocks rather than calling…
Anthropic on June 3 published a year-long study of how malicious actors have used Claude, examining 832 accounts the company banned for cyber abuse between March 2025 and March 2026. The headline…
On June 2, 2026, Anthropic disclosed that it is significantly expanding Project Glasswing, its program to harden critical-infrastructure software with frontier-model-assisted security analysis.…
Hugging Face users running the Transformers library below version 5.3.0 are exposed to a critical remote-code-execution vulnerability that turns a routine model load into attacker-controlled Python…
On June 2, 2026 a team at the University of Toronto and the Vector Institute released a paper on arXiv showing that an AI agent can power a computer worm capable of writing fresh exploits as it…
Anthropic's engineering team published a retrospective on 2026-05-25 describing how it contains Claude agents across three products — claude.ai, Claude Code, and Claude Cowork — with a focus on…
At CVPR 2026, NVIDIA Research presented three new physical-AI foundation models built around a common thesis: train at sufficient scale, and systems generalize across embodiments, scenarios and…
In the GPT-5.5 system card (April 23, 2026; updated April 24 for API safeguards), OpenAI classifies the model's biological/chemical and cybersecurity capabilities as "High" under its Preparedness…