Full KeyRank table

Every qualifying story, ranked. How KeyRank works

#StoryPillarKeyRankSource
1OpenAI Releases GPT-5.6 (Sol, Terra, Luna): A Three-Tier Model Family With Programmatic Tool Calling in the Responses APIFrontier92MarkTechPost
2Open-weight models surge to 29% of volume, price per token flattensFrontier88Vercel Blog
3Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberFrontier85Google DeepMind Blog
4Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligenceFrontier85The Decoder
5Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus PricingFrontier85MarkTechPost
6OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"Frontier85The Decoder
7The new GPT-5.6 family: Luna, Terra, SolFrontier85Simon Willison
8OpenAI's GPT-5.6 and ChatGPT Work aim to beat Anthropic on price, speed, and productivityFrontier85ZDNet AI
9OpenAI Releases GPT-Live and GPT-Live-1 mini: Full-Duplex Voice Models That Delegate Deeper Reasoning to GPT-5.5Frontier85MarkTechPost
10Tencent Releases Hy3: An Open 295B Mixture-of-Experts (MoE) Model with 21B Active Parameters and 256K ContextFrontier85MarkTechPost
11DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilitiesFrontier82Vercel Blog
12Anthropic says its Mythos model found vulnerabilities in cryptographic algorithms that secure the internetFrontier78The Decoder
13Microsoft launches MAI-Cyber-1-Flash cybersecurity AI model - qz.comFrontier78Reuters Technology
14Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model raceFrontier78The Decoder
15Exclusive: Cogent Security debuts VR-1, a frontier model built to prove attack pathsFrontier78SiliconAngle
16Why China is giving away its best AI modelsFrontier78The Verge AI
17Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action PredictionFrontier78MarkTechPost
18Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarksFrontier78The Decoder
19Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain whyFrontier78The Decoder
20Claude Opus 5 arrives with near Fable performance at half the priceFrontier78ZDNet AI
21Anthropic's new AI model rivals Fable 5 and is cheaper as businesses fret about costsFrontier78CNBC Technology
22Google CEO Pichai says Gemini's next leap depends on building "much larger base models"Frontier78The Decoder
23Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above its sizeFrontier78The Decoder
24Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest LabsFrontier78The Decoder
25Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI - VentureBeatFrontier78Reuters Technology
26Google ships three new Gemini Flash models but its frontier 3.5 Pro remains lost in trainingFrontier78The Decoder
27Google’s Gemini 3.6 Flash targets enterprise agent token costsFrontier78AI News
28Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic WorkloadsFrontier78MarkTechPost
29Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real CodebasesFrontier78MarkTechPost
30China delivers a one-two punch to America’s AI dominanceFrontier78The Verge AI
31Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex mathFrontier78The Decoder
32Kimi K3 open-weight model: China’s biggest AI is a bet on memory, not computeFrontier78AI News
33Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantageFrontier78The Decoder
34Ex-OpenAI CTO Murati's Thinking Machines drops Inkling, a 975B parameter model that leads US labs but trails ChinaFrontier78The Decoder
35Nvidia launches Cosmos 3 Edge model and expands its physical AI push in JapanFrontier78SiliconAngle
36China’s Moonshot throws down the gauntlet with Kimi K3, the world’s largest open-weights modelFrontier78SiliconAngle
37OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt InjectionFrontier78MarkTechPost
38NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEBFrontier78MarkTechPost
39Chinese AI start-up Moonshot launches model challenging Anthropic’s leadFrontier78Financial Times Technology
40GPT-5.6 Sol reportedly disproves a 30-year-old statistics conjecture in 90 minutes after humans couldn't crack itFrontier78The Decoder
41Meet GPT-Red: an LLM super-hacker OpenAI built to make its models saferFrontier78MIT Technology Review
42Anthropic Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8: Agentic Coding Benchmarks, API Pricing, and Cost-Performance Tradeoffs ComparedFrontier78MarkTechPost
43China's Orca world model matches specialized robotics systems without ever seeing a single action labelFrontier78The Decoder
44OpenAI finds roughly 30 percent of popular AI coding test is brokenFrontier78The Decoder
45Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server ThroughputFrontier78MarkTechPost
46Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Trillion Minutes of Sensor DataFrontier78MarkTechPost
47OpenAI's newest AI model is 54% more token efficient on agentic coding, Altman tells CNBCFrontier78CNBC Technology
48Anthropic's Claude Fable 5 dominates new industry benchmarks at a steep premiumFrontier78The Decoder
49Grok 4.5 is so cheap compared to Fable 5 and GPT 5.5 that benchmark gaps may not matter muchFrontier78The Decoder
50NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User ThroughputFrontier78MarkTechPost
51Grok 4.5Frontier78Hacker News
52OpenAI launches GPT-Live voice models that listen and speak simultaneously - ReutersFrontier78Reuters Technology
53Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian LensFrontier78The Decoder
54GPT-4's dominance lasted a year while today's top models barely survive seven weeks at the topFrontier78The Decoder
55Mistral's open-source Leanstral 1.5 aces formal math benchmarks and catches real bugs in codeFrontier78The Decoder
56Mistral AI Releases Leanstral 1.5: An Apache-2.0 Lean 4 Code Agent Model Solving 587 of 672 PutnamBench ProblemsFrontier78MarkTechPost
57Leanstral 1.5: Proof abundance for allFrontier78Hacker News
58How GPT-5.6 fuses frontier intelligence with frontier efficiencyFrontier75OpenAI Blog
59NVIDIA Releases Cosmos 3 Edge: A 4B-Parameter Open World Model That Reasons and Generates Robot Actions On-DeviceFrontier75MarkTechPost
60Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving CostFrontier75MarkTechPost
61Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AIFrontier75The Decoder
62OpenAI's AI beats every human at AtCoder, a top competitive programming contestFrontier75The Decoder
63GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the costFrontier75The Decoder
64Meta Superintelligence Labs Releases Muse Spark 1.1: A Multimodal Reasoning Model for Agentic Tasks on Meta Model APIFrontier75MarkTechPost
65SpaceXAI Releases Grok 4.5, a Cursor-Trained Model for Coding, Agentic Tasks, and Knowledge Work at $2/M InputFrontier75MarkTechPost
66OpenAI's GPT-5.6 launches Thursday after a delay forced by the U.S. governmentFrontier75The Decoder
67Thinking Machines bets on efficiency over size with its second model, Inkling SmallFrontier72The Decoder
68Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoidsFrontier72The Decoder
69DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding GainsFrontier72MarkTechPost
70PolyAI launches new real-time voice conversation model to make AI-driven calls more humanFrontier72SiliconAngle
71Google DeepMind debuts Gemini Robotics 2 model series for humanoid robotsFrontier72SiliconAngle
72Microsoft AI bets on cheap specialist models instead of chasing the frontierFrontier72The Decoder
73Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut itFrontier72The Decoder
74Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 ModelsFrontier72MarkTechPost
75Google DeepMind Ships Three Physical AI Models For Whole Body Control, Dexterity And Multi Robot CollaborationFrontier72MarkTechPost
76PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And ResponseFrontier72MarkTechPost
77China’s open-weight model lead exposes America’s AI blind spotFrontier72CNBC Technology
78Gemini Robotics 2 Brings Google's AI Into the Physical WorldFrontier72Wired AI
79Google DeepMind’s new AI model can control a robot’s entire bodyFrontier72The Verge AI
80Gemini Robotics 2 brings whole body intelligence to robotsFrontier72Google DeepMind Blog
81Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationFrontier72Google DeepMind Blog
82Microsoft is openly competing with OpenAI, Anthropic more than everFrontier72TechCrunch AI
83It’s Frighteningly Easy to Jailbreak Some Frontier AI ModelsFrontier72Wired AI
84ByteDance’s plan to dominate AIFrontier72Financial Times Technology
85How enabling two settings tripled our scores on the ARC-AGI-3 benchmarkFrontier72OpenAI Blog
86Amazon reportedly scales back its Nova AI models and bets on a new Frontier research teamFrontier72The Decoder
87moonshotai/Kimi-K3Frontier72Simon Willison
88Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasksFrontier72The Decoder
89Kimi AI and kvcache-ai Open Sources ‘AgentENV’: A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3Frontier72MarkTechPost
90Microsoft AI Releases MAI-Cyber-1-Flash: A 5B-Active-Parameter Cyber Model That Pushes MDASH to 95.95% on CyberGymFrontier72MarkTechPost
91Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hackerFrontier72Import AI (Blog)
92The AI giants’ new problem: open AI - The VergeFrontier72Reuters Technology
93Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining RunFrontier72MarkTechPost
94KwaiKAT Team Releases KAT-Coder-V2.5: An Agentic Coding Model Trained on 100,000+ Verifiable Repository EnvironmentsFrontier72MarkTechPost
95Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agentsFrontier72The Decoder
96Meet Open Dreamer: A JAX/Flax Reproduction of the Dreamer 4 World Model Pipeline, With the Full Training Recipe PublishedFrontier72MarkTechPost
97Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALMFrontier72MarkTechPost
98FAIRChem v2 UMA for Multidomain Atomistic Simulation across Molecules, Catalysts, Materials, Vibrations, and Molecular DynamicsFrontier72MarkTechPost
99Introducing Claude Opus 5Frontier72Simon Willison
100Anthropic launches Claude Opus 5 with efficiency, safety improvementsFrontier72SiliconAngle