| 1 | Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence | Frontier | 85 | The Decoder |
| 2 | DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities | Frontier | 82 | Vercel Blog |
| 3 | Anthropic says its Mythos model found vulnerabilities in cryptographic algorithms that secure the internet | Frontier | 78 | The Decoder |
| 4 | Microsoft launches MAI-Cyber-1-Flash cybersecurity AI model - qz.com | Frontier | 78 | Reuters Technology |
| 5 | Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race | Frontier | 78 | The Decoder |
| 6 | Exclusive: Cogent Security debuts VR-1, a frontier model built to prove attack paths | Frontier | 78 | SiliconAngle |
| 7 | Why China is giving away its best AI models | Frontier | 78 | The Verge AI |
| 8 | Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction | Frontier | 78 | MarkTechPost |
| 9 | Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks | Frontier | 78 | The Decoder |
| 10 | How GPT-5.6 fuses frontier intelligence with frontier efficiency | Frontier | 75 | OpenAI Blog |
| 11 | Thinking Machines bets on efficiency over size with its second model, Inkling Small | Frontier | 72 | The Decoder |
| 12 | Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids | Frontier | 72 | The Decoder |
| 13 | DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains | Frontier | 72 | MarkTechPost |
| 14 | PolyAI launches new real-time voice conversation model to make AI-driven calls more human | Frontier | 72 | SiliconAngle |
| 15 | Google DeepMind debuts Gemini Robotics 2 model series for humanoid robots | Frontier | 72 | SiliconAngle |
| 16 | Microsoft AI bets on cheap specialist models instead of chasing the frontier | Frontier | 72 | The Decoder |
| 17 | Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it | Frontier | 72 | The Decoder |
| 18 | Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models | Frontier | 72 | MarkTechPost |
| 19 | Google DeepMind Ships Three Physical AI Models For Whole Body Control, Dexterity And Multi Robot Collaboration | Frontier | 72 | MarkTechPost |
| 20 | PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response | Frontier | 72 | MarkTechPost |
| 21 | China’s open-weight model lead exposes America’s AI blind spot | Frontier | 72 | CNBC Technology |
| 22 | Gemini Robotics 2 Brings Google's AI Into the Physical World | Frontier | 72 | Wired AI |
| 23 | Google DeepMind’s new AI model can control a robot’s entire body | Frontier | 72 | The Verge AI |
| 24 | Gemini Robotics 2 brings whole body intelligence to robots | Frontier | 72 | Google DeepMind Blog |
| 25 | Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration | Frontier | 72 | Google DeepMind Blog |
| 26 | Microsoft is openly competing with OpenAI, Anthropic more than ever | Frontier | 72 | TechCrunch AI |
| 27 | It’s Frighteningly Easy to Jailbreak Some Frontier AI Models | Frontier | 72 | Wired AI |
| 28 | ByteDance’s plan to dominate AI | Frontier | 72 | Financial Times Technology |
| 29 | How enabling two settings tripled our scores on the ARC-AGI-3 benchmark | Frontier | 72 | OpenAI Blog |
| 30 | Amazon reportedly scales back its Nova AI models and bets on a new Frontier research team | Frontier | 72 | The Decoder |
| 31 | moonshotai/Kimi-K3 | Frontier | 72 | Simon Willison |
| 32 | Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks | Frontier | 72 | The Decoder |
| 33 | Kimi AI and kvcache-ai Open Sources ‘AgentENV’: A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3 | Frontier | 72 | MarkTechPost |
| 34 | Microsoft AI Releases MAI-Cyber-1-Flash: A 5B-Active-Parameter Cyber Model That Pushes MDASH to 95.95% on CyberGym | Frontier | 72 | MarkTechPost |
| 35 | Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker | Frontier | 72 | Import AI (Blog) |
| 36 | The AI giants’ new problem: open AI - The Verge | Frontier | 72 | Reuters Technology |
| 37 | Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run | Frontier | 72 | MarkTechPost |
| 38 | KwaiKAT Team Releases KAT-Coder-V2.5: An Agentic Coding Model Trained on 100,000+ Verifiable Repository Environments | Frontier | 72 | MarkTechPost |
| 39 | Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents | Frontier | 72 | The Decoder |
| 40 | Meet Open Dreamer: A JAX/Flax Reproduction of the Dreamer 4 World Model Pipeline, With the Full Training Recipe Published | Frontier | 72 | MarkTechPost |
| 41 | Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM | Frontier | 72 | MarkTechPost |
| 42 | FAIRChem v2 UMA for Multidomain Atomistic Simulation across Molecules, Catalysts, Materials, Vibrations, and Molecular Dynamics | Frontier | 72 | MarkTechPost |
| 43 | Introducing Claude Opus 5 | Frontier | 72 | Simon Willison |
| 44 | Anthropic launches Claude Opus 5 with efficiency, safety improvements | Frontier | 72 | SiliconAngle |
| 45 | Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown | Frontier | 72 | MarkTechPost |
| 46 | Open weights vs. closed: An AI civil war's afoot, and the stakes are existential | Frontier | 68 | ZDNet AI |
| 47 | Language models can't spark scientific revolutions, but world models might | Frontier | 62 | The Decoder |
| 48 | GPT Transcribe improves on its predecessor but can't catch ElevenLabs, Google, or Mistral on error rates | Frontier | 62 | The Decoder |
| 49 | Liquid AI Releases LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: Bidirectional Encoders That Stay Fast at 8K Context on CPU | Frontier | 62 | MarkTechPost |
| 50 | Moonshot AI Open-Sources MoonEP: A Perfectly Balanced Expert Parallelism Library for MoE Training | Frontier | 62 | MarkTechPost |
| 51 | LFM2.5-Encoders for Fast Long-Context Inference on CPU | Frontier | 62 | Hugging Face Blog |
| 52 | EvoLib: Turning experience into evolving knowledge | Frontier | 55 | Microsoft Research |
| 53 | smevals - a small eval suite for evaluating models, prompts, and harnesses | Frontier | 52 | Simon Willison |
| 54 | Nvidia’s Open Source Alliance Is Missing Some Key Names: OpenAI and Anthropic | Frontier | 52 | Wired AI |
| 55 | MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization | Frontier | 45 | Apple Machine Learning |
| 56 | Everyone Is Freaking Out About OpenAI and Anthropic’s Race for Dominance | Frontier | 45 | Wired AI |
| 57 | OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings | Frontier | 45 | The Decoder |
| 58 | Latest AI Uses Tabular Foundation Models To Turn Columnar Data Into Vital Insights | Frontier | 45 | Forbes Innovation |
| 59 | How controllers from industrial machinery can coordinate multitask machine learning | Frontier | 42 | Amazon Science |