| 1 | OpenAI Releases GPT-5.6 (Sol, Terra, Luna): A Three-Tier Model Family With Programmatic Tool Calling in the Responses API | Frontier | 92 | MarkTechPost |
| 2 | Open-weight models surge to 29% of volume, price per token flattens | Frontier | 88 | Vercel Blog |
| 3 | Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber | Frontier | 85 | Google DeepMind Blog |
| 4 | Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence | Frontier | 85 | The Decoder |
| 5 | Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing | Frontier | 85 | MarkTechPost |
| 6 | OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt" | Frontier | 85 | The Decoder |
| 7 | The new GPT-5.6 family: Luna, Terra, Sol | Frontier | 85 | Simon Willison |
| 8 | OpenAI's GPT-5.6 and ChatGPT Work aim to beat Anthropic on price, speed, and productivity | Frontier | 85 | ZDNet AI |
| 9 | OpenAI Releases GPT-Live and GPT-Live-1 mini: Full-Duplex Voice Models That Delegate Deeper Reasoning to GPT-5.5 | Frontier | 85 | MarkTechPost |
| 10 | Tencent Releases Hy3: An Open 295B Mixture-of-Experts (MoE) Model with 21B Active Parameters and 256K Context | Frontier | 85 | MarkTechPost |
| 11 | DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities | Frontier | 82 | Vercel Blog |
| 12 | Anthropic says its Mythos model found vulnerabilities in cryptographic algorithms that secure the internet | Frontier | 78 | The Decoder |
| 13 | Microsoft launches MAI-Cyber-1-Flash cybersecurity AI model - qz.com | Frontier | 78 | Reuters Technology |
| 14 | Moonshot AI releases Kimi K3 open weights and infrastructure after shaking up the frontier model race | Frontier | 78 | The Decoder |
| 15 | Exclusive: Cogent Security debuts VR-1, a frontier model built to prove attack paths | Frontier | 78 | SiliconAngle |
| 16 | Why China is giving away its best AI models | Frontier | 78 | The Verge AI |
| 17 | Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction | Frontier | 78 | MarkTechPost |
| 18 | Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks | Frontier | 78 | The Decoder |
| 19 | Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why | Frontier | 78 | The Decoder |
| 20 | Claude Opus 5 arrives with near Fable performance at half the price | Frontier | 78 | ZDNet AI |
| 21 | Anthropic's new AI model rivals Fable 5 and is cheaper as businesses fret about costs | Frontier | 78 | CNBC Technology |
| 22 | Google CEO Pichai says Gemini's next leap depends on building "much larger base models" | Frontier | 78 | The Decoder |
| 23 | Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above its size | Frontier | 78 | The Decoder |
| 24 | Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs | Frontier | 78 | The Decoder |
| 25 | Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI - VentureBeat | Frontier | 78 | Reuters Technology |
| 26 | Google ships three new Gemini Flash models but its frontier 3.5 Pro remains lost in training | Frontier | 78 | The Decoder |
| 27 | Google’s Gemini 3.6 Flash targets enterprise agent token costs | Frontier | 78 | AI News |
| 28 | Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads | Frontier | 78 | MarkTechPost |
| 29 | Cisco Foundation AI Releases Antares: 350M and 1B Open-Weight Models That Localize Known Vulnerabilities Inside Real Codebases | Frontier | 78 | MarkTechPost |
| 30 | China delivers a one-two punch to America’s AI dominance | Frontier | 78 | The Verge AI |
| 31 | Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math | Frontier | 78 | The Decoder |
| 32 | Kimi K3 open-weight model: China’s biggest AI is a bet on memory, not compute | Frontier | 78 | AI News |
| 33 | Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage | Frontier | 78 | The Decoder |
| 34 | Ex-OpenAI CTO Murati's Thinking Machines drops Inkling, a 975B parameter model that leads US labs but trails China | Frontier | 78 | The Decoder |
| 35 | Nvidia launches Cosmos 3 Edge model and expands its physical AI push in Japan | Frontier | 78 | SiliconAngle |
| 36 | China’s Moonshot throws down the gauntlet with Kimi K3, the world’s largest open-weights model | Frontier | 78 | SiliconAngle |
| 37 | OpenAI Details GPT-Red: An Internal Automated Red-Teaming Model That Beat Human Red-Teamers 84% To 13% On Prompt Injection | Frontier | 78 | MarkTechPost |
| 38 | NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB | Frontier | 78 | MarkTechPost |
| 39 | Chinese AI start-up Moonshot launches model challenging Anthropic’s lead | Frontier | 78 | Financial Times Technology |
| 40 | GPT-5.6 Sol reportedly disproves a 30-year-old statistics conjecture in 90 minutes after humans couldn't crack it | Frontier | 78 | The Decoder |
| 41 | Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer | Frontier | 78 | MIT Technology Review |
| 42 | Anthropic Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8: Agentic Coding Benchmarks, API Pricing, and Cost-Performance Tradeoffs Compared | Frontier | 78 | MarkTechPost |
| 43 | China's Orca world model matches specialized robotics systems without ever seeing a single action label | Frontier | 78 | The Decoder |
| 44 | OpenAI finds roughly 30 percent of popular AI coding test is broken | Frontier | 78 | The Decoder |
| 45 | Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput | Frontier | 78 | MarkTechPost |
| 46 | Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Trillion Minutes of Sensor Data | Frontier | 78 | MarkTechPost |
| 47 | OpenAI's newest AI model is 54% more token efficient on agentic coding, Altman tells CNBC | Frontier | 78 | CNBC Technology |
| 48 | Anthropic's Claude Fable 5 dominates new industry benchmarks at a steep premium | Frontier | 78 | The Decoder |
| 49 | Grok 4.5 is so cheap compared to Fable 5 and GPT 5.5 that benchmark gaps may not matter much | Frontier | 78 | The Decoder |
| 50 | NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput | Frontier | 78 | MarkTechPost |
| 51 | Grok 4.5 | Frontier | 78 | Hacker News |
| 52 | OpenAI launches GPT-Live voice models that listen and speak simultaneously - Reuters | Frontier | 78 | Reuters Technology |
| 53 | Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens | Frontier | 78 | The Decoder |
| 54 | GPT-4's dominance lasted a year while today's top models barely survive seven weeks at the top | Frontier | 78 | The Decoder |
| 55 | Mistral's open-source Leanstral 1.5 aces formal math benchmarks and catches real bugs in code | Frontier | 78 | The Decoder |
| 56 | Mistral AI Releases Leanstral 1.5: An Apache-2.0 Lean 4 Code Agent Model Solving 587 of 672 PutnamBench Problems | Frontier | 78 | MarkTechPost |
| 57 | Leanstral 1.5: Proof abundance for all | Frontier | 78 | Hacker News |
| 58 | How GPT-5.6 fuses frontier intelligence with frontier efficiency | Frontier | 75 | OpenAI Blog |
| 59 | NVIDIA Releases Cosmos 3 Edge: A 4B-Parameter Open World Model That Reasons and Generates Robot Actions On-Device | Frontier | 75 | MarkTechPost |
| 60 | Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost | Frontier | 75 | MarkTechPost |
| 61 | Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI | Frontier | 75 | The Decoder |
| 62 | OpenAI's AI beats every human at AtCoder, a top competitive programming contest | Frontier | 75 | The Decoder |
| 63 | GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost | Frontier | 75 | The Decoder |
| 64 | Meta Superintelligence Labs Releases Muse Spark 1.1: A Multimodal Reasoning Model for Agentic Tasks on Meta Model API | Frontier | 75 | MarkTechPost |
| 65 | SpaceXAI Releases Grok 4.5, a Cursor-Trained Model for Coding, Agentic Tasks, and Knowledge Work at $2/M Input | Frontier | 75 | MarkTechPost |
| 66 | OpenAI's GPT-5.6 launches Thursday after a delay forced by the U.S. government | Frontier | 75 | The Decoder |
| 67 | Thinking Machines bets on efficiency over size with its second model, Inkling Small | Frontier | 72 | The Decoder |
| 68 | Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids | Frontier | 72 | The Decoder |
| 69 | DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains | Frontier | 72 | MarkTechPost |
| 70 | PolyAI launches new real-time voice conversation model to make AI-driven calls more human | Frontier | 72 | SiliconAngle |
| 71 | Google DeepMind debuts Gemini Robotics 2 model series for humanoid robots | Frontier | 72 | SiliconAngle |
| 72 | Microsoft AI bets on cheap specialist models instead of chasing the frontier | Frontier | 72 | The Decoder |
| 73 | Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it | Frontier | 72 | The Decoder |
| 74 | Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models | Frontier | 72 | MarkTechPost |
| 75 | Google DeepMind Ships Three Physical AI Models For Whole Body Control, Dexterity And Multi Robot Collaboration | Frontier | 72 | MarkTechPost |
| 76 | PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response | Frontier | 72 | MarkTechPost |
| 77 | China’s open-weight model lead exposes America’s AI blind spot | Frontier | 72 | CNBC Technology |
| 78 | Gemini Robotics 2 Brings Google's AI Into the Physical World | Frontier | 72 | Wired AI |
| 79 | Google DeepMind’s new AI model can control a robot’s entire body | Frontier | 72 | The Verge AI |
| 80 | Gemini Robotics 2 brings whole body intelligence to robots | Frontier | 72 | Google DeepMind Blog |
| 81 | Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration | Frontier | 72 | Google DeepMind Blog |
| 82 | Microsoft is openly competing with OpenAI, Anthropic more than ever | Frontier | 72 | TechCrunch AI |
| 83 | It’s Frighteningly Easy to Jailbreak Some Frontier AI Models | Frontier | 72 | Wired AI |
| 84 | ByteDance’s plan to dominate AI | Frontier | 72 | Financial Times Technology |
| 85 | How enabling two settings tripled our scores on the ARC-AGI-3 benchmark | Frontier | 72 | OpenAI Blog |
| 86 | Amazon reportedly scales back its Nova AI models and bets on a new Frontier research team | Frontier | 72 | The Decoder |
| 87 | moonshotai/Kimi-K3 | Frontier | 72 | Simon Willison |
| 88 | Microsoft launches its own cybersecurity model MAI-Cyber-1-Flash but still depends on OpenAI for the toughest tasks | Frontier | 72 | The Decoder |
| 89 | Kimi AI and kvcache-ai Open Sources ‘AgentENV’: A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3 | Frontier | 72 | MarkTechPost |
| 90 | Microsoft AI Releases MAI-Cyber-1-Flash: A 5B-Active-Parameter Cyber Model That Pushes MDASH to 95.95% on CyberGym | Frontier | 72 | MarkTechPost |
| 91 | Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker | Frontier | 72 | Import AI (Blog) |
| 92 | The AI giants’ new problem: open AI - The Verge | Frontier | 72 | Reuters Technology |
| 93 | Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run | Frontier | 72 | MarkTechPost |
| 94 | KwaiKAT Team Releases KAT-Coder-V2.5: An Agentic Coding Model Trained on 100,000+ Verifiable Repository Environments | Frontier | 72 | MarkTechPost |
| 95 | Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents | Frontier | 72 | The Decoder |
| 96 | Meet Open Dreamer: A JAX/Flax Reproduction of the Dreamer 4 World Model Pipeline, With the Full Training Recipe Published | Frontier | 72 | MarkTechPost |
| 97 | Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM | Frontier | 72 | MarkTechPost |
| 98 | FAIRChem v2 UMA for Multidomain Atomistic Simulation across Molecules, Catalysts, Materials, Vibrations, and Molecular Dynamics | Frontier | 72 | MarkTechPost |
| 99 | Introducing Claude Opus 5 | Frontier | 72 | Simon Willison |
| 100 | Anthropic launches Claude Opus 5 with efficiency, safety improvements | Frontier | 72 | SiliconAngle |