SpecializedCognition

SWE-2

Context

200K tokens

Pricing

Available via Devin subscription and Cognition API; enterprise pricing on request

Modalities

text, code

Released

Sep 2025

Overview
SWE-2 is Cognition's second-generation autonomous software engineering model, designed to plan, write, debug, and iterate on code across entire codebases with minimal human intervention. Unlike copilot-style tools that assist developers line by line, SWE-2 operates as a full agent capable of completing multi-step engineering tasks end to end. It powers Devin, Cognition's AI software engineer product, and is benchmarked against real-world GitHub issue resolution and enterprise engineering workflows.
Why it matters
SWE-2 arrives at the inflection point where the market is deciding whether autonomous coding agents are a productivity multiplier or a wholesale replacement for junior engineering roles — and Cognition's $48B valuation suggests investors have already made their bet. For CTOs, the model represents a credible path to shrinking headcount on routine engineering tasks: bug fixes, test generation, refactoring, and greenfield feature work within well-scoped repositories. The competitive pressure is real: OpenAI's GPT-6 Astra, Anthropic's Claude Opus 4, and Google's Gemini 2.5 Pro are all competing for the same autonomous coding workflows, meaning SWE-2's window to establish defensible benchmark leadership is narrow. If Cognition can sustain accuracy on SWE-bench at scale without the hallucination and context-loss failures that plague longer-horizon tasks, it becomes a serious enterprise procurement decision, not just a developer toy. The $48B valuation stakes are high enough that any capability regression or security incident — like the supply-chain exploit attributed to OpenAI agents in May 2025 — could materially damage adoption momentum.

Key strengths

  • End-to-end autonomous task completion on real GitHub repositories, not just code completion
  • Multi-step planning and self-correction across long-horizon engineering workflows
  • State-of-the-art SWE-bench performance targeting enterprise-grade issue resolution
  • Designed for agentic orchestration — can spawn sub-tasks, run tests, and iterate without human checkpoints
  • Integrated with Devin's enterprise deployment harness, including sandboxed execution environments

THE FRIDAY BRIEFING

We cover ai models every week.

Subscribe free →

Know the terms. Know the moves.

ONE BRIEFING · EVERY FRIDAY · FREE

Free. Unsubscribe anytime.