FrontierThe story, in brief

Beyond 100K Tokens: Evaluating AI Agents in Long-Context Software Engineering

Beyond 100K tokens. Salesforce just released the benchmark that proves whether AI agents can actually code at enterprise scale.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

As codebases balloon to millions of lines, context-window capability becomes the differentiator between toy models and production-ready agents. LoCoBench-Agent gives founders and CTOs a standardized way to measure which models can handle real software engineering work.

The key facts

6 to know
  1. Benchmark name: LoCoBench-Agent

  2. Focus: AI agent evaluation in long-context software engineering tasks

  3. Context range: Beyond 100K tokens

  4. Use case: Coding assistants on large, million-line codebases

  5. Source: Salesforce Research

  6. Category: Model capability evaluation for code reasoning and generation

Go to the source

Salesforce Blogsalesforce.com

Publisher excerpt: As codebases grow to millions of lines of code, can AI agents still understand, reason, and code effectively? LoCoBench-Agent delivers the answer: a comprehensive benchmark for evaluating AI coding assistants across contexts ranging…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier