AgentsAugust 31, 2026via Vercel Blog

How our agents build on-brand pages with design.md

Why it matters

This is a production blueprint for making agents reliable at a specific, repeatable task. Vercel's eval methodology—frozen scenarios, blind A/B comparisons, deterministic checks layered with prose guidance—is a reusable pattern for any team deploying agents to generate consistent artifacts. The 57% failure reduction is modest but real, and the framework shows how to measure and iterate.

Key signals

  • design.md is a public file any agent can load to generate on-brand pages
  • Seven fixed eval scenarios tested against Claude Opus 4.8 and GPT-5.5
  • 200+ runs to build the guidance file through iterative evaluation
  • 57% reduction in known failures (39 failures with design.md vs. 91 without, on 6 test pages)
  • Three-part system: prose guidance in design.md, CSS stylesheet with constrained primitives, deterministic mechanical checks
  • @design-agent Slack bot uses design.md in production to build pages, reports, proposals
  • Weekly feedback loop: Slack threads + GitHub reviews + Figma comments aggregated to identify recurring complaints
  • Each correction encoded in narrowest applicable layer: prose guidance, stylesheet, or code check

The hook

Vercel shipped a public design.md file that cuts agent design failures by 57%. Here's how they built the eval loop that made it work.

Across Vercel, we use coding agents to design and build pages that have to look and feel like Vercel. The typography, color, and composition all need to carry the same judgment that we put into the pages we already ship ourselves. , our skill that teaches agents how we design when they work in our

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.