ProText: A Benchmark Dataset for Measuring (Mis)gendering in Long-Form Texts
Apple just released ProText—a benchmark that exposes how LLMs systematically misgender people in long-form text. Your model might be failing this test.

Why it matters
AI bias in language models is moving beyond simple pronoun resolution to real-world text tasks like summarization and rewriting. ProText gives teams a rigorous way to measure and fix gendering failures before deployment.
The key facts
9 to knowProText dataset measures gendering/misgendering across three dimensions: theme nouns (names, occupations, titles, kinship), theme category (male/female/neutral stereotypes), pronoun category (masculine/feminine/neutral/none)
Designed to evaluate state-of-the-art LLMs on text transformations: summarization and rewrites
Extends beyond traditional pronoun resolution benchmarks to long-form, stylistically diverse English texts
Released by Apple Machine Learning Research
ProText dataset measures gendering/misgendering across three dimensions: Theme nouns, Theme category (male/female/neutral), Pronoun category
Extends beyond traditional pronoun resolution benchmarks to long-form text transformations
Published by Apple Machine Learning Research on March 31, 2026
Designed to probe state-of-the-art LLMs for gender bias in summarization and rewrite tasks
Spans stylistically diverse English texts
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: We introduce ProText, a dataset for measuring gendering and misgendering in stylistically diverse long-form English texts. ProText spans three dimensions: Theme nouns (names, occupations, titles, kinship terms), Theme category (stereotypically male, stereotypically female,…