The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
How much tree-structured reasoning actually survives when language models turn logic into words? Apple Research measured it.

Why it matters
Apple researchers quantified a fundamental limit in how well LLMs preserve compositional structure when serializing reasoning into natural language and back. The round-trip protocol (generate → extract → verify via symbolic equivalence) provides an empirical oracle for understanding degradation across sixteen models—relevant to practitioners building chain-of-thought systems or evaluating model reasoning fidelity.
The key facts
10 to knowRound-trip protocol: procedurally generated arithmetic expressions → word problems → recovered expressions → symbolic equivalence verification
Evaluated all pairwise combinations of 16 models
Tests tree-structured compositional content preservation through serialization bottleneck
Exact oracle via symbolic equivalence (not proxy metrics)
Published Apple Research, September 2026
Round-trip protocol: generator converts procedurally generated arithmetic expressions into word problems; extractor recovers expressions from text alone
Symbolic equivalence provides exact oracle for measurement
Evaluated all pairwise combinations of sixteen models
Published by Apple ML Research, September 2026
Focuses on tree-structured compositional content survival through natural-language serialization
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for…