Apple researchers have published a study identifying a significant communication bottleneck when language models use natural language to exchange or reason through tree-structured information.
Communication Bottleneck: Apple Study Finds LLMs Struggle to Preserve Hierarchical Structure via Natural Language
The study, "The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models," demonstrates that when models convert procedural arithmetic expressions into word problems, a substantial amount of structural information is lost during the translation process. The researchers used a round-trip protocol—converting expressions to text and then back to expressions—to evaluate how sixteen different models perform this task.
Key findings from the evaluation include:
- Asymmetric and Lossy Channels: The accuracy of the round-trip process varies significantly depending on which model is used for generation versus extraction. In some cases, swapping the roles of the models shifted accuracy by up to 60.4 points. Interestingly, using different models for generation and extraction yielded better results than using the same model for both tasks.
- Generation as the Primary Failure Point: At least 73.6% of the round-trip failures were attributed to the generation phase. The difficulty of the task was driven more by the complexity of the tree structure—such as operator count, depth, and right-branching—than by the specific model family used.
- Trainability through Fine-Tuning: The study found that this communication gap can be mitigated through training. Using approximately 3,600 fine-tuning examples, researchers were able to lift the performance of all tested open-weight models above an untrained Gemini-3.1-Pro.
The results suggest that the necessity of serializing hierarchical structures into linear natural language is a primary limiting factor for models attempting to communicate complex, structured reasoning.
Sources
- The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models (Apple Machine Learning, 2026-09-29)
- arXiv