English

Model ReleasesApple

Apple Researchers Propose RLTL;DR to Enable Self-Improvement Through Self-Generated Feedback

Apple Machine Learning researchers have introduced RLTL;DR, a new paradigm designed to enable self-improvement in reinforcement learning when tasks are too difficult for agents to succeed through standard methods.

The common approach, reinforcement learning with verifiable rewards (RLVR), relies on agents making successful attempts to optimize behavior. However, this becomes problematic when tasks are so difficult that the agent has little to no chance of success, or when no teacher models are available.

RLTL;DR addresses this by having the policy observe a verifier's output after a failed attempt and then generating its own feedback in the form of a single "TL;DR" insight. Subsequent attempts are conditioned on these previous insights. The method also enables backpropagation on these in-context insights to internalize the mapping from task to insight.

In tests on challenging tool-calling and coding datasets, where standard GRPO training with a Qwen 3.5 9B Thinking policy failed to achieve significant success (0% to 1% Pass@1), RLTL;DR achieved a Pass@1 of 14–31% during training. Crucially, it maintained a 12–13% Pass@1 at evaluation time even without insights in context. The researchers noted that the internalization of task-to-insight mappings was a key factor in this performance breakthrough.

Sources

  1. RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback (Apple Machine Learning, 2026-10-01)
  2. arXiv