Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Apple researchers turn failed agent attempts into training insights with RLTL;DR

Apple researchers introduced RLTL;DR, a method that turns verifier feedback from failed agent attempts into short insights for later rollouts and model updates.

D
Oct 1, 2026 · 3 min read

Apple researchers introduced RLTL;DR, a training method for AI agents facing tasks where repeated attempts yield no successful examples. In experiments on selected tool-calling and coding tasks, the researchers reported that the method raised first-attempt success from the 0% to 1% range under a standard reinforcement-learning baseline to 12% to 13% when evaluated without an insight in the prompt.

The method addresses a failure mode in reinforcement learning with verifiable rewards. A verifier can score several agent attempts using signals such as unit tests or the final state of an environment, and group-based methods such as Group Relative Policy Optimization, or GRPO, can reinforce the better attempts. When every sampled attempt fails, however, the group has no successful outcome to learn from.

RLTL;DR changes both exploration and model updating. According to the paper, the policy tries the same task sequentially. After a failed attempt, it reviews the environment and verifier output, then writes a one-sentence insight about what to improve. Later attempts receive the accumulated insights in context when the running success rate is no more than 50%.

The researchers also train the policy to predict those earlier insights from the task description. They implement this as a supervised next-token loss on insight tokens already in the training context, combined with the GRPO objective at a default weight of 0.5. Their aim is to encode the short reminders in the model’s weights so it can apply them on a first attempt without receiving the insight at evaluation time.

The main experiments used Qwen 3.5 9B Thinking and focused on tasks the model had failed in all 128 initial attempts: 458 tasks from Apple’s proprietary Synthetic-API set, 34 AppWorld tasks and 123 LeetCode problems. AppWorld, also used in separate research on AI-agent consistency, is a simulated environment for interactive coding agents with programmatic checks of the resulting application state.

Across those filtered tasks, the authors reported that standard GRPO remained at 0% to 1% Pass@1. RLTL;DR reached 14% to 31% Pass@1 when insights were available in context during training and 12% to 13% on evaluation rollouts without insights. The results have not been independently reproduced.

The paper also reports held-out Pass@1 results of 91.1% on Synthetic-API, 61.7% on AppWorld and 49.1% on LeetCode, compared with 57.8%, 33.7% and 55.1% for GRPO. The authors explicitly caution that these are not benchmark scores because training used only subsets of the tasks and included AppWorld’s test-normal split. They reported no comparable generalization on LeetCode and said RLTL;DR and a self-distillation baseline appeared to overfit there.

A simplified variant, SFTL;DR, updates the model only on task-and-insight pairs instead of full rollout tokens. In one Synthetic-API ablation, the researchers reported that 4,000 unique pairs recovered most of the performance of RLTL;DR and conventional supervised fine-tuning on full rollouts. They measured 4.6 million forward tokens and 68,000 backward tokens for the deduplicated variant, versus 400 million forward and 6.6 million backward tokens for supervised fine-tuning on 3,500 rollouts.

Those figures cover only the model-update stage. The authors said collecting the rollouts needed to generate insights still consumed about one billion tokens across the compared approaches. Sequential sampling also made the sampling phase 4.5 times slower than parallel GRPO at eight attempts in their asynchronous setup, with overall wall-clock time 1.5 times higher.

The evidence is limited to one 9-billion-parameter policy and the tool-calling and coding settings studied. Synthetic-API is proprietary, and the authors said the dynamics that let training on short insights transfer to unaided rollout generation remain unknown.

More news