AWS details multi-turn RL workflow for SageMaker search agents
AWS has published a managed SageMaker AI workflow for training search agents across multi-step trajectories, reporting nDCG@10 gains on three of four held-out datasets but a small regression on FreshStack.
AWS published a SageMaker AI workflow for fine-tuning search agents with multi-turn reinforcement learning. In its own evaluation runs, the fine-tuned model recorded higher retrieval scores on three of four held-out datasets and lower failure rates on three, while FreshStack showed a small decline in retrieval quality.
For SageMaker users, the workflow brings trajectory-level optimization to managed SageMaker AI training jobs, allowing a search agent to be optimized across a complete sequence of tool calls rather than training each response in isolation.
The experiment used Qwen3.6-27B in AWS’s us-west-2 region. The agent could call BM25 lexical search and vector search, and its number of interactions was capped.
The SageMaker AI documentation describes multi-turn RL as training a policy across sequences of states, actions and rewards to maximize cumulative reward over an episode. It lists Qwen3.6-27B as a supported model in us-west-2 and says usage is charged across prefill, sample and train-token dimensions.
AWS used nDCG@10—the normalized discounted cumulative gain of the top 10 retrieved results—as the reward for the full trajectory. A run received a reward of minus one if the agent reached either the maximum number of turns or the per-turn sampling-token limit. The example changed three job settings: one maximum epoch, a global batch size of 128 and rollout concurrency of 32. The algorithm, advantage estimator and off-policy staleness settings remained at service defaults.
AWS said its training mix included FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique and MLQA, with 5% of each dataset reserved for validation. It evaluated the resulting model on four held-out benchmarks: FreshStack for recent technical-document retrieval, WixQA for enterprise support knowledge bases, BrowseComp-Plus for deep-research retrieval over a fixed corpus, and WANDS for product-search relevance.
In AWS’s table, nDCG@10 rose from 0.5725 to 0.6781 on WixQA, from 0.5762 to 0.6112 on WANDS and from 0.5136 to 0.6354 on BrowseComp-Plus. AWS characterized those changes as relative gains of 18.4%, 6% and 23.7%, respectively. FreshStack moved in the opposite direction, slipping from 0.4112 to 0.4089.
Reported failure rates fell from 0.67% to 0.17% on WixQA, from 0.20% to 0.05% on FreshStack and from 22.89% to 0.68% on BrowseComp-Plus. WANDS stayed at 0% for both the base and fine-tuned models.
Average turn counts moved in both directions. They increased from 4.3 to 4.5 on WixQA and from 2.2 to 2.9 on WANDS, while decreasing from 3.1 to 2.8 on FreshStack and from 7.0 to 6.3 on BrowseComp-Plus. The table therefore does not support a claim that fine-tuning reduced average turns on every held-out dataset.
The benchmark maintainers’ materials establish what the four datasets measure, but they do not validate AWS’s SageMaker results. AWS did not identify an independent evaluator or an independently reproduced run. Its post also did not report confidence intervals, variance, random-seed counts or statistical-significance tests, making the small FreshStack regression and WANDS gain difficult to assess.
AWS provided a short configuration example but did not link a complete experiment repository, trained checkpoint, processed datasets, prompts, exact turn and token caps, retriever configuration or the full default RL settings needed to reproduce the scores. It also omitted total training-token counts, elapsed training time and full experiment cost.
More news

MIT and Sakana AI propose SIFT to prioritize coding-agent benchmark runs

Google outlines verifiable private federated learning system for Gboard

NetApp puts approval gates around Workload Factory AI agents
