Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

ServiceNow CoreAI introduces AutoSynthData for enterprise-agent training

ServiceNow says AutoSynthData turns model failures and teacher successes into verified synthetic tasks, reporting gains in two EnterpriseOps-Gym domains.

D
Oct 2, 2026 · 2 min read

ServiceNow CoreAI has introduced AutoSynthData, a pipeline that uses a target model’s failures and a stronger teacher model’s successes to generate and validate training tasks for enterprise-agent environments.

In two EnterpriseOps-Gym domains, the researchers reported gains after supervised fine-tuning. Mean Pass@1 improved by 7.2 percentage points in Hybrid and rose from 18.77% to 27.18% in IT service management. These are ServiceNow’s controlled findings, not an independent replication, and they do not establish performance across the benchmark’s other six domains or in production systems.

The process starts by running the target and teacher models on diagnostic tasks. AutoSynthData distills the resulting capability gaps into sanitized specification cards that omit the original prompts, entities, trajectories and verifier details. A generator uses those cards to produce new tasks with different prompts, states and solution paths. As post-training changes the target model, later rounds can focus on failures that remain.

Each task combines a system specification, a user request and a verifier. In ServiceNow’s configuration, the pipeline favors tasks the target solves in no more than one of three trials and the stronger solver completes in at least two of three. It creates and validates core samples in a target phase, then expands them into variants in a multiply phase. Multiplied samples cannot seed further variants.

Candidates are tested inside the target environment. A positive gate replays a reference trajectory and checks the resulting state; a negative gate mutates expected outcomes to confirm that incorrect states fail. Rejected candidates can enter a critique-and-repair loop with a fixed retry limit before the relevant checks run again. At the batch level, the pipeline also looks for repeated task families, missing capability dimensions, systematic failures and redundant samples.

For the Hybrid experiment, ServiceNow said AutoSynthData generated 2,000 samples in about 18 hours, with Gemma-4-26B-A4B-it as the target and Qwen3.8-27B as the teacher. The researchers selected the epoch-five checkpoint and said verifier success increased from 63.01% to 68.55%, while the fine-tuned model closed 59% of the original Pass@1 gap to their reference model. They described the 7.2-point Pass@1 increase as a 35% relative gain.

For IT service management, the researchers paired the same Gemma target with DeepSeek-V4.1-Flash as the teacher. They said the pipeline produced 1,994 samples in 66 hours, attributing the longer run to the larger teacher and to conducting the experiment before later pipeline optimizations.

The experiments are limited to the Hybrid and ITSM sections of EnterpriseOps-Gym, a containerized, resettable benchmark with 1,150 expert-curated tasks across eight domains, 512 tools and 164 database tables. The ServiceNow repository says tasks operate against live MCP servers, while SQL verifiers score the final environment state rather than require one prescribed action sequence.

The published experiments use supervised fine-tuning. ServiceNow’s authors said the moving, difficulty-calibrated generation process could also support reinforcement learning and that they plan to test it beyond SFT. They did not report reinforcement-learning results.

More news