ServiceNow reports improved benchmark results after targeting synthetic tasks at an agent’s hardest enterprise workflows.
ServiceNow CoreAI describes AutoSynthData as a post-training system for enterprise AI agents. In a post published on Hugging Face, the company says it uses a target model’s failures and a stronger teacher’s successful behavior to decide what the model should learn next.
The system creates executable tasks made of three parts: a system specification, a user prompt, and a verifier. The verifier checks whether the agent reached a valid result, rather than requiring one exact sequence of actions.
AutoSynthData tests each candidate in the target environment. ServiceNow says candidates go through validation, execution, solver evaluation, and repair before acceptance. Batch reviews then check whether the training set covers varied skills instead of repeating the same examples.
That matters because enterprise agents often fail for environment-specific reasons, such as misusing tools or mishandling a workflow. The method aims to turn those failures into many new practice tasks while avoiding the original evaluation prompts.
What ServiceNow reported
In EnterpriseOps Gym’s Hybrid environment, ServiceNow used Gemma-4-26B-A4B-it as the target model and Qwen3.8-27B as the teacher. AutoSynthData produced 2,000 synthetic samples in about 18 hours.
ServiceNow reports a 7.2-percentage-point improvement in mean Pass@1, a first-attempt success measure. Verifier success rose from 63.01% to 68.55%.
In a separate ITSM experiment, the system generated 1,994 samples in 66 hours, using DeepSeek-V4.1-Flash as teacher. Mean Pass@1 rose from 18.77% to 27.18%.
These results come from ServiceNow’s controlled tests in two EnterpriseOps Gym environments. Applying the approach elsewhere would require an executable environment and reliable checks for whether the agent actually completed the work.
