OFICIAL Hugging Face Blog

AutoSynthData: Generating Training Data for Enterprise Agents

What happened
Based on Hugging Face Blog · Oct 02, 2026

ServiceNow CoreAI introduces AutoSynthData, a framework that converts agent failures into targeted training tasks to improve enterprise AI performance in specific environments.

AutoSynthData: Generating Training Data for Enterprise Agents
Hugging Face Blog — Hugging Face
Key points
·
AutoSynthData converts agent failures into targeted training tasks using a stronger teacher model to identify solvable capability gaps
·
Generated tasks must be feasible, realistic, and difficult to provide useful training signal for the target model
·
The EnterpriseOps Gym experiment validates tasks through execution, solver evaluation, and repair before inclusion in training datasets

AutoSynthData addresses a core challenge in enterprise AI: turning observed weaknesses of an agent into effective training data. The framework evaluates a target model’s failures in a defined environment, then uses a stronger teacher model to identify solvable tasks that expose similar capability gaps. By generating new tasks that are feasible, realistic, and appropriately challenging, AutoSynthData creates a curriculum that evolves as the model improves, focusing training on unresolved weaknesses.

The system specification defines the operational constraints of the agent, including system instructions, policies, and environment state, while the user prompt describes the task the agent must complete. A generated task must meet three criteria: feasibility (a valid solution exists in the environment), realism (the task resembles plausible user requests), and difficulty (the task targets a current weakness of the model). The verifier ensures solutions meet the task requirements without encoding a single reference trajectory.

In the EnterpriseOps Gym experiment, AutoSynthData first identifies capability gaps by running both the target model and a stronger teacher on evaluation tasks. These findings are distilled into sanitized capability specification cards, which guide the generation of new tasks with varied prompts, states, and solution paths. The generator creates multiple independent tasks in parallel, each validated through execution, solver evaluation, and repair before acceptance.

The framework operates in two phases: core sample generation followed by expansion into novel variants. Core samples are created from capability specifications, with each candidate undergoing validation, execution checks, and repair. The result is a vetted dataset tailored to the target model’s needs, enabling supervised fine-tuning that improves performance in the specific enterprise environment.

Original source → Deals on Clipraptor.com →