OFICIAL Google Cloud Blog

From weeks to minutes: The new agentic era of data pipelines

What happened
Based on Google Cloud Blog · Aug 31, 2026

Google Cloud introduces the Data Agent Kit, an open-source toolkit that integrates orchestration pipelines directly into IDEs and CLIs, enabling data teams to build and manage workflows using natural language prompts instead of complex Python code.

From weeks to minutes: The new agentic era of data pipelines
Google Cloud Blog — Google
Key points
·
Data pipelines are the backbone of the modern enterprise, yet a barrier to entry exists for orchestrating them, making this critical capability unavailable to many data professionals.
·
Following its announcements at Google Cloud NEXT ’26, where we introduced the Orchestration Pipelines framework, the company is fundamentally changing this dynamic.
·
The Data Agent Kit seamlessly embeds the Orchestration Pipelines framework into your workflow in two distinct ways.
·
First, it provides a dedicated Data Engineering tab for comprehensive pipeline management.

Data pipelines are essential for modern enterprises, yet many data professionals struggle to orchestrate them due to technical barriers. Google Cloud’s new Data Agent Kit addresses this by providing a freely available, open-source collection of tools that integrate directly into popular development environments like VS Code and Claude Code. The kit embeds the Orchestration Pipelines framework into workflows through a dedicated Data Engineering tab and a specialized agentic skill for creating, deploying, and troubleshooting Apache Airflow DAGs using natural language, eliminating the need for extensive Python boilerplate.

The framework decouples high-level orchestration logic from compute execution, making advanced MLOps capabilities accessible to all data roles, from analysts to machine learning engineers. Users can set up the extension in under two minutes by installing it in their preferred IDE or CLI and authenticating with their Google Cloud account. The toolkit provides deep contextual knowledge of pipeline syntax, variable substitution, secret management, and automated incident diagnosis for Airflow runs, enabling faster and more intuitive pipeline development.

A practical example demonstrates the toolkit’s impact in the logistics and retail sector, where accurate delivery estimates are critical to customer satisfaction. By predicting transit times based on warehouse and customer locations, operations teams can proactively notify customers or adjust shipping tiers before service level agreements are breached. The demonstration uses the public BigQuery dataset *bigquery-public-data.thelook_ecommerce* to build an end-to-end MLOps architecture that handles training, daily batch inference, and model drift evaluation.

With the Data Agent Kit configured, users can bypass Python DAG authoring entirely by issuing natural language prompts in VS Code. The toolkit generated PySpark scripts, dbt configurations, and three declarative YAML pipelines within minutes, showcasing its ability to automate complex workflows. While responses may vary based on model versions and context, follow-up prompts can refine outputs, demonstrating a shift toward faster, more accessible data pipeline development.

Original source → Deals on Clipraptor.com →