Skip to content

Pipelines overview

In Registry, Pipelines turns a Python file into a scheduled, monitored, multi-task workflow. Decorate functions with @dr.task, connect them inside a @dr.pipeline, then upload and run the pipeline. DataRobot runs each task in its own container, executes independent branches in parallel, and records status, logs, and results for the run.

Iterate in Draft, promote to a Locked version for reproducible runs, attach reusable inputs and images, and run on demand or on a schedule. View the directed acyclic graph (DAG) to inspect the pipeline.

What happens when a pipeline runs

When a pipeline runs, DataRobot:

  1. Parses the file and resolves the task-call graph into a DAG.
  2. Starts one isolated container per task, in dependency order, with parallel branches running concurrently.
  3. Passes each task's return value to its downstream consumers.
  4. Records status, logs, and results per task. The pipeline reaches Run completed when every task succeeds.

Tasks never share filesystems or memory. Data moves only through return values and parameters.

Pipelines concepts

The following table describes the building blocks of Pipelines.

Concept Description
@dr.task The unit of work. A Python function that runs in its own isolated container. Tasks do not share filesystems or memory; data flows through return values.
@dr.pipeline The function that calls tasks and defines the flow and dependencies. DataRobot infers concurrency and parallelism from this call graph.
Image The container blueprint that carries third-party Python packages (for example, scikit-learn, pandas, or scipy) and is reused across runs.
Input An immutable, auditable YAML payload of parameters injected into the @dr.pipeline entrypoint at run time.
Run A single run of a pipeline with a defined input. Status, logs, and outputs are captured and persisted per task.
Schedule A cron trigger that starts a run on a cadence and timezone.

Draft and Locked pipelines

A pipeline, its inputs, and image bindings follow a draft-to-locked lifecycle. The following table describes each state.

State Description
Draft A mutable pipeline for iteration. Source, inputs, and image binding can change.
Locked An immutable, versioned snapshot. Locked pipelines are eligible for reproducible runs and recurring schedules. Promotion from Draft to Locked is one-way.

Recurring schedules require a Locked pipeline and a locked input. For how to promote a pipeline and attach a cadence, see Lock and schedule.

Work with Pipelines in Registry

In Registry, click the Pipelines tile to open the list. The suggested path is:

  1. Add a pipeline: Provide Python source (@dr.task and @dr.pipeline), inputs, and a container image.
  2. View and manage pipelines: Open the pipeline to inspect the DAG, configuration, and run history.
  3. Run and schedule pipelines: Start a run, lock the pipeline, and optionally attach a schedule.
  4. Manage pipeline images: Create, rebuild, and reuse images across pipelines.

A run cannot start until the selected image is ready. If an image was created while adding the pipeline, wait for the build to finish on the Images tab before running.

Next steps

After reviewing these concepts, add a pipeline or inspect one that already exists.