Pipelines overview¶
In Registry, Pipelines turns a Python file into a scheduled, monitored, multi-task workflow. Decorate functions with @dr.task, connect them inside a @dr.pipeline, then upload and run the pipeline. DataRobot runs each task in its own container, executes independent branches in parallel, and records status, logs, and results for the run.
Iterate in Draft, promote to a Locked version for reproducible runs, attach reusable inputs and images, and run on demand or on a schedule. View the directed acyclic graph (DAG) to inspect the pipeline.
What happens when a pipeline runs¶
When a pipeline runs, DataRobot:
- Parses the file and resolves the task-call graph into a DAG.
- Starts one isolated container per task, in dependency order, with parallel branches running concurrently.
- Passes each task's return value to its downstream consumers.
- Records status, logs, and results per task. The pipeline reaches Run completed when every task succeeds.
Tasks never share filesystems or memory. Data moves only through return values and parameters.
Pipelines concepts¶
The following table describes the building blocks of Pipelines.
| Concept | Description |
|---|---|
@dr.task |
The unit of work. A Python function that runs in its own isolated container. Tasks do not share filesystems or memory; data flows through return values. |
@dr.pipeline |
The function that calls tasks and defines the flow and dependencies. DataRobot infers concurrency and parallelism from this call graph. |
| Image | The container blueprint that carries third-party Python packages (for example, scikit-learn, pandas, or scipy) and is reused across runs. |
| Input | An immutable, auditable YAML payload of parameters injected into the @dr.pipeline entrypoint at run time. |
| Run | A single run of a pipeline with a defined input. Status, logs, and outputs are captured and persisted per task. |
| Schedule | A cron trigger that starts a run on a cadence and timezone. |
Draft and Locked pipelines¶
A pipeline, its inputs, and image bindings follow a draft-to-locked lifecycle. The following table describes each state.
| State | Description |
|---|---|
| Draft | A mutable pipeline for iteration. Source, inputs, and image binding can change. |
| Locked | An immutable, versioned snapshot. Locked pipelines are eligible for reproducible runs and recurring schedules. Promotion from Draft to Locked is one-way. |
Recurring schedules require a Locked pipeline and a locked input. For how to promote a pipeline and attach a cadence, see Lock and schedule.
Work with Pipelines in Registry¶
In Registry, click the Pipelines tile to open the list. The suggested path is:
- Add a pipeline: Provide Python source (
@dr.taskand@dr.pipeline), inputs, and a container image. - View and manage pipelines: Open the pipeline to inspect the DAG, configuration, and run history.
- Run and schedule pipelines: Start a run, lock the pipeline, and optionally attach a schedule.
- Manage pipeline images: Create, rebuild, and reuse images across pipelines.
A run cannot start until the selected image is ready. If an image was created while adding the pipeline, wait for the build to finish on the Images tab before running.
Next steps¶
After reviewing these concepts, add a pipeline or inspect one that already exists.
- Try the tutorial: Build a pipeline using included scripts to learn the tools of the Pipelines interface.
- Add a pipeline: Create a pipeline from Python source, inputs, and an image.
- View and manage pipelines: Inspect the workflow DAG, details, and runs.