Skip to content

Troubleshooting

Diagnose the control plane first, then image builds, then an individual dispatch.

Start with the Deployment, its migration init container, and the rendered settings:

kubectl get deploy pipelines-api-server -n <namespace>
kubectl logs deploy/pipelines-api-server -c db-migrate -n <namespace>
kubectl get configmap pipelines-api-env-config -n <namespace> -o yaml

Service (control plane)

Symptom Likely cause Resolution
pipelines-api-server pods CrashLoopBackOff on the db-migrate init container Alembic migration failed, or the database is unreachable. Inspect the db-migrate init-container logs for the failing revision; confirm global.postgresql.* and the ExternalSecret database password. See Migrations.
Pods Running but every dispatch validation fails with an unknown resource bundle The resource-bundle catalog did not load. Confirm the pipelines-api-resource-bundles ConfigMap is present and populated. See Resource bundles.
401 / 403 on public routes Requests are not arriving with gateway identity headers, or a token failed validation. Confirm traffic reaches the service through the gateway (not directly) and that the Hydra auth config (config.auth.*) is correct. See Networking & security.
Callbacks from task pods rejected with 401 Clock skew, an expired HMAC token, or a rotated signing key. Confirm PIPELINES_API_CALLBACK_SIGNING_KEY is present and consistent; a Deployment restart briefly rejects in-flight tokens. See Secrets.

Migrations

The db-migrate init container runs the database migrations before the app starts. If it fails, the Deployment never becomes ready. Read its logs (the kubectl logs command above) to identify the failing revision, resolve the database condition, and let the Deployment retry.

Resource bundles

At startup the service loads the resource-bundle catalog from the pipelines-api-resource-bundles ConfigMap the chart renders. If the catalog is missing or empty, dispatches that reference a bundle fail validation. Verify the ConfigMap is present and populated:

kubectl get configmap pipelines-api-resource-bundles -n <namespace> -o jsonpath='{.data}'

See Resource bundles for how the catalog is defined.

Image builds

Symptom Likely cause Resolution
Image build fails immediately with INVALID_IMAGE_REGISTRY global.imageRegistry is unset or points at Docker Hub, so the service has no registry to push built pipeline images to. Set global.imageRegistry to your private container registry. See Deploy and enable.
Build never starts / status stuck The Image Build Service is unreachable, or the build-context bucket is misconfigured. Confirm config.buildService.url and config.ibs.buildContextBucket. See Prerequisites.
Air-gapped build fails while resolving pipelines-api-runtime-env:<version>-py<X.Y> (NAME_UNKNOWN, not found, or manifest unknown in the build log) The runtime image for that Python version is not in the private registry. Mirrors that copy an image list are missing one of the four tags; a pull-through proxy registry does not allow the datarobotdev/pipelines-api-runtime-env repository. Mirror all four pipelines-api-runtime-env tags for the chart's runtimeImageVersion, or add the repository to the pull-through allow-list. The four images are listed in the pipelines-api chart's datarobot.com/images annotation and are delivered with the release image set.
Creating an image with pythonVersion other than 3.12 returns 422 with "preprovisioned runtime environment that provides Python 3.12" config.ibs.runtimeImage is set to a single image reference without the {pythonVersion} placeholder, which pins every build to that interpreter. Remove the override so the chart derives the per-version reference, or keep the {pythonVersion} placeholder in it.

Dispatches

A dispatch moves through PREPARING (prep Job traces the graph) → RUNNING (per-task Jobs execute) → COMPLETED or FAILED.

Symptom Likely cause Resolution
Dispatch stuck in PREPARING, then FAILED The prep Job could not start or timed out (bad image, missing input, or admission policy rejection). Inspect the prep Job pod in the dispatch namespace; check image availability and any admission-policy denials.
Tasks never leave the ready queue Jobs cannot be scheduled (ServiceAccount RBAC, namespace, or quota). Confirm the pipelines-api-service-account RBAC and that the dispatch namespace has capacity; check the task-runner ServiceAccount's cloud identity (IRSA / Workload Identity), which is configured at the platform level. See Deployed Kubernetes components.
Task pods run but callbacks never arrive; the reconciler later marks them terminal Task pods cannot reach the internal callback endpoint. Confirm config.standaloneDispatch.callbackUrlBase resolves from dispatch pods and the /api/v2/internal/pipelines/** route is reachable in-cluster. See Deployed Kubernetes components.
Artifacts not encrypted with the tenant key (AWS, CMEK enabled) Tenant CMEK fell back to bucket-default SSE-S3 on a KMS error. Expected fail-open behavior; check KMS reachability and the tenant key configuration. See Encryption at rest.

Where to find dispatch logs

Per-task stdout and stderr are written to object storage and are retrievable through the task's .../tasks/{task_id}/logs endpoint. The control-plane logs (structured JSON) carry the dispatch_id and request_id for correlation; the reconciler polls task Jobs about every 10 seconds to catch lost callbacks.