Runbook: DecisionModelCandidateTimeout¶
What it means¶
A candidate revision has stayed in Starting or Evaluating past the (longer)
candidate window: the model-ready gate or the evaluation is not completing.
How to confirm¶
kubectl describe dm <name> -n <ns> # candidateRevision, ModelReady / Evaluated conditions
kubectl get pods -n <ns> -l decisionmodel.io/name=<name>
kubectl logs -n <ns> <candidate-pod>
ModelReady condition reason: DigestMismatch, DeviceMismatch,
NotPinned, or a probe error.
Common causes¶
- Serving container healthy to Kubernetes but
/api/psnever shows the expected digest/device (silent CPU fallback on a CUDA request, or a store with the wrong model). See DecisionModelModelMismatch. - The runtime keeps dropping the pin (
NotPinned): warmup succeeds but the model is evicted before the gate flips. - Evaluation runs but the candidate Pod is slow/unhealthy, or the dataset is very large.
Fix¶
- Fix the image/device so
/api/psreports the pinned digest on the expected device. - For
NotPinned, check the runtime'skeep_alivehandling and node memory pressure. - The candidate rolls back to the stable revision on its progress deadline; the stable keeps serving throughout. Correct the cause and re-apply the spec change.