Runbook: DecisionModelDegraded¶
What it means¶
The DecisionModel is in Degraded: the stable revision is serving on a subset (or none)
of its replicas, or the model store cannot be shared for the requested replica count.
How to confirm¶
kubectl describe dm <name> -n <ns> # Degraded condition reason
kubectl get pods -n <ns> -l decisionmodel.io/name=<name> -o wide
Common causes¶
CacheNotShareable:replicas > 1across nodes butspec.cache.accessModesisReadWriteOnce— a candidate scheduled on another node cannot attach the store.- Nodes lost / evicted Pods; not enough schedulable capacity (CPU, memory, GPU).
- Model-ready gate failing on some replicas (digest/device mismatch) — see DecisionModelModelMismatch.
Fix¶
- For
CacheNotShareable, setspec.cache.accessModes: [ReadWriteMany]with an RWX StorageClass, or keepreplicas: 1/ pin to one node. - Restore capacity or fix scheduling constraints (
spec.scheduling). - Fix any per-replica model mismatch. The stable revision keeps serving on its ready
replicas while Degraded; the goal is to get
modelReady == desired.