Operator metrics¶
The controller exports custom Prometheus metrics on controller-runtime's metrics
registry, served by the manager's metrics endpoint (--metrics-bind-address,
default :8443 HTTPS in the shipped config; protected by authn/authz). They are
registered automatically at startup — no extra flags are required.
All per-DecisionModel series carry only the namespace and name labels (never
the revision hash), plus a small fixed set of enum labels, so cardinality stays
bounded. When a DecisionModel is deleted, every series for it is removed, so a
churn of short-lived objects does not leak series.
Per-DecisionModel gauges and counters are updated only after a status write the API server accepted (the same rule the operator uses for Events). A conflicting or failed status write emits no metric mutation, so metrics never advertise a phase, replica count, or outcome that was not persisted.
Metrics¶
| Metric | Type | Labels | Description |
|---|---|---|---|
decisionmodel_phase |
gauge | namespace, name, phase |
1 for the DecisionModel's current phase and 0 for every other phase (one-hot), so a dashboard never shows a stale phase as active. phase is one of Pending, Resolving, Caching, Starting, Evaluating, Degraded, Promoting, Ready, Failed, RolledBack. |
decisionmodel_model_ready_replicas |
gauge | namespace, name |
Number of serving replicas that passed the model-ready gate (status.replicas.modelReady). |
decisionmodel_desired_replicas |
gauge | namespace, name |
Desired number of serving replicas (status.replicas.desired). |
decisionmodel_rollouts_total |
counter | namespace, name, result |
Rollout outcomes. result is promoted (a revision switch was promoted), rolled_back (a candidate failed but a stable revision kept serving), or failed (a candidate failed with no stable revision to fall back to). |
decisionmodel_phase_duration_seconds |
histogram | namespace, name, phase |
Time spent in a phase, observed when that phase is left. Buckets span 5s to 1h (5, 15, 30, 60, 120, 300, 600, 1200, 1800, 3600). |
decisionmodel_evaluation_accuracy |
gauge | namespace, name |
Accuracy of the last completed eval-gated evaluation (decimal in [0,1]). Only set once an evaluation has completed. |
decisionmodel_evaluation_ece |
gauge | namespace, name |
Expected Calibration Error of the last completed evaluation (decimal in [0,1]). |
decisionmodel_probe_results_total |
counter | namespace, name, result |
Model-readiness probe outcomes, counted on a gate transition (not re-counted for a steady-state Pod each requeue). result is ready, digest_mismatch, device_mismatch, not_pinned, or error. |
decisionmodel_registry_resolve_duration_seconds |
histogram | (none) | Duration of a model registry tag→digest resolution. No per-DecisionModel labels (bounded by design); default Prometheus buckets. |
Notes:
digest_mismatch/device_mismatchare the model-aware readiness signals: a Pod reportedContainersReadybut/api/psshowed the wrong digest or ran on the wrong device (e.g. a silent CPU fallback). A rising rate here means Pods look healthy to Kubernetes but are not actually serving the pinned model.not_pinnedmeans/api/psshowed the right model on the right device but unpinned (the runtime'skeep_alivewould evict it), so the Pod is held out of the Service until it is pinned again. Usually transient (the operator re-pins via warmup); a sustained rate means the runtime keeps dropping the pin.decisionmodel_registry_resolve_duration_secondsis only observed on an actual registry call; reconciles that reuse an already-recorded digest do not add a sample.
Example alerts¶
The expressions below are illustrative PromQL for Prometheus alerting rules; tune
thresholds and for windows to your environment.
Stuck in Starting for more than 15 minutes¶
A revision whose Pods never become model-ready (e.g. bad image, digest mismatch,
GPU never allocated) sits in Starting.
- alert: DecisionModelStuckStarting
expr: |
max by (namespace, name) (decisionmodel_phase{phase="Starting"}) == 1
for: 15m
labels:
severity: warning
annotations:
summary: "DecisionModel {{ $labels.namespace }}/{{ $labels.name }} stuck in Starting"
description: "Serving Pods have not become model-ready for 15m."
The same pattern catches a model wedged in Caching:
Elevated rollback / failure rate¶
Repeated rollbacks or failures indicate a bad candidate or a flaky eval/probe.
- alert: DecisionModelRolloutFailures
expr: |
sum by (namespace, name) (
rate(decisionmodel_rollouts_total{result=~"rolled_back|failed"}[15m])
) > 0
for: 15m
labels:
severity: warning
annotations:
summary: "DecisionModel {{ $labels.namespace }}/{{ $labels.name }} failing rollouts"
description: "One or more rollbacks/failures in the last 15m."
Share of rollouts that did not promote, over the last hour:
sum by (namespace, name) (increase(decisionmodel_rollouts_total{result=~"rolled_back|failed"}[1h]))
/
sum by (namespace, name) (increase(decisionmodel_rollouts_total[1h]))
Degraded¶
Degraded means the stable revision is serving on a subset (or none) of its
replicas, or the model store cannot be shared for the requested replica count.
- alert: DecisionModelDegraded
expr: |
max by (namespace, name) (decisionmodel_phase{phase="Degraded"}) == 1
for: 10m
labels:
severity: warning
annotations:
summary: "DecisionModel {{ $labels.namespace }}/{{ $labels.name }} is Degraded"
description: "Not all desired replicas are model-ready, or the model store is not shareable."
A tighter, phase-independent variant that fires whenever replicas fall short:
Silent model mismatch¶
Pods pass Kubernetes readiness but load the wrong digest/device.
sum by (namespace, name) (
rate(decisionmodel_probe_results_total{result=~"digest_mismatch|device_mismatch"}[15m])
) > 0
Poor calibration after an evaluation¶
Shipped alerts and runbooks¶
The chart can ship these as a PrometheusRule (requires the Prometheus Operator) and a
Grafana dashboard ConfigMap, both disabled by default:
prometheusRule:
enabled: true
labels: { release: kube-prometheus-stack } # match your Prometheus ruleSelector
grafanaDashboard:
enabled: true
metrics:
serviceMonitor:
enabled: true # scrape the metrics Service (HTTPS, bearer token)
labels: { release: kube-prometheus-stack } # match your Prometheus serviceMonitorSelector
The rules fire only on the metrics listed above plus controller-runtime built-ins
(controller_runtime_reconcile_errors_total, controller_runtime_reconcile_total).
Each alert links a runbook in docs/runbooks/:
| Alert | Signal | Runbook |
|---|---|---|
DecisionModelRolloutStuck |
decisionmodel_phase in a non-terminal phase |
RolloutStuck |
DecisionModelCandidateTimeout |
decisionmodel_phase Starting/Evaluating |
CandidateTimeout |
DecisionModelRolloutFailures |
decisionmodel_rollouts_total rolled_back/failed |
RolloutFailures |
DecisionModelEvaluationPoor |
decisionmodel_evaluation_accuracy / _ece |
EvaluationPoor |
DecisionModelDegraded |
decisionmodel_phase{phase="Degraded"} |
Degraded |
DecisionModelModelMismatch |
decisionmodel_probe_results_total digest/device mismatch |
ModelMismatch |
DecisionModelOperatorReconcileErrors |
controller_runtime_reconcile_errors_total |
OperatorReconcileErrors |
DecisionModelOperatorDown |
absent(controller_runtime_reconcile_total) |
OperatorDown |
Thresholds and for windows are illustrative — tune them (chart prometheusRule.windows
and the expressions) to your environment. There is no dedicated metric for a store-recovery
failure today; it surfaces via Degraded/Failed phase and the PrefetchFailed condition.