Eval-gated rollout¶
When spec.rollout.evaluation is set, a new revision is not promoted just
because its Pods became model-ready — it must first pass an evaluation against a
golden dataset. This catches a model that is accurate on average but
mis-calibrated (over-confident), and a candidate that regressed against the
current stable revision.
Golden dataset (JSONL)¶
The dataset is JSON Lines: one JSON object per line, each a case with state,
questions (the same shapes you POST to /v1/systemone), and expected (the
correct answer per question id).
{"state": "I was charged twice, please refund the second charge.", "questions": {"dept": {"type": "choice", "criteria": {"billing": "payments and refunds", "technical": "bugs"}}}, "expected": {"dept": "billing"}}
{"state": "The app crashes on login.", "questions": {"dept": {"type": "choice", "criteria": {"billing": "payments and refunds", "technical": "bugs"}}}, "expected": {"dept": "technical"}}
expectedvalues match the answer type: a label string forchoice, a boolean fornoul, an integer level forscore.- Score questions. Ollaya answers a
scorequestion with an expected-value levelscore = Σ i·pᵢ(a float, usually between integer levels) plus a per-levelprobabilitiesmap and aconfidence(docs/api.md). A case counts as correct when|predicted − expected| ≤ tolerance(levels). The tolerance is, in order: a per-case"tolerance": {"<qid>": 0.5}map, elsespec.rollout.evaluation.scoreTolerance, else the default0.5(which means "rounds to the expected level"). A score answer contributes to ECE/Brier only when the runtime returns a usable probability distribution (keys0..n-1summing to ~1); otherwise it counts toward accuracy only. A score question is no longer silently scored as wrong (it was before this release).
{"state": "…", "questions": {"sev": {"type": "score"}}, "expected": {"sev": 3}, "tolerance": {"sev": 1}}
DatasetInvalid; phase
RolledBack if a stable revision is serving, else Failed). An invalid dataset
is a permanent failure: the revision is recorded in status.failedRevision
and not retried automatically. The operator does not watch the ConfigMap /
Secret holding the dataset, so after you fix it you must either change a
model-affecting spec field or bump the decisionmodel.io/retry annotation (see
Retrying a failed rollout) for the operator to pick
up the corrected dataset. (A dataset object that is simply missing is treated
as transient and retried on its own, in case it is created shortly after.)
- Only the first maxCases lines are used (default 500).
Store it in a ConfigMap (or Secret) and point datasetRef at the key:
Secret capability labels¶
A Secret the operator reads must carry the label for the capability it is used for — one label grants exactly one capability:
| Capability | Label | Allows |
|---|---|---|
| Engine API key | decisionmodel.io/api-key: "true" |
spec.auth.apiKeySecretRef |
| Eval dataset | decisionmodel.io/eval-dataset: "true" |
a datasetRef.secretRef dataset |
| Download token | decisionmodel.io/download-token: "true" |
spec.cache.downloadTokenSecretRef |
A ConfigMap dataset needs no label. A dataset Secret must carry
decisionmodel.io/eval-dataset: "true".
Transition (one release). A dataset Secret labelled only with the old
decisionmodel.io/api-key: "true"is still accepted, but the operator emits aDeprecatedSecretLabelWarning Event naming the label to add. Adddecisionmodel.io/eval-dataset: "true"— it will become required. A dataset Secret with neither label is held (not failed): phaseEvaluating,Evaluated=FalsereasonSecretNotAllowed, until the label is added or the evaluation timeout expires.
apiVersion: v1
kind: ConfigMap
metadata:
name: support-router-eval
data:
cases.jsonl: |
{"state": "I was charged twice, please refund.", "questions": {"dept": {"type": "choice", "criteria": {"billing": "payments and refunds", "technical": "bugs"}}}, "expected": {"dept": "billing"}}
{"state": "The app crashes on login.", "questions": {"dept": {"type": "choice", "criteria": {"billing": "payments and refunds", "technical": "bugs"}}}, "expected": {"dept": "technical"}}
Configure the gate¶
apiVersion: decisionmodel.io/v1alpha1
kind: DecisionModel
metadata:
name: support-router
spec:
model: laya:en
device: cpu
rollout:
evaluation:
datasetRef:
configMapRef:
name: support-router-eval
key: cases.jsonl
minAccuracy: "0.90" # candidate must reach at least this accuracy
maxAccuracyDrop: "0.02" # and not drop more than this vs the stable baseline (optional)
maxECE: "0.10" # absolute calibration gate (optional)
maxECEIncrease: "0.05" # calibration must not worsen this much vs baseline (optional)
maxCases: 500 # cap on dataset lines used (optional, default 500)
scoreTolerance: "0.5" # correctness band for score questions, in levels (optional, default 0.5)
Gate fields (all decimal strings in [0,1]):
| Field | Meaning | If omitted |
|---|---|---|
minAccuracy |
Minimum candidate accuracy to promote. | required |
maxAccuracyDrop |
Max accuracy drop vs the stable baseline. | no drop constraint |
maxECE |
Max absolute Expected Calibration Error. | absolute ECE gate off |
maxECEIncrease |
Max ECE increase vs the stable baseline. | relative ECE gate off |
maxCases |
First N dataset lines scored. | 500 |
scoreTolerance |
Correctness band for score questions, in levels (not in [0,1]). |
0.5 |
Baseline is the current stable revision scored on the same dataset; the
*Drop / *Increase gates compare the candidate against it. The baseline is run
only when a relative gate is configured (maxAccuracyDrop or
maxECEIncrease); with only absolute gates (minAccuracy, maxECE) the baseline
is not run at all, so it never delays promotion or causes an EvaluationTimeout.
On a first rollout (no stable revision) there is no baseline either, so only the
absolute gates apply.
Known limitation — one API key per DecisionModel. The operator uses a single
spec.auth.apiKeySecretReffor every revision, including the baseline run against the stable revision. If you changespec.model(a new candidate) and the API-key value in the same edit, the stable revision's Pods still run with the old key, so the baseline scoring request gets401and the baseline run fails. With a relative gate configured (maxAccuracyDrop/maxECEIncrease) that missing baseline rolls the candidate back. Change the model and the API key in two separate steps (rotate the key first, let it roll — see Rotating the API key — then change the model), or temporarily drop the relative gates for the key-rotation rollout. Per-revision keys are planned for a later version (ARCHITECTURE §9).
The API-key Secret must contain the referenced key: apiKeySecretRef.optional is
not supported (rejected by CEL), and a labelled Secret missing the key is reported
as Degraded (reason APIKeyInvalid) with no workloads changed, so the prober and
the serving Pod never read different keys.
ECE (expected calibration error) and the Brier score come from the calibration metrics: they measure whether the model's stated confidence matches how often it is right. Accuracy alone does not catch an over-confident model; the ECE gates do.
Reading the result¶
The evaluation outcome is recorded in status.evaluation:
{
"revision": "a1b2c3d4e5",
"accuracy": "0.9400",
"ece": "0.0700",
"brier": "0.0500",
"baselineAccuracy": "0.9300",
"baselineEce": "0.0600",
"cases": 500,
"failedCases": 30,
"minAccuracy": "0.90",
"maxAccuracyDrop": "0.02",
"maxEce": "0.10",
"maxEceIncrease": "0.02",
"result": "Passed",
"reason": "",
"datasetDigest": "9f2c…",
"approvalId": "a1b2c3d4e5f6",
"completedAt": "2026-09-28T00:00:00Z"
}
result is Passed or Failed; on a failure reason carries the gate message
(e.g. accuracy 0.8700 < minAccuracy 0.9200). The minAccuracy / maxAccuracyDrop
/ maxEce / maxEceIncrease fields echo the gate that was applied, so the
recorded result is self-describing.
- Pass -> the candidate is promoted (phase goes
Evaluating->Ready;status.lastPromotionTimeis stamped and thePromotedcondition isTrue/Promoted). - Fail -> the candidate is rejected and the stable revision keeps serving. The
phase ends
RolledBack(orFailedwhen there was no stable revision), theReadycondition staysTruewith reasonCandidateRejected, and theDegradedcondition isTruewith reasonEvaluationFailed. The rejected revision is recorded instatus.failedRevisionwith itsreason,messageandfailedAt. AWarningevent is emitted too.
Other evaluation reasons: DatasetInvalid (bad JSONL / missing dataset),
EvaluationTimeout (scoring exceeded the deadline), EvaluationUnsupported
(the engine has no Decide capability). See the condition-reason table below and
the phase table in quickstart.md.
Missing or unreadable dataset¶
A dataset that is not yet usable does not fail the revision immediately — the
candidate holds in Evaluating (no traffic moves, its Deployment is kept) and
resumes on its own once the dataset becomes readable, with no spec change or
retry annotation needed:
| Situation | Evaluated reason while held |
Resolves when |
|---|---|---|
| dataset ConfigMap/Secret does not exist | DatasetNotFound |
the object is created |
| the referenced key is absent | DatasetKeyNotFound |
the key is added |
dataset Secret missing the decisionmodel.io/api-key: "true" label |
SecretNotAllowed |
the label is added |
A Warning event is emitted once per cause while held. A hold is bounded by the
Evaluating timeout (spec.rollout.timeouts.evaluating, default 10m): if the
dataset is still unusable when it expires, the candidate is rolled back (the
stable revision keeps serving). A transient API error on the read (timeout, 5xx,
throttling, RBAC) is retried with workqueue backoff and never records a failed
revision; it is bounded by the same timeout.
Only a genuinely invalid dataset (malformed JSONL) is a permanent
DatasetInvalid failure. Fixing the content in place is not auto-detected (the
dataset is not watched); set the decisionmodel.io/retry annotation to re-run.
Watching a rollout from the CLI¶
kubectl get dm shows the stable (Active) and in-flight (Candidate) models,
the last evaluation Accuracy, the Phase, and the Ready condition Reason:
$ kubectl get dm
NAME ACTIVE CANDIDATE ACCURACY PHASE REASON AGE
support-router laya:en kev:en 0.8700 RolledBack CandidateRejected 6d
billing-router laya:en 0.9400 Ready Promoted 20d
-o wide adds the active digest (short), device, model-ready replicas and the
last promotion time.
Conditions and reasons¶
The condition types are stable; the reasons tell the story and are part of the API (safe to alert on):
| Type | Reason | Meaning |
|---|---|---|
Resolved |
Resolved |
tag resolved to a digest |
Resolved |
ResolveFailed / ModelNotFound |
registry resolution failed |
Cached |
Cached |
weights present in the store |
Cached |
Caching / PrefetchFailed / CacheTimeout |
prefetch in progress or failed |
ModelReady |
ModelReady |
all serving Pods report the expected digest on the expected device |
ModelReady |
DigestMismatch / DeviceMismatch / NotPinned / ProbeError |
the readiness gate failed (e.g. silent CPU fallback) |
Evaluated |
EvaluationRunning |
golden-dataset scoring is in progress |
Evaluated |
EvaluationPassed |
the candidate met every configured gate |
Evaluated |
EvaluationSkipped |
no evaluation configured; promotion follows readiness |
Evaluated |
BaselineUnavailable |
a relative gate needs a stable baseline that is not reachable yet |
Evaluated |
DatasetNotFound / DatasetKeyNotFound / SecretNotAllowed |
the golden dataset is not yet readable; the candidate holds (see "Missing or unreadable dataset") |
Evaluated |
DatasetChanged |
the dataset content changed while awaiting promotion; re-evaluating |
Promoted |
PromotionPending |
a Manual-policy candidate passed its gate and waits for approval |
Promoted |
Promoted |
the candidate was promoted and is serving |
Stabilizing |
Stabilizing |
the new stable is in its post-promotion stabilization window (previous revision kept) |
Stabilizing |
StableRolling |
a quorum shortfall is ignored because the stable Deployment is mid-rollout (replicas scale-up, key rotation, scheduling change, restart) |
Stabilizing |
PostPromotionUnhealthy |
the new stable looks unhealthy in the window (debouncing before rollback) |
Ready |
Ready |
the stable revision is serving and healthy |
Ready |
CandidateRejected |
a candidate was rejected (or rolled back after promotion); the surviving revision keeps serving |
Degraded |
EvaluationFailed |
the candidate failed a gate (message names it) |
Degraded |
PostPromotionUnhealthy |
the new stable was rolled back to the previous revision during the stabilization window |
Degraded |
CacheNotShareable / StoreTerminating / StoreLost / StorePrefetchFailed |
a store issue that does not stop serving |
Degraded |
SecretNotAllowed / DownloadTokenInvalid |
a referenced Secret is not usable |
Events¶
Events are human-readable and name the model and short digest instead of a bare
revision hash, so kubectl describe dm reads like a changelog:
Normal EvaluationStarted evaluating revision a1b2c3d4e5 against 500 cases
Warning EvaluationFailed candidate kev:en@sha256:1a2b3c4d failed evaluation: accuracy 0.8700 < minAccuracy 0.9200; laya:en@sha256:c305a927 keeps serving
Warning RolledBack candidate kev:en@sha256:1a2b3c4d rejected (EvaluationFailed): ...; laya:en@sha256:c305a927 keeps serving
Normal Promoted promoted laya:en@sha256:c305a927 (accuracy 0.9400, baseline 0.9300)
The event reason (the first column after the type) is stable and safe to key alerts on; the message wording may change.
Promotion policy¶
spec.rollout.promotion selects how a model-ready candidate is promoted:
| Value | Behaviour |
|---|---|
Automatic |
promote as soon as the candidate is model-ready (and, when evaluation is set, has passed the gate) |
EvaluationGated |
requires rollout.evaluation; the gate decides promotion (rejected by CEL without evaluation) |
Manual |
hold the passed candidate in AwaitingPromotion until a human approves it |
When promotion is unset the effective policy is EvaluationGated if
rollout.evaluation is set, else Automatic.
The boolean
spec.rollout.manualPromotion: trueis deprecated but still works as an alias forpromotion: Manual. SettingpromotionandmanualPromotion: trueto disagreeing values is rejected by CEL.
Manual promotion (hold for approval)¶
Set spec.rollout.promotion: Manual to hold a candidate that passed its gate
until a human approves it, instead of promoting automatically. It works with or
without spec.rollout.evaluation: the candidate is built, becomes model-ready,
passes the gate (if configured), and then waits.
spec:
rollout:
promotion: Manual
evaluation: # optional; Manual works without it too
datasetRef: { configMapRef: { name: support-router-eval, key: cases.jsonl } }
minAccuracy: "0.90"
While waiting:
- The phase is
AwaitingPromotionand the conditionPromotedisFalsewith reasonPromotionPending; its message names the annotation and the candidate hash. The stable revision keeps serving on the Service the whole time. - There is no timeout — a parked candidate waits indefinitely.
promotionis not part of the revision hash, so changing it does not create a new revision; switching it toAutomaticreleases a parked candidate to promote.
Approve the candidate by annotating the DecisionModel with its approvalID
(shown in status.evaluation.approvalId, in the AwaitingPromotion Event, and in
the Promoted condition message):
id=$(kubectl get dm support-router -o jsonpath='{.status.evaluation.approvalId}')
kubectl annotate dm support-router decisionmodel.io/promote="$id" --overwrite
- The approvalID is
sha256(revisionHash + policyHash + datasetDigest)[:12], so it identifies this revision with this evaluation result. It changes whenever the revision, the gate policy, or the dataset content changes — an approval cannot carry over to a re-evaluation. Withoutrollout.evaluation(a manual-only hold) it is derived from the revision hash alone. - A value that is neither the current approvalID nor (without evaluation) the bare
revision hash is a stale approval: it is ignored with a
StaleApprovalWarning that names the current approvalID. - With
rollout.evaluationconfigured, only the approvalID promotes. A barestatus.candidateRevision.hashis not honoured — it would otherwise carry an approval given for an old result across a re-evaluation of the same revision — and is ignored with aDeprecatedApprovalWarning naming the current approvalID. - Without evaluation (a manual-only hold) the bare revision hash still works for
one release and emits a
DeprecatedApprovalWarning; prefer the approvalID. - On approval the controller promotes the candidate (
Promoted=True/Promoted), then removes the annotation once the promotion is durable. - Anyone who can
updatethe DecisionModel can approve — it is the same trust level as editing the spec, not a separate permission.
Caveats:
- The first revision of a DecisionModel (no stable yet) is not held — it promotes without approval, because there is nothing serving to protect.
- While a candidate is parked, changing the evaluation policy — any
rollout.evaluationthreshold (minAccuracy,maxAccuracyDrop,maxECE,maxECEIncrease,scoreTolerance), thedatasetRef, ormaxCases— or editing the dataset content in place (same ConfigMap/Secret) makes the recorded result stale, so the operator leavesAwaitingPromotionand re-evaluates rather than promoting on the old result. A dataset content change is detected by re-reading its bytes on each reconcile while parked (reasonDatasetChanged, Warning Event); there is no watch, so the detection delay is at most one regate interval (~60s). Re-evaluation rotates the approvalID, so an approval set for the old identity does not carry over — re-approve with the new approvalID. - Otherwise, to force a fresh candidate, change a model-affecting field (model,
digest, device, resources, image, scheduling) or switch
promotiontoAutomaticand back toManual. - A
promoteannotation whose hash never became stable (e.g. you approved, then reverted the spec) stays on the object and would approve a later revision that happens to resolve to that same hash. Remove a stale approval by hand:kubectl annotate dm support-router decisionmodel.io/promote-.
How Argo CD and Flux surface AwaitingPromotion is described in
gitops.md (Argo CD reports it Suspended; a Flux healthCheckExprs
check can treat it as in-progress).
Post-promotion stabilization¶
A promotion is not the end of the risk: a new stable can look model-ready at the
instant of promotion and fall over a minute later. spec.rollout.stabilization
(a Go duration, default 5m, 0 disables it) keeps the previous revision's
Deployment running — scaled to its replicas but out of the Service — for a window
after promotion, so the operator can switch traffic back instantly if the new
stable turns out unhealthy.
During the window (condition Stabilizing=True), the operator rolls back to the
previous revision if the new stable:
- drops below the model-ready quorum
max(1, ceil(replicas/2))for longer than a short debounce (30s) — a brief dip during a rolling restart does not trip it; or - reports a readiness-gate
DigestMismatch/DeviceMismatch(wrong model loaded), or a container inCrashLoopBackOff— these roll back immediately, no debounce.
A quorum shortfall caused by an intentional in-place rollout of the stable
(replicas scale-up, API-key rotation, scheduling change, or kubectl rollout
restart) is not treated as a failure: while the stable Deployment is
mid-rollout — generation not yet observed, or the Progressing condition is
True with a reason other than NewReplicaSetAvailable (e.g. ReplicaSetUpdated)
— the condition is Stabilizing=True with reason StableRolling and the
debounce does not advance. A completed rollout settles to Progressing=True,
NewReplicaSetAvailable and stays there even if its Pods later go unready, so a
post-rollout health failure is not masked and still rolls back after the
debounce. Immediate failures (gate mismatch, crash loop) always roll back.
Protection stays bounded: a rollout that exceeds its Deployment
progressDeadlineSeconds (Progressing=False, ProgressDeadlineExceeded) is counted
as a failure and rolled back. If the stabilization window elapses while a rollout
is still in progress, the window is extended (the previous revision is kept) until
the rollout settles healthy or trips the progress deadline.
On rollback the Service switches back to the previous revision (its Pods are still
running, so there is no cold start), the new revision is recorded in
status.failedRevision with reason PostPromotionUnhealthy, the phase is
RolledBack, a Warning Event RolledBackAfterPromotion is emitted, and
decisionmodel_rollouts_total{result="rolled_back_after_promotion"} is
incremented. The rolled-back revision is not retried automatically (a spec change
or the decisionmodel.io/retry annotation is required).
If the window passes healthy, the operator emits Stabilized, drops the
Stabilizing condition, and garbage-collects the previous revision as usual.
Notes:
- The check runs on each reconcile from persisted status
(
status.previousRevision) and live cluster state, so it is correct across an operator restart and uses no in-memory timers. - A spec change during the window starts a new candidate as usual and ends the window (the previous revision is then collected normally).
stabilization: 0restores the previous behaviour: the previous revision is removed after a short endpoint-gap grace, with no rollback window.- The previous revision's store PVC lives through the window (so the footprint is ~2× the model during it), the same as during a rollout.
Rollout timeouts¶
Each rollout phase has a progress timeout; if the phase does not complete in
time the rollout fails (and rolls back to the stable revision if there is one).
Override them under spec.rollout.timeouts for large models that need longer on
a cold node:
spec:
rollout:
timeouts:
caching: "2h" # Caching phase (and the prefetch Job's activeDeadlineSeconds)
starting: "30m" # Starting phase (and the per-Pod model warmup bound)
evaluating: "20m" # Evaluating phase (and each golden-dataset run)
| Field | Bounds | Default |
|---|---|---|
caching |
The Caching phase. Also set as the prefetch Job's activeDeadlineSeconds. |
30m |
starting |
The Starting phase. Also bounds the background model warmup on each Pod. | 10m |
evaluating |
The Evaluating phase and each golden-dataset evaluation run. | 10m |
- Values are Go durations (
90m,2h,1h30m) and must be between 1m and 24h. timeoutsis not part of the revision hash, so changing it does not create a new revision; the new phase timeouts apply on the next reconcile.- An existing prefetch Job keeps the
activeDeadlineSecondsit was created with — a changedcachingdoes not retroactively extend a pull that is already running. To apply a new caching timeout to an in-flight pull, retry with thedecisionmodel.io/retryannotation (see below) so the Job is recreated.
Suggested starting points for large GGUF models (e.g. the ~19 GB clef:flash
/ 9B-class models) — estimates, tune to your registry bandwidth and node
disk/CPU:
spec:
rollout:
timeouts:
caching: "2h" # ~19 GB pull + sha256 verify over a slow link
starting: "30m" # loading tens of GB into memory on a cold node
evaluating: "30m" # larger model = slower per-case decisions
Retrying a failed rollout¶
A failed revision is not retried automatically. Change the spec (a new revision)
or bump the decisionmodel.io/retry annotation to clear the failure once — see
quickstart.md.