GitOps: Argo CD and Flux¶
Keep your DecisionModel manifests in Git and let Argo CD or Flux apply them.
Both tools need to know when a DecisionModel is healthy, progressing or
degraded — otherwise a blue-green rollout that is rolling back can show up as
"Healthy" and hide a regression.
A DecisionModel reports its state through status.phase, the standard
status.conditions (Ready, Degraded, …) and status.observedGeneration.
The relevant phases are Pending, Resolving, Caching, Starting,
Evaluating, Promoting, AwaitingPromotion, Ready, Degraded,
RolledBack and Failed (see quickstart.md for what each
means).
What goes in Git¶
Put only the DecisionModel (and its ConfigMaps/Secrets, e.g. a golden
dataset or an API key) in Git. The operator creates and owns the Deployments,
Jobs, Service and PVCs for each revision — do not commit those; Argo CD/Flux
would fight the controller over them. Point the Service, replica counts, cache
and resources through the DecisionModel spec, not through hand-written child
objects.
You do not need ignoreDifferences for the DecisionModel itself: status
is a Kubernetes subresource on this CRD (subresources.status: {}), so neither
tool treats status as part of the desired manifest and there is nothing to
ignore there.
Argo CD — custom health check¶
Argo CD does not know this CRD, so without a health check it reports a
DecisionModel as Healthy as soon as it exists. Add a Lua health check that maps
our phases and conditions. Put this under resource.customizations in the
argocd-cm ConfigMap:
# argocd-cm ConfigMap
data:
resource.customizations.health.decisionmodel.io_DecisionModel: |
local hs = {}
hs.status = "Progressing"
hs.message = "Waiting for DecisionModel status"
if obj.status == nil then
return hs
end
-- Spec changed but not yet observed by the controller: still progressing.
if obj.metadata ~= nil and obj.metadata.generation ~= nil then
if obj.status.observedGeneration == nil or obj.status.observedGeneration < obj.metadata.generation then
hs.status = "Progressing"
hs.message = "Waiting for the operator to observe generation " .. tostring(obj.metadata.generation)
return hs
end
end
-- Prefer a condition message when present (richer than the phase alone).
-- Use the Ready message normally; use the Degraded message only when the
-- Degraded condition is actually True (a stale Degraded=False must not show).
local readyMsg = nil
local degradedMsg = nil
if obj.status.conditions ~= nil then
for _, c in ipairs(obj.status.conditions) do
if c.type == "Ready" and c.message ~= nil and c.message ~= "" then
readyMsg = c.message
end
if c.type == "Degraded" and c.status == "True" and c.message ~= nil and c.message ~= "" then
degradedMsg = c.message
end
end
end
local phase = obj.status.phase
if phase == "Ready" then
hs.status = "Healthy"
hs.message = readyMsg or "Serving"
return hs
elseif phase == "AwaitingPromotion" then
-- Candidate passed its gate and is waiting for a human to approve it
-- (spec.rollout.manualPromotion); the stable keeps serving. Surface it as
-- Suspended so a sync does not sit in Progressing forever waiting on a
-- human, and Argo CD shows it as "paused for input".
hs.status = "Suspended"
hs.message = "Awaiting manual promotion (set the decisionmodel.io/promote annotation)"
return hs
elseif phase == "Pending" or phase == "Resolving" or phase == "Caching"
or phase == "Starting" or phase == "Evaluating" or phase == "Promoting" then
hs.status = "Progressing"
hs.message = "Rolling out: " .. phase
return hs
elseif phase == "Degraded" or phase == "RolledBack" or phase == "Failed" then
hs.status = "Degraded"
hs.message = degradedMsg or readyMsg or phase
return hs
end
hs.status = "Progressing"
hs.message = "Unknown phase: " .. tostring(phase)
return hs
Behaviour:
| DecisionModel state | Argo CD health |
|---|---|
Ready |
Healthy |
Pending / Resolving / Caching / Starting / Evaluating / Promoting |
Progressing |
AwaitingPromotion (candidate passed its gate, waiting for human approval) |
Suspended |
Degraded (partial or zero readiness) |
Degraded (with the Degraded condition message) |
RolledBack (candidate failed, stable still serving) |
Degraded |
Failed |
Degraded |
observedGeneration < generation (spec not yet seen) |
Progressing |
AwaitingPromotion and its Promoted condition (False/AwaitingApproval while
a candidate waits, True after promotion) exist only when
spec.rollout.manualPromotion is set; otherwise the phase never appears.
Unlike kstatus (below), the Argo CD script reads status.phase, so it reports
RolledBack and Failed as Degraded even though the stable's Ready condition
is still True.
This script was tested with gopher-lua (the Lua runtime Argo CD embeds) against
sample status JSON for every phase above, including the two Degraded sub-cases
and the stale-generation case — all mapped as in the table.
Flux — kstatus and healthCheckExprs¶
Flux's default health evaluation uses kstatus (fluxcd/cli-utils, a fork
of kubernetes-sigs/cli-utils),
which has no rule specific to this CRD and applies its generic rules:
- If
metadata.generation != status.observedGeneration, the resource is InProgress. The operator always setsobservedGeneration, so this works. - If the standard
Reconciling/Stalledconditions are present, kstatus uses them. The operator sets neither. - Otherwise kstatus falls back to the
Readycondition:Ready=True→ Current (healthy);Ready=False→ InProgress.
Default kstatus misreports this CRD — do not rely on it alone. The problem is
rule 3 together with how the controller keeps the stable serving during trouble
(verified in internal/controller/decisionmodel_controller.go):
RolledBackreportsCurrent(healthy). On a rollback the controller runsapplyStableReadinessand setsReady=Truewhen the stable revision is fully ready, then sets the phase toRolledBack(the failed candidate was rolled back; the stable keeps serving). TheReadycondition isTrue, so kstatus reports Current — a failed rollout looks healthy. ADegraded=Truecondition with the real reason is also set, but kstatus does not look at it.- A candidate rolling out over a healthy stable reports
Currenttoo. While a new candidate isStarting/Evaluating/AwaitingPromotion, the controller updatesModelReadybut does not touchReady, soReadykeeps the stable's lastTrue. kstatus seesReady=True→ Current, even though a rollout is in flight. (A first rollout, with no prior stable, has noReady=Trueyet, so that case does show InProgress.) - The partial-
Degradedcase reportsCurrent. When some but not all replicas are model-ready, the phase isDegradedbutReadyis stillTrue(the model answers on the ready replicas), so kstatus reports Current.
So with default kstatus, a Flux wait returns "ready" during rollbacks, in-flight
rollouts over a healthy stable, and partial degradation. Use a custom check instead.
Use healthCheckExprs (CEL on status.phase)¶
Flux v2.5+ (Feb 2025) supports custom health checks via
spec.healthCheckExprs (Flux CEL cheatsheet).
Key the check on status.phase so it reflects rollouts rather than the Ready
condition. The input to each expression is the resource object; a resource is
InProgress while neither current nor failed matches:
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: decision-models
namespace: flux-system
spec:
interval: 10m
path: ./clusters/prod/decision-models
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 10m
wait: true
healthChecks:
- apiVersion: decisionmodel.io/v1alpha1
kind: DecisionModel
name: support-router
namespace: default
healthCheckExprs:
- apiVersion: decisionmodel.io/v1alpha1
kind: DecisionModel
# Healthy only when fully Ready and the controller has seen the latest spec.
current: >-
has(status.phase) && status.phase == 'Ready' &&
status.observedGeneration == metadata.generation
# Fail fast on a failed/rolled-back/degraded rollout instead of waiting for the timeout.
failed: >-
has(status.phase) &&
(status.phase == 'Failed' || status.phase == 'RolledBack' || status.phase == 'Degraded')
With this, Pending / Resolving / Caching / Starting / Evaluating /
Promoting / AwaitingPromotion (none matching current or failed) are
InProgress, Ready is Current, and Failed / RolledBack / Degraded
are Failed so wait: true fails fast. Treat AwaitingPromotion as InProgress
here; if you use manual promotion and do not want Flux to block on it, add
status.phase == 'AwaitingPromotion' to the current expression instead.
Not tested on a live cluster in this change: the
healthCheckExprsschema, the v2.5 availability and the CEL field names (current/failed, thehas()macro) are from the Flux docs linked above; the CEL expressions were not run through a Flux controller here. The kstatus misreport above is derived from the controller code. Validate the CEL on your Flux version (the CEL Playground accepts the resource object as input) before relying on it.