Quickstart¶
Deploy a laya:en decision model and send it a request. About 10 minutes on a
laptop-scale cluster.
Prerequisites¶
- A Kubernetes cluster (v1.29+; CI tests 1.35 and 1.37). kind works:
kind create cluster. kubectlpointed at that cluster.- A default StorageClass that supports
ReadWriteOnce(kind ships one). Forreplicas > 1across nodes you needReadWriteMany— see Scaling. - Egress from the prefetch Job to
https://ollaya.dev(the model registry: manifests) and tohttps://huggingface.coplus the Hugging Face CDN it redirects to (*.hf.co: the weights). Air-gapped clusters can mirror both, see Supported models.
Install the operator¶
From a release (recommended)¶
Replace <tag> with a released version (e.g. v0.3.0):
kubectl apply -f https://github.com/maks3201/decision-model-operator/releases/download/<tag>/install.yaml
This installs the CRD, RBAC, and the operator Deployment into the
decision-model-operator-system namespace.
From source¶
Verify the operator is running:
kubectl -n decision-model-operator-system get deploy
kubectl -n decision-model-operator-system logs deploy/decision-model-operator-controller-manager
With Helm¶
The chart is published on GHCR as an OCI artifact (and listed on
Artifact Hub).
Replace <version> with a chart version without the v (e.g. 0.3.0); the operator
image defaults to the matching v<version> tag:
helm install dmo oci://ghcr.io/maks3201/charts/decision-model-operator --version <version> \
--namespace decision-model-operator-system --create-namespace
From a source checkout, use charts/decision-model-operator instead of the OCI reference.
The chart installs the CRD (from its crds/ directory), RBAC, the manager
Deployment and the metrics Service. Common overrides:
helm install dmo oci://ghcr.io/maks3201/charts/decision-model-operator --version <version> \
--namespace decision-model-operator-system --create-namespace \
--set replicaCount=1 \
--set manager.allowedRegistries="ollaya.dev\,registry.example.com" \
--set manager.maxConcurrentReconciles=4
Values of note: image.repository/tag/pullPolicy, replicaCount, resources,
nodeSelector/tolerations/affinity, and the manager policy flags
manager.allowedRegistries, manager.allowInsecureRegistries,
manager.allowImageOverride, manager.maxConcurrentReconciles,
manager.ollayaRegistry. See charts/decision-model-operator/values.yaml.
Helm does not upgrade or delete CRDs in crds/; manage CRD upgrades out of band.
Upgrade the operator with helm upgrade, uninstall with
helm uninstall dmo -n decision-model-operator-system (the CRD stays).
Behind an HTTP proxy¶
Corporate clusters often reach the model registry only through an HTTP proxy. Set the proxy on the operator and it is applied to the operator's own egress and copied into every prefetch Job (the only workload that pulls weights):
helm install dmo oci://ghcr.io/maks3201/charts/decision-model-operator --version <version> \
--namespace decision-model-operator-system --create-namespace \
--set proxy.httpProxy=http://proxy.corp:3128 \
--set proxy.httpsProxy=http://proxy.corp:3128 \
--set 'proxy.noProxy=10.96.0.0/12\,.svc\,.cluster.local'
These render as HTTP_PROXY / HTTPS_PROXY / NO_PROXY env on the manager
container (nothing is rendered for an empty value).
The proxy applies to the operator's registry calls (resolving a tag's digest)
and is copied to the prefetch Job that pulls weights. It
does not apply to the operator's in-cluster calls to serving Pods: the
readiness prober and evaluator connect to Pod IPs with a separate, proxy-less
HTTP client (so the Authorization: Bearer <api key> header never traverses the
proxy and readiness no longer depends on noProxy). You therefore do not
need the pod CIDR in noProxy. Still set noProxy for the registry/API path:
- the service CIDR — so in-cluster Services (e.g. the Kubernetes API server) are reached directly, not via the proxy;
.svc,.cluster.local— in-cluster DNS names.
The service CIDR is the API server's --service-cluster-ip-range (often the
ClusterIP range of the kubernetes Service in default). Example:
--set 'proxy.noProxy=10.96.0.0/12\,.svc\,.cluster.local'.
If your proxy URL carries credentials, keep it out of the Helm values: create a
Secret with keys httpProxy / httpsProxy / noProxy and set
--set proxy.existingSecret=<name>. The manager env then sources each var from
the Secret (valueFrom.secretKeyRef, each key optional); the plaintext
proxy.* values are ignored when existingSecret is set.
A credentialed proxy URL (user:pass@…) is used by the operator only.
The operator does not copy a credentialed proxy URL into the prefetch Jobs —
a Job's spec is readable by anyone who can get jobs in that (tenant) namespace,
so the password would leak. It copies NO_PROXY/no_proxy to the Jobs always,
and logs one startup warning naming the skipped variables (names only, never
the value). A value it cannot parse, or one containing @, is treated as
credentialed and skipped (fail-closed).
Consequence: with a credentialed-only proxy configuration the prefetch pull
has no proxy, so in a proxy-only network it fails visibly as PrefetchFailed
(phase Failed/RolledBack, a Degraded condition) rather than leaking the
credential. For the Job's own egress to the registry, use an IP-allowlisted
proxy that needs no credentials, or a node-level / transparent proxy, set
via a plain (credential-free) HTTP_PROXY the operator can forward. A proxy URL
without credentials is copied to the prefetch Jobs as normal.
Serving Pods need no proxy — they do not pull.
Create a DecisionModel¶
# dm.yaml
apiVersion: decisionmodel.io/v1alpha1
kind: DecisionModel
metadata:
name: support-router
spec:
engine: ollaya
model: laya:en # an explicit ":tag" is required (a bare name resolves to :latest, a heavier artifact)
device: cpu
replicas: 1
resources:
requests:
cpu: "1"
memory: 3500Mi
limits:
memory: 4Gi
dm is the short name for DecisionModel. The columns show MODEL, DEVICE,
PHASE, and READY (model-ready replicas).
What the phases mean¶
The operator drives a level-triggered state machine
(status.phase):
| Phase | Meaning |
|---|---|
Pending |
Accepted, not yet processed. |
Resolving |
Resolving the tag to an immutable digest from the registry. |
Caching |
A prefetch Job is downloading the weights into the model-store PVC and verifying the digest. |
Starting |
Serving Pods are created; waiting for the model to load and pass the readiness gate. |
Evaluating |
(rollout with spec.rollout.evaluation) scoring the candidate against a golden dataset. |
Promoting |
Switching the Service to the new revision. |
Ready |
The stable revision is serving and all desired replicas are model-ready. |
Degraded |
The stable revision is serving but not all desired replicas are model-ready: 0 < modelReady < desired (Ready=True, ModelReady=False, reason ReplicasNotModelReady), or modelReady == 0 (Ready=False, reason NoModelReadyPods). The model keeps serving on whatever replicas are ready. |
Failed |
The rollout failed with no healthy stable revision (e.g. tag not found on a first rollout, resolve failure, or a policy violation). |
RolledBack |
A candidate failed and traffic stayed on the previous stable revision, which keeps serving. |
Failed vs RolledBack: both mean the latest revision did not roll out. If a
stable revision was already serving, its traffic is preserved and the phase is
RolledBack; if there was nothing healthy to fall back to, it is Failed.
If you change the spec rapidly back and forth so a revision's model-store PVC is
being re-created while its previous PVC is still Terminating, the operator
waits (requeues every ~5 s in Caching) for the old PVC to finish deleting
rather than failing — the rollout resumes once the store PVC is available.
Degraded is a phase, not just a condition: since a healthy stable revision
keeps serving, partial (or zero) replica readiness moves status.phase to
Degraded rather than Ready. The Ready condition can still be True in the
partial case (0 < modelReady < desired), because the model is answering on the
ready replicas; when modelReady == 0 the Ready condition is False. A
separate Degraded=True condition (see below) also accompanies cache-sharing
warnings and rejected candidates without necessarily changing the phase.
Key conditions (kubectl describe dm support-router → Conditions):
Resolved— the tag resolved to a digest (recorded instatus.stableRevision.digest).Cached— weights are present in the PVC and the digest matched.ModelReady— enough Pods report the expected digest loaded on the expected device.Ready— the model is serving.Degraded— a policy or runtime problem; read the reason and message.
Condition reasons you may see¶
Each reason resolves to a specific phase and condition. The Phase / condition
column states where the reason actually lands (kubectl describe dm → Conditions,
or -o jsonpath='{.status.conditions}'):
| Reason | Phase / condition | Meaning | Fix |
|---|---|---|---|
NoModelReadyPods |
phase Degraded; Ready=False, ModelReady=False, Degraded=True |
No serving Pod reports the expected digest/device (modelReady == 0). |
Check Pod logs / device; a CPU pod set to device: cuda never loads. |
ReplicasNotModelReady |
phase Degraded; Ready=True, ModelReady=False, Degraded=True |
Some, but not all, replicas are model-ready (0 < modelReady < desired). |
Wait, or check node capacity / the store PVC access mode. |
DigestMismatch |
Pod readiness gate decisionmodel.io/model-ready=False (keeps the Pod out of the Service; feeds the counts above) |
A Pod loaded a different digest than expected. | The tag moved upstream mid-rollout; re-resolve (change spec) or pin spec.digest. |
DeviceMismatch |
Pod readiness gate decisionmodel.io/model-ready=False |
A Pod loaded on the wrong device. | Fix spec.device / node GPU availability. |
NotPinned |
Pod readiness gate decisionmodel.io/model-ready=False |
The right model on the right device, but it is not pinned in the runtime (keep_alive), so the runtime may evict it and silently stop serving. |
Usually transient — the operator's warmup pins it (keep_alive=-1); the gate recovers once /api/ps shows it pinned. Persisting suggests the runtime is dropping the pin. |
CacheNotShareable |
Degraded=True condition; phase unchanged (the model still serves) |
replicas > 1 with a ReadWriteOnce store. |
Set spec.cache.accessModes: ["ReadWriteMany"] and an RWX storageClass. |
CacheSpecImmutable |
Degraded=True condition; phase unchanged |
A spec.cache change the operator cannot apply to the existing PVC: an accessModes change (access modes are immutable), a size shrink (PVCs cannot shrink), a storageClassName change, or a size grow the StorageClass refused (it does not allow volume expansion). The message names which. A size grow on an expansion-capable StorageClass is applied in place and does not set this. |
For a shrink / class / access-mode change, delete the DM (and its model-store PVC) and recreate it, or trigger a new revision with a model-affecting change (which provisions a fresh PVC from the current spec.cache); for a refused grow, use a StorageClass with allowVolumeExpansion: true. |
RegistryNotAllowed |
phase Failed; Ready=False (permanent — not retried) |
spec.model host is not in --allowed-registries (or is http:// without --allow-insecure-registries). |
Use an allowed registry, or have the operator configured to allow it. |
ImageOverrideNotAllowed |
phase Failed; Ready=False (permanent) |
spec.image set but the operator forbids overrides. |
Remove spec.image, or run the operator with --allow-image-override. |
SecretNotAllowed |
phase Degraded; Ready=False |
The API-key Secret lacks the required label. | Label it decisionmodel.io/api-key: "true" (see below); the operator picks it up within ~60 s. |
InvalidModelName |
phase Failed; Ready=False (permanent) |
spec.model did not parse. |
Use [host/][namespace/]model:tag with an explicit tag. |
ModelNotFound |
phase Failed; Resolved=False (requeued and re-resolved) |
The tag does not exist in the registry. | Fix the tag; a bare name resolves to :latest. |
DatasetInvalid |
phase RolledBack (stable exists) or Failed; Degraded=True |
Eval dataset missing or malformed JSONL. | See evaluation.md. |
EvaluationFailed |
phase RolledBack (stable exists) or Failed; Degraded=True |
Candidate did not meet the eval gates. | Inspect status.evaluation; see evaluation.md. |
EvaluationTimeout |
phase RolledBack (stable exists) or Failed; Degraded=True |
Evaluation exceeded its deadline. | Reduce maxCases or check the model. |
ResourceConflict |
Ready=False, Degraded=True reason ResourceConflict |
An object the operator needs (Service, Deployment, Job or PVC) already exists under the DecisionModel's name but is not owned by it (different or absent controller owner). The operator never modifies or deletes a foreign object; a foreign Service blocks promotion, so traffic cannot reach the model. | Rename or remove the conflicting object (it was not created by this operator), then the next reconcile proceeds. |
StoreLost |
Degraded=True reason StoreLost; stable keeps whatever Pods are ready |
The stable revision's model-store PVC went missing (e.g. deleted out of band). The operator recreates it (annotated decisionmodel.io/store-recovering) and re-runs the prefetch before letting Pods serve on it; it will not serve an empty store. |
Usually none — recovery is automatic. Check StorageClass/quota if it stays here. |
StorePrefetchFailed |
Degraded=True reason StorePrefetchFailed; held without churn |
The recovery prefetch Job exhausted its retries (bad registry, digest mismatch, disk). | Fix the underlying cause, then bump decisionmodel.io/retry to re-run recovery. |
StoreTerminating |
Degraded=True reason StoreTerminating; the stable keeps serving |
The stable store PVC was deleted but is still Terminating because the stable Pods still mount it (the kubernetes.io/pvc-protection finalizer). The operator does not stop the stable automatically (a PVC delete cannot be undone). |
Delete the stable Pods so the PVC can finish deleting and recovery can run (StoreLost → recreate + prefetch). Expect a short serving gap while the Pods restart. |
For candidate failures (DatasetInvalid, EvaluationFailed, EvaluationTimeout,
and SecretNotAllowed when reading an eval dataset), the phase is RolledBack
when a stable revision is still serving, and Failed only when there is no
stable revision to fall back to.
The StoreLost / StorePrefetchFailed / StoreTerminating reasons concern the
stable revision's model store; the stable keeps serving on whatever Pods are
ready throughout (Ready can stay True), and the operator recovers the store in
place. StoreTerminating is the only one that needs a human action — see its row.
While a candidate is mid-rollout the operator also keeps the stable healthy:
it restores a deleted stable Deployment, re-applies an API-key rotation, keeps the
stable's PodDisruptionBudget, and re-checks the gate on the stable Pods — all
rendered from status.stableRevision, never the in-flight candidate's spec.
Retrying a failed rollout¶
A failed revision is not retried automatically (a transient registry blip should
not loop). To retry without changing the spec, set/change the
decisionmodel.io/retry annotation to any new value:
The controller clears the failed revision once when this value differs from the one it last consumed. Returning the spec to the last stable revision also clears it.
Prefetch failures: transient vs permanent¶
The prefetch Job tells a transient download failure from a permanent one and reacts differently:
- Transient (network error, registry/Hugging Face
5xx, timeout): the Job retries up to itsbackoffLimit(4) within the Caching deadline, so a short outage recovers on its own. - Permanent — the Job fails immediately (no wasted retries):
ModelNotFound— the registry has no such tag (spec.modelis wrong, or a bare name resolved to a different artifact). Fix the tag; a digest pin (spec.digest) protects against a tag moving.DigestMismatch— the pulled manifest's sha256 does not match the pinned digest (the tag moved upstream mid-rollout, or the content is corrupt). Re-resolve by changing the spec, or pinspec.digest.
The reason is shown on the failed revision's condition/Event. A permanent failure will not clear on its own — fix the cause, then retry as above.
Secret capability labels¶
A Secret the operator reads must carry the label for the capability it is used for — one label grants exactly one capability:
| Capability | Label | Granted to |
|---|---|---|
| Engine API key | decisionmodel.io/api-key: "true" |
spec.auth.apiKeySecretRef |
| Eval dataset | decisionmodel.io/eval-dataset: "true" |
a datasetRef.secretRef dataset |
| Download token | decisionmodel.io/download-token: "true" |
spec.cache.downloadTokenSecretRef |
A Secret without the required label is refused: the API key and download token
report phase Degraded, Ready=False, reason SecretNotAllowed; a dataset
Secret is held in Evaluating (Evaluated=False, reason SecretNotAllowed)
until the label is added. This guards against a DecisionModel editor pointing the
operator at an arbitrary Secret.
kubectl label secret my-ollaya-key decisionmodel.io/api-key=true
kubectl label secret my-golden-dataset decisionmodel.io/eval-dataset=true
Adding the label to an existing Secret is enough to recover: the operator re-reads the Secret on its next reconcile and clears the condition within ~60 s — no spec change and no annotation are required.
Transition. A dataset Secret labelled only with the older
decisionmodel.io/api-key: "true"still works for one release but emits aDeprecatedSecretLabelWarning; switch it todecisionmodel.io/eval-dataset.
The API-key Secret must contain the referenced key: apiKeySecretRef.optional is
rejected (CEL), and a labelled Secret missing the key is reported as Degraded,
reason APIKeyInvalid, with no workloads changed.
Rotating the API key¶
You can rotate the key in place by changing the value inside the same Secret
(no spec change). The operator does not watch Secrets (they stay get-only
RBAC); instead it reads the Secret on each reconcile and records a hash of the
value in the decisionmodel.io/apikey-checksum annotation on the serving
Deployment. When the value changes, the new checksum lands on the Pod template and
the Deployment rolls per the rollout strategy.
What to expect after you update the Secret value:
- Pickup is on the next reconcile, within ~60 s (a
ReadyDecisionModel re-checks about once a minute; there is no Secret watch, so it is not instant). - A brief 401 window. The still-running old Pods keep the previous key in their
env until they are replaced, so the operator's readiness prober gets
401from them. The gate tolerates this for a bounded window (up to ~3 min: 3 re-checks at ~60 s) rather than flapping the Pod out immediately; the roll normally completes within it. A runtime that keeps returning401past that window has its gate flippedFalse(reasonProbeError). - Downtime follows the strategy table. On the default
ReadWriteOncestore withreplicas: 1the roll isRecreate, so there is a short outage while the new Pod starts and reloads the model. With an RWX store andreplicas > 1the rolling update keeps other replicas serving. - On upgrade, the first reconcile only records the checksum (it does not roll a Pod that is otherwise unchanged); a roll happens only when the value actually changes afterwards.
Model-aware readiness is the point of this operator: a Pod only joins the
Service once the runtime's /api/ps reports the expected digest on the
expected device, via the Pod readiness gate decisionmodel.io/model-ready.
A Pod that silently fell back to CPU, or loaded the wrong weights, never gets
traffic.
The gate is also re-verified every 60 s (an /api/ps Inspect only, no
re-warm): if a Pod later unloads the model, loads a different digest, ends up on
the wrong device, or loses its pin, the gate flips back to False (reason
ProbeError / DigestMismatch / DeviceMismatch / NotPinned) and the Pod
leaves the Service. A Ready DecisionModel therefore reconciles about once a
minute; when nothing changed the re-check writes nothing (no Pod or status
update), so it is cheap.
Traffic moves only after the candidate's Pods are genuinely Ready — the
model-ready gate is True and the Pod's Ready condition (kubelet /
PodReady) is True, so the Pod is actually in the Service's EndpointSlice. The
operator persists status.stableRevision before it switches the Service
selector, so a crash mid-promotion cannot leave the Service pointing at a
revision the status does not record (and GC never deletes the revision the live
Service selects).
Operator flags¶
The manager (controller) accepts these policy flags (set via make deploy
args, or Helm manager.* values):
| Flag | Default | Meaning |
|---|---|---|
--allowed-registries |
ollaya.dev |
Comma-separated registry hosts a spec.model may resolve from (SSRF guard). |
--allow-insecure-registries |
false |
Permit http:// model registries (dev only). |
--allow-image-override |
false |
Permit spec.image to override the engine's default image. |
--max-concurrent-reconciles |
4 |
Max DecisionModels reconciled concurrently. |
--runtime-version-policy |
Pinned |
How an unset spec.runtimeVersion resolves: Pinned reuses the stable revision's recorded runtime version (an operator upgrade starts no rollout) or FollowOperator uses the engine default. |
--max-concurrent-rollouts |
0 |
Max DecisionModels rolling out at once across the watched scope (0 = unlimited). Others wait in phase Pending (reason RolloutQueued), FIFO. |
--ollaya-registry |
(empty) | Default registry base URL for host-less model names; empty = OLLAYA_REGISTRY env, else https://ollaya.dev. |
--ollaya-hf-endpoint |
(empty) | Base URL for model-weight downloads (a Hugging Face mirror or enterprise endpoint; the registry still serves manifests). Empty = OLLAYA_HF_ENDPOINT env, else Hugging Face. Needs the 0.10.0+ runtime. |
The security defaults are deliberately strict: only ollaya.dev, no http://,
and no image override unless you opt in.
Upgrading the operator¶
A new operator release may ship a newer default engine runtime image. What
happens to your DecisionModels on helm upgrade depends on
--runtime-version-policy:
- Pinned (default). A DecisionModel that does not set
spec.runtimeVersionkeeps the runtime version its stable revision was promoted with, so an operator upgrade alone starts no rollout. The new default is surfaced as theRuntimeUpdateAvailablecondition and a one-off Event; adopt it deliberately by settingspec.runtimeVersion(e.g.0.10.0) on each DecisionModel, which starts a normal blue-green rollout for that one. - FollowOperator. An unset
spec.runtimeVersiontracks the operator default, so an upgrade rolls every DecisionModel onto the new runtime at once. Bound the blast radius with--max-concurrent-rollouts(e.g.2–3): DecisionModels beyond the budget wait in phasePending(reasonRolloutQueued) and roll out in FIFO order as slots free up.
Pin a specific runtime regardless of policy with spec.runtimeVersion (mutually
exclusive with spec.image):
Eval-gated rollout¶
For blue-green promotion gated on a golden-dataset accuracy and calibration
check, set spec.rollout.evaluation — see evaluation.md.
Inspect¶
kubectl describe dm support-router # spec, status, conditions, recent Events
kubectl get dm support-router -o yaml # full status, including status.stableRevision.digest
kubectl get events --field-selector involvedObject.name=support-router
Send a request¶
The operator creates a Service named after the DecisionModel
(support-router) in the same namespace, exposing port http (11435).
Port-forward it and POST to /v1/systemone:
curl -sS -X POST http://127.0.0.1:11435/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary '{
"model": "laya:en",
"state": "I was charged twice for my subscription this month. Please refund the second charge.",
"questions": {
"department": {
"type": "choice",
"criteria": {
"billing": "billing, payments, charges, refunds",
"technical": "technical bugs and errors",
"sales": "pricing and plans",
"other": "anything else"
}
}
}
}'
The response contains one answer per question id. For a choice question you
get choice, confidence and probabilities:
{"answers": {"department": {"type": "choice", "choice": "billing",
"confidence": 0.97, "probabilities": {"billing": 0.98, "technical": 0.01, "sales": 0.005, "other": 0.005}}}}
Question types: choice (pick one label), noul (probability a statement
holds), score (expected level over an ordered scale).
Supported models¶
There is no built-in allow-list of model names: any model in an Ollaya registry can be referenced. What has been verified so far:
| Models | Status |
|---|---|
laya:en |
E2E on every change (kind, CPU), and once on a T4 GPU (EKS, runtime 0.7.3) |
laya:multilingual, nli:latest, gliclass:latest |
Pulled, served and measured (see sizing) |
gliclass:latest, nli:latest |
Nightly E2E (kind, CPU) |
jevk5:latest (GGUF, Q8_0, llama.cpp runner) |
Reached Ready on kind; the gate checks digest, device and pinning, not precision, so quantized models pass (spike 007) |
Other GGUF models (winnow, jeb, cygnet, …) |
Same runner as jevk5, not run individually; 8–13 GB, raise spec.cache.size |
Large models (nimble, clef, jeeves, ~18–19 GB) |
Not verified: size spec.cache.size and the GPU yourself |
Reference a model by name:tag:
- An explicit tag is required. A bare name (
laya) resolves to:latest, a different and heavier artifact; the operator rejects a model without a tag (InvalidModelName). Pinlaya:en,nli:latest, etc. - Namespaced and host-qualified names are supported:
[host/][namespace/]model:tag, e.g.acme/triage:v3orregistry.example.com/acme/triage:v3. - Internal mirror / air-gapped: set the operator's default registry for
host-less names with
--ollaya-registry(Helmmanager.ollayaRegistry, or theOLLAYA_REGISTRYenv). This covers both steps that reach a registry: the operator's digest resolve and the prefetch Job'sollaya pull(the Job is givenOLLAYA_REGISTRYso the pull targets the same mirror). The mirror must speak the Ollaya/Ollama registry API —GET <base>/v2/<namespace>/<model>/manifests/<tag>returning the v2 manifest whosesha256(manifest bytes)equals the pinned digest, plus the blob endpointsollaya pullfetches; the default namespace islibrary(e.g.library/layaforlaya:en). Anhttp://mirror base (dev only) makes the pull non-TLS automatically. Alternatively, write the host directly into the model name (mirror.corp/library/laya:en); an explicit host in the name overrides--ollaya-registryfor that model, in both resolve and pull. When--ollaya-registrypoints at a mirror, the operator also setsOLLAYA_REGISTRYon the serving Pods (not just the prefetch Job), soollaya servefinds the model in the store it was pulled into. The on-disk store path is keyed by the registry authority with the port's:written as_— a mirror atmirror.corp:8080stores the manifest undermanifests/mirror.corp_8080/library/laya/en— which the operator computes consistently for the pull, the serve lookup and the Job's digest check. -
Registry allow-list: the operator only resolves from hosts in
--allowed-registries(Helmmanager.allowedRegistries, defaultollaya.dev) as an SSRF guard; add your mirror there.http://registries need--allow-insecure-registries(dev only). -
Weight-download mirror (Hugging Face endpoint): the prefetch Job contacts two distinct hosts: the model registry (default
ollaya.dev, set via--ollaya-registry) for manifests and blob digests, and the weight host (default Hugging Face) for the model-weight blobs themselves. In an air-gapped cluster you may need mirrors for both. Set the weight-download mirror with--ollaya-hf-endpoint(Helmmanager.ollayaHFEndpoint); the runtime reads it asOLLAYA_HF_ENDPOINT(requires runtime v0.10.0+; older runtimes ignore it harmlessly). Serving Pods never receive this variable — they mount the store read-only and never download. When both mirrors are behind a corporate proxy, the proxy env already covers both: the controller copiesHTTP_PROXY/HTTPS_PROXY/NO_PROXYto the prefetch container, so any request from the CLI — registry or weight host — goes through the proxy.
helm install dmo oci://ghcr.io/maks3201/charts/decision-model-operator --version <version> \
--namespace decision-model-operator-system --create-namespace \
--set manager.ollayaRegistry=https://registry.internal \
--set manager.ollayaHFEndpoint=https://hf-mirror.internal \
--set 'manager.allowedRegistries=registry.internal'
- Download token (private / gated models): if the weight host (Hugging Face or the mirror) requires authentication, create a Secret with the token and reference it from the DecisionModel:
The Secret must carry the label decisionmodel.io/download-token: "true",
or the operator refuses to read it (phase Degraded, reason
SecretNotAllowed):
kubectl create secret generic my-hf-token --from-literal=token=hf_...
kubectl label secret my-hf-token decisionmodel.io/download-token=true
The token is injected into the prefetch Job as OLLAYA_HF_TOKEN and sent
only to the weight-download endpoint, never to the model registry and never
to serving Pods.
See sizing.md for measured per-model resources and PVC sizes.
What changes trigger a new revision¶
The operator rolls out some spec changes blue-green (a new revision: a second Deployment and its own model-store PVC, verified by the readiness gate before traffic switches) and applies others in place on the running Deployment.
| Spec change | How it is applied |
|---|---|
engine, model, digest, device, resources, engine image |
New revision (re-resolve/prefetch as needed). |
scheduling (nodeSelector, tolerations, affinity, runtimeClassName) |
New revision. A placement change re-downloads into a second PVC and, for device: cuda, needs a second GPU for the duration of the rollout. An unschedulable candidate ends as RolledBack (reason StartTimeout) with the stable still serving — it does not hang silently. |
replicas |
In place (scale the stable Deployment). |
cache (size / accessModes / class) |
In place. A size grow is applied by expanding the PVC when the StorageClass allows volume expansion; a size shrink, a storageClassName change, or an accessModes change is rejected with CacheSpecImmutable (Degraded=True) and the PVC is left as is (delete the DM and its PVC to change those). accessModes: ["ReadOnlyMany"] is rejected at admission — the prefetch Job must write the store. Note: because each revision gets its own store PVC (named by revision hash) provisioned from the current spec.cache, a changed class or a shrink also takes effect automatically the next time a model-affecting change (model, digest, device, resources, image, scheduling) rolls a new revision — not only via delete-and-recreate. |
auth.apiKeySecretRef |
In place — changing which Secret/key is referenced updates the stable Deployment's env and it rolls (see the downtime note below). Rotating the value inside the same Secret is also picked up — see Rotating the API key. |
Engine resource defaults (the built-in per-model memory/cpu requests, see sizing.md) only apply to a new revision. A running stable keeps the resources it was created with; upgrading the operator does not re-request a running stable.
Rollout strategy and downtime¶
The serving Deployment's update strategy depends on the store's access mode and the device, so an in-place change (or an unavoidable operator-driven roll) does not deadlock on the volume:
| Store access mode | Device | Strategy |
|---|---|---|
ReadWriteOnce (default) |
any | Recreate (old Pod stops before the new one starts) |
ReadWriteMany |
cuda |
RollingUpdate, maxSurge: 0, maxUnavailable: 1 (no second GPU needed) |
ReadWriteMany |
cpu |
RollingUpdate, maxSurge: 25%, maxUnavailable: 25% |
For replicas: 1 on an RWO store (the default) or on cuda, any in-place
change that alters the Pod template causes a short outage: the old Pod stops,
the new Pod starts, loads the model, and passes the gate (for a large model that
can be minutes). For replicas > 1 on an RWX store the no-surge rolling update
keeps N-1 Pods serving; on RWX + cpu the default surge means no downtime (the
gate still protects the Service). After the placement/resource freeze, the only
things that roll a running stable in place are an auth.apiKeySecretRef change
and an operator upgrade that changes a non-placement/non-resource field of the
rendered Pod (rare).
Scaling¶
- Always pin an explicit tag (
laya:en), never a bare name — a bare name resolves to:latest, a different and heavier artifact. - For
replicas > 1spread across nodes, the model-store PVC must beReadWriteMany; the defaultReadWriteOncebinds the PVC to one node. (ReadOnlyManyis not accepted — the prefetch Job must write the store.) Set it viaspec.cache:
PVC access modes are immutable once created.
See sizing.md for per-model resources and PVC sizing.