decision-model-operator — architecture¶
Status: alpha (v0.3). Module github.com/maks3201/decision-model-operator,
API group decisionmodel.io/v1alpha1 (the domain is not registered; the group may change before v1).
Engine contract: internal/engine/engine.go.
1. Goal¶
Safe production rollout of decision models: a new model version is evaluated against a golden dataset (and against production) before it receives traffic; a regression is blocked or rolled back. Running the model server is a means, not the product.
DecisionModel CRD
│
▼
Lifecycle / rollout controller (internal/controller: state machine, revisions, status)
├── Runtime adapter (internal/engine: contract; internal/engine/ollaya: Ollaya)
├── Evaluation engine (internal/controller/evaluator.go, evaluate_flow.go; internal/eval/calibration)
└── Promotion / rollback (internal/controller/promotion.go: Service switch, failed revisions)
The runtime adapter serves the model (workloads, health, what is loaded). The controller owns the lifecycle. The evaluation engine decides quality. Promotion decides whether a candidate becomes active. Only the adapter knows the runtime.
Non-goals: generic model serving (KServe, KubeAI, Ollama operators), GPU scheduling, autoscaling (stay compatible with HPA/KEDA), a model registry, experiment tracking, LLM text generation, multi-cluster, distributed inference.
2. Facts the design is built on¶
Verified against code, docs and local runs on 2026-09-27 (details in docs/spikes/).
| Fact | Source | Design consequence |
|---|---|---|
laya-serve: FastAPI, /v1/systemone, /health; one inference at a time per Pod |
laya/serve.py |
concurrency 1 → scale out; autoscale on in-flight/queue |
Laya does not pin the HF revision (snapshot_download without revision) |
laya/agent.py |
the operator must resolve and pin versions itself |
Laya silently falls back to CPU on OOM / missing CUDA; /health does not show it |
laya/agent.py, serve.py |
the operator needs its own check of the actual device |
Ollaya (ollaya-dev/ollaya, 0.10.0, Rust): "Ollama for decision models". Laya, Kev, JevK5, decider, NLI, GLiClass behind one API; weights pinned to commit + sha256; /v1/systemone, /api/pull, /api/ps, /api/decide (load/unload, keep_alive), 503 QUEUE_FULL |
README, docs/api.md |
primary runtime: the multi-model abstraction already exists |
Ollaya has no /metrics; liveness is GET / (200 even with an API key); no implicit pull |
docs/api.md |
readiness and prefetch are the operator's job |
Image ghcr.io/ollaya-dev/ollaya:0.10.0 (native arm64 + amd64; :0.10.0-cuda amd64 only): ollaya serve, 0.0.0.0:11435, UID 1000 |
spikes 001, 006 | near-default PodSpec |
/api/ps reports name, digest, device, precision, expires_at |
spike 001 | Inspect() for the readiness gate |
POST /api/decide {"model":…,"keep_alive":-1} loads and pins a model |
spike 001 | Warmup() is one request |
Serving from a read-only store works; pull on a RO store fails |
spike 001 | serving Pods mount the PVC read-only; only the prefetch Job writes |
Model digest = sha256(manifest bytes) from GET https://ollaya.dev/v2/library/<name>/manifests/<tag>; equals /api/ps digest |
spike 001 | resolver needs no daemon, plain HTTP |
pull accepts a tag only, never @sha256: |
spike 001 | pull by tag, then verify the manifest sha256 in the Job |
A bare name (laya) resolves to latest, a different, heavier artifact |
spike 001 | the API requires an explicit tag |
OLLAYA_DEVICE=cuda on the CPU image fails at model load, not at startup; the port is open |
spike 001 | a port/HTTP probe is not enough; gate on /api/ps |
Sizes (CPU, F32, cgroup anon, re-measured 2026-10-02): laya:en 3.09 GiB, laya:multilingual 1.85 GiB, nli 3.62 GiB, gliclass 2.42 GiB |
spike 002 (corrected), docs/sizing.md |
built-in default requests per measured name:tag |
Ollama 0.35 and llama.cpp also serve /v1/systemone. Ollama: /api/ps digest = sha256 of the registry manifest, no pull by digest, RO store works; but no device field (only size_vram), no server-side API key, different request schema |
spike 004 | second engine possible later, not in the next stage |
| OCI image volumes (KEP-4639, GA in k8s 1.36) can carry an Ollaya store; gate unaffected; Ollaya registry is not OCI, so the model must be repackaged | spike 005 | opt-in cache backend candidate (later) |
| No existing Kubernetes operator for Laya / Ollaya / Jev-style models found | GitHub + web search | the niche is open (re-check before release) |
3. Positioning¶
| Helm chart | KServe | ollama-operator-style | this operator | |
|---|---|---|---|---|
| Deployment / Service / GPU | yes | yes | yes | yes |
| Model pinned to a digest | manual | storageUri | partial | yes, in status |
| Prefetch / cache | no | LocalModelCache | pull at start | Job → PVC before rollout |
| "Model loaded on the expected device" | no | no | no | yes (readiness gate) |
| Eval-gated rollout | no | no | no | yes |
| Knows the Jev / System-1 contract | no | no | no | yes |
The value is in the last three rows. Without them the project is a Helm chart.
4. Components¶
Kubernetes API
│ watch DecisionModel, owned objects, Pods
▼
┌──────────────────────── controller-manager ────────────────────────┐
│ DecisionModel reconciler (state machine, §6) │
│ ├── Resolver name:tag → pinned digest (registry HTTP) │
│ ├── Cache PVC + prefetch Job (`ollaya pull` + sha256) │
│ ├── Engine ollaya → PodSpec, JobSpec, HTTP to Pods │
│ ├── Revisions one Deployment per revision (stable/candidate) │
│ ├── Prober warmup + /api/ps → Pod readiness gate │
│ ├── Evaluator golden set via /v1/systemone (async) │
│ └── Status conditions, events, observedGeneration │
└────────────────────────────────────────────────────────────────────┘
│ owns
▼
PVC <name>-store-<rev> Job (prefetch) Deployment <name>-<rev> PDB Service <name>
│ Pod: ollaya serve
│ readinessGate: decisionmodel.io/model-ready
▼
POST /v1/systemone (Jev-compatible)
Engine contract¶
internal/engine/engine.go is normative and does not import CRD types: the controller maps
DecisionModel.spec → engine.Params; the engine maps Params → PodSpec / JobSpec and talks
to Pods over HTTP. Methods: Resolve, ServingPodSpec (store RO), PrefetchJobSpec (store RW,
digest check), ServicePort, Inspect (/api/ps), Warmup (keep_alive=-1). Optional
capabilities detected by type assertion: Decider (/v1/systemone, used by the evaluator) and
CanonicalName.
Implementations: ollaya. A laya-serve engine is possible later; no other abstractions until
a second real engine exists.
5. API: DecisionModel (v1alpha1)¶
apiVersion: decisionmodel.io/v1alpha1
kind: DecisionModel
metadata:
name: support-router
spec:
engine: ollaya # only ollaya today
model: laya:en # explicit tag required (CEL)
# digest: <64 hex> # optional hard pin; otherwise resolved and recorded in status
replicas: 2
device: cpu # cpu | cuda
resources: {} # PodSpec resources; cuda adds nvidia.com/gpu: 1 if absent
scheduling: {} # nodeSelector / tolerations / affinity passthrough
cache:
size: 10Gi
accessModes: [ReadWriteOnce] # RWX needed for replicas > 1 across nodes (§7)
auth:
apiKeySecretRef: { name: support-router-key, key: token }
rollout:
evaluation: # optional; without it promotion follows ModelReady
datasetRef: { configMapRef: { name: support-router-golden, key: cases.jsonl } }
minAccuracy: "0.90"
maxAccuracyDrop: "0.02" # vs the stable revision on the same dataset
status:
phase: Ready
phaseTransitionTime: "…"
observedGeneration: 3
endpoint: http://support-router.default.svc:11435/v1/systemone
stableRevision: { hash: 7c9f…, model: laya:en, digest: c305…, device: cpu, precision: F32 }
candidateRevision: null
failedRevision: null
evaluation: { accuracy: "0.9400", baselineAccuracy: "0.9300", cases: 500 }
replicas: { desired: 2, modelReady: 2 }
conditions: [Resolved, Cached, ModelReady, Evaluated, Promoted, Ready, Degraded]
Golden dataset: JSONL, one case per line:
{"state": …, "questions": {…}, "expected": {"<qid>": "<choice>" | true | false}}.
Revision hash = hash of engine, model, digest, device, resources, engine image and placement
(scheduling; omitted when empty so older hashes are unchanged). A scheduling change is a
blue-green candidate. replicas, cache and auth are in-place updates of the stable Deployment.
In-place rules:
- A running stable is rendered from its recorded identity; its placement and resources are frozen
from the live Deployment, so an operator upgrade (new default resources, GPU toleration, arch
pin) does not roll it. Stables recorded by an older operator are adopted without a roll.
- Every serving Deployment has an explicit strategy: RWO store → Recreate; shareable store +
cuda → RollingUpdate{maxSurge: 0, maxUnavailable: 1}; otherwise 25%/25%. For replicas: 1
any in-place change is a short outage; that is the price of no Multi-Attach / GPU deadlock.
- Known limit: an auth secret-ref change rolls the stable in place, and the prober uses the
current key for every revision; a per-revision key is a v0.2 item.
Manual promotion: spec.rollout.manualPromotion: true parks a candidate that passed its
gate in AwaitingPromotion until decisionmodel.io/promote=<candidate hash> is set; no timeout
while waiting; the first revision is never held.
6. Reconcile state machine¶
spec changed (new revision)
│
▼
┌──────────┐ ┌──────────┐ ┌─────────────┐ ┌────────────┐
│ Resolving├───►│ Caching ├───►│ Starting ├───►│ Evaluating │
└────┬─────┘ └────┬─────┘ │ (candidate │ └──┬──────┬──┘
│ fail │ fail │ Deployment,│ pass│ │fail
▼ ▼ │ model-ready│ ▼ ▼
Failed / Failed / │ gate) │ Promoting RolledBack
RolledBack RolledBack └──────┬──────┘ │ (candidate deleted,
│ timeout ▼ stable keeps serving)
▼ Ready
RolledBack (old revision deleted)
- Resolving:
model→ digest via the registry. Stored instatus.candidateRevision. - Caching: the prefetch Job pulls the tag into the PVC and verifies the manifest sha256 against the resolved digest. Idempotent.
- Starting:
Deployment <name>-<rev>is created without traffic (the Service selectsdecisionmodel.io/revision=<stable>). Pods carry the readiness gatedecisionmodel.io/model-ready; the prober runsWarmup→Inspectand sets the gate True only if digest and device match. Mismatch (e.g. silent CPU fallback) → gate False with reasonDigestMismatch/DeviceMismatch, an Event, and no promotion. - Evaluating: the golden set runs asynchronously against one candidate Pod (and the stable
revision for a baseline, if missing). Skipped without
rollout.evaluation. - Promoting: record the candidate as
status.stableRevision, switch the Service selector, then delete the old revision after a short grace. Blue-green, not rolling: evaluation must pass before traffic moves. - Failure: rollback if a stable revision exists, otherwise
Failed. The failed revision is recorded instatus.failedRevisionand not retried automatically. Retry with thedecisionmodel.io/retryannotation.
Other phases: AwaitingPromotion (manual promotion, §5) and Degraded (not all stable replicas
are model-ready, or a referenced Secret is not allowed). The Degraded condition additionally
flags issues that do not stop serving, e.g. CacheNotShareable or StoreTerminating. Promoting
is reserved and never observed: promotion persists status.stableRevision and then moves the
Service in one reconcile.
Rules:
- Everything is owned via OwnerReferences; no finalizers (no external resources).
- Level-triggered: phase is derived from cluster state on every reconcile.
- Progress timeouts on Caching / Starting / Evaluating, measured from phaseTransitionTime.
- metav1.Condition everywhere, observedGeneration always set.
- Watches: DecisionModel (generation changes), owned Deployment / Job / Service / PVC / PDB, Pods by
label decisionmodel.io/name.
7. Model cache¶
One store PVC per revision, <dm>-store-<rev>, mounted at OLLAYA_MODELS=/models. The
prefetch Job mounts it RW, that revision's serving Pods RO; serving Pods therefore cannot pull or
delete models. fsGroup: 1000 (fsGroupChangePolicy: OnRootMismatch) on both.
- Why per revision: blue-green runs stable and candidate at the same time. With one shared RWO PVC a
candidate scheduled on another node (e.g. one GPU per node) cannot attach it → Multi-Attach
deadlock (reproduced on EKS). A revision's PVC is garbage-collected with its Deployment/Job (never
the stable or candidate; the just-demoted previous revision only within the promotion grace). Two
PVCs exist during a rollout, so the store footprint ≈ 2× the model while it runs. Use a
StorageClass with
reclaimPolicy: Delete, or deleted stores keep their volumes. - Migration: a stable promoted before per-revision stores keeps mounting the shared
<dm>-storeuntil its next promotion; the shared PVC is deleted once no workload references it. - The stable revision is rendered from its recorded identity in
status.stableRevision(engine, model, digest, device, image, resources), never from the live spec, so starting — or failing — a candidate never rewrites the running stable Pod template. - SELinux: on enforcing nodes whose CSI driver lacks SELinux mount support the runtime relabels the volume to the mounting Pod's MCS level; the operator sets one stable level per DecisionModel on the prefetch and serving Pods so they can share the store.
replicas: 1or single node: RWO is enough.replicas > 1across nodes needs RWX (spec.cache.accessModes); otherwise the operator setsDegradedreasonCacheNotShareable.- Later: node-local cache or OCI image volumes (model as an image), similar to KServe LocalModelCache / modelcars.
- Air-gapped:
--ollaya-registry/OLLAYA_REGISTRY→ internal mirror.
Namespace-scoped mode¶
By default the operator watches all namespaces. --watch-namespaces=ns1,ns2 (chart value
watchNamespaces) scopes the manager cache with cache.Options.DefaultNamespaces, and the
reconciler refuses any DecisionModel outside the set (defence in depth). The operator then makes no
cluster-scoped API calls at runtime and runs with one Role per watched namespace; the CRD remains
cluster-scoped and is installed by a cluster admin.
8. Resources and scheduling¶
device: cuda→nvidia.com/gpu: 1,OLLAYA_DEVICE=cuda,:0.10.0-cudaimage. GPU behaviour verified on EKS (g4dn.xlarge, T4, Bottlerocket, 0.7.3-cuda; not yet re-verified on 0.10.0-cuda):/api/psreportscuda:0, which the gate treats as classcuda(only the exactcuda:<n>form; anything else fails the gate). A blue-green rollout needs a second GPU for the candidate. Models are small (hundreds of MB to a few GB); recommend GPU time-slicing / MIG in the docs, the operator does not automate it.device: cpu→ CPU image; requests from the sizing table in spike 002.schedulingis a passthrough.
9. Security¶
- Pods: non-root (UID 1000),
readOnlyRootFilesystem, drop ALL,RuntimeDefaultseccomp, no service account token. - Engine API key from a Secret (
OLLAYA_API_KEY); the prober and evaluator use the same key. Secrets and ConfigMaps are read through the uncached API reader withget-only RBAC. - Multi-tenant guards:
--allowed-registries(defaultollaya.dev),--allow-insecure-registries(default false),--allow-image-override(default false), and Secrets referenced by a DecisionModel must carry the labeldecisionmodel.io/api-key: "true". - Trust boundary is the namespace: anyone who can create Pods in it can mount any Secret there. Ownership checks before probing/GC are defence in depth, not a tenant boundary.
- Calls to serving Pods (
Inspect,Warmup,Decide, which carry the API key) use a client without a proxy; only the registry client honoursHTTP(S)_PROXY. Proxy URLs with credentials are never copied into tenant prefetch Jobs. - Service is
ClusterIPonly; an example NetworkPolicy ships in the docs. - The prefetch Job is the only component with egress: to the registry (manifests, graphs) and to
Hugging Face (
huggingface.co, redirected to*.hf.co) for the weights. Both can be mirrored (--ollaya-registry,--ollaya-hf-endpoint); a download token comes from a labelled Secret and reaches the prefetch Job only. - A golden dataset may contain real data; store it in a Secret if sensitive.
- Repository: gitleaks pre-commit hook and CI scan.
10. Observability¶
- Conditions and Events (events.k8s.io/v1) on the CR are the primary UX; printer columns:
Model, Device, Phase, Ready, Age (Digest with
-o wide). - Operator metrics: controller-runtime defaults plus
decisionmodel_*series (phase, ready replicas, rollouts, phase durations, eval accuracy / ECE, probe results, registry latency); seedocs/metrics.md. - Ollaya has no inference metrics → later: upstream
/metricsor a sidecar.
11. Stack¶
- Go (module minimum 1.26, built with 1.27), controller-runtime v0.25, k8s libraries v0.37, operator-sdk / kubebuilder layout.
- Tests: envtest (reconcile), kind e2e with the CPU Ollaya image and
laya:en(GitHub Actions, amd64). - Distribution: multi-arch image in GHCR and
install.yamlattached to a GitHub Release on tagv*; Helm chart as an OCI artifact (oci://ghcr.io/maks3201/charts/decision-model-operator). OLM later.
12. Roadmap¶
Stages, not version numbers; the CHANGELOG records what each release ships.
| Stage | Scope |
|---|---|
| Released | DecisionModel, ollaya engine, CPU + CUDA, per-revision PVC store + prefetch Job, readiness gate (digest + device + pinned), blue-green switch, eval-gated and manual promotion (accuracy, ECE/Brier), security guards, namespace-scoped mode, conditions, Events, e2e on kind |
| Next | per-revision API key, autoscaling (KEDA on in-flight / queue), metrics, node-local cache, GPU e2e |
| Later | shadow traffic (agreement rate), confidence cascade / fallback to an LLM |
13. Open questions¶
- CUDA on a real GPU:
/api/psreportscuda:0(EKS T4, answered); behaviour when the GPU disappears at runtime is still unverified. - GGUF / llama.cpp runner models (
jevk5,winnow): not verified (smallest is 4.5 GB). - Model licenses differ per model: what to state in the docs.
- Final project name and API group domain.