API Reference¶
Packages¶
decisionmodel.io/v1alpha1¶
Package v1alpha1 contains API Schema definitions for the decisionmodel v1alpha1 API group.
Resource Types¶
AuthSpec¶
AuthSpec configures the engine API key.
Appears in: - DecisionModelSpec
| Field | Description | Default | Validation |
|---|---|---|---|
apiKeySecretRef SecretKeySelector |
APIKeySecretRef references a Secret key holding the engine API key. The Secret must carry the label decisionmodel.io/api-key: "true" (opt-in), or the operator reports Degraded (SecretNotAllowed) and creates no workloads. The same key is used by the operator's prober and evaluator. The referenced key must exist (optional is not supported: the prober and the serving Pod would otherwise disagree); a missing key is reported as Degraded (APIKeyInvalid). |
Optional: {} |
CacheSpec¶
CacheSpec configures the model store PVC.
Appears in: - DecisionModelSpec
| Field | Description | Default | Validation |
|---|---|---|---|
size Quantity |
Size is the requested PVC size for the model store. | 10Gi | Optional: {} |
storageClassName string |
StorageClassName selects the PVC storage class. RWX is required for replicas>1 spread across nodes. |
Optional: {} |
|
accessModes PersistentVolumeAccessMode array |
AccessModes for the model store PVC. Defaults to ReadWriteOnce. Set ReadWriteMany when running replicas>1 spread across nodes. ReadOnlyMany is rejected: the prefetch Job must write the store. PVC access modes are immutable once created. |
Optional: {} |
|
downloadTokenSecretRef SecretKeySelector |
DownloadTokenSecretRef references a Secret key holding a credential for weight downloads (e.g. a Hugging Face token for private or gated repositories, or a mirror). It is injected into the prefetch Job only, never into serving Pods. The Secret must carry the label decisionmodel.io/download-token: "true" (opt-in guard), or the operator reports Degraded (SecretNotAllowed) and does not create a prefetch Job. The optional field is not supported: when a ref is set the Secret must exist and contain the referenced key, or the operator reports Degraded (DownloadTokenInvalid). |
Optional: {} |
DatasetKeyRef¶
DatasetKeyRef names an object and a key within it.
Appears in: - DatasetRef
| Field | Description | Default | Validation |
|---|---|---|---|
name string |
Name of the ConfigMap or Secret. | MinLength: 1 |
|
key string |
Key holding the JSONL content. | MinLength: 1 |
DatasetRef¶
DatasetRef references a golden dataset (JSONL) in a ConfigMap or Secret. Exactly one of configMapRef / secretRef must be set.
Appears in: - EvaluationSpec
| Field | Description | Default | Validation |
|---|---|---|---|
configMapRef DatasetKeyRef |
ConfigMapRef selects a key in a ConfigMap holding the JSONL dataset. | Optional: {} |
|
secretRef DatasetKeyRef |
SecretRef selects a key in a Secret holding the JSONL dataset. The Secret must carry the label decisionmodel.io/eval-dataset: "true" (opt-in guard). For one release a Secret labelled only decisionmodel.io/api-key: "true" is still accepted (deprecated; a Warning Event names the label to add). |
Optional: {} |
DecisionModel¶
DecisionModel is the Schema for the decisionmodels API.
metadata.name must be a DNS-1035 label of at most 43 characters (see MaxNameLength): the operator derives a Service named after it (no dots) and Job/PVC names that must stay within 63 characters.
| Field | Description | Default | Validation |
|---|---|---|---|
apiVersion string |
decisionmodel.io/v1alpha1 |
||
kind string |
DecisionModel |
||
metadata ObjectMeta |
Refer to Kubernetes API documentation for fields of metadata. |
||
spec DecisionModelSpec |
|||
status DecisionModelStatus |
DecisionModelPhase¶
Underlying type: string
DecisionModelPhase enumerates the high-level lifecycle phase of a DecisionModel.
Appears in: - DecisionModelStatus
| Field | Description |
|---|---|
Pending |
|
Resolving |
|
Caching |
|
Starting |
|
Evaluating |
|
Degraded |
|
Promoting |
PhasePromoting is reserved and never set. Promotion is a single reconcile that persists status.stableRevision first and only then moves the Service, so there is no observable intermediate phase. It is kept, not removed, because the documented phase set and copied health checks (e.g. GitOps/Argo Lua) name it. |
Ready |
|
AwaitingPromotion |
PhaseAwaitingPromotion means the candidate passed its gate and waits for a human to approve it (spec.rollout.manualPromotion). The stable keeps serving. |
Failed |
|
RolledBack |
DecisionModelSpec¶
DecisionModelSpec defines the desired state of DecisionModel.
Appears in: - DecisionModel
| Field | Description | Default | Validation |
|---|---|---|---|
engine string |
Engine is the serving runtime for the model. | ollaya | Enum: [ollaya] Optional: {} |
model string |
Model is the model name in engine terms, e.g. "laya:en". The last path segment must include an explicit ":tag" (a bare name resolves to :latest, a different and heavier artifact — see spike 001 §9). |
MinLength: 1 Required: {} |
|
digest string |
Digest optionally pins the model to an immutable digest (bare hex sha256). When empty, the operator resolves the tag and records the digest in status. |
Pattern: ^[a-f0-9]\{64\}$ Optional: {} |
|
replicas integer |
Replicas is the number of serving Pods for the stable revision (default 1). With replicas > 1 spread across nodes the model store must be shareable — see cache.accessModes; otherwise the operator reports Degraded (CacheNotShareable). |
1 | Minimum: 1 Optional: {} |
device string |
Device selects the target compute device for serving. "cuda" also requests an nvidia.com/gpu and uses the engine's CUDA image. |
cpu | Enum: [cpu cuda] Optional: {} |
image string |
Image overrides the engine's default container image. Rejected unless the operator is started with --allow-image-override (a DecisionModel editor could otherwise run an arbitrary image under the operator's Pod template). Mutually exclusive with runtimeVersion. |
Optional: {} |
|
runtimeVersion string |
RuntimeVersion pins the engine runtime release (e.g. "0.10.0", plain MAJOR.MINOR.PATCH) independently of the operator's default. Pinning it means an operator upgrade that changes the default runtime image does not start a rollout for this DecisionModel. Mutually exclusive with image. When unset the effective version follows --runtime-version-policy. |
Pattern: ^[0-9]+\.[0-9]+\.[0-9]+$ Optional: {} |
|
resources ResourceRequirements |
Resources are the compute resource requirements for the serving container. | Optional: {} |
|
scheduling SchedulingSpec |
Scheduling is a passthrough of nodeSelector/tolerations/affinity for serving Pods. | Optional: {} |
|
cache CacheSpec |
Cache configures the model store PVC. | Optional: {} |
|
auth AuthSpec |
Auth configures the engine API key. | Optional: {} |
|
rollout RolloutSpec |
Rollout configures rollout behaviour, including eval-gated promotion. | Optional: {} |
DecisionModelStatus¶
DecisionModelStatus defines the observed state of DecisionModel.
Appears in: - DecisionModel
| Field | Description | Default | Validation |
|---|---|---|---|
phase DecisionModelPhase |
Phase is the high-level lifecycle phase. | Optional: {} |
|
phaseTransitionTime Time |
PhaseTransitionTime is when Phase last changed. Used for progress timeouts. | Optional: {} |
|
lastPromotionTime Time |
LastPromotionTime is when a candidate was most recently promoted to stable. | Optional: {} |
|
observedGeneration integer |
ObservedGeneration is the generation last processed by the controller. | Optional: {} |
|
endpoint string |
Endpoint is the in-cluster serving endpoint URL. | Optional: {} |
|
stableRevision RevisionStatus |
StableRevision is the revision currently receiving traffic. | Optional: {} |
|
candidateRevision RevisionStatus |
CandidateRevision is the revision being rolled out, if any. | Optional: {} |
|
failedRevision RevisionStatus |
FailedRevision is a revision that failed to roll out. The controller does not automatically retry it; a spec change (new revision) is required, or setting the decisionmodel.io/retry annotation to a new token re-attempts the same revision. |
Optional: {} |
|
previousRevision PreviousRevisionStatus |
PreviousRevision is the revision that was stable immediately before the most recent promotion. It may linger for a short grace period after promotedAt so the new revision's endpoints populate before it is removed. |
Optional: {} |
|
lastRetryToken string |
LastRetryToken is the value of the decisionmodel.io/retry annotation the controller last consumed to clear a failed revision. |
Optional: {} |
|
storeRecovery StoreRecoveryStatus |
StoreRecovery tracks the bounded retries of lost-store recovery for the stable revision. It is persisted so the attempt bound survives an operator restart (an in-memory counter would reset and allow more recreations than the documented limit). |
Optional: {} |
|
replicas ReplicaStatus |
Replicas reports desired and model-ready replica counts. | Optional: {} |
|
evaluation EvaluationStatus |
Evaluation records the most recent eval-gated rollout result. | Optional: {} |
|
conditions Condition array |
Conditions represent the latest available observations of the object's state. | Optional: {} |
EvaluationResult¶
Underlying type: string
EvaluationResult is the outcome of applying the evaluation gate.
Appears in: - EvaluationStatus
| Field | Description |
|---|---|
Passed |
EvaluationPassed means the candidate met every configured gate. |
Failed |
EvaluationFailed means the candidate failed at least one gate. |
EvaluationSpec¶
EvaluationSpec configures the eval-gated rollout.
Appears in: - RolloutSpec
| Field | Description | Default | Validation |
|---|---|---|---|
datasetRef DatasetRef |
DatasetRef points at the golden dataset (JSONL). | ||
minAccuracy string |
MinAccuracy is the minimum accuracy (decimal string in [0,1], e.g. "0.90") the candidate must reach to be promoted. |
Pattern: ^(0(\.[0-9]+)?\|1(\.0+)?)$ |
|
maxAccuracyDrop string |
MaxAccuracyDrop is the maximum tolerated accuracy drop vs the stable baseline (decimal string, e.g. "0.02"). Only enforced when a stable revision exists to provide a baseline; empty means no drop constraint. |
Pattern: ^(0(\.[0-9]+)?\|1(\.0+)?)$ Optional: {} |
|
maxCases integer |
MaxCases caps how many dataset cases are used (first N). | 500 | Maximum: 5000 Minimum: 1 Optional: {} |
maxECE string |
MaxECE is the maximum tolerated Expected Calibration Error (decimal string in [0,1], e.g. "0.10"). Empty disables the absolute ECE gate. |
Pattern: ^(0(\.[0-9]+)?\|1(\.0+)?)$ Optional: {} |
|
maxECEIncrease string |
MaxECEIncrease is the maximum tolerated ECE increase vs the stable baseline (decimal string). Only enforced when a stable revision exists to provide a baseline; empty disables the relative ECE gate. |
Pattern: ^(0(\.[0-9]+)?\|1(\.0+)?)$ Optional: {} |
|
scoreTolerance string |
ScoreTolerance is the correctness band for "score" questions (decimal string): a predicted expected-value level counts as correct when it is within this many levels of the golden integer level. A case may override it per question via a "tolerance" map. Empty means the default (0.5, "rounds to the expected level"). |
Pattern: ^[0-9]+(\.[0-9]+)?$ Optional: {} |
EvaluationStatus¶
EvaluationStatus records the outcome of an eval-gated rollout evaluation.
Appears in: - DecisionModelStatus
| Field | Description | Default | Validation |
|---|---|---|---|
revision string |
Revision is the revision hash the evaluation was run against. | ||
accuracy string |
Accuracy is the candidate accuracy (decimal string). | ||
baselineAccuracy string |
BaselineAccuracy is the stable revision's accuracy for this dataset, if known. | ||
cases integer |
Cases is the number of questions scored. | ||
failedCases integer |
FailedCases is the number of questions that were wrong or unanswerable. | ||
calibratedCases integer |
CalibratedCases is the number of scored questions that yielded a usable probability distribution and therefore contributed to ECE/Brier. A calibration gate (maxEce / maxEceIncrease) does not pass when this is 0: ECE over an empty set is 0, which would otherwise look perfectly calibrated. |
Optional: {} |
|
ece string |
ECE is the candidate's expected calibration error (decimal string). | ||
brier string |
Brier is the candidate's Brier score (decimal string). | ||
baselineEce string |
BaselineECE is the stable revision's ECE for this dataset, if known. | ||
minAccuracy string |
MinAccuracy echoes the accuracy floor the gate applied, if any. | Optional: {} |
|
maxAccuracyDrop string |
MaxAccuracyDrop echoes the accuracy-drop limit the gate applied, if any. | Optional: {} |
|
maxEce string |
MaxECE echoes the absolute ECE limit the gate applied, if any. | Optional: {} |
|
maxEceIncrease string |
MaxECEIncrease echoes the relative ECE-increase limit the gate applied, if any. | Optional: {} |
|
result EvaluationResult |
Result is the gate outcome: Passed or Failed. | Enum: [Passed Failed] Optional: {} |
|
reason string |
Reason is the gate message (e.g. the failing comparison), human-readable. | Optional: {} |
|
policyHash string |
PolicyHash is a hash of the effective evaluation policy (thresholds, datasetRef, maxCases) this result was produced under. A parked candidate in AwaitingPromotion whose current policy hash differs is re-evaluated rather than promoted on the stale result. |
Optional: {} |
|
scorerVersion integer |
ScorerVersion is the scoring/calibration implementation version this result was produced by. It is part of policyHash (and so approvalId), so an operator upgrade that changes scoring re-evaluates a parked candidate. |
Optional: {} |
|
datasetDigest string |
DatasetDigest is the sha256 (bare hex) of the dataset bytes this result was produced from. Together with policyHash it is the result's identity: if the dataset content changes (same ConfigMap/Secret, edited in place) a parked candidate is re-evaluated rather than promoted on the stale result. |
Optional: {} |
|
approvalId string |
ApprovalID is the identity a manual approval must name: the first 12 hex of sha256(revisionHash + policyHash + datasetDigest). Set the annotation decisionmodel.io/promote to this value to promote. It changes whenever the revision, policy or dataset changes, so an approval cannot carry over a re-evaluation. |
Optional: {} |
|
completedAt Time |
CompletedAt is when the evaluation finished. | Optional: {} |
PreviousRevisionStatus¶
PreviousRevisionStatus records the revision demoted by the latest promotion and when the promotion happened, so its workloads may linger for a grace period before being garbage-collected.
Appears in: - DecisionModelStatus
| Field | Description | Default | Validation |
|---|---|---|---|
hash string |
Hash is the demoted revision's hash. | ||
promotedAt Time |
PromotedAt is when the newer revision was promoted (this one demoted). | Optional: {} |
|
revision RevisionStatus |
Revision is the demoted revision's full recorded identity (engine, model, digest, device, image, resources, placement). It is retained so that, if the new stable turns out unhealthy during the stabilization window, the operator can promote this revision back to stable and render its workloads from its own recorded state rather than from a live Deployment that may already be gone. |
Optional: {} |
PromotionPolicy¶
Underlying type: string
PromotionPolicy selects how a model-ready candidate is promoted.
Appears in: - RolloutSpec
| Field | Description |
|---|---|
Automatic |
PromotionAutomatic promotes a candidate as soon as it is model-ready (and, when evaluation is configured, has passed the gate). |
EvaluationGated |
PromotionEvaluationGated requires rollout.evaluation and lets the gate decide promotion. Rejected by CEL when evaluation is unset. |
Manual |
PromotionManual holds a passed candidate in AwaitingPromotion until a human approves it via the decisionmodel.io/promote annotation. |
ReplicaStatus¶
ReplicaStatus reports desired and model-ready replica counts.
Appears in: - DecisionModelStatus
| Field | Description | Default | Validation |
|---|---|---|---|
desired integer |
Desired is the desired number of serving replicas. | Optional: {} |
|
modelReady integer |
ModelReady is the number of replicas that passed the model readiness gate. | Optional: {} |
RevisionStatus¶
RevisionStatus records the resolved, model-affecting identity of a revision.
Model-affecting fields (engine, model, digest, device, image, resources) together reproduce the revision hash, so a stable revision can be rendered from its own recorded state rather than the current spec (which may already describe a different, pending candidate).
Appears in: - DecisionModelStatus - PreviousRevisionStatus
| Field | Description | Default | Validation |
|---|---|---|---|
hash string |
Hash is the short revision hash of the model-affecting spec fields. | ||
engine string |
Engine is the serving runtime for this revision. | Optional: {} |
|
model string |
Model is the model name for this revision. | ||
digest string |
Digest is the resolved immutable digest for this revision. | ||
device string |
Device is the target device for this revision. | ||
image string |
Image is the resolved serving container image for this revision. | Optional: {} |
|
runtimeVersion string |
RuntimeVersion is the resolved engine runtime version for this revision (e.g. "0.10.0"). Empty when the image was user-set (spec.image), because the version is then unknown. Recorded so a stable revision keeps its runtime version across operator upgrades under --runtime-version-policy=Pinned. |
Optional: {} |
|
resources ResourceRequirements |
Resources are the compute resource requirements of this revision's serving container. |
Optional: {} |
|
precision string |
Precision is the quantization level reported by the engine (e.g. F32, F16). | ||
placement string |
Placement is a short hash of spec.scheduling (nodeSelector, tolerations, affinity, runtimeClassName) as it was when this revision was created, or "none" when no scheduling was set. Empty only on a revision recorded by an older operator version, which the controller adopts on its first reconcile. Placement is part of a revision's identity: a change of scheduling starts a new revision (blue-green, own store) instead of rolling the running Deployment in place. A hash is recorded, not the spec, because Affinity is a very large schema and this type appears three times in the CRD. |
Optional: {} |
|
reason string |
Reason is a short machine reason for why this revision failed. Set only on status.failedRevision; empty on stable/candidate. |
Optional: {} |
|
message string |
Message is a human-readable explanation of a failure. Set only on status.failedRevision; empty on stable/candidate. |
Optional: {} |
|
failedAt Time |
FailedAt is when this revision was recorded as failed. Set only on status.failedRevision; nil on stable/candidate. |
Optional: {} |
RolloutSpec¶
RolloutSpec configures rollout behaviour.
Appears in: - DecisionModelSpec
| Field | Description | Default | Validation |
|---|---|---|---|
evaluation EvaluationSpec |
Evaluation gates promotion on a golden-dataset accuracy check. When unset, a candidate is promoted as soon as all its Pods are model-ready. |
Optional: {} |
|
stabilization Duration |
Stabilization keeps the previous revision's Deployment running (scaled to its replicas, out of the Service) for this long after a promotion, so the operator can switch traffic back instantly if the new stable turns out unhealthy during the window. Default 5m when unset; "0" disables it (the previous revision is removed after a short endpoint-gap grace, as before). Must be between 1m and 24h when set to a non-zero value. |
MaxLength: 32 Pattern: ^([0-9]+(\.[0-9]+)?(ns\|us\|ms\|s\|m\|h))+$ Type: string Optional: {} |
|
promotion PromotionPolicy |
Promotion selects how a candidate that passed ModelReady is promoted: - Automatic: promote as soon as the candidate is model-ready (and, when evaluation is configured, has passed the gate). - EvaluationGated: like Automatic but requires rollout.evaluation to be set (rejected by CEL otherwise); the gate decides promotion. - Manual: hold the candidate in AwaitingPromotion until a human sets the annotation decisionmodel.io/promote to status.evaluation.approvalId (or, without evaluation, the candidate's revision hash for one release) (evaluation still runs when configured). When unset the effective policy is EvaluationGated if rollout.evaluation is set, else Automatic. The deprecated manualPromotion:true is an alias for Manual; setting both promotion and manualPromotion:true to disagreeing values is rejected by CEL. |
Enum: [Automatic EvaluationGated Manual] Optional: {} |
|
manualPromotion boolean |
ManualPromotion holds a candidate that passed its gate (model-ready, plus evaluation when configured) in phase AwaitingPromotion until a human approves it by setting the annotation decisionmodel.io/promote to status.evaluation.approvalId (or, without evaluation, the candidate's revision hash for one release). The stable revision keeps serving meanwhile. There is no progress timeout while waiting. The very first revision of a DecisionModel (no stable revision yet) is promoted without approval, since there is no traffic to protect. The approval annotation is removed once the promotion has been persisted. Deprecated: use promotion: Manual. manualPromotion:true keeps working as an alias for promotion: Manual. |
false | Optional: {} |
timeouts RolloutTimeouts |
Timeouts overrides the progress timeouts of a rollout. Unset fields keep the built-in defaults. Large models (tens of GB) need more than the defaults on a cold node. Not part of the revision hash; a change applies to the phase timeouts immediately, but a prefetch Job that already exists keeps the deadline it was created with. |
Optional: {} |
RolloutTimeouts¶
RolloutTimeouts overrides the progress timeouts of a rollout. Each value must be between 1m and 24h.
Appears in: - RolloutSpec
| Field | Description | Default | Validation |
|---|---|---|---|
caching Duration |
Caching bounds the Caching phase and is also used as the prefetch Job's activeDeadlineSeconds. Default 30m. |
MaxLength: 32 Pattern: ^([0-9]+(\.[0-9]+)?(ns\|us\|ms\|s\|m\|h))+$ Type: string Optional: {} |
|
starting Duration |
Starting bounds the Starting phase (candidate Pods becoming model-ready) and, when set, also the background model warmup on each Pod. When unset the phase timeout is 10m and the warmup bound stays 2m. Default 10m. |
MaxLength: 32 Pattern: ^([0-9]+(\.[0-9]+)?(ns\|us\|ms\|s\|m\|h))+$ Type: string Optional: {} |
|
evaluating Duration |
Evaluating bounds the Evaluating phase and each golden-dataset run. Default 10m. |
MaxLength: 32 Pattern: ^([0-9]+(\.[0-9]+)?(ns\|us\|ms\|s\|m\|h))+$ Type: string Optional: {} |
SchedulingSpec¶
SchedulingSpec is a passthrough of standard Pod scheduling controls.
Appears in: - DecisionModelSpec
| Field | Description | Default | Validation |
|---|---|---|---|
nodeSelector object (keys:string, values:string) |
NodeSelector is a selector which must be true for the Pod to fit on a node. | Optional: {} |
|
tolerations Toleration array |
Tolerations allow the Pod to schedule onto nodes with matching taints. | Optional: {} |
|
affinity Affinity |
Affinity constrains Pod scheduling. | Optional: {} |
|
runtimeClassName string |
RuntimeClassName selects the RuntimeClass for serving Pods, e.g. "nvidia" when the NVIDIA GPU Operator does not make it the default runtime. It is applied to serving Pods only (the prefetch Job needs no GPU runtime). Because it is part of spec.scheduling, changing it changes the revision's placement hash and therefore starts a new revision (blue-green), like any other scheduling change. |
MaxLength: 253 MinLength: 1 Pattern: ^[a-z0-9]([-a-z0-9.]*[a-z0-9])?$ Optional: {} |
StoreRecoveryStatus¶
StoreRecoveryStatus is the bounded-retry bookkeeping for lost-store recovery of the stable revision. It survives an operator restart so the attempt bound is honoured across restarts.
Appears in: - DecisionModelStatus
| Field | Description | Default | Validation |
|---|---|---|---|
attempts integer |
Attempts is the number of failed recovery prefetch Jobs counted so far toward the retry bound. |
Optional: {} |
|
lastFailedJob string |
LastFailedJob is the UID of the last failed prefetch Job already counted, so a repeated observation of the same Job (e.g. a stale cache read) is not double-counted. |
Optional: {} |
|
exhausted boolean |
Exhausted is true once Attempts reached the bound: recovery stopped with Degraded=StorePrefetchFailed and will not recreate the prefetch Job again until a new revision or a decisionmodel.io/retry token resets it. |
Optional: {} |