Spike 007 — GGUF (llama.cpp runner) models under the operator¶
Tested 2026-10-06 on macOS arm64 (OrbStack, 11.7 GiB Docker) with
ghcr.io/ollaya-dev/ollaya:0.10.0 (CPU). Resolves ARCHITECTURE §13 open question 2:
do ONNX-based facts (readiness gate, digest, device) hold for GGUF models served by
Ollaya's llama.cpp runner?
Model chosen¶
jevk5:latest — the smallest GGUF model in the default registry. Manifest layer sizes
probed via GET https://ollaya.dev/v2/library/<model>/manifests/<tag>:
| model | GGUF weights layer | note |
|---|---|---|
| jevk5:latest | 4.48 GB | smallest; chosen |
| winnow:e4b | 8.01 GB | |
| jeb:latest | 9.79 GB | |
| cygnet:latest | 12.67 GB |
No smaller GGUF tag exists (probed winnow:e2b/e1b/1b/q4, jevk5:1b/2b/3b/mini/small/q4/e2b,
jeb:mini/1b/small/q4/e2b, cygnet:small — all 404; the registry does not expose a tag list).
All GGUF manifests carry the same layer set: application/vnd.ollaya.weights.gguf plus tiny
.decision, .calibration, .license layers. jevk5:latest at 4.48 GB fits a 16 GB CI runner
(ubuntu-latest) with room for the operator and control plane.
Facts (jevk5:latest, Q8_0, CPU, 0.10.0)¶
| Fact | Result |
|---|---|
| Pull into RW store | exit 0, ~374 s (4.48 GB), 4.2 GB on disk |
| Manifest digest on disk == registry sha256 | yes — 58b5d6516410070d5f165d77357ed7bfefe5ffedde63e6ae75ab5101d4fe0221 (1767-byte manifest) |
Serve from RO store + writable .ollaya emptyDir + readOnlyRootFilesystem |
works, GET / → 200 (same layout as ONNX; spike 006) |
Warmup /api/decide keep_alive:-1 |
done_reason:"load", cold load ~18 s |
/api/ps name |
jevk5:latest |
/api/ps digest |
58b5d651…0221 (== manifest sha256, bare hex) |
/api/ps device |
"cpu" |
/api/ps precision (details.quantization_level) |
"Q8_0" (quantization, not F32/F16) |
/api/ps expires_at after keep_alive:-1 |
null → Pinned=True |
/api/ps details.format |
"gguf" |
Memory (cgroup anon, loaded) |
5 207 502 848 B = 4966 MiB ≈ 4.85 GiB |
/v1/systemone (evaluator path) |
works — a choice question returned choice:"cancel", confidence 0.9785, full probabilities |
Does our gate accept it?¶
The readiness gate (internal/controller/readiness.go, with Inspect in
internal/engine/ollaya/ollaya.go) checks exactly three things against the recorded revision:
loaded.Digest != rev.Digest→DigestMismatch. GGUF digest = sha256(manifest), same scheme as ONNX → passes.loaded.Device != rev.Device→DeviceMismatch. GGUF reports"cpu"(normalized bynormalizeDevice, which only collapsescuda:<n>→cuda) → passes on a CPU revision.!loaded.Pinned(i.e.expires_at != null) →NotPinned. keep_alive:-1 setsexpires_at:null→ passes.
Precision is recorded but NOT gated. Inspect copies details.quantization_level into
Loaded.Precision and the controller stores it in status.stableRevision.precision, but no
condition compares it. So Q8_0 (or any GGUF quant) never fails the gate. This is the key
finding: the gate is precision-agnostic and therefore correct for GGUF out of the box.
CUDA fallback (CPU image, OLLAYA_DEVICE=cuda)¶
OLLAYA_DEVICE=cuda on the CPU image with jevk5:latest: server starts (GET / → 200), but the
model fails to load — MODEL_LOAD_FAILED: llama.cpp finds no device CUDA0 (devices: []); the
CUDA libraries are not installed. /api/ps → {"models":[]}. The model never appears, so the
gate never sees a loaded entry → the Pod's readiness gate stays False and the probe eventually
times out. Same safe outcome as the ONNX CUDA-misconfig case (spike 001): a naive TCP probe
would pass, the /api/ps gate does not. (The failure here is a hard load error rather than a
silent CPU fallback, so there is no wrong-device entry to flag; the empty /api/ps is enough.)
End-to-end on kind¶
Deployed the operator (kustomize, default image now 0.10.0 from the runtime-bump merge) and
applied a DecisionModel jevk5:latest, device: cpu, replicas: 1, resources.requests
.memory: 5120Mi/limits.memory: 6Gi, cache.size: 10Gi:
- phase: Resolving → Caching ~4 min (prefetch Job pulled 4.48 GB, completed in 4m14s) → Starting ~1.5 min (cold load) → Ready at ~5m18s total.
- conditions:
Resolved=True, Cached=True, ModelReady=True, Ready=True, Degraded=False. status.stableRevision: digest58b5d651…0221, devicecpu, precisionQ8_0, imageghcr.io/ollaya-dev/ollaya:0.10.0.- owned objects: Deployment 1/1, Service :11435, PVC 10Gi Bound, prefetch Job Complete.
- Default phase timeouts (Caching 30m, Starting 10m) were comfortably sufficient for this
4.48 GB model; a larger GGUF (winnow/jeb/cygnet at 8–13 GB, or the 9B-class ~18–19 GB models)
would need the raised
spec.rollout.timeoutsalready documented in sizing.md.
Verdict¶
| Aspect | Verdict |
|---|---|
| Readiness gate (digest/device/pinned) | works as-is — no change needed; the gate is precision-agnostic |
| Digest scheme (sha256 of manifest) | works as-is — identical to ONNX |
Serving from RO store, readOnlyRootFilesystem |
works as-is |
/v1/systemone evaluator path |
works as-is |
| CUDA misconfig on CPU image | safe as-is — model fails to load, /api/ps empty, gate stays False |
Memory default in resources.go |
docs only — jevk5 is not in the measured table; add a sizing row (below). No code change required: an unmeasured model simply gets no built-in default, and the user sets spec.resources (as this spike did). A jevk5:latest row could optionally be added to cpuDefaults. |
GGUF models work under the operator with no gate or engine change. The only follow-up is a
documentation sizing row (and optionally a built-in resources.go default entry).
Proposed docs/sizing.md row¶
| Model | Family | Download | Memory (anon, loaded) |
|---|---|---|---|
jevk5:latest |
jevk5 (GGUF, Q8_0) | ~4.48 GB | ~4.85 GiB (4966 MiB) |
Recommended resources (rule of thumb: request ≈ anon rounded up to 512Mi, limit ≈ anon × 1.4):
# jevk5:latest (GGUF Q8_0, anon ~4.85 GiB)
resources:
requests: { cpu: "1", memory: "5120Mi" }
limits: { cpu: "2", memory: "7Gi" }
cache:
size: 6Gi # 4.48 GB download + headroom
Environment¶
- macOS arm64, OrbStack (11.7 GiB Docker).
ghcr.io/ollaya-dev/ollaya:0.10.0(arm64, CPU).- kind v0.33.0, k8s v1.37.0 for the end-to-end run.
- All containers, volumes, images, and the kind cluster removed after testing;
config/manager/kustomization.yaml(mutated bymake deploy) restored.