Spike 002 — Second/third model families + resource sizing¶
Goal: prove the Ollaya engine (internal/engine/ollaya) is not Laya-specific, and give users
sizing numbers per model.
Environment¶
- Host: macOS (arm64 / Apple Silicon), Docker via OrbStack. Image
ghcr.io/ollaya-dev/ollaya:0.7.3(native arm64). Serving env identical toServingPodSpec(OLLAYA_HOST=0.0.0.0:11435,OLLAYA_MODELS=/models,OLLAYA_DEVICE=cpu,OLLAYA_KEEP_ALIVE=-1), read-only model store at/models, writable emptyDir at/home/ollaya/.ollaya,--user 1000:1000,fsGroup:1000simulated by chowning the volumes. - Each model: prefetch (the
PrefetchJobSpecscript —ollaya pullthensha256sumof$OLLAYA_MODELS/$MANIFEST_PATHvs the registry digest), then serve read-only, thenResolve/Inspect/Decidevia the engine, thendocker stats, then cleanup. - Device: CPU only (no GPU on this host); precision therefore F32 for all (matches spike 001).
Model selection (all < 3 GB download; sizes = sum of manifest layer bytes)¶
Probed the registry manifests before pulling (no wasted downloads):
| Model | Family | Runner (layer types) | Download |
|---|---|---|---|
laya:multilingual |
laya | ONNX (graph.onnx) |
683 MB |
nli:latest |
nli | ONNX (graph.onnx) |
884 MB |
gliclass:latest |
gliclass | ONNX (graph.onnx) |
1768 MB |
Picked these three (distinct families, all under the 3 GB cap). The only GGUF / llama.cpp-runner
candidate found, jevk5:latest (layer graph.gguf), is 4482 MB > 3 GB, so it was skipped per
the task; kev:0.8b/decider:0.8b turned out to be ONNX, not GGUF. So the GGUF runner path
remains unverified here (no small GGUF model exists in the public registry today) — flagged for a
later spike if a small GGUF model ships.
Results (per model, CPU / F32)¶
All three worked with the engine unchanged. /api/ps field shapes are identical across
families (name, digest, device, details.quantization_level, expires_at); no per-family
differences.
Data correction (2026-10-02): the "RSS (docker stats)" column below over-reports memory because
docker statscounts the page cache left by the model pull. Re-measured by cgroupanon(/sys/fs/cgroup/memory.stat, loaded withkeep_alive:-1):laya:en~3.09 GiB,laya:multilingual~1.85 GiB,nli:latest~3.62 GiB,gliclass:latest~2.42 GiB. Use docs/sizing.md for the authoritative numbers and the method; the figures in the table below are kept as originally recorded.
| Model | Prefetch digest verify | Resolve == on-disk | /api/ps name / device / precision / pinned | Decide (choice+noul) | Load time | RSS (docker stats) | Decide latency (avg of 5, warm) |
|---|---|---|---|---|---|---|---|
laya:multilingual |
ok (2840…84eb) |
match | laya:multilingual / cpu / F32 / true |
mapped ok | ~2.8 s | ~1.85 GiB | ~72 ms |
nli:latest |
ok (b331…e205) |
match | nli:latest / cpu / F32 / true |
mapped ok | ~5.8 s | ~3.75 GiB | ~242 ms |
gliclass:latest |
ok (8c4c…e5a6) |
match | gliclass:latest / cpu / F32 / true |
mapped ok | ~4.9 s | ~2.66 GiB | ~131 ms |
Notes:
- MANIFEST_PATH computed by the engine (manifests/ollaya.dev/library/<model>/<tag>) matched
the on-disk path for every model; the prefetch digest check passed each time.
- Decide mapped choice (Choice + Probabilities + Confidence) and noul (Noul, no confidence)
correctly for all three; the answer JSON shape is the same across families.
- RSS is peak resident set while a single model is loaded and warm (CPU, F32).
Engine incompatibilities found¶
None. Resolve, Inspect (/api/ps mapping), and Decide (/v1/systemone mapping) all work
across the laya / nli / gliclass families with no code change. No new unit test was required
(the existing table-driven tests already cover the field shapes these models return).
Recommended resources per model¶
Guidance for a CPU / F32 serving Pod. Memory limit ≈ observed RSS + ~40% headroom (model load plus request working set); requests set to the observed RSS. GPU is out of scope here (unverified on this host). PVC size ≥ download size with margin for the writable emptyDir logs.
# laya:multilingual (download ~683 MB, RSS ~1.85 GiB)
resources:
requests: { cpu: "1", memory: "2Gi" }
limits: { cpu: "2", memory: "3Gi" }
# PVC: 2Gi
# gliclass:latest (download ~1.77 GB, RSS ~2.66 GiB)
resources:
requests: { cpu: "1", memory: "3Gi" }
limits: { cpu: "2", memory: "4Gi" }
# PVC: 4Gi
# nli:latest (download ~884 MB, RSS ~3.75 GiB)
resources:
requests: { cpu: "1", memory: "4Gi" }
limits: { cpu: "2", memory: "6Gi" }
# PVC: 2Gi
Rule of thumb for a new ONNX model: memory request ≈ RSS, memory limit ≈ RSS × 1.4,
cpu request 1 / limit 2 for low-latency single-request serving; size the PVC to the download
plus a little for logs. RSS does not track download size — nli downloads less than gliclass
but loads a larger graph into memory, so measure RSS per model rather than inferring it from the
download.
Caveats / unverified¶
- CPU / F32 only; GPU and F16 unverified (no GPU on this arm64 host, as in spike 001).
- GGUF / llama.cpp runner unverified: the only GGUF model in the registry (
jevk5:latest) exceeds the 3 GB cap. Re-run this spike for a small GGUF model when one is available to confirm the engine handles a non-ONNX runner (expected to be transparent — the engine only speaks HTTP). - RSS/latency measured under OrbStack on Apple Silicon; absolute numbers will differ on cluster nodes, but the relative ordering and the sizing rule of thumb should hold.