From 881b118158e82895f0209e0f79c3213913e14d77 Mon Sep 17 00:00:00 2001 From: "Claude (noether)" Date: Sat, 11 Jul 2026 14:58:50 +0200 Subject: [PATCH] Add fix plan: cleanup_frequency=10 is the accumulation trigger; consumer-side bounding is the primary lever Co-Authored-By: Claude Opus 4.8 --- docs/fix-plan.md | 71 ++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 71 insertions(+) create mode 100644 docs/fix-plan.md diff --git a/docs/fix-plan.md b/docs/fix-plan.md new file mode 100644 index 0000000..43cbee6 --- /dev/null +++ b/docs/fix-plan.md @@ -0,0 +1,71 @@ +# Fix plan — panvk priv-VA-heap exhaustion (ub>64 OOM) + +## Confirmed root cause (no longer a guess) +`vk::CommandBuffer::end: ErrorOutOfDeviceMemory` at ub>64 is **cross-command-buffer accumulation +in panvk's device-wide 4 GB priv VA heap**: +- llama.cpp `ggml-vulkan.cpp:2616` uses `cleanup_frequency = 10` — it lets **10 command buffers + pile up** (each submitted, none freed) before `resetCommandPool` frees their panvk allocations. +- Each cmdbuf retains, in the device `tls` pool (VA from `priv_heap = [32 MB, 4 GB)`): + - WLS (per-dispatch scratch) — **fixed** via grow-only reuse (patch 0002, verified 644/46). + - **TLS stack** (register spill), one per cmdbuf at `emit_tls()`/`EndCommandBuffer`: **304 MB** for + the L-variant matmul shader (`tls_size` 5328 B/thread) vs 38 MB at ub64 (M shader). +- ~10 cmdbufs × ~304 MB TLS (+ WLS) ≈ >4 GB → OOM. ub64 uses the M shader (38 MB) so 10 cmdbufs ≈ + 400 MB — fits. That's the exact ub64 cliff. +- Compounded by a ~5× overallocation: WLS/TLS sized `× core_id_range = 19` (sparse mask `0x50005`, + 4 real cores). + +## Strategy: two complementary levers +- **Lever A (primary, consumer side, easy): bound cmdbuf accumulation.** Reset the pool before its + retained VA can exceed the heap. This directly removes the mechanism. +- **Lever B (secondary, driver side, headroom): shrink per-cmdbuf footprint.** WLS reuse (done); + optionally the ×19 core overallocation. + +## Phased plan (reboots into panthor are free — instrument-driven) + +### Phase 1 — Validate Lever A + measure the heap (1 panthor boot) +1. Instrument panvk `panvk_as_alloc`(priv_heap): log running reserved bytes + BO count per alloc, + and dump peak at the failing alloc. Confirms peak VA and #live cmdbufs at OOM. +2. Patch llama.cpp: make `cleanup_frequency` an env (`GGML_VK_CMD_CLEANUP_FREQ`), default unchanged. + Rebuild llama.cpp (one TU, fast). +3. Test: `GGML_VK_CMD_CLEANUP_FREQ=3` (+ WLS reuse) at ub128/256/512. **Expected: no crash.** + Sweep 2/3/4/6 to find the max that fits. +- **Exit criterion:** ub≥128 runs crash-free with a small cleanup freq. If not, the accumulation + model is wrong — fall back to Phase 4 (structural panvk). + +### Phase 2 — Make the fix principled (not a magic constant) +- Drive cleanup by **accumulated VA**, not buffer count: track per-cmdbuf estimated tls-pool bytes; + trigger `resetCommandPool` when the running sum would exceed a safe fraction (e.g. 60%) of the + device priv-heap size (query it). Auto-engages only on small-VA devices (panvk); mainline GPUs + with huge VA keep the count-based path. Env override for tuning. +- Keep patch 0002 (WLS reuse) for headroom. + +### Phase 3 — Measure the payoff + correctness +- With ub128/256/512 working, measure prefill t/s at large ubatch. **Hypothesis:** large ubatch is + far less overhead-bound than ub64, so throughput rises well above the 21 t/s ub64 ceiling. Sweep + ub to find the sweet spot; compare vs the i8mm-enabled CPU. +- Correctness: run a real completion (not just llama-bench) through llama-cli on the GPU; verify + coherent output (not garbage) — the WLS/TLS reuse must not corrupt scratch. +- Regression: ub64 still works; run on a second (larger) model to confirm generality. + +### Phase 4 — Deeper panvk fixes (only if Phase 1 fails or headroom too tight) +- **×19 → dense core-id remap:** biggest single lever (304 MB→64 MB). Investigate whether panvk can + index WLS/TLS by a dense 0..N-1 core number via the affinity/core mask instead of the sparse HW + id range. Hard (HW-indexed) — spike first, drop if infeasible. +- **Post-submit pool BO recycling** in panvk: free big `tls`-pool BOs once a cmdbuf's fence signals, + instead of waiting for reset. Structural; upstream-worthy. + +### Phase 5 — Second-model (Fable) review +- Review the final patch set (llama.cpp cleanup + panvk WLS reuse) for correctness, VA-safety + (no dangling scratch across submits), and upstream-readiness. Non-skippable. + +### Phase 6 — Package, document, upstream +- Update `rocky-vulkan-llama`: final patches, benchmarks, build recipe. +- Co-installable artifact (model the `mesa-panvk-bifrost` PKGBUILD): patched llama.cpp + patched + panvk to `/usr/lib/panvk-rocky/`, opt-in ICD. Deploy to ampere (same G610) + boltzmann. +- Upstream candidates: WLS-reuse (closes a real Mesa TODO); cleanup-by-VA (legit small-VA-device + fix for llama.cpp Vulkan on Mali/Adreno — the unsolved ggml #9464). + +## Effort / risk +- Phase 1–3 (the likely-complete fix): ~1–2 panthor boots, mostly llama.cpp-side + the existing + panvk WLS patch. Low risk, high confidence given the confirmed mechanism. +- Phase 4 is the fallback if consumer-side bounding proves insufficient — higher effort, deferred.