881b118158
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
4.6 KiB
4.6 KiB
Fix plan — panvk priv-VA-heap exhaustion (ub>64 OOM)
Confirmed root cause (no longer a guess)
vk::CommandBuffer::end: ErrorOutOfDeviceMemory at ub>64 is cross-command-buffer accumulation
in panvk's device-wide 4 GB priv VA heap:
- llama.cpp
ggml-vulkan.cpp:2616usescleanup_frequency = 10— it lets 10 command buffers pile up (each submitted, none freed) beforeresetCommandPoolfrees their panvk allocations. - Each cmdbuf retains, in the device
tlspool (VA frompriv_heap = [32 MB, 4 GB)):- WLS (per-dispatch scratch) — fixed via grow-only reuse (patch 0002, verified 644/46).
- TLS stack (register spill), one per cmdbuf at
emit_tls()/EndCommandBuffer: 304 MB for the L-variant matmul shader (tls_size5328 B/thread) vs 38 MB at ub64 (M shader).
- ~10 cmdbufs × ~304 MB TLS (+ WLS) ≈ >4 GB → OOM. ub64 uses the M shader (38 MB) so 10 cmdbufs ≈ 400 MB — fits. That's the exact ub64 cliff.
- Compounded by a ~5× overallocation: WLS/TLS sized
× core_id_range = 19(sparse mask0x50005, 4 real cores).
Strategy: two complementary levers
- Lever A (primary, consumer side, easy): bound cmdbuf accumulation. Reset the pool before its retained VA can exceed the heap. This directly removes the mechanism.
- Lever B (secondary, driver side, headroom): shrink per-cmdbuf footprint. WLS reuse (done); optionally the ×19 core overallocation.
Phased plan (reboots into panthor are free — instrument-driven)
Phase 1 — Validate Lever A + measure the heap (1 panthor boot)
- Instrument panvk
panvk_as_alloc(priv_heap): log running reserved bytes + BO count per alloc, and dump peak at the failing alloc. Confirms peak VA and #live cmdbufs at OOM. - Patch llama.cpp: make
cleanup_frequencyan env (GGML_VK_CMD_CLEANUP_FREQ), default unchanged. Rebuild llama.cpp (one TU, fast). - Test:
GGML_VK_CMD_CLEANUP_FREQ=3(+ WLS reuse) at ub128/256/512. Expected: no crash. Sweep 2/3/4/6 to find the max that fits.
- Exit criterion: ub≥128 runs crash-free with a small cleanup freq. If not, the accumulation model is wrong — fall back to Phase 4 (structural panvk).
Phase 2 — Make the fix principled (not a magic constant)
- Drive cleanup by accumulated VA, not buffer count: track per-cmdbuf estimated tls-pool bytes;
trigger
resetCommandPoolwhen the running sum would exceed a safe fraction (e.g. 60%) of the device priv-heap size (query it). Auto-engages only on small-VA devices (panvk); mainline GPUs with huge VA keep the count-based path. Env override for tuning. - Keep patch 0002 (WLS reuse) for headroom.
Phase 3 — Measure the payoff + correctness
- With ub128/256/512 working, measure prefill t/s at large ubatch. Hypothesis: large ubatch is far less overhead-bound than ub64, so throughput rises well above the 21 t/s ub64 ceiling. Sweep ub to find the sweet spot; compare vs the i8mm-enabled CPU.
- Correctness: run a real completion (not just llama-bench) through llama-cli on the GPU; verify coherent output (not garbage) — the WLS/TLS reuse must not corrupt scratch.
- Regression: ub64 still works; run on a second (larger) model to confirm generality.
Phase 4 — Deeper panvk fixes (only if Phase 1 fails or headroom too tight)
- ×19 → dense core-id remap: biggest single lever (304 MB→64 MB). Investigate whether panvk can index WLS/TLS by a dense 0..N-1 core number via the affinity/core mask instead of the sparse HW id range. Hard (HW-indexed) — spike first, drop if infeasible.
- Post-submit pool BO recycling in panvk: free big
tls-pool BOs once a cmdbuf's fence signals, instead of waiting for reset. Structural; upstream-worthy.
Phase 5 — Second-model (Fable) review
- Review the final patch set (llama.cpp cleanup + panvk WLS reuse) for correctness, VA-safety (no dangling scratch across submits), and upstream-readiness. Non-skippable.
Phase 6 — Package, document, upstream
- Update
rocky-vulkan-llama: final patches, benchmarks, build recipe. - Co-installable artifact (model the
mesa-panvk-bifrostPKGBUILD): patched llama.cpp + patched panvk to/usr/lib/panvk-rocky/, opt-in ICD. Deploy to ampere (same G610) + boltzmann. - Upstream candidates: WLS-reuse (closes a real Mesa TODO); cleanup-by-VA (legit small-VA-device fix for llama.cpp Vulkan on Mali/Adreno — the unsolved ggml #9464).
Effort / risk
- Phase 1–3 (the likely-complete fix): ~1–2 panthor boots, mostly llama.cpp-side + the existing panvk WLS patch. Low risk, high confidence given the confirmed mechanism.
- Phase 4 is the fallback if consumer-side bounding proves insufficient — higher effort, deferred.