panvk WLS fix v1 + confirmed mechanism (partial: reuse halves but doesn't eliminate)
Instrument-first panthor boot confirms Fable's diagnosis: ~326 per-dispatch WLS allocs exhaust the sub-4GB priv VA heap at ub>=128. Grow-only single-max reuse (fix v1) halves the pressure (326->715 dispatches) but WLS sizes vary 128x (155KB-19MB) so a single buffer is insufficient. Refined fix pending 2nd review. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,30 @@
|
||||
# Benchmarks & WLS instrumentation (panthor boot 2026-07-11)
|
||||
|
||||
Patched panvk (WLS-reuse, env-gated) loaded via `VK_ICD_FILENAMES`; Qwen2.5-3B-Q4_K_M, pp512.
|
||||
|
||||
## Confirmed: the mechanism is per-dispatch WLS exhausting panvk's sub-4 GB priv VA heap
|
||||
| Case | ubatch | WLS allocs/graph | Result |
|
||||
|------|--------|------------------|--------|
|
||||
| no-reuse | 64 | (freed per reset) | 21.0 t/s ✓ |
|
||||
| no-reuse | 128 | **326** before OOM | crash `CommandBuffer::end: OutOfDeviceMemory` |
|
||||
| **reuse v1** | 128 | 715 before OOM | **still crash** (survives ~2× longer) |
|
||||
| reuse v1 | 256 | 690 | still crash |
|
||||
| reuse v1 | 512 | 432 | still crash |
|
||||
|
||||
WLS per-dispatch sizes vary 128× (155 KB … 19 MB; dominant 9.5 MB), summing ~2 GB/graph without
|
||||
reuse. driver_set/tsd sibling allocs are tiny (descriptors) — not contributors.
|
||||
|
||||
## Status of the fix
|
||||
- **fix v1** (grow-only single-max WLS buffer, `patches/mesa-panvk-0001-*`): confirmed the
|
||||
mechanism and halved the pressure, but **insufficient** — the 128×-varying WLS sizes defeat a
|
||||
single-max buffer (the original TODO warned about exactly this). Refined fix in progress
|
||||
(2nd Fable pass).
|
||||
- ub64 remains the only crash-free mode at 21 t/s (overhead-bound; ties an i8mm-off CPU).
|
||||
|
||||
## How to reproduce (on panthor kernel, patched panvk)
|
||||
```
|
||||
export XDG_RUNTIME_DIR=/run/user/$(id -u) VK_ICD_FILENAMES=$HOME/rocky-panvk-icd.json
|
||||
cd ~/src/llama.cpp/build-vulkan/bin
|
||||
PANVK_WLS_LOG=1 ./llama-bench -m MODEL -p 512 -n 0 -ngl 99 -b 512 -ub 128 # crash, ~326 allocs
|
||||
PANVK_WLS_REUSE=1 PANVK_WLS_LOG=1 ./llama-bench -m MODEL -p 512 -n 0 -ngl 99 -b 512 -ub 128 # v1: still crash
|
||||
```
|
||||
@@ -0,0 +1,49 @@
|
||||
# Fable review — CONFIRMED root cause & fix (2026-07-11)
|
||||
|
||||
Second-model (Fable) review of `root-cause.md`, verified against Mesa 26.1.3 source. It corrected
|
||||
my WLS-*size* hypothesis: right allocation, wrong mechanism.
|
||||
|
||||
## Confirmed mechanism: sub-4 GB priv-VA-heap exhaustion from per-dispatch WLS
|
||||
|
||||
CSF panvk allocates a **fresh WLS (workgroup-local-storage) buffer per `vkCmdDispatch` and never
|
||||
reuses it** — there is a literal TODO admitting this at `csf/panvk_vX_cmd_dispatch.c:99`. The JM
|
||||
backend (`jm/panvk_vX_cmd_dispatch.c:48-57`) and gallium (`pan_job.c:438`) both reuse one WLS per
|
||||
batch; **CSF is the outlier.**
|
||||
|
||||
- On CSF (`PAN_ARCH >= 10`) the VM is created **without AUTO_VA**; all internal/priv BOs draw VA
|
||||
from `priv_heap = [PANVK_VA_RESERVE_BOTTOM (32 MB), 1ull<<32)` ≈ **4064 MB**
|
||||
(`csf/panvk_vX_device.c:466-476`). User `VkDeviceMemory` uses a *separate* ≥4 GB heap — which is
|
||||
why every `GGML_VK_*` data-buffer knob changed nothing.
|
||||
- Pool BOs are freed only at cmdbuf **reset** (`csf/panvk_vX_cmd_buffer.c:818-820`), not at submit.
|
||||
So "submit every 4 nodes still crashes" did **not** prove single-dispatch — the accumulation
|
||||
lives in the driver pool, which submits don't drain. That supports the aggregate diagnosis.
|
||||
- Failure chain: `util_vma_heap_alloc` → 0 → `panvk_as_alloc` (`panvk_device.h:188`) →
|
||||
`panvk_priv_bo.c:62-70` → `VK_ERROR_OUT_OF_DEVICE_MEMORY` → caught by
|
||||
`panvk_catch_indirect_alloc_failure` (`panvk_cmd_alloc.h:24-29`) → deferred to
|
||||
`vkEndCommandBuffer`. Exactly the observed signature, 31 GB RAM untouched.
|
||||
|
||||
## Why exactly 64 → 72
|
||||
`ggml_vk_guess_matmul_pipeline` (ggml-vulkan.cpp:7267-7273): `n≤64` → **M** warptile, `n>64` → **L**.
|
||||
Shared mem: M-mmq 128×33×2 = 8448 B → pow2 → **16 KB**; L-mmq 256×33×2 = 16896 B → **32 KB**.
|
||||
Per-dispatch WLS = `pow2(size) × instances × core_id_range`, with `core_id_range = 19` (live core
|
||||
mask `0x50005` → sparse, highest core ID 18 → ×19, ~5× overallocation by construction). ~10-20 MB
|
||||
(M) vs ~20-40 MB (L) per dispatch. A Qwen-3B ubatch graph records ~300 dispatches: ub64 ≈ 3-3.5 GB
|
||||
(just under 4064 MB); ub72 flips to L → ~2× → over the heap. The shared-mem-fit theory is ruled
|
||||
out — llama.cpp gates L/M on `maxComputeSharedMemorySize` and would silently downgrade, not crash.
|
||||
|
||||
## The fix (implemented): reuse a grow-only WLS buffer per cmdbuf
|
||||
Mirrors JM/gallium. Adds `wls_ptr`/`wls_size` to `struct panvk_tls_state`; a dispatch reuses the
|
||||
buffer when `wls_total_size <= state.tls.wls_size`, else allocates and remembers it. Collapses ~300
|
||||
allocations to ~1 (a few if the size grows M→L). Safe on reset (`panvk_reset_cmdbuf` memsets state).
|
||||
Gated behind `PANVK_WLS_REUSE` for A/B testing; `PANVK_WLS_LOG` traces per-dispatch sizes.
|
||||
See `patches/mesa-panvk-0001-csf-reuse-wls-across-dispatches.patch`.
|
||||
|
||||
Rejected alternatives: (b) lowering advertised `maxComputeSharedMemorySize` — collateral damage to
|
||||
flash-attn/other ops, and the L shader fits 32 KB fine; growing the priv heap — riskier (WLS may
|
||||
not cross a 4 GB boundary, `pan_desc.c:933-935`).
|
||||
|
||||
## Predictions to verify on the panthor boot (instrument-first)
|
||||
1. `PANVK_WLS_LOG` off-reuse: ~300 allocations of ~20-40 MB; failure when running total ≈ 3.9 GB.
|
||||
2. Reuse on: **one** WLS allocation; ub128/256/512 no longer crash.
|
||||
3. Free checks: a **larger model at ub64** should also crash without reuse (more dispatches); ub64
|
||||
without reuse sits measurably close to the 4 GB edge.
|
||||
@@ -0,0 +1,55 @@
|
||||
--- a/b/src/panfrost/vulkan/csf/panvk_cmd_buffer.h 2026-06-18 08:44:14.000000000 +0200
|
||||
+++ b/src/panfrost/vulkan/csf/panvk_cmd_buffer.h 2026-07-11 13:50:15.332900002 +0200
|
||||
@@ -409,6 +409,11 @@
|
||||
struct pan_ptr desc;
|
||||
struct pan_tls_info info;
|
||||
unsigned max_wg_count;
|
||||
+
|
||||
+ /* rocky-vulkan-llama: grow-only WLS buffer reused across dispatch
|
||||
+ * commands (PANVK_WLS_REUSE) instead of one BO per dispatch. */
|
||||
+ uint64_t wls_ptr;
|
||||
+ unsigned wls_size;
|
||||
};
|
||||
|
||||
struct panvk_cond_render_state {
|
||||
--- a/b/src/panfrost/vulkan/csf/panvk_vX_cmd_dispatch.c 2026-06-18 08:44:14.000000000 +0200
|
||||
+++ b/src/panfrost/vulkan/csf/panvk_vX_cmd_dispatch.c 2026-07-11 13:50:15.336233209 +0200
|
||||
@@ -13,6 +13,7 @@
|
||||
#include "genxml/gen_macros.h"
|
||||
|
||||
#include "panvk_buffer.h"
|
||||
+#include <stdio.h>
|
||||
#include "panvk_cmd_alloc.h"
|
||||
#include "panvk_cmd_buffer.h"
|
||||
#include "panvk_cmd_desc_state.h"
|
||||
@@ -100,10 +101,26 @@
|
||||
* instance count) might differ significantly between dispatch commands,
|
||||
* rather than track a single maximum size, we might want to consider
|
||||
* multiple allocations for different size buckets. */
|
||||
- tlsinfo.wls.ptr =
|
||||
- panvk_cmd_alloc_dev_mem(cmdbuf, tls, wls_total_size, 4096).gpu;
|
||||
- if (!tlsinfo.wls.ptr)
|
||||
- return 0;
|
||||
+ /* rocky-vulkan-llama fix (addresses the TODO above): reuse a grow-only
|
||||
+ * WLS buffer across dispatches. CSF allocated a fresh WLS BO per
|
||||
+ * dispatch, exhausting the sub-4GB priv VA heap over a ~300-dispatch
|
||||
+ * LLM prefill graph -> OOM at ub>64. Gated by PANVK_WLS_REUSE (A/B). */
|
||||
+ const bool wls_reuse = getenv("PANVK_WLS_REUSE") != NULL;
|
||||
+ if (wls_reuse && wls_total_size <= cmdbuf->state.tls.wls_size) {
|
||||
+ tlsinfo.wls.ptr = cmdbuf->state.tls.wls_ptr;
|
||||
+ } else {
|
||||
+ tlsinfo.wls.ptr =
|
||||
+ panvk_cmd_alloc_dev_mem(cmdbuf, tls, wls_total_size, 4096).gpu;
|
||||
+ if (!tlsinfo.wls.ptr)
|
||||
+ return 0;
|
||||
+ if (wls_reuse) {
|
||||
+ cmdbuf->state.tls.wls_ptr = tlsinfo.wls.ptr;
|
||||
+ cmdbuf->state.tls.wls_size = wls_total_size;
|
||||
+ }
|
||||
+ }
|
||||
+ if (getenv("PANVK_WLS_LOG"))
|
||||
+ fprintf(stderr, "[panvk-wls] total_size=%u reuse=%d\n",
|
||||
+ wls_total_size, wls_reuse);
|
||||
}
|
||||
|
||||
cmdbuf->state.tls.info.tls.size =
|
||||
Reference in New Issue
Block a user