# rocky-vulkan-llama GPU-accelerated **LLM prefill** on the RK3588 **Mali-G610** via the open Vulkan stack (panthor kernel driver + Mesa **panvk** + llama.cpp Vulkan backend). Goal: make the Mali GPU a useful **prompt-processing (prefill) accelerator** for local agents on RK3588 boards (boltzmann = Radxa ROCK 5 ITX+; ampere = CoolPi CM5 GenBook), so a ~5k-token agent context doesn't cost ~150–250 s of CPU prefill before the model starts generating. > Decode (token generation) is memory-bandwidth-bound (~28 GB/s shared LPDDR5) and the GPU can't > help it. **Prefill is compute-bound** — that's the only place the GPU's FLOPS matter. This repo > is about prefill. ## Dependency stack | Layer | Requirement | Notes | |-------|-------------|-------| | Kernel | **panthor** DRM driver (mainline ≥ 6.x), GPU node bound to panthor, **not** the vendor `mali` kbase blob | On boltzmann: boot `linux-rk3588-marfrit-A1` (mainline 7.0.0-rc3). The vendor BSP kernel (`6.1.75-npu-port`) binds `mali` → panvk cannot enumerate. | | CSF firmware | `/lib/firmware/arm/mali/arch10.10/mali_csffw.bin` | Present on boltzmann. | | Userspace GL/Vulkan | **Mesa panvk** (`libvulkan_panfrost.so`), Mesa **26.1.3** | `panfrost_icd.json` ICD. Enumerates `Mali-G610 MC4 (panvk)` with `uma:1 fp16:1 int dot:1 matrix cores:none`. | | Runtime env | `XDG_RUNTIME_DIR` must be set | panvk init fails without it. | | Inference | **llama.cpp** Vulkan backend (`-DGGML_VULKAN=ON`), commit `389ff61` | Build tools: `glslc`, `libvulkan`, cmake. | ## Status (2026-07-11) - ✅ Full open stack works: panthor binds the GPU, panvk enumerates it, llama.cpp runs on it. - ⚠️ **Blocker:** panvk throws `vk::CommandBuffer::end: ErrorOutOfDeviceMemory` on any prefill dispatch above **ubatch 64**. Only `-ub 64` runs; at ub64 prefill is **21 t/s**, which is *overhead-bound* (ub16/32 far slower) — i.e. 21 t/s is the "forced-tiny-chunk" ceiling, not the Mali GEMM ceiling. It ties an (i8mm-disabled) CPU. - 🔬 Root cause narrowed to a panvk driver allocation that scales past ub64; see [docs/root-cause.md](docs/root-cause.md). The **fix is a Mesa panvk patch** (this repo). ## Contents - `patches/` — collected patches (llama.cpp + Mesa panvk). - `docs/` — root-cause analysis, benchmark data, build/boot procedure. - `mesa/` — Mesa 26.1.3 source (gitignored tarball) for the panvk build. ## Benchmarks (Qwen2.5-3B-Q4_K_M, boltzmann, pp512) | Backend | t/s | Notes | |---------|-----|-------| | CPU (plain build, 8 threads) | 22.4 | `+noi8mm` handicap; i8mm would be higher | | GPU (panvk Vulkan, `-ub 64`) | 21.0 | ub≥72 → OOM crash | | GPU (panvk, ub128+) | crash | `vk::CommandBuffer::end: ErrorOutOfDeviceMemory` | ## Prior art (why the Mali is "framebuffer-only") Arm chose no matrix cores (int8 dotprod on vector ALUs); panvk is graphics-first (Mesa 26.1 sprint = Proton/gaming; Vulkan 1.2 conformance only 2025); llama.cpp Vulkan is ~15× slower than CPU on Mali+Adreno and unsolved upstream ([ggml #9464]). But **MLC-LLM/TVM** (auto-tuned OpenCL kernels) get decent Mali perf → the hardware is capable; the gap is tuned kernels + driver maturity. [ggml #9464]: https://github.com/ggml-org/llama.cpp/discussions/9464