Add REPLICATION.md: reproduce rocket-NPU 1GHz LLM accel on vanilla kernel+userspace

Full recipe: DT (enable rknn cores, OPP table 300-1000MHz, delete assigned-clock-rates),
driver (OPP config + dev_pm_opp_set_rate 1GHz, autosuspend SError mitigation), ggml-rocket
build/run, results (pp512 46.7 vs vendor 51.4, DDR-bandwidth-bound), correctness gate.
Links rk-llama.cpp @ rocket-npu-backend and linux-rk3588-marfrit @ npuclk-opp.
This commit is contained in:
marfrit
2026-07-14 23:41:44 +02:00
parent a2244675b1
commit 27df00a201
2 changed files with 193 additions and 0 deletions
+9
View File
@@ -112,3 +112,12 @@ Headline: mainline driver name is **`rocket`** (not `rknpu`); it's
already shipped in boltzmann's kernel as a built module. Phase-2
unblock is small (DT enable + IOMMU v1.0 mitigation + modprobe),
not a driver port.
## Replicating this work
See **[REPLICATION.md](REPLICATION.md)** for the full recipe (vanilla kernel + userspace):
kernel DT/driver changes for 1 GHz NPU DVFS, the power-domain SError caveat, and
building/running the `ggml-rocket` backend.
- Userspace: [marfrit/rk-llama.cpp](https://git.reauktion.de/marfrit/rk-llama.cpp) branch `rocket-npu-backend` (K>8192: `rocket-ksplit-wip`)
- Kernel: [marfrit/linux-rk3588-marfrit](https://git.reauktion.de/marfrit/linux-rk3588-marfrit) branch `npuclk-opp`
+184
View File
@@ -0,0 +1,184 @@
# Replicating the RK3588 Rocket-NPU LLM acceleration
Run `llama.cpp` LLM inference on the **RK3588 NPU via the mainline `rocket`
DRM-accel driver** — no vendor `librknnrt` blob — at 1 GHz. Measured on
`boltzmann` (Radxa Rock 5 ITX+, RK3588, 32 GB): gemma-4-E2B-it-Q8_0
**pp512 ≈ 46.7 t/s**, within ~9 % of the vendor closed stack (51.4), both
DDR-bandwidth-bound.
## Repositories
| Component | Repo | Branch | Commit |
|---|---|---|---|
| Userspace — `ggml-rocket` backend | https://git.reauktion.de/marfrit/rk-llama.cpp | `rocket-npu-backend` | `288d70413` |
| Userspace — K>8192 K-segmentation (optional) | https://git.reauktion.de/marfrit/rk-llama.cpp | `rocket-ksplit-wip` | `6d5a075a6` |
| Kernel — 1 GHz OPP DVFS + SError mitigation | https://git.reauktion.de/marfrit/linux-rk3588-marfrit | `npuclk-opp` | `a811e6cc` |
---
## 1. Kernel — from a vanilla RK3588 tree
Base: a mainline-class RK3588 kernel that already carries `drivers/accel/rocket/`
(the upstream Rocket driver by Tomeu Vizoso; present in recent mainline /
Collabora rockchip trees). Reference here: 7.0.0-rc3 (`linux-rk3588-marfrit`),
`CONFIG_DRM_ACCEL_ROCKET=m`, `CONFIG_PM_OPP=y`.
Upstream `rocket` runs the NPU **DT-pinned at 200 MHz, no DVFS**. Two changes
take it to 1 GHz.
### 1a. Device tree (board DTS)
The 3 NPU cores are defined `disabled` in `rk3588-base.dtsi` as
`rknn_core_0/1/2`. In your board DTS, enable them (Jaguar-style override), wire
the supply, attach an OPP table, and **delete `assigned-clock-rates`**:
```dts
&rknn_core_0 {
npu-supply = <&vdd_npu_s0>;
sram-supply = <&vdd_npu_s0>;
operating-points-v2 = <&npu_opp_table>;
/delete-property/ assigned-clock-rates;
status = "okay";
};
/* identical &rknn_core_1, &rknn_core_2 */
&rknn_mmu_0 { status = "okay"; }; /* + _1, _2 */
/ {
npu_opp_table: opp-table-npu {
compatible = "operating-points-v2";
opp-300000000 { opp-hz = /bits/ 64 < 300000000>; opp-microvolt = <700000>; };
opp-400000000 { opp-hz = /bits/ 64 < 400000000>; opp-microvolt = <700000>; };
opp-500000000 { opp-hz = /bits/ 64 < 500000000>; opp-microvolt = <700000>; };
opp-600000000 { opp-hz = /bits/ 64 < 600000000>; opp-microvolt = <700000>; };
opp-700000000 { opp-hz = /bits/ 64 < 700000000>; opp-microvolt = <700000>; };
opp-800000000 { opp-hz = /bits/ 64 < 800000000>; opp-microvolt = <750000>; };
opp-900000000 { opp-hz = /bits/ 64 < 900000000>; opp-microvolt = <800000>; };
opp-1000000000 { opp-hz = /bits/ 64 <1000000000>; opp-microvolt = <850000>; };
};
};
```
Notes:
- The **`npu`-named clock** (`<&scmi_clk SCMI_CLK_NPU>`) is the scalable one;
`aclk/hclk/pclk` are fixed AXI/AHB/APB bus clocks (do not scale with it).
- `vdd_npu_s0` is the board's NPU rail (`rk8602@42`, 0.550.95 V); voltages above
are the vendor OPP values.
- **Deleting `assigned-clock-rates` is mandatory.** Left in, `of_clk_set_defaults()`
forces the clock rate at device-creation — *before* the driver probes and
*before* the i2c PMIC reaches voltage — an undervolt-at-speed hang. Driving the
rate through OPP (below) sequences voltage-then-frequency.
### 1b. Driver — `drivers/accel/rocket/rocket_core.c`, in `rocket_core_init()`
After `devm_clk_bulk_get(...)`, register the OPP config + table and pin 1 GHz:
```c
struct dev_pm_opp_config opp_cfg = {
.clk_names = (const char * const []){ "npu", NULL },
.regulator_names = (const char * const []){ "npu", NULL },
};
devm_pm_opp_set_config(dev, &opp_cfg);
devm_pm_opp_of_add_table(dev);
/* ... after pm_runtime_resume_and_get(dev) succeeds: */
dev_pm_opp_set_rate(dev, 1000000000); /* voltage-first via OPP */
```
And raise the runtime-PM autosuspend delay (SError mitigation — see caveat):
```c
pm_runtime_set_autosuspend_delay(dev, 3600000); /* was 50 */
```
### 1c. Build / verify
Build kernel + `rocket.ko` + your board DTB, install, boot. Check:
```sh
cat /sys/kernel/debug/clk/scmi_clk_npu/clk_rate # 1000000000
cat /sys/class/regulator/.../microvolts # 850000 on vdd_npu_s0
ls /dev/accel/ # accel0
dmesg | grep -i 'rocket.*core .* version' # 3 cores probe
```
### ⚠ Caveat — power-domain SError under load
At 1 GHz, the driver's **per-job runtime-PM autosuspend (50 ms)** rapidly
power-gates the NPU genpd domain between fast tiles. That rapid gating races the
Rockchip power-domain controller into a **fatal asynchronous SError panic** under
sustained compute (`rocket_job_run → __pm_runtime_resume → genpd_… →
rockchip_pd_power_off → regmap_mmio_read → SError`). The `autosuspend_delay =
3600000` above is the **working mitigation** (the domain stays powered while a
client is active, so it never rapid-gates). A cleaner session-scoped PM hold is
WIP. Symptom if you skip it: hard SoC hang (no console) under a real workload.
---
## 2. Userspace — the `ggml-rocket` backend
```sh
git clone https://git.reauktion.de/marfrit/rk-llama.cpp
cd rk-llama.cpp
git checkout rocket-npu-backend
cmake -B build -DGGML_ROCKET=ON -DGGML_RKNPU2=OFF \
-DGGML_NATIVE=ON -DGGML_OPENMP=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-bench llama-cli -j"$(nproc)"
```
Backend lives in `ggml/src/ggml-rocket/`: `librocket.c` (thin DRM ioctl wrapper
over `/dev/accel`) + `rkt_*.c` (the INT8 GEMM regcmd builder — NVDLA-style CBUF
tiling, ported from Mesa Teflon's Rocket Gallium driver, golden-byte-verified
against a Mesa oracle; see `userspace/npu-probe/verify/`). It offloads `MUL_MAT`
(attention + FFN GEMMs) to the NPU INT8 path; dequant / RoPE / softmax /
sampling / embedding stay on the A76 NEON cores.
Run (requires `/dev/accel/accel0` from part 1):
```sh
ROCKET_NFD=3 ./build/bin/llama-bench -m <model-Q8_0.gguf> -p 512 -n 128 -ngl 99
```
Knobs:
- `ROCKET_NFD=3` — one fd per NPU core (a single fd caps at 2 cores via the
DRM-sched tie-break; use 3 for all three).
- `GGML_ROCKET_DEBUG=1` — per-op routing (which MUL_MATs offload vs fall to CPU).
- `GGML_ROCKET_PROF=1` — setup / submit / npu-wait phase timing.
- Compile-time: `ROCKET_TILE_M/N`, `ROCKET_K_MAX=8192`, `ROCKET_MIN_BATCH=32`.
Optional: K>8192 matmuls (e.g. gemma `ffn_down` K=12288) fall back to CPU by
default. Branch `rocket-ksplit-wip` adds NPU K-segmentation (correct — coherent
output — but pp512-neutral, being bandwidth-bound).
---
## 3. Results & findings (boltzmann, gemma-4-E2B-it-Q8_0, NPU @ 1 GHz)
| backend | pp512 (t/s) | tg128 |
|---|---|---|
| **rocket (mainline, this work)** | **46.7** | 6.5 |
| vendor RKNPU2 (`librknnrt`) | 51.4 | 8.3 |
~9 % gap, and it is **DDR-bandwidth-bound**, not compute- or core-bound.
Levers explored that did **not** close it (all empirically measured):
- **M-chunk tiling** (`ROCKET_TILE_M` up): regressed — CBUF input-bank pressure.
- **K>8192 segmentation**: correct, pp512-neutral.
- **CPU/NPU double-buffer** (overlap CPU prep with NPU compute on the idle cores):
bandwidth-bound. CPU memory throughput drops 3344 % the moment the NPU streams
weights (idle cores ≠ idle bandwidth). Matches the earlier rknpu2 "E8" result.
Corollaries: the NPU compute clock is **not** the bottleneck (pp512 is flat
800 MHz → 1 GHz); LPDDR5 bandwidth is the wall (no `dmc` devfreq headroom on this
board). The vendor stack hits the same wall (its own notes call prefill
"NPU-bound near the compute ceiling").
---
## 4. Correctness gate
There's no perplexity binary and `llama-completion` crashes on some chat
templates, so validate changes with a **deterministic piped-`llama-cli` greedy
A/B**: temp 0, fixed seed, feed the prompt + `/exit`, filter the
`[ Prompt: … t/s ]` line, diff the generated content across the change. A
corrupted matmul yields garbage tokens; a numerically-fine change keeps coherent
output. Note: `test-backend-ops`' default `MUL_MAT` shapes are all `n=1`, which
rocket declines (`M < MIN_BATCH`) — add `M ≥ 32`, large-`K` shapes if you use it.