Full recipe: DT (enable rknn cores, OPP table 300-1000MHz, delete assigned-clock-rates), driver (OPP config + dev_pm_opp_set_rate 1GHz, autosuspend SError mitigation), ggml-rocket build/run, results (pp512 46.7 vs vendor 51.4, DDR-bandwidth-bound), correctness gate. Links rk-llama.cpp @ rocket-npu-backend and linux-rk3588-marfrit @ npuclk-opp.
7.9 KiB
Replicating the RK3588 Rocket-NPU LLM acceleration
Run llama.cpp LLM inference on the RK3588 NPU via the mainline rocket
DRM-accel driver — no vendor librknnrt blob — at 1 GHz. Measured on
boltzmann (Radxa Rock 5 ITX+, RK3588, 32 GB): gemma-4-E2B-it-Q8_0
pp512 ≈ 46.7 t/s, within ~9 % of the vendor closed stack (51.4), both
DDR-bandwidth-bound.
Repositories
| Component | Repo | Branch | Commit |
|---|---|---|---|
Userspace — ggml-rocket backend |
https://git.reauktion.de/marfrit/rk-llama.cpp | rocket-npu-backend |
288d70413 |
| Userspace — K>8192 K-segmentation (optional) | https://git.reauktion.de/marfrit/rk-llama.cpp | rocket-ksplit-wip |
6d5a075a6 |
| Kernel — 1 GHz OPP DVFS + SError mitigation | https://git.reauktion.de/marfrit/linux-rk3588-marfrit | npuclk-opp |
a811e6cc |
1. Kernel — from a vanilla RK3588 tree
Base: a mainline-class RK3588 kernel that already carries drivers/accel/rocket/
(the upstream Rocket driver by Tomeu Vizoso; present in recent mainline /
Collabora rockchip trees). Reference here: 7.0.0-rc3 (linux-rk3588-marfrit),
CONFIG_DRM_ACCEL_ROCKET=m, CONFIG_PM_OPP=y.
Upstream rocket runs the NPU DT-pinned at 200 MHz, no DVFS. Two changes
take it to 1 GHz.
1a. Device tree (board DTS)
The 3 NPU cores are defined disabled in rk3588-base.dtsi as
rknn_core_0/1/2. In your board DTS, enable them (Jaguar-style override), wire
the supply, attach an OPP table, and delete assigned-clock-rates:
&rknn_core_0 {
npu-supply = <&vdd_npu_s0>;
sram-supply = <&vdd_npu_s0>;
operating-points-v2 = <&npu_opp_table>;
/delete-property/ assigned-clock-rates;
status = "okay";
};
/* identical &rknn_core_1, &rknn_core_2 */
&rknn_mmu_0 { status = "okay"; }; /* + _1, _2 */
/ {
npu_opp_table: opp-table-npu {
compatible = "operating-points-v2";
opp-300000000 { opp-hz = /bits/ 64 < 300000000>; opp-microvolt = <700000>; };
opp-400000000 { opp-hz = /bits/ 64 < 400000000>; opp-microvolt = <700000>; };
opp-500000000 { opp-hz = /bits/ 64 < 500000000>; opp-microvolt = <700000>; };
opp-600000000 { opp-hz = /bits/ 64 < 600000000>; opp-microvolt = <700000>; };
opp-700000000 { opp-hz = /bits/ 64 < 700000000>; opp-microvolt = <700000>; };
opp-800000000 { opp-hz = /bits/ 64 < 800000000>; opp-microvolt = <750000>; };
opp-900000000 { opp-hz = /bits/ 64 < 900000000>; opp-microvolt = <800000>; };
opp-1000000000 { opp-hz = /bits/ 64 <1000000000>; opp-microvolt = <850000>; };
};
};
Notes:
- The
npu-named clock (<&scmi_clk SCMI_CLK_NPU>) is the scalable one;aclk/hclk/pclkare fixed AXI/AHB/APB bus clocks (do not scale with it). vdd_npu_s0is the board's NPU rail (rk8602@42, 0.55–0.95 V); voltages above are the vendor OPP values.- Deleting
assigned-clock-ratesis mandatory. Left in,of_clk_set_defaults()forces the clock rate at device-creation — before the driver probes and before the i2c PMIC reaches voltage — an undervolt-at-speed hang. Driving the rate through OPP (below) sequences voltage-then-frequency.
1b. Driver — drivers/accel/rocket/rocket_core.c, in rocket_core_init()
After devm_clk_bulk_get(...), register the OPP config + table and pin 1 GHz:
struct dev_pm_opp_config opp_cfg = {
.clk_names = (const char * const []){ "npu", NULL },
.regulator_names = (const char * const []){ "npu", NULL },
};
devm_pm_opp_set_config(dev, &opp_cfg);
devm_pm_opp_of_add_table(dev);
/* ... after pm_runtime_resume_and_get(dev) succeeds: */
dev_pm_opp_set_rate(dev, 1000000000); /* voltage-first via OPP */
And raise the runtime-PM autosuspend delay (SError mitigation — see caveat):
pm_runtime_set_autosuspend_delay(dev, 3600000); /* was 50 */
1c. Build / verify
Build kernel + rocket.ko + your board DTB, install, boot. Check:
cat /sys/kernel/debug/clk/scmi_clk_npu/clk_rate # 1000000000
cat /sys/class/regulator/.../microvolts # 850000 on vdd_npu_s0
ls /dev/accel/ # accel0
dmesg | grep -i 'rocket.*core .* version' # 3 cores probe
⚠ Caveat — power-domain SError under load
At 1 GHz, the driver's per-job runtime-PM autosuspend (50 ms) rapidly
power-gates the NPU genpd domain between fast tiles. That rapid gating races the
Rockchip power-domain controller into a fatal asynchronous SError panic under
sustained compute (rocket_job_run → __pm_runtime_resume → genpd_… → rockchip_pd_power_off → regmap_mmio_read → SError). The autosuspend_delay = 3600000 above is the working mitigation (the domain stays powered while a
client is active, so it never rapid-gates). A cleaner session-scoped PM hold is
WIP. Symptom if you skip it: hard SoC hang (no console) under a real workload.
2. Userspace — the ggml-rocket backend
git clone https://git.reauktion.de/marfrit/rk-llama.cpp
cd rk-llama.cpp
git checkout rocket-npu-backend
cmake -B build -DGGML_ROCKET=ON -DGGML_RKNPU2=OFF \
-DGGML_NATIVE=ON -DGGML_OPENMP=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-bench llama-cli -j"$(nproc)"
Backend lives in ggml/src/ggml-rocket/: librocket.c (thin DRM ioctl wrapper
over /dev/accel) + rkt_*.c (the INT8 GEMM regcmd builder — NVDLA-style CBUF
tiling, ported from Mesa Teflon's Rocket Gallium driver, golden-byte-verified
against a Mesa oracle; see userspace/npu-probe/verify/). It offloads MUL_MAT
(attention + FFN GEMMs) to the NPU INT8 path; dequant / RoPE / softmax /
sampling / embedding stay on the A76 NEON cores.
Run (requires /dev/accel/accel0 from part 1):
ROCKET_NFD=3 ./build/bin/llama-bench -m <model-Q8_0.gguf> -p 512 -n 128 -ngl 99
Knobs:
ROCKET_NFD=3— one fd per NPU core (a single fd caps at 2 cores via the DRM-sched tie-break; use 3 for all three).GGML_ROCKET_DEBUG=1— per-op routing (which MUL_MATs offload vs fall to CPU).GGML_ROCKET_PROF=1— setup / submit / npu-wait phase timing.- Compile-time:
ROCKET_TILE_M/N,ROCKET_K_MAX=8192,ROCKET_MIN_BATCH=32.
Optional: K>8192 matmuls (e.g. gemma ffn_down K=12288) fall back to CPU by
default. Branch rocket-ksplit-wip adds NPU K-segmentation (correct — coherent
output — but pp512-neutral, being bandwidth-bound).
3. Results & findings (boltzmann, gemma-4-E2B-it-Q8_0, NPU @ 1 GHz)
| backend | pp512 (t/s) | tg128 |
|---|---|---|
| rocket (mainline, this work) | 46.7 | 6.5 |
vendor RKNPU2 (librknnrt) |
51.4 | 8.3 |
~9 % gap, and it is DDR-bandwidth-bound, not compute- or core-bound. Levers explored that did not close it (all empirically measured):
- M-chunk tiling (
ROCKET_TILE_Mup): regressed — CBUF input-bank pressure. - K>8192 segmentation: correct, pp512-neutral.
- CPU/NPU double-buffer (overlap CPU prep with NPU compute on the idle cores): bandwidth-bound. CPU memory throughput drops 33–44 % the moment the NPU streams weights (idle cores ≠ idle bandwidth). Matches the earlier rknpu2 "E8" result.
Corollaries: the NPU compute clock is not the bottleneck (pp512 is flat
800 MHz → 1 GHz); LPDDR5 bandwidth is the wall (no dmc devfreq headroom on this
board). The vendor stack hits the same wall (its own notes call prefill
"NPU-bound near the compute ceiling").
4. Correctness gate
There's no perplexity binary and llama-completion crashes on some chat
templates, so validate changes with a deterministic piped-llama-cli greedy
A/B: temp 0, fixed seed, feed the prompt + /exit, filter the
[ Prompt: … t/s ] line, diff the generated content across the change. A
corrupted matmul yields garbage tokens; a numerically-fine change keeps coherent
output. Note: test-backend-ops' default MUL_MAT shapes are all n=1, which
rocket declines (M < MIN_BATCH) — add M ≥ 32, large-K shapes if you use it.