Champion shader ported to July build (29x over stock GPU) BUT ties CPU prefill / loses decode -> no faster-than-CPU win on Mali (VA squeeze). Honest correction.
This commit is contained in:
@@ -118,3 +118,36 @@ still unknown but moot — use plain mainline). GPU wedges if you pkill a job mi
|
||||
|
||||
**Next:** port the champion shader onto `build-vulkan` and re-measure gemma-4-E2B prefill on GPU;
|
||||
the correctness question is now closed (works).
|
||||
|
||||
## CHAMPION SHADER PORTED TO JULY BUILD + honest CPU comparison (2026-07-12)
|
||||
|
||||
Ported the champion `mul_mmq` warptile onto stock July mainline via a REDUCED patch: the 4
|
||||
essential ggml-vulkan hunks (VK_VENDOR_ID_ARM, GGML_VK_CMD_CLEANUP_FREQ env, is_mali warptile
|
||||
override {256,64,64,...}, Mali wg_denoms/align) + the mul_mmq.comp BK_STEP spec-constant. Dropped
|
||||
the 2 submit-cadence hunks (they hit July's restructured `ggml_backend_vk_graph_compute`; secondary).
|
||||
July already had the BK_STEP scaffolding + `l_warptile_mmq_int` variants upstream, and warptile is
|
||||
`std::vector<uint32_t>` so the 12th element (BK_STEP=2) compiles + plumbs as spec-constant 11.
|
||||
Reduced patch applies clean; build OK; correctness preserved (gemma-E2B -> ' Paris.').
|
||||
|
||||
### Measured (gemma-4-E2B, 242-tok prompt, warm)
|
||||
| path | prefill | decode |
|
||||
|------|---------|--------|
|
||||
| GPU stock July (no champion) | 0.9 t/s | 3.2 t/s |
|
||||
| **GPU champion shader** | **26.5 t/s** | 4.3 t/s |
|
||||
| **CPU (4x A76)** | **28.1 t/s** | **23.7 t/s** |
|
||||
|
||||
### VERDICT — the champion shader works, but it does NOT beat CPU on this box
|
||||
- Champion vs stock GPU: **29x prefill** (0.9 -> 26.5). The shader port is a real success and hits
|
||||
its documented ~28 t/s. Bug reproduced-and-fixed on the July build.
|
||||
- Champion GPU vs CPU: **prefill is a TIE** (26.5 vs 28.1), and **CPU wins decode decisively**
|
||||
(23.7 vs 4.3, the unified-DDR bandwidth wall the GPU can't beat).
|
||||
- **Correction to an earlier optimistic claim:** the champion's documented "4.8x" is stock-GPU ->
|
||||
champion-GPU, NOT vs CPU. The real result is GPU reaches PARITY with CPU on prefill, not a win.
|
||||
- The only way GPU prefill would beat CPU is a model big enough for the tile efficiency to pull
|
||||
ahead — but those don't fit on the Mali (sub-4GB VA heaps). The models that FIT are too small
|
||||
for the GPU to win. Squeeze: fits => too small to win; big enough to win => doesn't fit.
|
||||
- Net on boltzmann: **CPU (llama.cpp, Q4 QAT) remains the right path** for gemma. The GPU champion
|
||||
is a correct, impressive shader that lands at prefill-parity / decode-loss — not an upgrade.
|
||||
|
||||
Op note: SIGKILL of even an IDLE GPU llama-server wedges the panvk context (next GPU init hangs
|
||||
at 0% CPU) -> reboot to clear. Confirmed twice.
|
||||
|
||||
Reference in New Issue
Block a user