ffmpeg-v4l2-request-fourier: revert ctx flip — PR #36 was a measurement artifact (0015) #105

Merged
marfrit merged 1 commits from claude-noether/marfrit-packages:noether/revert-h264-ctx-qpu-capable into main 2026-05-25 20:33:53 +00:00
Owner

Reverts the no_qpu → qpu-capable ctx flip that landed via patch 0014 (PR #104). Pairs with daedalus-fourier!37.

Why

PR #104 was justified by daedalus-fourier PR #36's QPU 4.30× faster than CPU NEON bench result. That number was a measurement artifact.

v3d_runner.read_spv() did a bare cwd-relative fopen() with no path search. The bench in PR #36 was run from the source dir, the SPVs at $builddir/v3d_*.spv were not found, every QPU dispatch returned -1 fast, the bench loop ignored the rc and timed the failure path (~1–5 µs per iteration). That's where the suspiciously-uniform-around-10ns/op QPU numbers came from.

daedalus-fourier PR #37 fixes the SPV search + bench preflight. Corrected numbers on hertz (Pi 5 V3D 7.1):

kernel CPU ns/op QPU ns/op winner
IDCT 4x4 luma 10.75 217.63 CPU 20.24×
IDCT 8x8 luma 29.69 785.94 CPU 26.47×
Deblock luma_v 17.63 467.42 CPU 26.51×
Deblock luma_h 38.30 498.53 CPU 13.02×
qpel mc20 (8x8) 30.17 1300.44 CPU 43.10×
qpel mc02 (8x8) 17.69 1363.40 CPU 77.08×
qpel mc22 (8x8) 71.60 1948.37 CPU 27.21×

1080p sum: CPU 5.57 ms vs QPU 123.54 ms — QPU 22× SLOWER than CPU.

What

New patch 0015-h264-ctx-revert-to-no-qpu.patch is a 2-line revert of patch 0014 across both H.264 TUs (h264_idct_daedalus.c, h264_qpel_daedalus.c). Wired into arch PKGBUILD + debian build-deb.sh. Both pkgrel bumped 11 → 12.

Until daedalus QPU dispatch overhead is actually competitive (multi-task effort on the daedalus-fourier side), libavcodec.so substitution must stay on daedalus_ctx_create_no_qpu() so host processes (firefox-fourier RDD, mpv-fourier, daedalus_v4l2_daemon) don't pessimize their H.264 decode path.

Note on alternatives

I considered just deleting patch 0014 outright vs. adding a 0015 that reverts it. Kept the additive-revert approach so the historical record stays intact (0014 + 0015 together = no net change; the commit messages on each carry the why).

Reverts the `no_qpu → qpu-capable` ctx flip that landed via patch 0014 (PR #104). Pairs with [daedalus-fourier!37](https://git.reauktion.de/marfrit/daedalus-fourier/pulls/37). ## Why PR #104 was justified by daedalus-fourier PR #36's `QPU 4.30× faster than CPU NEON` bench result. **That number was a measurement artifact.** `v3d_runner.read_spv()` did a bare cwd-relative `fopen()` with no path search. The bench in PR #36 was run from the source dir, the SPVs at `$builddir/v3d_*.spv` were not found, every QPU dispatch returned -1 fast, the bench loop ignored the rc and timed the failure path (~1–5 µs per iteration). That's where the suspiciously-uniform-around-10ns/op QPU numbers came from. daedalus-fourier PR #37 fixes the SPV search + bench preflight. Corrected numbers on hertz (Pi 5 V3D 7.1): | kernel | CPU ns/op | QPU ns/op | winner | |---|---|---|---| | IDCT 4x4 luma | **10.75** | 217.63 | CPU **20.24×** | | IDCT 8x8 luma | **29.69** | 785.94 | CPU **26.47×** | | Deblock luma_v | **17.63** | 467.42 | CPU **26.51×** | | Deblock luma_h | **38.30** | 498.53 | CPU **13.02×** | | qpel mc20 (8x8) | **30.17** | 1300.44 | CPU **43.10×** | | qpel mc02 (8x8) | **17.69** | 1363.40 | CPU **77.08×** | | qpel mc22 (8x8) | **71.60** | 1948.37 | CPU **27.21×** | **1080p sum**: CPU **5.57 ms** vs QPU **123.54 ms** — QPU **22× SLOWER** than CPU. ## What New patch `0015-h264-ctx-revert-to-no-qpu.patch` is a 2-line revert of patch 0014 across both H.264 TUs (`h264_idct_daedalus.c`, `h264_qpel_daedalus.c`). Wired into arch PKGBUILD + debian build-deb.sh. Both pkgrel bumped 11 → 12. Until daedalus QPU dispatch overhead is actually competitive (multi-task effort on the daedalus-fourier side), libavcodec.so substitution must stay on `daedalus_ctx_create_no_qpu()` so host processes (firefox-fourier RDD, mpv-fourier, daedalus_v4l2_daemon) don't pessimize their H.264 decode path. ## Note on alternatives I considered just deleting patch 0014 outright vs. adding a 0015 that reverts it. Kept the additive-revert approach so the historical record stays intact (0014 + 0015 together = no net change; the commit messages on each carry the why).
marfrit added 1 commit 2026-05-25 19:48:03 +00:00
Reverts the no_qpu → qpu-capable ctx flip that landed via patch 0014
(marfrit-packages PR #104).

PR #104 was justified by daedalus-fourier PR #36's "QPU 4.30x faster
than CPU NEON" bench result.  That number was a measurement artifact:
v3d_runner.read_spv() did a bare cwd-relative fopen() with no path
search, so when the bench was run from the source dir (as in PR #36),
the SPVs at $builddir/v3d_*.spv were not found, every QPU dispatch
returned -1 fast, and the bench loop timed the failure path.

daedalus-fourier PR #37 fixes the SPV search + bench preflight.
Corrected numbers on hertz (Pi 5 V3D 7.1):

  kernel             CPU ns/op  QPU ns/op  winner
  IDCT 4x4 luma          10.75     217.63  CPU 20.24x
  IDCT 8x8 luma          29.69     785.94  CPU 26.47x
  Deblock luma_v         17.63     467.42  CPU 26.51x
  Deblock luma_h         38.30     498.53  CPU 13.02x
  qpel mc20 (8x8)        30.17    1300.44  CPU 43.10x
  qpel mc02 (8x8)        17.69    1363.40  CPU 77.08x
  qpel mc22 (8x8)        71.60    1948.37  CPU 27.21x

  1080p sum: CPU 5.57 ms vs QPU 123.54 ms — QPU 22x SLOWER.

Until daedalus QPU dispatch overhead is actually competitive (separate
multi-task effort tracked on the daedalus-fourier side), libavcodec.so
substitution must stay on daedalus_ctx_create_no_qpu() so the host
processes (firefox-fourier RDD, mpv-fourier, daedalus_v4l2_daemon)
don't pessimize their H.264 decode path.

Adds 0015-h264-ctx-revert-to-no-qpu.patch (2-line revert of patch 0014)
to both arch PKGBUILD and debian build-deb.sh.  Both pkgrel bumped
11 → 12.  Refs reauktion/daedalus-fourier!37.
marfrit merged commit bdf3fffe2d into main 2026-05-25 20:33:53 +00:00
marfrit deleted branch noether/revert-h264-ctx-qpu-capable 2026-05-25 20:33:54 +00:00
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: marfrit/marfrit-packages#105