Compare commits

...

8 Commits

Author SHA1 Message Date
marfrit 9b0cb71370 Merge pull request 'ffmpeg-v4l2-request-fourier: export ff_h264_set_mb_inspect_cb (0016 amend, PKGREL 15)' (#108) from claude-noether/marfrit-packages:noether/h264-mb-inspect-export-symbol into main
build and publish packages / distcc-avahi-aarch64 (push) Successful in 6s
build and publish packages / mesa-panvk-bifrost-aarch64 (push) Successful in 4s
build and publish packages / mesa-panvk-bifrost-video-aarch64 (push) Successful in 5s
build and publish packages / lmcp-any (push) Successful in 4s
build and publish packages / lmcp-debian (push) Successful in 4s
build and publish packages / claude-his-any (push) Successful in 4s
build and publish packages / aish-any (push) Successful in 4s
build and publish packages / ffmpeg-v4l2-request-aarch64 (push) Successful in 12m2s
build and publish packages / claude-his-debian (push) Successful in 4s
build and publish packages / aish-debian (push) Successful in 5s
build and publish packages / libva-v4l2-request-fourier-aarch64 (push) Successful in 3s
build and publish packages / mpv-fourier-aarch64 (push) Successful in 6s
build and publish packages / ffmpeg-v4l2-request-debian (push) Successful in 7m15s
build and publish packages / daedalus-v4l2-debian (push) Successful in 6s
build and publish packages / libva-v4l2-request-fourier-debian (push) Successful in 6s
build and publish packages / mpv-fourier-debian (push) Successful in 4s
build and publish packages / daedalus-v4l2-dkms-debian (push) Successful in 4s
Reviewed-on: #108
2026-05-26 14:54:27 +00:00
marfrit 25610930ad ffmpeg-v4l2-request-fourier: export ff_h264_set_mb_inspect_cb (0016 amend, PKGREL 15)
The 0016 patch declared ff_h264_set_mb_inspect_cb in h264dec.h and
defined it in h264_mb.c, but didn't touch libavcodec/libavcodec.v.
FFmpeg's default version script exports only `av_*`, `avcodec_*`,
`avpriv_*`, and `avsubtitle_free`; everything else is hidden as LOCAL
behind a `*` glob.  Result: `nm -D libavcodec.so.62 | grep
ff_h264_set_mb_inspect_cb` returned nothing → dlsym() returned NULL.

Static-link CLI consumer (daedalus_decode_h264) was unaffected
because static linking doesn't care about symbol visibility.  The
daedalus-v4l2 daemon shadow_decoder path (PR-Q3a.1) dlopens
libavcodec.so.62 and resolves the callback via dlsym — that needs
the symbol exported.

Fix: add ff_h264_set_mb_inspect_cb to the global list in
libavcodec/libavcodec.v.  Single-line addition to the 0016 patch.
Mirrored across the arch/ + debian/ patch trees.

PKGREL bump 14 → 15, changelog entry added (debian side).  PKGBUILD
pkgrel bumped on arch side too.  No behaviour change to the decode
path: the callback is still opt-in via the H264Context function
pointer; only consumers that have explicitly installed a callback
pay the one-load-one-branch cost per MB.

dejavu-check: this is fixing the existing 0016 observation-hook to
actually work as a dlsym intercept (the architectural shape the
patch was designed for).  NOT adding new per-kernel substitution.
Same shape, same patch number, same intent.  Just hiding/exporting
plumbing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 15:38:03 +02:00
marfrit 368fcff41f Merge pull request 'ffmpeg-v4l2-request-fourier: preserve sl->mb for inspection callback (0017)' (#107) from claude-noether/marfrit-packages:noether/h264-mb-coeffs-side-buffer into main
build and publish packages / distcc-avahi-aarch64 (push) Successful in 16s
build and publish packages / mesa-panvk-bifrost-aarch64 (push) Successful in 2s
build and publish packages / mesa-panvk-bifrost-video-aarch64 (push) Successful in 3s
build and publish packages / lmcp-any (push) Successful in 4s
build and publish packages / lmcp-debian (push) Successful in 3s
build and publish packages / claude-his-any (push) Successful in 4s
build and publish packages / aish-any (push) Successful in 4s
build and publish packages / ffmpeg-v4l2-request-aarch64 (push) Failing after 2m20s
build and publish packages / libva-v4l2-request-fourier-aarch64 (push) Has been skipped
build and publish packages / mpv-fourier-aarch64 (push) Has been skipped
build and publish packages / ffmpeg-v4l2-request-debian (push) Has been skipped
build and publish packages / libva-v4l2-request-fourier-debian (push) Has been skipped
build and publish packages / mpv-fourier-debian (push) Has been skipped
build and publish packages / claude-his-debian (push) Successful in 14s
build and publish packages / daedalus-v4l2-debian (push) Successful in 13s
build and publish packages / daedalus-v4l2-dkms-debian (push) Successful in 3s
build and publish packages / aish-debian (push) Failing after 11m25s
Reviewed-on: #107
2026-05-26 07:48:47 +00:00
claude-noether ea99dc8e27 ffmpeg-v4l2-request-fourier: preserve sl->mb for inspection callback (0017)
Companion to 0016 (PR #106).  Adds a coefficient side buffer in
H264Context, populated at the start of ff_h264_hl_decode_mb with a
single memcpy from sl->mb BEFORE IDCT-add zeros it.  The existing
post-pixel-work callback (still in 0016) can now read:
  - h->mb_inspect_coeffs  = pre-IDCT coefficients (this patch)
  - h->cur_pic.f->data    = post-pixel-work pre-deblock reconstruction

and derive P = pixels − IDCT(C) for daedalus-decoder's frame-major
dispatch in PR-A3+.

Memcpy gated on (h->mb_inspect_cb != NULL).  Zero cost when no
consumer is registered.  Side buffer = 16 * 48 int16 = 1536 bytes
(matches the 8-bit half of sl->mb's int16_t[16 * 48 * 2] declared
size; high-bit-depth uses the upper half — not preserved here since
the daedalus-decoder consumer is 8-bit-only).

Single-threaded decode assumed at the consumer side
(avctx->thread_count = 1).  Multi-slice / multi-threaded streams
would race on the single side buffer — explicit limitation of the
inspection mechanism, future extension would put per-slice buffers
in H264SliceContext.

Verified: patches 0016 + 0017 apply cleanly and build in sequence
against the Kwiboo v4l2-request-n8.1 fork at the pinned commit
b57fbbe5.  ff_h264_set_mb_inspect_cb symbol exported as before.

Wired into arch PKGBUILD + debian build-deb.sh patch sequence.
pkgrel bumped 13 → 14.

Refs reauktion/daedalus-decoder!14 (PR-A2 callback wiring complete,
PR-A3 coefficient extraction is the next consumer).
2026-05-26 09:46:10 +02:00
marfrit 59901bceca Merge pull request 'ffmpeg-v4l2-request-fourier: per-MB inspection callback for H.264 (0016)' (#106) from claude-noether/marfrit-packages:noether/h264-mb-inspect-callback into main
build and publish packages / distcc-avahi-aarch64 (push) Successful in 5s
build and publish packages / mesa-panvk-bifrost-aarch64 (push) Successful in 2s
build and publish packages / mesa-panvk-bifrost-video-aarch64 (push) Successful in 4s
build and publish packages / lmcp-any (push) Successful in 4s
build and publish packages / lmcp-debian (push) Successful in 4s
build and publish packages / claude-his-any (push) Successful in 3s
build and publish packages / aish-any (push) Successful in 4s
build and publish packages / ffmpeg-v4l2-request-aarch64 (push) Failing after 4m54s
build and publish packages / libva-v4l2-request-fourier-aarch64 (push) Has been skipped
build and publish packages / mpv-fourier-aarch64 (push) Has been skipped
build and publish packages / ffmpeg-v4l2-request-debian (push) Has been skipped
build and publish packages / libva-v4l2-request-fourier-debian (push) Has been skipped
build and publish packages / mpv-fourier-debian (push) Has been skipped
build and publish packages / claude-his-debian (push) Successful in 7s
build and publish packages / daedalus-v4l2-debian (push) Successful in 5s
build and publish packages / aish-debian (push) Successful in 5s
build and publish packages / daedalus-v4l2-dkms-debian (push) Successful in 5s
Reviewed-on: #106
2026-05-26 04:05:57 +00:00
claude-noether 87cbb9b70a ffmpeg-v4l2-request-fourier: per-MB inspection callback for H.264 (0016)
Adds 0016-h264-mb-inspect-callback.patch to the FFmpeg fork.  Adds an
opt-in callback fired by ff_h264_hl_decode_mb after the existing
pixel work, for tools that need per-MB visibility into H.264 decode.

API:
  typedef void (*ff_h264_mb_inspect_cb)(void *opaque,
                                         const struct H264Context *h,
                                         int mb_x, int mb_y);
  void ff_h264_set_mb_inspect_cb(AVCodecContext *avctx,
                                  ff_h264_mb_inspect_cb cb, void *opaque);

Two new fields appended to H264Context (internal struct, declared in
h264dec.h not h264.h, no ABI surface to non-libavcodec callers).
Callback fires post-pixel-work for every MB in coded order; receives
const H264Context* so it can inspect any state (slice ctx via
h->slice_ctx, reconstructed pixels via h->cur_pic.f->data[plane],
etc.).

Default (cb==NULL): zero behaviour change, one load + one branch per
MB in the decoder hot path.

Shape distinction: per-MB observation, NOT per-kernel function-pointer
hijack (the 0003-0014 substitution-arc pattern that PR #105 reverted
+ daedalus-fourier PR #37's measurement-correction architecturally
retired).  Per-block synchronous Vulkan dispatch from libavcodec is
non-competitive; per-MB CPU-side observation feeding a per-frame
daedalus-decoder batch submit is the right shape (frame-major UMA
dispatch verdict, memory: dejavu).

Used by:
  - daedalus-decoder/tools/daedalus_decode_h264 (PR-A1b, follow-up)
  - future daedalus-v4l2 daemon refactor

Wired into arch PKGBUILD source[] + prepare() and debian build-deb.sh
patch sequence.  pkgrel bumped 12 → 13.

Refs reauktion/daedalus-decoder!12.
2026-05-26 05:57:49 +02:00
marfrit bdf3fffe2d Merge pull request 'ffmpeg-v4l2-request-fourier: revert ctx flip — PR #36 was a measurement artifact (0015)' (#105) from claude-noether/marfrit-packages:noether/revert-h264-ctx-qpu-capable into main
build and publish packages / distcc-avahi-aarch64 (push) Successful in 3s
build and publish packages / mesa-panvk-bifrost-aarch64 (push) Successful in 4s
build and publish packages / mesa-panvk-bifrost-video-aarch64 (push) Successful in 3s
build and publish packages / lmcp-any (push) Successful in 3s
build and publish packages / lmcp-debian (push) Successful in 3s
build and publish packages / claude-his-any (push) Successful in 2s
build and publish packages / aish-any (push) Successful in 3s
build and publish packages / ffmpeg-v4l2-request-aarch64 (push) Successful in 11m59s
build and publish packages / claude-his-debian (push) Successful in 4s
build and publish packages / aish-debian (push) Successful in 4s
build and publish packages / libva-v4l2-request-fourier-aarch64 (push) Successful in 2s
build and publish packages / mpv-fourier-aarch64 (push) Successful in 3s
build and publish packages / ffmpeg-v4l2-request-debian (push) Successful in 7m33s
build and publish packages / daedalus-v4l2-debian (push) Successful in 3s
build and publish packages / libva-v4l2-request-fourier-debian (push) Successful in 4s
build and publish packages / mpv-fourier-debian (push) Successful in 4s
build and publish packages / daedalus-v4l2-dkms-debian (push) Successful in 3s
Reviewed-on: #105
2026-05-25 20:33:52 +00:00
claude-noether f4047f3145 ffmpeg-v4l2-request-fourier: revert ctx flip — PR #36 was a measurement artifact (0015)
Reverts the no_qpu → qpu-capable ctx flip that landed via patch 0014
(marfrit-packages PR #104).

PR #104 was justified by daedalus-fourier PR #36's "QPU 4.30x faster
than CPU NEON" bench result.  That number was a measurement artifact:
v3d_runner.read_spv() did a bare cwd-relative fopen() with no path
search, so when the bench was run from the source dir (as in PR #36),
the SPVs at $builddir/v3d_*.spv were not found, every QPU dispatch
returned -1 fast, and the bench loop timed the failure path.

daedalus-fourier PR #37 fixes the SPV search + bench preflight.
Corrected numbers on hertz (Pi 5 V3D 7.1):

  kernel             CPU ns/op  QPU ns/op  winner
  IDCT 4x4 luma          10.75     217.63  CPU 20.24x
  IDCT 8x8 luma          29.69     785.94  CPU 26.47x
  Deblock luma_v         17.63     467.42  CPU 26.51x
  Deblock luma_h         38.30     498.53  CPU 13.02x
  qpel mc20 (8x8)        30.17    1300.44  CPU 43.10x
  qpel mc02 (8x8)        17.69    1363.40  CPU 77.08x
  qpel mc22 (8x8)        71.60    1948.37  CPU 27.21x

  1080p sum: CPU 5.57 ms vs QPU 123.54 ms — QPU 22x SLOWER.

Until daedalus QPU dispatch overhead is actually competitive (separate
multi-task effort tracked on the daedalus-fourier side), libavcodec.so
substitution must stay on daedalus_ctx_create_no_qpu() so the host
processes (firefox-fourier RDD, mpv-fourier, daedalus_v4l2_daemon)
don't pessimize their H.264 decode path.

Adds 0015-h264-ctx-revert-to-no-qpu.patch (2-line revert of patch 0014)
to both arch PKGBUILD and debian build-deb.sh.  Both pkgrel bumped
11 → 12.  Refs reauktion/daedalus-fourier!37.
2026-05-25 21:47:33 +02:00
9 changed files with 620 additions and 9 deletions
@@ -0,0 +1,73 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Markus Fritsche <mfritsche@reauktion.de>
Date: Mon, 25 May 2026 22:00:00 +0200
Subject: [PATCH] avcodec/aarch64/h264: revert ctx flip — daedalus-fourier PR
#36 was a measurement artifact
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Reverts the daedalus_ctx_create_no_qpu() → daedalus_ctx_create() flip
that landed in 0014-h264-ctx-qpu-capable.patch (marfrit-packages PR
#104). The flip was justified by daedalus-fourier PR #36 which
reported a 4.30x QPU-over-CPU win on the 1080p H.264 hot-path sum.
That number was a measurement artifact. The bench tool's
v3d_runner.read_spv() did a bare fopen() that resolved relative to
cwd; when run from the source directory (as in PR #36), the SPVs at
$builddir/v3d_*.spv were not found, every QPU dispatch returned -1
fast, and the loop timed the failure path. Daedalus-fourier PR #37
fixes the SPV search + bench preflight; corrected numbers from hertz
(Pi 5 V3D 7.1) show QPU is 12-77x SLOWER than CPU NEON at every
H.264 hot-path kernel:
kernel CPU ns/op QPU ns/op winner
IDCT 4x4 luma 10.75 217.63 CPU 20.24x
IDCT 8x8 luma 29.69 785.94 CPU 26.47x
Deblock luma_v 17.63 467.42 CPU 26.51x
Deblock luma_h 38.30 498.53 CPU 13.02x
qpel mc20 (8x8) 30.17 1300.44 CPU 43.10x
qpel mc02 (8x8) 17.69 1363.40 CPU 77.08x
qpel mc22 (8x8) 71.60 1948.37 CPU 27.21x
1080p sum: CPU 5.57 ms vs QPU 123.54 ms — QPU 22x slower.
Until the daedalus QPU dispatch overhead is actually competitive (a
multi-task effort tracked on the daedalus-fourier side), the
libavcodec.so substitution must stay on daedalus_ctx_create_no_qpu()
to avoid pessimizing every host process that loads it
(firefox-fourier RDD, mpv-fourier, daedalus_v4l2_daemon).
Both H.264 TUs (h264_idct_daedalus.c, h264_qpel_daedalus.c) are
reverted; the change is a 2-line revert of patch 0014.
Refs reauktion/daedalus-fourier!37 (the retraction PR).
---
libavcodec/aarch64/h264_idct_daedalus.c | 2 +-
libavcodec/aarch64/h264_qpel_daedalus.c | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/libavcodec/aarch64/h264_idct_daedalus.c b/libavcodec/aarch64/h264_idct_daedalus.c
--- a/libavcodec/aarch64/h264_idct_daedalus.c
+++ b/libavcodec/aarch64/h264_idct_daedalus.c
@@ -32,7 +32,7 @@ static pthread_once_t g_dctx_once = PTHREAD_ONCE_INIT;
static void daedalus_ctx_init_once(void)
{
- g_dctx = daedalus_ctx_create();
+ g_dctx = daedalus_ctx_create_no_qpu();
}
void ff_h264_idct_add_daedalus(uint8_t *dst, int16_t *block, int stride);
diff --git a/libavcodec/aarch64/h264_qpel_daedalus.c b/libavcodec/aarch64/h264_qpel_daedalus.c
--- a/libavcodec/aarch64/h264_qpel_daedalus.c
+++ b/libavcodec/aarch64/h264_qpel_daedalus.c
@@ -38,7 +38,7 @@ static pthread_once_t g_dctx_once = PTHREAD_ONCE_INIT;
static void daedalus_ctx_init_once(void)
{
- g_dctx = daedalus_ctx_create();
+ g_dctx = daedalus_ctx_create_no_qpu();
}
void ff_put_h264_qpel8_mc20_daedalus(uint8_t *dst, const uint8_t *src, ptrdiff_t stride);
@@ -0,0 +1,132 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Markus Fritsche <mfritsche@reauktion.de>
Date: Tue, 26 May 2026 06:00:00 +0200
Subject: [PATCH] avcodec/h264: per-MB inspection callback (daedalus-decoder
hook)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Adds an opt-in callback fired in ff_h264_hl_decode_mb after the
existing pixel work, used by tools that need per-MB visibility into
the H.264 decode. Initially driven by daedalus-decoder's CLI test
harness (tools/daedalus_decode_h264) which shadows libavcodec's
decode with a frame-major daedalus-decoder run for byte-exact diff
on real H.264 streams; later target is a daedalus-v4l2 daemon
refactor that drives daedalus_decoder_append_mb directly from the
callback instead of letting libavcodec do per-MB pixel work.
Shape: ONE inspection point per MB. Distinct from the per-kernel
function-pointer-hijack pattern that used to live in 0003-0014
patches (now reverted via 0015 for ctx, and architecturally retired
per daedalus-fourier PR #37's measurement-correction). Per-block
synchronous Vulkan dispatch from libavcodec was structurally non-
competitive; per-MB CPU-side observation feeding a per-frame batch
submit is the right shape.
Two new fields in H264Context (appended at end of struct; no ABI
surface visible to non-libavcodec callers since H264Context is
internal — declared in h264dec.h, not h264.h). One new exported
function ff_h264_set_mb_inspect_cb to set them.
Zero behaviour change when cb == NULL (the default): one load +
one branch per MB in the decoder hot path, both branch-predicted
to fall through.
Used by:
- daedalus-decoder/tools/daedalus_decode_h264 (PR-A1b)
- daedalus-v4l2 daemon shadow-mode path (PR-Q3a.1+)
The CLI static-links libavcodec.a so symbol visibility doesn't matter
there. The daemon dlopens libavcodec.so.62 and resolves the callback
via dlsym, so the symbol MUST be exported — added to libavcodec.v
explicitly (FFmpeg's default version script hides every `ff_*` symbol
as LOCAL behind a glob).
Refs reauktion/daedalus-decoder!12 (Stage 2 PR-b complete).
---
libavcodec/h264_mb.c | 20 ++++++++++++++++++++
libavcodec/h264dec.h | 26 ++++++++++++++++++++++++++
libavcodec/libavcodec.v | 1 +
3 files changed, 47 insertions(+)
--- a/libavcodec/h264dec.h
+++ b/libavcodec/h264dec.h
@@ -334,6 +334,16 @@
int pic_order_cnt_bit_size;
} H264SliceContext;
+/* Per-MB inspection callback type — see ff_h264_set_mb_inspect_cb()
+ * below. Fired by ff_h264_hl_decode_mb after the existing pixel work
+ * for every macroblock in coded order. Receives a const H264Context*
+ * so the callback can inspect any slice/picture state (h->slice_ctx
+ * for current slice, h->cur_pic.f->data[plane] for reconstructed
+ * samples, etc.). */
+typedef void (*ff_h264_mb_inspect_cb)(void *opaque,
+ const struct H264Context *h,
+ int mb_x, int mb_y);
+
/**
* H264Context
*/
@@ -579,6 +589,10 @@
int non_gray; ///< Did we encounter a intra frame after a gray gap frame
int noref_gray;
int skip_gray;
+
+ /* Per-MB inspection hook — set via ff_h264_set_mb_inspect_cb. */
+ ff_h264_mb_inspect_cb mb_inspect_cb;
+ void *mb_inspect_opaque;
} H264Context;
extern const uint16_t ff_h264_mb_sizes[4];
@@ -607,6 +621,16 @@
const H2645NAL *nal, void *logctx);
void ff_h264_hl_decode_mb(const H264Context *h, H264SliceContext *sl);
+
+/**
+ * Install an opt-in per-MB inspection callback that fires from
+ * ff_h264_hl_decode_mb after each macroblock's pixel work. Default
+ * is NULL (no callback installed); the check is a single branch on
+ * the decoder hot path. See ff_h264_mb_inspect_cb for signature.
+ */
+void ff_h264_set_mb_inspect_cb(AVCodecContext *avctx,
+ ff_h264_mb_inspect_cb cb, void *opaque);
+
void ff_h264_decode_init_vlc(void);
/**
--- a/libavcodec/h264_mb.c
+++ b/libavcodec/h264_mb.c
@@ -815,4 +815,20 @@
hl_decode_mb_simple_16(h, sl);
} else
hl_decode_mb_simple_8(h, sl);
+
+ /* Per-MB inspection callback (opt-in via ff_h264_set_mb_inspect_cb).
+ * Fired AFTER pixel work — reconstructed samples are in
+ * h->cur_pic.f->data[plane] at the MB's raster position by the
+ * time this runs. Callback may inspect slice context via
+ * h->slice_ctx + sl->mb_xy, coeffs via sl->mb, etc. */
+ if (h->mb_inspect_cb)
+ h->mb_inspect_cb(h->mb_inspect_opaque, h, sl->mb_x, sl->mb_y);
+}
+
+void ff_h264_set_mb_inspect_cb(AVCodecContext *avctx,
+ ff_h264_mb_inspect_cb cb, void *opaque)
+{
+ H264Context *h = avctx->priv_data;
+ h->mb_inspect_cb = cb;
+ h->mb_inspect_opaque = opaque;
}
--- a/libavcodec/libavcodec.v
+++ b/libavcodec/libavcodec.v
@@ -3,6 +3,7 @@
av_*;
avcodec_*;
avpriv_*;
+ ff_h264_set_mb_inspect_cb;
avsubtitle_free;
local:
*;
@@ -0,0 +1,88 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Markus Fritsche <mfritsche@reauktion.de>
Date: Tue, 26 May 2026 07:30:00 +0200
Subject: [PATCH] avcodec/h264: preserve sl->mb coefficients for the inspection
callback (companion to 0016)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Patch 0016 adds a per-MB inspection callback fired at the end of
ff_h264_hl_decode_mb. By that time the IDCT-add path has already
zeroed sl->mb (FFmpeg's convention — see ff_h264_idct_add_neon and
friends), so consumers reading coefficients from the callback get
zeros.
Add a coefficient side buffer in H264Context, populated at the
START of ff_h264_hl_decode_mb (before any IDCT runs) with a single
memcpy from sl->mb. The post-pixel-work callback (still in 0016)
can then read both:
- the side-buffer coefficients (= just-entropy-decoded, pre-IDCT)
- the reconstructed pixels in h->cur_pic.f->data (= P + IDCT(C),
pre-deblock for this MB)
and the consumer can derive P = pixels IDCT(C) for daedalus-
decoder's frame-major dispatch.
Memcpy is gated on (h->mb_inspect_cb != NULL) — zero overhead when
no consumer is registered. Buffer size = sizeof(int16_t) * 16 * 48
= 1536 bytes per H264Context (fits in one cache line family;
allocated once at H264Context lifetime, reused per MB).
8-bit path only. High-bit-depth H.264 uses the upper half of
sl->mb (int16_t[16 * 48 * 2] declared; the * 2 reserves space for
the high-depth case); preserving the high-depth coefficients
correctly would need a wider side buffer. Punted for now — the
daedalus-decoder consumer is 8-bit-only.
Single-threaded decode assumed at the consumer side (avctx->
thread_count = 1). Multi-slice / multi-threaded streams would
race on the single side buffer — that's an explicit limitation of
the inspection mechanism, documented in 0016's comment block.
Future extension: per-H264SliceContext side buffers.
Used by:
- daedalus-decoder/tools/daedalus_decode_h264 PR-A3+ (CLI test
harness extracts coefficients here for daedalus-decoder
IDCT validation on real H.264 streams).
Refs reauktion/daedalus-decoder!14 (PR-A2 callback wiring).
---
libavcodec/h264_mb.c | 9 +++++++++
libavcodec/h264dec.h | 8 ++++++++
2 files changed, 17 insertions(+)
--- a/libavcodec/h264dec.h
+++ b/libavcodec/h264dec.h
@@ -593,6 +593,14 @@
/* Per-MB inspection hook — set via ff_h264_set_mb_inspect_cb. */
ff_h264_mb_inspect_cb mb_inspect_cb;
void *mb_inspect_opaque;
+
+ /* Per-MB coefficient side buffer — populated at the start of
+ * ff_h264_hl_decode_mb so the post-pixel-work inspection callback
+ * can read the just-entropy-decoded coefficients before IDCT-add
+ * zeros sl->mb. 16 blocks × 48 int16 = libavcodec sl->mb size
+ * (matches DECLARE_ALIGNED(16, int16_t, mb)[16 * 48 * 2] for the
+ * 8-bit half; high-bit-depth paths skip this — see h264_mb.c). */
+ DECLARE_ALIGNED(16, int16_t, mb_inspect_coeffs)[16 * 48];
} H264Context;
extern const uint16_t ff_h264_mb_sizes[4];
--- a/libavcodec/h264_mb.c
+++ b/libavcodec/h264_mb.c
@@ -801,6 +801,15 @@
{
const int mb_xy = sl->mb_xy;
const int mb_type = h->cur_pic.mb_type[mb_xy];
+
+ /* Snapshot just-entropy-decoded coefficients before IDCT-add
+ * destroys them. Only when an inspection callback is registered
+ * — zero cost otherwise. 8-bit path only (high-bit-depth uses
+ * the upper half of sl->mb which we don't preserve here). */
+ if (h->mb_inspect_cb && !h->pixel_shift)
+ memcpy((int16_t *) (uintptr_t) h->mb_inspect_coeffs, sl->mb,
+ sizeof(((H264Context *) NULL)->mb_inspect_coeffs));
+
int is_complex = CONFIG_SMALL || sl->is_complex ||
IS_INTRA_PCM(mb_type) || sl->qscale == 0;
+9 -3
View File
@@ -24,7 +24,7 @@ _srcname=FFmpeg
_version='8.1'
_commit='b57fbbe50c9b2656fad86a1a7eeabfd2b2a50935' # v4l2-request-n8.1 tip 2026-04-24
pkgver=8.1.r123329.b57fbbe
pkgrel=11 # pkgrel=11 — libavcodec.so daedalus ctx flipped no_qpu → qpu-capable (PR #36 bench: QPU 4.30x on 1080p hot-path sum, 2026-05-25)
pkgrel=15 # pkgrel=15 export ff_h264_set_mb_inspect_cb via libavcodec.v so dlsym consumers (daedalus-v4l2 daemon shadow_decoder, PR-Q3a.1) can resolve the symbol; static-link CLI was unaffected. No behaviour change to existing decode path. (2026-05-26)
epoch=2
# daedalus-fourier pin. 209a421 = PR #2 merge (Phase 8c — public API
@@ -101,8 +101,11 @@ source=("git+https://github.com/Kwiboo/FFmpeg.git#commit=${_commit}"
'0011-h264-chroma-dc-hadamard-daedalus-fourier.patch'
'0012-h264-qpel-rest-daedalus-fourier.patch'
'0013-h264-deblock-chroma-intra-daedalus-fourier.patch'
'0014-h264-ctx-qpu-capable.patch')
sha256sums=('SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP')
'0014-h264-ctx-qpu-capable.patch'
'0015-h264-ctx-revert-to-no-qpu.patch'
'0016-h264-mb-inspect-callback.patch'
'0017-h264-mb-coeffs-side-buffer.patch')
sha256sums=('SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP' 'SKIP')
pkgver() {
cd "${_srcname}"
@@ -127,6 +130,9 @@ prepare() {
patch -Np1 -i "${srcdir}/0012-h264-qpel-rest-daedalus-fourier.patch"
patch -Np1 -i "${srcdir}/0013-h264-deblock-chroma-intra-daedalus-fourier.patch"
patch -Np1 -i "${srcdir}/0014-h264-ctx-qpu-capable.patch"
patch -Np1 -i "${srcdir}/0015-h264-ctx-revert-to-no-qpu.patch"
patch -Np1 -i "${srcdir}/0016-h264-mb-inspect-callback.patch"
patch -Np1 -i "${srcdir}/0017-h264-mb-coeffs-side-buffer.patch"
}
build() {
@@ -0,0 +1,73 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Markus Fritsche <mfritsche@reauktion.de>
Date: Mon, 25 May 2026 22:00:00 +0200
Subject: [PATCH] avcodec/aarch64/h264: revert ctx flip — daedalus-fourier PR
#36 was a measurement artifact
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Reverts the daedalus_ctx_create_no_qpu() → daedalus_ctx_create() flip
that landed in 0014-h264-ctx-qpu-capable.patch (marfrit-packages PR
#104). The flip was justified by daedalus-fourier PR #36 which
reported a 4.30x QPU-over-CPU win on the 1080p H.264 hot-path sum.
That number was a measurement artifact. The bench tool's
v3d_runner.read_spv() did a bare fopen() that resolved relative to
cwd; when run from the source directory (as in PR #36), the SPVs at
$builddir/v3d_*.spv were not found, every QPU dispatch returned -1
fast, and the loop timed the failure path. Daedalus-fourier PR #37
fixes the SPV search + bench preflight; corrected numbers from hertz
(Pi 5 V3D 7.1) show QPU is 12-77x SLOWER than CPU NEON at every
H.264 hot-path kernel:
kernel CPU ns/op QPU ns/op winner
IDCT 4x4 luma 10.75 217.63 CPU 20.24x
IDCT 8x8 luma 29.69 785.94 CPU 26.47x
Deblock luma_v 17.63 467.42 CPU 26.51x
Deblock luma_h 38.30 498.53 CPU 13.02x
qpel mc20 (8x8) 30.17 1300.44 CPU 43.10x
qpel mc02 (8x8) 17.69 1363.40 CPU 77.08x
qpel mc22 (8x8) 71.60 1948.37 CPU 27.21x
1080p sum: CPU 5.57 ms vs QPU 123.54 ms — QPU 22x slower.
Until the daedalus QPU dispatch overhead is actually competitive (a
multi-task effort tracked on the daedalus-fourier side), the
libavcodec.so substitution must stay on daedalus_ctx_create_no_qpu()
to avoid pessimizing every host process that loads it
(firefox-fourier RDD, mpv-fourier, daedalus_v4l2_daemon).
Both H.264 TUs (h264_idct_daedalus.c, h264_qpel_daedalus.c) are
reverted; the change is a 2-line revert of patch 0014.
Refs reauktion/daedalus-fourier!37 (the retraction PR).
---
libavcodec/aarch64/h264_idct_daedalus.c | 2 +-
libavcodec/aarch64/h264_qpel_daedalus.c | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/libavcodec/aarch64/h264_idct_daedalus.c b/libavcodec/aarch64/h264_idct_daedalus.c
--- a/libavcodec/aarch64/h264_idct_daedalus.c
+++ b/libavcodec/aarch64/h264_idct_daedalus.c
@@ -32,7 +32,7 @@ static pthread_once_t g_dctx_once = PTHREAD_ONCE_INIT;
static void daedalus_ctx_init_once(void)
{
- g_dctx = daedalus_ctx_create();
+ g_dctx = daedalus_ctx_create_no_qpu();
}
void ff_h264_idct_add_daedalus(uint8_t *dst, int16_t *block, int stride);
diff --git a/libavcodec/aarch64/h264_qpel_daedalus.c b/libavcodec/aarch64/h264_qpel_daedalus.c
--- a/libavcodec/aarch64/h264_qpel_daedalus.c
+++ b/libavcodec/aarch64/h264_qpel_daedalus.c
@@ -38,7 +38,7 @@ static pthread_once_t g_dctx_once = PTHREAD_ONCE_INIT;
static void daedalus_ctx_init_once(void)
{
- g_dctx = daedalus_ctx_create();
+ g_dctx = daedalus_ctx_create_no_qpu();
}
void ff_put_h264_qpel8_mc20_daedalus(uint8_t *dst, const uint8_t *src, ptrdiff_t stride);
@@ -0,0 +1,132 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Markus Fritsche <mfritsche@reauktion.de>
Date: Tue, 26 May 2026 06:00:00 +0200
Subject: [PATCH] avcodec/h264: per-MB inspection callback (daedalus-decoder
hook)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Adds an opt-in callback fired in ff_h264_hl_decode_mb after the
existing pixel work, used by tools that need per-MB visibility into
the H.264 decode. Initially driven by daedalus-decoder's CLI test
harness (tools/daedalus_decode_h264) which shadows libavcodec's
decode with a frame-major daedalus-decoder run for byte-exact diff
on real H.264 streams; later target is a daedalus-v4l2 daemon
refactor that drives daedalus_decoder_append_mb directly from the
callback instead of letting libavcodec do per-MB pixel work.
Shape: ONE inspection point per MB. Distinct from the per-kernel
function-pointer-hijack pattern that used to live in 0003-0014
patches (now reverted via 0015 for ctx, and architecturally retired
per daedalus-fourier PR #37's measurement-correction). Per-block
synchronous Vulkan dispatch from libavcodec was structurally non-
competitive; per-MB CPU-side observation feeding a per-frame batch
submit is the right shape.
Two new fields in H264Context (appended at end of struct; no ABI
surface visible to non-libavcodec callers since H264Context is
internal — declared in h264dec.h, not h264.h). One new exported
function ff_h264_set_mb_inspect_cb to set them.
Zero behaviour change when cb == NULL (the default): one load +
one branch per MB in the decoder hot path, both branch-predicted
to fall through.
Used by:
- daedalus-decoder/tools/daedalus_decode_h264 (PR-A1b)
- daedalus-v4l2 daemon shadow-mode path (PR-Q3a.1+)
The CLI static-links libavcodec.a so symbol visibility doesn't matter
there. The daemon dlopens libavcodec.so.62 and resolves the callback
via dlsym, so the symbol MUST be exported — added to libavcodec.v
explicitly (FFmpeg's default version script hides every `ff_*` symbol
as LOCAL behind a glob).
Refs reauktion/daedalus-decoder!12 (Stage 2 PR-b complete).
---
libavcodec/h264_mb.c | 20 ++++++++++++++++++++
libavcodec/h264dec.h | 26 ++++++++++++++++++++++++++
libavcodec/libavcodec.v | 1 +
3 files changed, 47 insertions(+)
--- a/libavcodec/h264dec.h
+++ b/libavcodec/h264dec.h
@@ -334,6 +334,16 @@
int pic_order_cnt_bit_size;
} H264SliceContext;
+/* Per-MB inspection callback type — see ff_h264_set_mb_inspect_cb()
+ * below. Fired by ff_h264_hl_decode_mb after the existing pixel work
+ * for every macroblock in coded order. Receives a const H264Context*
+ * so the callback can inspect any slice/picture state (h->slice_ctx
+ * for current slice, h->cur_pic.f->data[plane] for reconstructed
+ * samples, etc.). */
+typedef void (*ff_h264_mb_inspect_cb)(void *opaque,
+ const struct H264Context *h,
+ int mb_x, int mb_y);
+
/**
* H264Context
*/
@@ -579,6 +589,10 @@
int non_gray; ///< Did we encounter a intra frame after a gray gap frame
int noref_gray;
int skip_gray;
+
+ /* Per-MB inspection hook — set via ff_h264_set_mb_inspect_cb. */
+ ff_h264_mb_inspect_cb mb_inspect_cb;
+ void *mb_inspect_opaque;
} H264Context;
extern const uint16_t ff_h264_mb_sizes[4];
@@ -607,6 +621,16 @@
const H2645NAL *nal, void *logctx);
void ff_h264_hl_decode_mb(const H264Context *h, H264SliceContext *sl);
+
+/**
+ * Install an opt-in per-MB inspection callback that fires from
+ * ff_h264_hl_decode_mb after each macroblock's pixel work. Default
+ * is NULL (no callback installed); the check is a single branch on
+ * the decoder hot path. See ff_h264_mb_inspect_cb for signature.
+ */
+void ff_h264_set_mb_inspect_cb(AVCodecContext *avctx,
+ ff_h264_mb_inspect_cb cb, void *opaque);
+
void ff_h264_decode_init_vlc(void);
/**
--- a/libavcodec/h264_mb.c
+++ b/libavcodec/h264_mb.c
@@ -815,4 +815,20 @@
hl_decode_mb_simple_16(h, sl);
} else
hl_decode_mb_simple_8(h, sl);
+
+ /* Per-MB inspection callback (opt-in via ff_h264_set_mb_inspect_cb).
+ * Fired AFTER pixel work — reconstructed samples are in
+ * h->cur_pic.f->data[plane] at the MB's raster position by the
+ * time this runs. Callback may inspect slice context via
+ * h->slice_ctx + sl->mb_xy, coeffs via sl->mb, etc. */
+ if (h->mb_inspect_cb)
+ h->mb_inspect_cb(h->mb_inspect_opaque, h, sl->mb_x, sl->mb_y);
+}
+
+void ff_h264_set_mb_inspect_cb(AVCodecContext *avctx,
+ ff_h264_mb_inspect_cb cb, void *opaque)
+{
+ H264Context *h = avctx->priv_data;
+ h->mb_inspect_cb = cb;
+ h->mb_inspect_opaque = opaque;
}
--- a/libavcodec/libavcodec.v
+++ b/libavcodec/libavcodec.v
@@ -3,6 +3,7 @@
av_*;
avcodec_*;
avpriv_*;
+ ff_h264_set_mb_inspect_cb;
avsubtitle_free;
local:
*;
@@ -0,0 +1,88 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Markus Fritsche <mfritsche@reauktion.de>
Date: Tue, 26 May 2026 07:30:00 +0200
Subject: [PATCH] avcodec/h264: preserve sl->mb coefficients for the inspection
callback (companion to 0016)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Patch 0016 adds a per-MB inspection callback fired at the end of
ff_h264_hl_decode_mb. By that time the IDCT-add path has already
zeroed sl->mb (FFmpeg's convention — see ff_h264_idct_add_neon and
friends), so consumers reading coefficients from the callback get
zeros.
Add a coefficient side buffer in H264Context, populated at the
START of ff_h264_hl_decode_mb (before any IDCT runs) with a single
memcpy from sl->mb. The post-pixel-work callback (still in 0016)
can then read both:
- the side-buffer coefficients (= just-entropy-decoded, pre-IDCT)
- the reconstructed pixels in h->cur_pic.f->data (= P + IDCT(C),
pre-deblock for this MB)
and the consumer can derive P = pixels IDCT(C) for daedalus-
decoder's frame-major dispatch.
Memcpy is gated on (h->mb_inspect_cb != NULL) — zero overhead when
no consumer is registered. Buffer size = sizeof(int16_t) * 16 * 48
= 1536 bytes per H264Context (fits in one cache line family;
allocated once at H264Context lifetime, reused per MB).
8-bit path only. High-bit-depth H.264 uses the upper half of
sl->mb (int16_t[16 * 48 * 2] declared; the * 2 reserves space for
the high-depth case); preserving the high-depth coefficients
correctly would need a wider side buffer. Punted for now — the
daedalus-decoder consumer is 8-bit-only.
Single-threaded decode assumed at the consumer side (avctx->
thread_count = 1). Multi-slice / multi-threaded streams would
race on the single side buffer — that's an explicit limitation of
the inspection mechanism, documented in 0016's comment block.
Future extension: per-H264SliceContext side buffers.
Used by:
- daedalus-decoder/tools/daedalus_decode_h264 PR-A3+ (CLI test
harness extracts coefficients here for daedalus-decoder
IDCT validation on real H.264 streams).
Refs reauktion/daedalus-decoder!14 (PR-A2 callback wiring).
---
libavcodec/h264_mb.c | 9 +++++++++
libavcodec/h264dec.h | 8 ++++++++
2 files changed, 17 insertions(+)
--- a/libavcodec/h264dec.h
+++ b/libavcodec/h264dec.h
@@ -593,6 +593,14 @@
/* Per-MB inspection hook — set via ff_h264_set_mb_inspect_cb. */
ff_h264_mb_inspect_cb mb_inspect_cb;
void *mb_inspect_opaque;
+
+ /* Per-MB coefficient side buffer — populated at the start of
+ * ff_h264_hl_decode_mb so the post-pixel-work inspection callback
+ * can read the just-entropy-decoded coefficients before IDCT-add
+ * zeros sl->mb. 16 blocks × 48 int16 = libavcodec sl->mb size
+ * (matches DECLARE_ALIGNED(16, int16_t, mb)[16 * 48 * 2] for the
+ * 8-bit half; high-bit-depth paths skip this — see h264_mb.c). */
+ DECLARE_ALIGNED(16, int16_t, mb_inspect_coeffs)[16 * 48];
} H264Context;
extern const uint16_t ff_h264_mb_sizes[4];
--- a/libavcodec/h264_mb.c
+++ b/libavcodec/h264_mb.c
@@ -801,6 +801,15 @@
{
const int mb_xy = sl->mb_xy;
const int mb_type = h->cur_pic.mb_type[mb_xy];
+
+ /* Snapshot just-entropy-decoded coefficients before IDCT-add
+ * destroys them. Only when an inspection callback is registered
+ * — zero cost otherwise. 8-bit path only (high-bit-depth uses
+ * the upper half of sl->mb which we don't preserve here). */
+ if (h->mb_inspect_cb && !h->pixel_shift)
+ memcpy((int16_t *) (uintptr_t) h->mb_inspect_coeffs, sl->mb,
+ sizeof(((H264Context *) NULL)->mb_inspect_coeffs));
+
int is_complex = CONFIG_SMALL || sl->is_complex ||
IS_INTRA_PCM(mb_type) || sl->qscale == 0;
+7 -6
View File
@@ -33,12 +33,10 @@ FFMPEG_VERSION=8.1
# epoch 2 matches Debian's stock ffmpeg (currently 7:7.1.x in trixie);
# +rfourier suffix to avoid colliding with upstream/Debian rebuilds.
PKGVER=2:${FFMPEG_VERSION}+rfourier+gb57fbbe
PKGREL=11 # pkgrel=11libavcodec.so daedalus ctx flipped no_qpu → qpu-capable (PR #36 bench: QPU 4.30x on 1080p hot-path sum, 2026-05-25)
# (cycle 9 of the daedalus-v4l2#11 step 2 substitution arc; closes
# the libavcodec.so substitution sequence 6 IDCT4 / 7 IDCT8 /
# 8 luma-v deblock / 9 qpel mc20). Pulls daedalus-fourier PR #2
# which extends the public API with
# daedalus_recipe_dispatch_h264_qpel_mc20. (2026-05-23)
PKGREL=15 # pkgrel=15export ff_h264_set_mb_inspect_cb via libavcodec.v so
# dlsym consumers (daedalus-v4l2 daemon shadow_decoder, PR-Q3a.1)
# can resolve the symbol; static-link CLI was unaffected. No
# behaviour change to existing decode path. (2026-05-26)
# daedalus-fourier pin. 209a421 = daedalus-fourier PR #2 merge — public
# API now exposes daedalus_recipe_dispatch_h264_qpel_mc20 +
@@ -81,6 +79,9 @@ patch -Np1 -i "$HERE/0011-h264-chroma-dc-hadamard-daedalus-fourier.patch"
patch -Np1 -i "$HERE/0012-h264-qpel-rest-daedalus-fourier.patch"
patch -Np1 -i "$HERE/0013-h264-deblock-chroma-intra-daedalus-fourier.patch"
patch -Np1 -i "$HERE/0014-h264-ctx-qpu-capable.patch"
patch -Np1 -i "$HERE/0015-h264-ctx-revert-to-no-qpu.patch"
patch -Np1 -i "$HERE/0016-h264-mb-inspect-callback.patch"
patch -Np1 -i "$HERE/0017-h264-mb-coeffs-side-buffer.patch"
# --- daedalus-fourier: fetch + build static .a with PIC, install to a
# per-build prefix; libavcodec.so links it into the shared object so
+18
View File
@@ -1,3 +1,21 @@
ffmpeg-v4l2-request-fourier (2:8.1+rfourier+gb57fbbe-15) bookworm trixie; urgency=medium
* Amend 0016-h264-mb-inspect-callback.patch to also add
ff_h264_set_mb_inspect_cb to libavcodec/libavcodec.v so the
symbol is exported (GLOBAL) on the shipped libavcodec.so.62.
Without this, FFmpeg's default version script hides every ff_*
symbol behind a glob → LOCAL → dlsym() returns NULL. The CLI
consumer (daedalus_decode_h264) was unaffected because it
static-links libavcodec.a; the daedalus-v4l2 daemon (PR-Q3a.1
shadow_decoder path) dlopens libavcodec.so.62 and needs the
symbol resolvable at runtime.
* No behaviour change to existing decode path. Callback is still
opt-in via the function pointer (NULL default), so paying the
one-load-one-branch cost only when a consumer has explicitly
installed an inspection callback.
-- Markus Fritsche <mfritsche@reauktion.de> Tue, 26 May 2026 15:00:00 +0200
ffmpeg-v4l2-request-fourier (2:8.1+rfourier+gb57fbbe-10) bookworm trixie; urgency=medium
* Add 0007-h264-qpel-mc20-daedalus-fourier.patch —