Skip to content

Commit a9e2541

Browse files
alexarjeclaude
andcommitted
The reference distance, recovered from the cadence source never carried
FFmpeg's source is a sign, plus or minus one whatever the distance, so a P-frame after a run of B-frames reported its whole span's displacement as one frame's --- three to four times over on B-heavy encodes. The gap back to the previous reference frame is recoverable from the picture-type cadence while decoding in display order, and past-referencing vectors now divide by it, in both the data loop and the grid. Test-first with a forced I B B P cadence: 6.0 before, 2.0 after, on a block moving 2 px a frame. The residuals that remain --- multi-reference blocks and future-referencing B vectors --- are documented where the old false claim about source's magnitude stood. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GFjgT8igZuQu4EPMvSqK6u
1 parent 8a668d7 commit a9e2541

4 files changed

Lines changed: 113 additions & 45 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
88
## [Unreleased]
99

1010
### Changed
11+
- **Motion-vector displacements are corrected for the reference cadence.** FFmpeg's
12+
`source` carries only the sign of a vector's reference, so a P-frame following a run
13+
of B-frames reported its whole multi-frame displacement as one frame's — 3× to 4× on
14+
typical B-heavy encodes. Past-referencing vectors are now divided by the
15+
display-order gap to the previous reference frame, recovered from the picture-type
16+
cadence while decoding. This changes every motion-vector number on B-frame encodes:
17+
displacements, `magnitude`, the grams and everything built on them. Two residuals
18+
remain and are documented: multi-reference blocks reaching older frames, and
19+
future-referencing B-frame vectors, which default filtering excludes.
1120
- **The motion-vector motiongrams now use the classic orientations, named for the
1221
picture.** The vertical motiongram is the tall image — the frame's width across,
1322
time running downward — and carries sideways travel; the horizontal motiongram is

‎musicalgestures/_motionvectors.py‎

Lines changed: 56 additions & 22 deletions
Original file line numberDiff line numberDiff line change
@@ -104,16 +104,20 @@ def mg_motionvectordata(self: "musicalgestures.MgVideo") -> "MgMotionVectorData"
104104
predicted. Without it a B-frame's vectors point backwards half the time and averaging
105105
them gives roughly nothing.
106106
107-
**What that normalisation cannot fix is the reference distance.** `source` records
108-
only the direction, plus or minus one, never how many frames away the reference was.
109-
An encoder with multiple reference frames will predict some blocks from two or four
110-
frames back, and those read as two or four times the per-frame displacement. Measured
111-
on a block moving 4 pixels per frame, 53 per cent of vectors came back as 4 and the
112-
rest as 8 or 16, with `source` at plus or minus one throughout. So `median_dx` and
113-
`median_dy` are medians for a reason and are robust to it; `magnitude` is a sum and is
114-
not, and inherits the over-count. Treat `magnitude` as a quantity to correlate against
115-
itself over time, which is what it was validated for, rather than as pixels per
116-
second.
107+
**The reference distance is corrected from the cadence, with two residuals.**
108+
`source` records only the direction, plus or minus one, never how many frames away
109+
the reference was, so the distance is recovered instead from the picture-type
110+
cadence: past-referencing vectors are divided by the display-order gap to the
111+
previous reference frame, which makes a P-frame after a run of B-frames read
112+
per-frame displacement rather than the whole span's. What the cadence cannot see:
113+
blocks reaching an OLDER reference through multi-reference prediction --- measured
114+
on a block moving 4 pixels per frame, such blocks read as 8 or 16 with `source` at
115+
plus or minus one throughout --- and B-frame vectors pointing at a future
116+
reference, which default filtering excludes. So `median_dx` and `median_dy` are
117+
medians for a reason and are robust to the residuals; `magnitude` is a sum and
118+
inherits what remains of the over-count. Treat `magnitude` as a quantity to
119+
correlate against itself over time, which is what it was validated for, rather
120+
than as pixels per second.
117121
118122
Returns:
119123
MgMotionVectorData: `time` in seconds, `picture_type` as 'I'/'P'/'B',
@@ -152,10 +156,27 @@ def mg_motionvectordata(self: "musicalgestures.MgVideo") -> "MgMotionVectorData"
152156
magnitude: list[float] = []
153157
med_dx: list[float] = []
154158
med_dy: list[float] = []
159+
#: The reference distance, recovered from the cadence. `source` carries only the
160+
#: SIGN of the reference --- measured on a real encode it is plus or minus one
161+
#: throughout, whatever the distance --- so a P-frame that follows a run of
162+
#: B-frames reports its whole multi-frame displacement as one frame's. What IS
163+
#: recoverable while decoding in display order is the gap back to the previous
164+
#: reference frame (I or P), and dividing the past-referencing vectors by that gap
165+
#: corrects the cadence part exactly. Two residuals remain and are documented
166+
#: rather than hidden: blocks reaching an OLDER reference through multi-reference
167+
#: prediction, and B-frame vectors pointing at a future reference --- the toolbox
168+
#: excludes B-frames by default, and their vectors were already the worse signal.
169+
display_idx = -1
170+
last_ref_idx = None
155171
for frame in container.decode(stream):
172+
display_idx += 1
173+
kind = _PICTURE_TYPES.get(int(frame.pict_type), "?")
174+
past_gap = 1.0 if last_ref_idx is None else max(1.0, display_idx - last_ref_idx)
175+
if kind in ("I", "P"):
176+
last_ref_idx = display_idx
156177
time.append(float(frame.pts * stream.time_base) if frame.pts is not None
157178
else (time[-1] if time else 0.0))
158-
kinds.append(_PICTURE_TYPES.get(int(frame.pict_type), "?"))
179+
kinds.append(kind)
159180
vectors = frame.side_data.get("MOTION_VECTORS")
160181
table = vectors.to_ndarray() if vectors is not None else None
161182
if table is None or len(table) == 0:
@@ -165,20 +186,20 @@ def mg_motionvectordata(self: "musicalgestures.MgVideo") -> "MgMotionVectorData"
165186
med_dy.append(0.0)
166187
continue
167188
scale = np.maximum(table["motion_scale"].astype(np.float64), 1)
168-
#: Two corrections in one division, and both are needed to get a number that
169-
#: means "which way did this move, per frame".
170-
#:
171-
#: ffmpeg defines the vector as `src = dst + motion / motion_scale`, so it
172-
#: points from where the block is now back to where it came from: content
173-
#: travelling right carries a NEGATIVE motion_x. And `source` is negative when
174-
#: the reference is an earlier frame, positive when it is a later one, with its
175-
#: magnitude the distance in frames. Dividing by `source` undoes both --- the
176-
#: backwards sense of the vector and the backwards sense of a future reference
177-
#: cancel --- and scales a multi-frame prediction down to one frame.
189+
#: Two corrections in one division: ffmpeg defines the vector as
190+
#: `src = dst + motion / motion_scale`, so it points from where the block is
191+
#: now back to where it came from --- content travelling right carries a
192+
#: NEGATIVE motion_x --- and `source` is negative for a past reference,
193+
#: positive for a future one. Dividing by `source` undoes both senses; the
194+
#: cadence gap above then scales the past-referencing multi-frame predictions
195+
#: down to one frame.
178196
source = table["source"].astype(np.float64)
179197
source[source == 0] = -1
180198
dx = table["motion_x"].astype(np.float64) / scale / source
181199
dy = table["motion_y"].astype(np.float64) / scale / source
200+
past = source < 0
201+
dx[past] /= past_gap
202+
dy[past] /= past_gap
182203
area = table["w"].astype(np.float64) * table["h"]
183204
counts.append(len(table))
184205
magnitude.append(float((np.hypot(dx, dy) * area).sum()))
@@ -282,15 +303,25 @@ def motion_vector_grid(filename, deterministic=False, threshold=0.0):
282303
rows = max(1, -(-height // _GRID))
283304

284305
previous_time = 0.0
306+
display_idx = -1
307+
last_ref_idx = None
285308
try:
286309
for frame in container.decode(stream):
310+
display_idx += 1
311+
kind = _PICTURE_TYPES.get(int(frame.pict_type), "?")
312+
#: The cadence gap to the previous reference frame; see motionvectordata
313+
#: for why `source` cannot provide the distance itself.
314+
past_gap = (1.0 if last_ref_idx is None
315+
else max(1.0, display_idx - last_ref_idx))
316+
if kind in ("I", "P"):
317+
last_ref_idx = display_idx
287318
vx = np.zeros((rows, cols), np.float64)
288319
vy = np.zeros((rows, cols), np.float64)
289320
w = np.zeros((rows, cols), np.float64)
290321
t = (float(frame.pts * stream.time_base)
291322
if frame.pts is not None else previous_time)
292323
previous_time = t
293-
is_p = _PICTURE_TYPES.get(int(frame.pict_type), "?") == "P"
324+
is_p = kind == "P"
294325
mvs = frame.side_data.get("MOTION_VECTORS")
295326
table = mvs.to_ndarray() if mvs is not None else None
296327
if table is not None and len(table):
@@ -299,6 +330,9 @@ def motion_vector_grid(filename, deterministic=False, threshold=0.0):
299330
source[source == 0] = -1
300331
dx = table["motion_x"].astype(np.float64) / scale / source
301332
dy = table["motion_y"].astype(np.float64) / scale / source
333+
past = source < 0
334+
dx[past] /= past_gap
335+
dy[past] /= past_gap
302336
area = (table["w"].astype(np.float64) * table["h"])
303337
if threshold > 0:
304338
keep = np.hypot(dx, dy) >= threshold

‎tests/_synth.py‎

Lines changed: 12 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -84,7 +84,7 @@ def decaying_tone(t60, sr=SR, f=440.0, dur=None, onset=0.05):
8484

8585

8686
def moving_block_video(path, dx=4, dy=0, frames=40, size=(320, 240), block=48,
87-
fps=25, noise=0):
87+
fps=25, noise=0, x264opts=None):
8888
"""An H.264 clip of one textured block translating by an exact (dx, dy) per frame.
8989
9090
Motion vectors are a claim about displacement, so testing them needs footage whose
@@ -128,10 +128,14 @@ def moving_block_video(path, dx=4, dy=0, frames=40, size=(320, 240), block=48,
128128
frame[y:y + block, x:x + block] = texture
129129
raw += frame.tobytes()
130130

131+
#: `x264opts` pins the encoder's structure where a test needs it deterministic ---
132+
#: for example "bframes=2:b-adapt=0:ref=1" for a fixed I B B P cadence with single
133+
#: reference, so every P-frame's reference is exactly 3 display frames back.
134+
extra = ["-x264opts", x264opts] if x264opts else []
131135
subprocess.run(
132136
["ffmpeg", "-y", "-loglevel", "error", "-f", "rawvideo", "-pix_fmt", "gray",
133137
"-s", f"{W}x{H}", "-r", str(fps), "-i", "pipe:0",
134-
"-c:v", "libx264", "-pix_fmt", "yuv420p", "-g", "12", str(path)],
138+
"-c:v", "libx264", "-pix_fmt", "yuv420p", "-g", "12", *extra, str(path)],
135139
input=bytes(raw), check=True)
136140
return str(path)
137141

@@ -154,7 +158,7 @@ def intra_only_video(path, frames=12, size=(160, 120), fps=25):
154158

155159

156160
def oscillating_block_video(path, amplitude=40, period=12, frames=96, size=(320, 240),
157-
block=48, fps=25):
161+
block=48, fps=25, x264opts=None):
158162
"""A block that goes back and forth, so its direction cancels but its motion does not.
159163
160164
The counterpart to `moving_block_video`: both leave the same amount of movement in
@@ -176,9 +180,13 @@ def oscillating_block_video(path, amplitude=40, period=12, frames=96, size=(320,
176180
x = max(0, min(W - block, x))
177181
frame[y0:y0 + block, x:x + block] = texture
178182
raw += frame.tobytes()
183+
#: `x264opts` pins the encoder's structure where a test needs it deterministic ---
184+
#: for example "bframes=2:b-adapt=0:ref=1" for a fixed I B B P cadence with single
185+
#: reference, so every P-frame's reference is exactly 3 display frames back.
186+
extra = ["-x264opts", x264opts] if x264opts else []
179187
subprocess.run(
180188
["ffmpeg", "-y", "-loglevel", "error", "-f", "rawvideo", "-pix_fmt", "gray",
181189
"-s", f"{W}x{H}", "-r", str(fps), "-i", "pipe:0",
182-
"-c:v", "libx264", "-pix_fmt", "yuv420p", "-g", "12", str(path)],
190+
"-c:v", "libx264", "-pix_fmt", "yuv420p", "-g", "12", *extra, str(path)],
183191
input=bytes(raw), check=True)
184192
return str(path)

‎tests/test_motionvectordata.py‎

Lines changed: 36 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -39,25 +39,14 @@ def test_recovers_the_horizontal_displacement_it_was_given(self, moving_right):
3939
assert np.median(moved) == pytest.approx(4, abs=1.0)
4040

4141
def test_recovers_the_vertical_displacement_it_was_given(self, moving_down):
42-
"""3 px a frame, within the spread that the reference distance introduces.
43-
44-
**The tolerance is wide for a reason, and the reason is a real limitation.**
45-
`motion_x`/`motion_y` are divided by `source`, and FFmpeg's `source` carries only
46-
the SIGN of the reference -- past or future -- not its DISTANCE. A vector
47-
referencing a frame two back therefore reports twice the per-frame displacement,
48-
and an encode mixing distance-1 and distance-2 references gives a median between
49-
the two.
50-
51-
That is not hypothetical: ffmpeg 6.1.1 emits only distance-1 references here and
52-
this reads exactly 3.00, while CI's newer build mixes them and reads 4.5 -- on
53-
all nine matrix jobs, so it is the encoder and not flakiness. The test was
54-
written against one encoder and had never run on another, because CI did not
55-
install PyAV until 1.21.0.
56-
57-
The narrow assertions worth keeping are elsewhere in this class: sign, and the
58-
absence of a component on the unmoved axis. Whether the reader should divide by
59-
the reference distance is a question about the measure rather than about this
60-
test.
42+
"""3 px a frame, within the spread the multi-reference residual introduces.
43+
44+
The cadence part of the reference distance is corrected now --- see
45+
Test_reference_distance --- but a block reaching an OLDER reference through
46+
multi-reference prediction still over-reports, and which blocks do that is the
47+
encoder's choice: ffmpeg 6.1.1 emits only distance-1 references here and reads
48+
exactly 3.00, while CI's newer build mixes references. The tolerance covers
49+
that residual, not the cadence, which is asserted tightly below.
6150
"""
6251
mv = musicalgestures.MgVideo(moving_down).motionvectordata()
6352
moved = mv.median_dy[mv.n_vectors > 0]
@@ -146,3 +135,31 @@ class Test_it_is_cheap:
146135
def test_reports_how_many_frames_carried_vectors(self, moving_right):
147136
mv = musicalgestures.MgVideo(moving_right).motionvectordata()
148137
assert 0 < int((mv.n_vectors > 0).sum()) <= len(mv.time)
138+
139+
140+
class Test_reference_distance:
141+
"""The cadence part of the reference distance is corrected, and says so in numbers.
142+
143+
FFmpeg's `source` carries only the sign of the reference, so a P-frame that follows
144+
a run of B-frames reports its whole multi-frame displacement as if it were one
145+
frame's. The distance to the previous REFERENCE frame is recoverable from the
146+
picture-type cadence while decoding, and that is the correction under test: with a
147+
forced I B B P cadence and single reference, every P-frame's vectors span exactly
148+
3 display frames, and a block moving 2 px a frame must read 2, not 6.
149+
150+
What this cannot correct, and is not tested: blocks reaching an older reference
151+
through multi-reference prediction (suppressed here with ref=1), and B-frame
152+
vectors that point at a future reference, which the toolbox excludes by default.
153+
"""
154+
155+
@pytest.fixture(scope="class")
156+
def cadenced(self, tmp_path_factory):
157+
return moving_block_video(tmp_path_factory.mktemp("mv") / "cadence.mp4",
158+
dx=2, dy=0, frames=60,
159+
x264opts="bframes=2:b-adapt=0:ref=1")
160+
161+
def test_a_p_frame_after_b_frames_reports_per_frame_displacement(self, cadenced):
162+
mv = musicalgestures.MgVideo(cadenced).motionvectordata()
163+
p = (mv.picture_type == "P") & (mv.n_vectors > 0) & (mv.median_dx != 0)
164+
assert p.sum() > 5
165+
assert np.median(mv.median_dx[p]) == pytest.approx(2, abs=0.75)

0 commit comments

Comments
 (0)