Skip to content

feature/adm: divide in the float decouple instead of estimating the reciprocal - #1655

Open
lusoris wants to merge 1 commit into
Netflix:masterfrom
VMAFx:fix/adm-float-division
Open

lusoris wants to merge 1 commit into
Netflix:masterfrom
VMAFx:fix/adm-float-division

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026 •

Copy link
Copy Markdown

Fixes #1654.

float_adm (-Denable_float=true) computed t / o in adm_decouple_s() as t * rcp_s(o), one Newton step on _mm_rcp_ss(), whenever ADM_OPT_RECIP_DIVISION and __SSE2__ were set. RCPSS is specified only by a relative error bound, so its low bits depend on the processor, and the result differed from builds that divide (MSVC, which does not define __SSE2__, and ARM). This removes the macro and the estimate path; DIVS() is a plain division on every target. The diff is adm_options.h (macro) and adm_tools.c (rcp_s, the __SSE2__ nesting, the <emmintrin.h> include).

What changes for users

  • Float ADM scores on x86 move by about 1e-7 and no longer depend on the processor's RCPSS. Measured on master 9e48141b, %.17g through the C API, one thread, the three Netflix pairs (54 frames): 30 frames change in at least one float_adm output (adm2, adm3, aim, adm_scale0..3), 18 of them in adm2; the largest change in any output is 1.39e-07 (adm_scale2), in adm2 5.8e-08. vmaf_float_v0.6.1 changes on 18 of 54 frames, by at most 1.2e-05 on the 0-100 scale (src01 mean 76.667439368 to 76.667441136). The XML and JSON writers print six decimals, so most of this is below what they show.
  • Integer adm (the default models, vmaf_v0.6.1 and the other integer models) is not affected: integer_adm.c does not include adm_tools.c and has its own DIVS (integer_adm.h:179). adm2, adm3, aim and the vmaf_v0.6.1 score are bit-identical before and after on all 54 frames. float_adm is built only with -Denable_float=true.
  • No other reciprocal estimate exists in libvmaf/src (grep for _mm_rcp, _mm256_rcp, rcp14, vrecpe, rsqrt: one hit, adm_tools.c:45). adm_tools.c.o has 3 rcpss instructions before and none after.

Validation

Master 9e48141b against this branch, x86-64 Linux, Ryzen 9 9950X3D, GCC 16.2.1, Meson 1.12.1, -Denable_float=true -Denable_checkasm=true:

Rebased on master 9e48141b (2026-10-02). The builds, test counts, the score comparison above and the Netflix pair comparison were re-run on it (same 18 frames in adm2, largest adm2 change 5.8e-08, largest change in any float_adm output 1.39e-07, vmaf_float_v0.6.1 18 of 54 frames, at most 1.2e-05; integer adm, adm3, aim and vmaf_v0.6.1 bit-identical on all 54 frames, and .engagement/golden_compare.py reports all three Netflix pairs identical for the default model); the Python tests and the timing table were not.

meson setup build libvmaf --buildtype release -Denable_float=true -Denable_checkasm=true
ninja -C build && meson test -C build
# master 25/25, branch 25/25

meson setup build-sanitize libvmaf --buildtype debug -Db_sanitize=address,undefined -Db_lto=false -Denable_float=true -Denable_checkasm=true
ninja -C build-sanitize && meson test -C build-sanitize
# master 22 passed 3 failed, branch 22 passed 3 failed; the same three abort on both: test_predict and test_pic_preallocation (LeakSanitizer), checkasm (heap-buffer-overflow in adm_dwt2_16, integer_adm.c:2603)

Upstream's Python tests, the way .github/workflows/libvmaf.yml runs them (pytest -m main under Python 3.11 with python/requirements.txt), on result_test, feature_assembler_test, quality_runner_test, local_explainer_test, vmafexec_test, feature_extractor_test, vmafexec_feature_extractor_test and routine_test (the files that use float_adm or vmaf_float_*), with the test resources the three Netflix pairs need: master 334 passed, 3 skipped; branch 334 passed, 3 skipped. No assertion was edited and none fails.

Cost, float_adm only, vmaf CLI, --threads 1, 480 frames of 576x324 and 30 frames of 1920x1080, median of 3, interleaved, ms/frame (the host was busy, load average about 30, so read these as "no slower", not as a speed-up):

size master branch
576x324 1.86 1.71
1920x1080 20.97 20.07

Master does not need extra flags at 9e48141b.

Not measured

  • A second processor. All numbers are from one Zen 5 machine; there is no Intel measurement, so the size of the difference between vendors is not known. Under QEMU 11.1.1 user-mode emulation (-cpu Haswell, whose RCPSS is an exact division) the raw instruction equals 1.0f / x for all 2^23 mantissas in [1, 2) and rcp_s differs from it for 1 287 375 of them, against 2 562 503 on Zen 5; that is a different implementation of the instruction, not another vendor.
  • 4K, other bit depths, and the ADM_OPT_AVOID_ATAN-off branch (not compiled).
  • Builds other than x86-64 GCC: MSVC and ARM already divided and are untouched by this diff.

@lusoris
lusoris force-pushed the fix/adm-float-division branch from f4775e6 to b42580b Compare October 2, 2026 05:24
…eciprocal

With ADM_OPT_RECIP_DIVISION defined and __SSE2__ set, adm_decouple_s()
computed t / o as t * rcp_s(o), where rcp_s() is one Newton step on
_mm_rcp_ss(). The instruction is specified only by a relative error bound
(1.5 * 2^-12) and its low bits differ between processors, so the float ADM
scores depended on the CPU, and on x86 they differed from builds that
already divide (MSVC, which does not define __SSE2__, and ARM).

Drop the macro and the estimate path; DIVS() is the IEEE division on every
target. Integer adm does not include adm_tools.c and has its own DIVS().

On a Ryzen 9 9950X3D, rcp_s(x) differs from 1.0f / x for 2562503 of the
8388608 mantissas in [1, 2), by up to 3 ulp. float_adm moves by at most
1.4e-7 per frame and score.

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
@lusoris
lusoris force-pushed the fix/adm-float-division branch from b42580b to b480a56 Compare October 2, 2026 18:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

float_adm: the decouple divides by an RCPSS estimate, so the scores depend on the processor

1 participant