Repository navigation
Conversation
lusoris
force-pushed
the
fix/adm-float-division
branch
from
October 2, 2026 05:24
f4775e6 to
b42580b
Compare
…eciprocal With ADM_OPT_RECIP_DIVISION defined and __SSE2__ set, adm_decouple_s() computed t / o as t * rcp_s(o), where rcp_s() is one Newton step on _mm_rcp_ss(). The instruction is specified only by a relative error bound (1.5 * 2^-12) and its low bits differ between processors, so the float ADM scores depended on the CPU, and on x86 they differed from builds that already divide (MSVC, which does not define __SSE2__, and ARM). Drop the macro and the estimate path; DIVS() is the IEEE division on every target. Integer adm does not include adm_tools.c and has its own DIVS(). On a Ryzen 9 9950X3D, rcp_s(x) differs from 1.0f / x for 2562503 of the 8388608 mantissas in [1, 2), by up to 3 ulp. float_adm moves by at most 1.4e-7 per frame and score. Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
lusoris
force-pushed
the
fix/adm-float-division
branch
from
October 2, 2026 18:46
b42580b to
b480a56
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1654.
float_adm(-Denable_float=true) computedt / oinadm_decouple_s()ast * rcp_s(o), one Newton step on_mm_rcp_ss(), wheneverADM_OPT_RECIP_DIVISIONand__SSE2__were set.RCPSSis specified only by a relative error bound, so its low bits depend on the processor, and the result differed from builds that divide (MSVC, which does not define__SSE2__, and ARM). This removes the macro and the estimate path;DIVS()is a plain division on every target. The diff isadm_options.h(macro) andadm_tools.c(rcp_s, the__SSE2__nesting, the<emmintrin.h>include).What changes for users
RCPSS. Measured on master9e48141b,%.17gthrough the C API, one thread, the three Netflix pairs (54 frames): 30 frames change in at least onefloat_admoutput (adm2,adm3,aim,adm_scale0..3), 18 of them inadm2; the largest change in any output is 1.39e-07 (adm_scale2), inadm25.8e-08.vmaf_float_v0.6.1changes on 18 of 54 frames, by at most 1.2e-05 on the 0-100 scale (src01mean 76.667439368 to 76.667441136). The XML and JSON writers print six decimals, so most of this is below what they show.adm(the default models,vmaf_v0.6.1and the other integer models) is not affected:integer_adm.cdoes not includeadm_tools.cand has its ownDIVS(integer_adm.h:179).adm2,adm3,aimand thevmaf_v0.6.1score are bit-identical before and after on all 54 frames.float_admis built only with-Denable_float=true.libvmaf/src(grepfor_mm_rcp,_mm256_rcp,rcp14,vrecpe,rsqrt: one hit,adm_tools.c:45).adm_tools.c.ohas 3rcpssinstructions before and none after.Validation
Master
9e48141bagainst this branch, x86-64 Linux, Ryzen 9 9950X3D, GCC 16.2.1, Meson 1.12.1,-Denable_float=true -Denable_checkasm=true:Rebased on master
9e48141b(2026-10-02). The builds, test counts, the score comparison above and the Netflix pair comparison were re-run on it (same 18 frames inadm2, largestadm2change 5.8e-08, largest change in anyfloat_admoutput 1.39e-07,vmaf_float_v0.6.118 of 54 frames, at most 1.2e-05; integeradm,adm3,aimandvmaf_v0.6.1bit-identical on all 54 frames, and.engagement/golden_compare.pyreports all three Netflix pairs identical for the default model); the Python tests and the timing table were not.Upstream's Python tests, the way
.github/workflows/libvmaf.ymlruns them (pytest -m mainunder Python 3.11 withpython/requirements.txt), onresult_test,feature_assembler_test,quality_runner_test,local_explainer_test,vmafexec_test,feature_extractor_test,vmafexec_feature_extractor_testandroutine_test(the files that usefloat_admorvmaf_float_*), with the test resources the three Netflix pairs need: master 334 passed, 3 skipped; branch 334 passed, 3 skipped. No assertion was edited and none fails.Cost,
float_admonly,vmafCLI,--threads 1, 480 frames of 576x324 and 30 frames of 1920x1080, median of 3, interleaved, ms/frame (the host was busy, load average about 30, so read these as "no slower", not as a speed-up):Master does not need extra flags at
9e48141b.Not measured
-cpu Haswell, whoseRCPSSis an exact division) the raw instruction equals1.0f / xfor all 2^23 mantissas in [1, 2) andrcp_sdiffers from it for 1 287 375 of them, against 2 562 503 on Zen 5; that is a different implementation of the instruction, not another vendor.ADM_OPT_AVOID_ATAN-off branch (not compiled).