Repository navigation
Conversation
lusoris
force-pushed
the
fix/cuda-buffer-alloc-oom
branch
from
October 2, 2026 05:24
88049ba to
9267326
Compare
…s out of memory vmaf_cuda_buffer_alloc() wraps cuMemAlloc() in CHECK_CUDA, which prints the error and calls assert(0). When another process holds the VRAM, cuMemAlloc() returns CUDA_ERROR_OUT_OF_MEMORY and a second run is killed with SIGABRT instead of failing: "vmaf_cuda_buffer_alloc: Assertion `0' failed". The function already returns negative errno values for its other failures. Call cuMemAlloc() directly, pop the context, free the half-built buffer and return -ENOMEM for CUDA_ERROR_OUT_OF_MEMORY and -EIO for any other error. *p_buf is now assigned only on success; before, it was set before the allocation. The ffnvcodec headers do not name CUDA_ERROR_OUT_OF_MEMORY, so the value 2 gets a local macro. All callers (the init functions of the motion, vif and adm extractors) already treat a non-zero return by jumping to their free_ref label, which releases only buffers that are not NULL; since *p_buf stays NULL on failure that holds. integer_motion_cuda.c leaked its write_score_parameters allocation on that path and now frees it. Reproduction: a second process holds all but 700, 1000 or 3000 MiB of the RTX 4090 and vmaf runs on two 4096x2160 8-bit frames with --gpumask 0. Unpatched master aborts with exit status 134 at 700 and 1000 MiB; with this change the same runs print "problem reading pictures" and exit with status 254; with 3000 MiB free both succeed. test_cuda_buffer_alloc requests SIZE_MAX / 2 bytes and checks for -ENOMEM, a NULL buffer and a working state afterwards; on unpatched master it aborts. This only converts this call site: CHECK_CUDA has 204 call sites under libvmaf/src on master (203 after this change) and the other sites still abort. Co-Authored-By: Claude Opus 5.5 <[email protected]>
lusoris
force-pushed
the
fix/cuda-buffer-alloc-oom
branch
from
October 2, 2026 18:45
9267326 to
d994c67
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
vmaf_cuda_buffer_alloc()aborts the process whencuMemAlloc()fails. With another process holding the VRAM, a secondvmafrun dies with SIGABRT instead of exiting with an error. This addresses #1420 for this call; it is not a general fix, see the last section.Cause
vmaf_cuda_buffer_alloc()(libvmaf/src/cuda/common.c) callscuMemAlloc()throughCHECK_CUDA, which prints the error and callsassert(0)(cuda_helper.cuh). When the device is out of memory the driver returnsCUDA_ERROR_OUT_OF_MEMORY, the macro asserts, and the assertion text in the report (common.c:166: vmaf_cuda_buffer_alloc: Assertion '0' failed) is this call. The function documents negative errno returns and already uses them for its argument checks and for a failedcalloc().Reproducer
Master
8e7a1ac4e, release build with-Denable_cuda=true, RTX 4090 (24 GiB), CUDA 13.4. A second process takes all but N MiB of the device and sleeps:and
vmafscores two 4096x2160 8-bit frames with--gpumask 0(-rand-dthe same file,--frame_cnt 2) while it runs:Assertion '0' failed, exit status 134 (SIGABRT)problem reading pictures, exit status 254Assertion '0' failed, exit status 134problem reading pictures, exit status 254On this branch the run also prints
libvmaf ERROR context could not be synchronizedandproblem flushing context.Fix
vmaf_cuda_buffer_alloc()callscuMemAlloc()directly, pops the context in every case, and on failure frees the half-built buffer and returns-ENOMEMforCUDA_ERROR_OUT_OF_MEMORYand-EIOfor any other error.*p_bufis now assigned only on success; before, it was set before the allocation. The ffnvcodec headers do not nameCUDA_ERROR_OUT_OF_MEMORY, so the value 2 gets a local macro.Callers:
vmaf_cuda_buffer_alloc()is called from the init functions of the motion, vif and adm extractors. All three already treat a non-zero return by jumping tofree_ref, which releases only buffers that are not NULL, and*p_bufstays NULL on failure, so none dereferences or frees a half-built buffer.integer_motion_cuda.cdid leak itswrite_score_parametersallocation on that path;free_refnow frees it.Tests
test_cuda_buffer_alloc(new,libvmaf/test/, needs a CUDA device) checks the-EINVALreturns, then requestsSIZE_MAX / 2bytes and requires-ENOMEM, a NULL buffer, and a working state afterwards (a 1 MiB allocation succeeds and is freed). On unpatched master the second test aborts with the assertion above; on this branch both pass. It does not need a second process. Out of memory with a real second process is not tested automatically; the reproducer above is the check.Validation
x86-64 Linux, GCC 16.2.1, nvcc 13.4, RTX 4090, release build on master
8e7a1ac4ewith-Denable_cuda=true.Rebased on master
9e48141b(2026-10-02). On that base, with-Denable_cuda=true -Denable_float=true -Denable_checkasm=true,meson testgives 28 of 29 (the failure istest_cuda_pic_preallocation, which also fails on unpatched9e48141b, 27 of 28 there), and the CPU-Db_sanitize=address,undefined -Db_lto=falsebuild gives 22 of 25 on master and here (test_predictandtest_pic_preallocationabort under LeakSanitizer andcheckasmaborts on a heap-buffer-overflow inadm_dwt2_16(integer_adm.c:2603); all three also abort on unpatched9e48141b). The CLI output of the three Netflix pairs (--gpumask 0,vmaf_v0.6.1) against unpatched9e48141b: identical on all three pairs. The sweeps and the other measurements below were taken on8e7a1ac4eand not repeated; the files and x86 code paths they depend on are unchanged since (the one upstream change in between is an arm64-only ADM kernel).meson test: 27 of 28 pass on master, 28 of 29 here (the new test). The failure on both istest_cuda_pic_preallocation, which segfaults in the host-pinned case on unpatched master; this change does not touch it.--gpumask 0,vmaf_v0.6.1) on the three Netflix pairs: identical to master, per frame and pooled.Not covered
This converts one call site.
CHECK_CUDAhas 204 call sites underlibvmaf/srcon master (203 after this change); a failure in any other driver call, including the other allocations (cuMemAllocPitchfor pictures,cuMemHostAlloc,cuModuleLoadData,cuStreamCreate), still asserts. Making the macro return an error would touch every caller and is a separate change. Wheninit_fex_cuda()fails, the streams, events and modules it created before the allocation are not released.Overlap: no open PR changes
common.c. #1612 also editsfree_refininteger_motion_cuda.c(it freessad_host);git merge-treeagainstpull/1612/headis clean and the mergedfree_reffrees each allocation once. #1614, #1553, #1613 and #1619 add tests tolibvmaf/test/meson.build: also clean.The workflow run on this PR needs a maintainer's approval.