Skip to content

cuda: return an error from vmaf_cuda_buffer_alloc() when the device is out of memory - #1646

Open
lusoris wants to merge 1 commit into
Netflix:masterfrom
VMAFx:fix/cuda-buffer-alloc-oom
Open

lusoris wants to merge 1 commit into
Netflix:masterfrom
VMAFx:fix/cuda-buffer-alloc-oom

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026 •

Copy link
Copy Markdown

vmaf_cuda_buffer_alloc() aborts the process when cuMemAlloc() fails. With another process holding the VRAM, a second vmaf run dies with SIGABRT instead of exiting with an error. This addresses #1420 for this call; it is not a general fix, see the last section.

Cause

vmaf_cuda_buffer_alloc() (libvmaf/src/cuda/common.c) calls cuMemAlloc() through CHECK_CUDA, which prints the error and calls assert(0) (cuda_helper.cuh). When the device is out of memory the driver returns CUDA_ERROR_OUT_OF_MEMORY, the macro asserts, and the assertion text in the report (common.c:166: vmaf_cuda_buffer_alloc: Assertion '0' failed) is this call. The function documents negative errno returns and already uses them for its argument checks and for a failed calloc().

Reproducer

Master 8e7a1ac4e, release build with -Denable_cuda=true, RTX 4090 (24 GiB), CUDA 13.4. A second process takes all but N MiB of the device and sleeps:

size_t free_b, tot; CUcontext c; CUdevice d; CUdeviceptr p;
cuInit(0); cuDeviceGet(&d, 0); cuDevicePrimaryCtxRetain(&c, d); cuCtxSetCurrent(c);
cuMemGetInfo(&free_b, &tot); cuMemAlloc(&p, free_b - ((size_t)N << 20)); sleep(25);

and vmaf scores two 4096x2160 8-bit frames with --gpumask 0 (-r and -d the same file, --frame_cnt 2) while it runs:

MiB left free master this branch
700 Assertion '0' failed, exit status 134 (SIGABRT) problem reading pictures, exit status 254
1000 Assertion '0' failed, exit status 134 problem reading pictures, exit status 254
3000 exit status 0 exit status 0

On this branch the run also prints libvmaf ERROR context could not be synchronized and problem flushing context.

Fix

vmaf_cuda_buffer_alloc() calls cuMemAlloc() directly, pops the context in every case, and on failure frees the half-built buffer and returns -ENOMEM for CUDA_ERROR_OUT_OF_MEMORY and -EIO for any other error. *p_buf is now assigned only on success; before, it was set before the allocation. The ffnvcodec headers do not name CUDA_ERROR_OUT_OF_MEMORY, so the value 2 gets a local macro.

Callers: vmaf_cuda_buffer_alloc() is called from the init functions of the motion, vif and adm extractors. All three already treat a non-zero return by jumping to free_ref, which releases only buffers that are not NULL, and *p_buf stays NULL on failure, so none dereferences or frees a half-built buffer. integer_motion_cuda.c did leak its write_score_parameters allocation on that path; free_ref now frees it.

Tests

test_cuda_buffer_alloc (new, libvmaf/test/, needs a CUDA device) checks the -EINVAL returns, then requests SIZE_MAX / 2 bytes and requires -ENOMEM, a NULL buffer, and a working state afterwards (a 1 MiB allocation succeeds and is freed). On unpatched master the second test aborts with the assertion above; on this branch both pass. It does not need a second process. Out of memory with a real second process is not tested automatically; the reproducer above is the check.

Validation

x86-64 Linux, GCC 16.2.1, nvcc 13.4, RTX 4090, release build on master 8e7a1ac4e with -Denable_cuda=true.

Rebased on master 9e48141b (2026-10-02). On that base, with -Denable_cuda=true -Denable_float=true -Denable_checkasm=true, meson test gives 28 of 29 (the failure is test_cuda_pic_preallocation, which also fails on unpatched 9e48141b, 27 of 28 there), and the CPU -Db_sanitize=address,undefined -Db_lto=false build gives 22 of 25 on master and here (test_predict and test_pic_preallocation abort under LeakSanitizer and checkasm aborts on a heap-buffer-overflow in adm_dwt2_16 (integer_adm.c:2603); all three also abort on unpatched 9e48141b). The CLI output of the three Netflix pairs (--gpumask 0, vmaf_v0.6.1) against unpatched 9e48141b: identical on all three pairs. The sweeps and the other measurements below were taken on 8e7a1ac4e and not repeated; the files and x86 code paths they depend on are unchanged since (the one upstream change in between is an arm64-only ADM kernel).

  • meson test: 27 of 28 pass on master, 28 of 29 here (the new test). The failure on both is test_cuda_pic_preallocation, which segfaults in the host-pinned case on unpatched master; this change does not touch it.
  • CUDA CLI output (--gpumask 0, vmaf_v0.6.1) on the three Netflix pairs: identical to master, per frame and pooled.
  • Not run: an ASan/UBSan build, a leak check of the failure path.

Not covered

This converts one call site. CHECK_CUDA has 204 call sites under libvmaf/src on master (203 after this change); a failure in any other driver call, including the other allocations (cuMemAllocPitch for pictures, cuMemHostAlloc, cuModuleLoadData, cuStreamCreate), still asserts. Making the macro return an error would touch every caller and is a separate change. When init_fex_cuda() fails, the streams, events and modules it created before the allocation are not released.

Overlap: no open PR changes common.c. #1612 also edits free_ref in integer_motion_cuda.c (it frees sad_host); git merge-tree against pull/1612/head is clean and the merged free_ref frees each allocation once. #1614, #1553, #1613 and #1619 add tests to libvmaf/test/meson.build: also clean.

The workflow run on this PR needs a maintainer's approval.

…s out of memory

vmaf_cuda_buffer_alloc() wraps cuMemAlloc() in CHECK_CUDA, which prints the error and calls assert(0). When another process holds the VRAM, cuMemAlloc() returns CUDA_ERROR_OUT_OF_MEMORY and a second run is killed with SIGABRT instead of failing: "vmaf_cuda_buffer_alloc: Assertion `0' failed". The function already returns negative errno values for its other failures.

Call cuMemAlloc() directly, pop the context, free the half-built buffer and return -ENOMEM for CUDA_ERROR_OUT_OF_MEMORY and -EIO for any other error. *p_buf is now assigned only on success; before, it was set before the allocation. The ffnvcodec headers do not name CUDA_ERROR_OUT_OF_MEMORY, so the value 2 gets a local macro.

All callers (the init functions of the motion, vif and adm extractors) already treat a non-zero return by jumping to their free_ref label, which releases only buffers that are not NULL; since *p_buf stays NULL on failure that holds. integer_motion_cuda.c leaked its write_score_parameters allocation on that path and now frees it.

Reproduction: a second process holds all but 700, 1000 or 3000 MiB of the RTX 4090 and vmaf runs on two 4096x2160 8-bit frames with --gpumask 0. Unpatched master aborts with exit status 134 at 700 and 1000 MiB; with this change the same runs print "problem reading pictures" and exit with status 254; with 3000 MiB free both succeed. test_cuda_buffer_alloc requests SIZE_MAX / 2 bytes and checks for -ENOMEM, a NULL buffer and a working state afterwards; on unpatched master it aborts. This only converts this call site: CHECK_CUDA has 204 call sites under libvmaf/src on master (203 after this change) and the other sites still abort.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@lusoris
lusoris force-pushed the fix/cuda-buffer-alloc-oom branch from 9267326 to d994c67 Compare October 2, 2026 18:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant