Skip to content

Bug: Support consumer RDNA2 (gfx1030) under ROCm/HIP backend for Ternary-Bonsai-27B #264

Description

@bumblebeejoe

Description
The custom ternary math execution pipeline inside the PrismML llama.cpp fork (prism branch) currently hard-limits execution on consumer-grade AMD RDNA2 cards (specifically the Radeon RX 6600, gfx1030) due to a strict hardware register assertion check inside the custom compiled Flash Attention kernels.

When attempting to load Ternary-Bonsai-2-27B-PQ2_0.gguf, the server crashes instantly during the graph compute phase with a core dump.

Steps to Reproduce

  1. Build the llama-prism fork using CMake with -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1030.
  2. Execute llama-server targeting the ternary layout with dynamic layer offloading (e.g., -ngl 35 -c 8192 --parallel 1 --kv-unified).
  3. Stream an initial prompt payload to the active network port.

Error Stack / Traceback
text
/home/llama-prism/ggml/src/ggml-cuda/template-instances/../fattn-common.cuh:1114: GGML_ASSERT(max_blocks_per_sm > 0) failed
Aborted (core dumped)

Rationale & Expected Behavior
While RDNA3 architectures (gfx1101) have been confirmed to run unpatched via the HIP backend, consumer RDNA2 platforms return 0 when the flash attention tile configuration queries max_blocks_per_sm inside fattn-common.cuh. Since the ternary format forces flash attention natively, the system cannot load or fall back cleanly.

We request either:

  1. A patch allowing standard matrix lookup structures to bypass the max_blocks_per_sm assertion for consumer cards.
  2. A formal Vulkan implementation of the custom type 142 fused Gated Delta Net layers so consumer AMD users can bypass the enterprise ROCm/HIP kernel constraints entirely.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions