Description
The custom ternary math execution pipeline inside the PrismML llama.cpp fork (prism branch) currently hard-limits execution on consumer-grade AMD RDNA2 cards (specifically the Radeon RX 6600, gfx1030) due to a strict hardware register assertion check inside the custom compiled Flash Attention kernels.
When attempting to load Ternary-Bonsai-2-27B-PQ2_0.gguf, the server crashes instantly during the graph compute phase with a core dump.
Steps to Reproduce
- Build the
llama-prism fork using CMake with -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1030.
- Execute
llama-server targeting the ternary layout with dynamic layer offloading (e.g., -ngl 35 -c 8192 --parallel 1 --kv-unified).
- Stream an initial prompt payload to the active network port.
Error Stack / Traceback
text
/home/llama-prism/ggml/src/ggml-cuda/template-instances/../fattn-common.cuh:1114: GGML_ASSERT(max_blocks_per_sm > 0) failed
Aborted (core dumped)
Rationale & Expected Behavior
While RDNA3 architectures (gfx1101) have been confirmed to run unpatched via the HIP backend, consumer RDNA2 platforms return 0 when the flash attention tile configuration queries max_blocks_per_sm inside fattn-common.cuh. Since the ternary format forces flash attention natively, the system cannot load or fall back cleanly.
We request either:
- A patch allowing standard matrix lookup structures to bypass the
max_blocks_per_sm assertion for consumer cards.
- A formal Vulkan implementation of the custom
type 142 fused Gated Delta Net layers so consumer AMD users can bypass the enterprise ROCm/HIP kernel constraints entirely.
Description
The custom ternary math execution pipeline inside the PrismML
llama.cppfork (prismbranch) currently hard-limits execution on consumer-grade AMD RDNA2 cards (specifically the Radeon RX 6600,gfx1030) due to a strict hardware register assertion check inside the custom compiled Flash Attention kernels.When attempting to load
Ternary-Bonsai-2-27B-PQ2_0.gguf, the server crashes instantly during the graph compute phase with a core dump.Steps to Reproduce
llama-prismfork using CMake with-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1030.llama-servertargeting the ternary layout with dynamic layer offloading (e.g.,-ngl 35 -c 8192 --parallel 1 --kv-unified).Error Stack / Traceback
text
/home/llama-prism/ggml/src/ggml-cuda/template-instances/../fattn-common.cuh:1114: GGML_ASSERT(max_blocks_per_sm > 0) failed
Aborted (core dumped)
Rationale & Expected Behavior
While RDNA3 architectures (
gfx1101) have been confirmed to run unpatched via the HIP backend, consumer RDNA2 platforms return 0 when the flash attention tile configuration queriesmax_blocks_per_sminsidefattn-common.cuh. Since the ternary format forces flash attention natively, the system cannot load or fall back cleanly.We request either:
max_blocks_per_smassertion for consumer cards.type 142fused Gated Delta Net layers so consumer AMD users can bypass the enterprise ROCm/HIP kernel constraints entirely.