Experimental Linux kernel fast-path patches for SRAM-based AI inference servers, targeting io_uring submission, registered buffers, CQ polling, wakeup attribution, and completion latency.
-
Updated
May 3, 2026 - C
Experimental Linux kernel fast-path patches for SRAM-based AI inference servers, targeting io_uring submission, registered buffers, CQ polling, wakeup attribution, and completion latency.
Experimental Linux kernel patchset and benchmark suite for semantic memory hints in inference workloads. Explores whether user-space intent (streaming vs reuse vs ephemeral memory) can influence reclaim behavior in Multi-Gen LRU (MGLRU).
SSD-resident INT2/INT4 KV cache with asynchronous staging and fused CUDA dequantizing decode attention for long-context LLM inference.
Exploratory AI infrastructure project modeling semantic KV-cache orchestration, memory tiering, and HBM/CXL movement tradeoffs for long-context LLM inference.
Publication-grade patent companion site and filing materials for a deterministic-gather architecture for hierarchically managed KV state in autoregressive neural network inference.
A curated map of AFD, PD disaggregation, KV-cache systems, MoE serving, and re-aggregation baselines for LLM serving.
|HACKATHON| This repository contains a CRNN (Convolutional Recurrent Neural Network) based solution for the R.O.A.D. Barbados Historic Handwriting Challenge. The project focuses on recognizing and transcribing historic handwritten text from Barbados documents using deep learning techniques.
Selected projects, runnable checks, and a published paper.
To associate your repository with the inference-systems topic, visit your repo's landing page and select "manage topics."