-
Updated
Oct 8, 2026 - Rust
activation-steering
Here are 137 public repositories matching this topic...
[ICLR 2025] General-purpose activation steering library
-
Updated
Oct 8, 2026 - Python
OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
-
Updated
Sep 23, 2026 - Python
KV Cache Steering for Controlling Frozen LLMs
-
Updated
Aug 18, 2026 - Python
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
-
Updated
Oct 1, 2026 - Python
[EMNLP 2026 Main] Steering Geometry: Validating Human Value Geometry in LLM Steering Space.
-
Updated
Sep 9, 2026 - Python
A bilingual awesome list for refusal suppression research: benchmarks, papers, tools, models, and ecosystem updates.
-
Updated
Jun 12, 2026
[ACL 2026] - Official repo for the paper: "Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection"
-
Updated
Sep 13, 2026 - Jupyter Notebook
Activation steering and trait monitoring for HuggingFace transformers
-
Updated
Sep 23, 2026 - Python
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
-
Updated
Sep 18, 2026 - Python
[ICLR 2026] ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
-
Updated
Sep 30, 2025 - Python
Runtime rank-1 refusal projection for DeepSeek-V4-Flash-0731: 757KB of directions instead of a 1.54GB weight overlay, lambda as a hot-swappable dial. Full A/B measurements on 2x DGX Spark.
-
Updated
Oct 2, 2026 - Python
Transformer interpretability research framework for circuit discovery, sparse features, and controlled residual-stream interventions.
-
Updated
Aug 16, 2026 - Python
Feature steering for open LLMs: find an interpretable SAE feature, steer the model, measure the causal effect
-
Updated
Jul 4, 2026 - Python
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
-
Updated
Mar 17, 2026 - Python
How Do Language Models Use Memory? Strategy readout from model activations and adaptive memory control with Duet.
-
Updated
Sep 26, 2026 - Python
Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions
-
Updated
Sep 4, 2026 - Python
MiniMax-H3 custom nodes: steering / cache / prompt director
-
Updated
Sep 13, 2026 - Python
GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
-
Updated
Jun 21, 2026 - Python
🏆[ICML 2026 Spotlight] Official implementation of "DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions"
-
Updated
Jun 21, 2026 - Python
Add this topic to your repo
To associate your repository with the activation-steering topic, visit your repo's landing page and select "manage topics."