Calibrated 151M Non-Autoregressive Decision Engine beating TypeSafe Jev & Laya on LocalLLaMA/typed-decisions (77.10% acc, 0.0636 Brier, 0.0144 ECE)
-
Updated
Sep 20, 2026 - Python
Calibrated 151M Non-Autoregressive Decision Engine beating TypeSafe Jev & Laya on LocalLLaMA/typed-decisions (77.10% acc, 0.0636 Brier, 0.0144 ECE)
Knowledge-graph RAG built from structured metadata instead of LLM extraction. 929M edges, zero LLM calls, 83.2% on held-out PubMedQA. Case study: the full PubMed 2026 baseline.
Calibrated abstention benchmark for drug–target interaction prediction, grounded in physical difficulty coordinates
Uncertainty based selection of compatible inputs
An editable, auditable 807K-param byte-level LLM: CRUD single facts with provable per-edit locality, and abstain when unsure instead of guessing. CPU, offline.
Awesome Jev: an evidence survey of Jev and Jev-like typed decision models — calibration, selective control and open implementations, with a searchable literature site.
A pluggable decision substrate for RAG pipelines; explicit, calibrated state → Decision → confidence → action gates, with Jev as the first swappable backend.
Behavioral Trust Clustering a thermodynamic governance layer that reduces LLM hallucination by 52% on HumanEval. Drop-in wrapper for any decoder. MIT.
Confidence-aware support intent classification with reproducible evaluation, selective prediction and a typed FastAPI inference boundary.
One causal attention head, three switches: softmax attention, unnormalized-kernel attention and the exact Markov path product as settings of one operator. Proved in Lean 4; every failure kept on the record.
Do LLM judges know when they're wrong? Three open-weight judges (Qwen2.5-7B, kev-8b, auto-j-13b) graded against MT-Bench human votes: calibration, position and padding attacks, and what auto-accepting confident verdicts would cost. Interactive site included.
We show that a model owner can artificially introduce uncertainty into their model and provide a corresponding detection mechanism.
Safer continual learning for PyTorch models — research alpha
Uncertainty-aware bean-leaf triage: safe image inputs, selective prediction, reproducible evaluation, FastAPI, observability, and CPU MLOps.
Empirical benchmark & framework evaluating Small Language Models (SLMs) on operational risk triage with temperature calibration, selective abstention, and physics telemetry.
A comprehensive library for uncertainty quantification in machine learning.
Project FEAR: asymmetric RLVR rewards for LLM calibration, selective abstention, and judge capability (GRPO, Qwen 2.5). Finding O: the reward sweep collapses to a one-dimensional abstention dial.
Selective RAG with conformal abstention: a hallucination detector that scores its own confidence and abstains, with a finite-sample precision guarantee.
A tiny, offline check for abstention in text-to-SQL and RAG-over-warehouse assistants: does it say it cannot answer instead of returning a confident wrong number? No model access, no network.
Explainable action ranking from historical sequences — a Rust library and CLI, not a contextual bandit or online-learning library.
To associate your repository with the selective-prediction topic, visit your repo's landing page and select "manage topics."