Skip to content
#

selective-prediction

Here are 117 public repositories matching this topic...

Calibrated 151M Non-Autoregressive Decision Engine beating TypeSafe Jev & Laya on LocalLLaMA/typed-decisions (77.10% acc, 0.0636 Brier, 0.0144 ECE)

  • Updated Sep 20, 2026
  • Python
judge-calibration

Do LLM judges know when they're wrong? Three open-weight judges (Qwen2.5-7B, kev-8b, auto-j-13b) graded against MT-Bench human votes: calibration, position and padding attacks, and what auto-accepting confident verdicts would cost. Interactive site included.

  • Updated Sep 25, 2026
  • Python

Add this topic to your repo

To associate your repository with the selective-prediction topic, visit your repo's landing page and select "manage topics."

Learn more