Skip to content
DavidGeeraertsPublic

Repository files navigation

📖 The Book of AI - Artificial Intelligence landscape

This is a work in progress --started 2025-02-- to document my exploration of Artificial Intelligence, Large Language Models (LLM), Natural Language Models (NLM), Natural Language Processing (NLP), etc.

🔥 Hotplate

📰 General AI vibe for October 2026

🧭 Building and refining open-source agent harnesses for long-running and autonomous work (LOOP for efficiency and duration, OpenHands and OpenCode for coding, goose for general MCP-native use), combined with heavy adoption of MCP tools, persistent memory systems, and self-hosted personal agents that live in messaging apps (OpenClaw and Hermes Agent leading). Strong bias toward tools you can actually run yourself, keep fully under your control, and host on your own hardware or cheap VPS (Virtual Private Server).


Letter Meaning
W Weights
A Artifacts
L Licenses
D Data
O Origin

OpenWALDO brings everything people expect from open source to AI: a community anyone can join; source code and training data anyone can contribute to and audit; open tools; public governance and trust earned through transparency; and models anyone can compose, train, validate, reproduce, extend, and improve.

Ajax is a fine-tuned (abliterated) Qwen 3.5 9B model, trained for Odysseus to be your always on agent. It handles your daily tasks from search to browse web to email to your calendar - all your daily tasks completely privately. Ajax's refusal has been ablated for a freer, less restricted AI experience. Please use responsibly.

Odysseus is a self-hosted AI workspace for chat, agents, deep research, email, calendar, and local/API model backends. It is local-first and privacy-first when you run local models on your own machine.

LLM Leaderboard.

Heretic removes restrictions (abliteration parameters) from language models, making sure they always follow your instructions.


AI Model Types

  • Large Language Models (LLM)
    • General-purpose language models
    • Code-specialized models {GitHub Copilot (Codex), Anthropic Claude Code, StarCoder, CodeLlama, DeepSeek Coder}
  • Image Generation Models
    • Text-to-Image {Stable Diffusion, Midjourney, DALL-E}
    • Image-to-Image
  • Multimodal Models (Text, Image, Audio, Video) {OpenAI Sora, Google Gemini, Meta Seamless, }

Prominent AI Organizations Creating Foundational AI Models

(Order by Name ↑)

Foundational AI Models

(Order by Name ↑)

Base Model Data

(Order by Name ↑)

Miscellaneous Tools

(Order by Name ↑)

Benchmarks

Benchmark Status in 2026 Strengths Weaknesses Still useful?
MMLU Saturated Historical reference for broad knowledge Frontier models cluster ~90%+; little differentiation; contamination risk Mostly no (use MMLU-Pro if needed)
GSM8K Saturated Simple math word problems Top models near ceiling (~95%+); many invalid questions noted No for frontier comparison
GPQA (esp. Diamond) Near-saturating Graduate-level “Google-proof” science reasoning Leaders pushing 90%+; small set → noise Still decent for science reasoning, but headroom shrinking
SWE-bench (Verified/Pro variants) Contested / partially saturating Real GitHub issue resolution; strong signal for coding agents Contamination concerns; scores climbing into high 70s–90s on easier versions Yes for coding/agentic work (prefer newer/harder variants like Pro)
GAIA Active but contested Multi-step agentic tasks (browsing, tools, files) that are easy for humans Large score gaps depending on scaffolding/tools; some saturation reports Yes for assistant/agent evaluation
LiveBench Highly regarded Monthly-refreshed questions from recent sources; objective ground-truth scoring; covers reasoning, coding, math, data analysis, language, instruction following Not human preference; not pure agentic long-horizon One of the strongest current objective capability benchmarks
LMArena (Chatbot/Arena) Highly regarded Large-scale blind human preference (Elo-style); real user prompts Style/verbosity bias; overlapping confidence intervals; not verifiable correctness Best for “which model feels best in open-ended use”
BenchLM Aggregator Combines many sources into overall indices Depends on underlying benchmarks; not a primary eval itself Useful meta-view, not a standalone benchmark
Mensa Norway IQ Test Niche Attempts IQ-style scoring Narrow, not standard in frontier AI evaluation Rarely used for serious model comparison
WebDev Arena Specialized Web development tasks Narrow domain Useful only if you care specifically about web-dev agents

Running AI Models Locally

Aspect LM Studio Ollama
Primary focus Polished desktop GUI for discovery, chat & experimentation CLI + background service for scripting, APIs & integrations
License Proprietary (free for personal + internal business use) Fully open-source (MIT)
Interface Excellent GUI + CLI (lms) + headless daemon (llmster) CLI-first + local API; has added a basic desktop app
Model discovery Built-in Hugging Face browser with size/VRAM estimates Curated library + simple ollama pull model:tag
API OpenAI- + Anthropic-compatible (port 1234) OpenAI-compatible + native endpoints (port 11434)
Best hardware edge Stronger MLX performance on Apple Silicon Lower overhead, often slightly faster on NVIDIA
Headless/server Yes (llmster since early 2026) Native from the start (systemd/Docker-friendly)
Extras Built-in document chat (RAG), MCP client, LM Link (remote device sharing), continuous batching Excellent ecosystem integrations, Modelfiles, official Docker image, optional cloud models

When Ollama is usually betterYou want a fully open-source, auditable tool.

  • You are a developer building scripts, agents, or apps (many frameworks and coding tools assume Ollama by default).
  • You need Docker, Kubernetes, or long-running headless servers with minimal footprint.
  • You prefer the lowest idle RAM and fastest cold-start times (especially on NVIDIA GPUs).
  • You want the absolute simplest CLI workflow (ollama run model).

Web/Online Models

(Order by Name ↑)

AI [blog] Resources

(Order by Name ↑)

Prominent AI People

(Sorted by Surname ↑)

AI Publication Websites

(Order by Name ↑)

AI Publications

(Sorted by Publication Date ↓)

Articles

(Sorted by Publication Date ↓)

Online Books

(Sorted by Publication Date ↓)

YouTube Videos & Channels

(Sorted by Publication Date ↓)

Releases

Packages

Contributors