This is a work in progress --started 2025-02-- to document my exploration of Artificial Intelligence, Large Language Models (LLM), Natural Language Models (NLM), Natural Language Processing (NLP), etc.
🧭 Building and refining open-source agent harnesses for long-running and autonomous work (LOOP for efficiency and duration, OpenHands and OpenCode for coding, goose for general MCP-native use), combined with heavy adoption of MCP tools, persistent memory systems, and self-hosted personal agents that live in messaging apps (OpenClaw and Hermes Agent leading). Strong bias toward tools you can actually run yourself, keep fully under your control, and host on your own hardware or cheap VPS (Virtual Private Server).
| Letter | Meaning |
|---|---|
| W | Weights |
| A | Artifacts |
| L | Licenses |
| D | Data |
| O | Origin |
OpenWALDO brings everything people expect from open source to AI: a community anyone can join; source code and training data anyone can contribute to and audit; open tools; public governance and trust earned through transparency; and models anyone can compose, train, validate, reproduce, extend, and improve.
Ajax is a fine-tuned (abliterated) Qwen 3.5 9B model, trained for Odysseus to be your always on agent. It handles your daily tasks from search to browse web to email to your calendar - all your daily tasks completely privately. Ajax's refusal has been ablated for a freer, less restricted AI experience. Please use responsibly.
Odysseus is a self-hosted AI workspace for chat, agents, deep research, email, calendar, and local/API model backends. It is local-first and privacy-first when you run local models on your own machine.
LLM Leaderboard.
Heretic removes restrictions (abliteration parameters) from language models, making sure they always follow your instructions.
- Large Language Models (LLM)
- General-purpose language models
- Code-specialized models {GitHub Copilot (Codex), Anthropic Claude Code, StarCoder, CodeLlama, DeepSeek Coder}
- Image Generation Models
- Text-to-Image {Stable Diffusion, Midjourney, DALL-E}
- Image-to-Image
- Multimodal Models (Text, Image, Audio, Video) {OpenAI Sora, Google Gemini, Meta Seamless, }
(Order by Name ↑)
- Alibaba
- Allen Institute for AI
- Amazon AI
- Anthropic
- xAI
- Cohere
- Deepseek
- Google AI
- Hugging Face
- IBM
- Meta AI
- Microsoft AI
- Mistral
- OpenAI
- Stability AI
(Order by Name ↑)
(Order by Name ↑)
- AWS Open Data
- Common Crawl
- FineWeb -- curated Internet dataset.
- Google Dataset Search
- Hugging Face Datasets
- Kaggle
- Microsoft Planetary Computer Data Catalog
- MNIST (Modified National Institute of Standards and Technology database)
- Open Images Dataset
- SNAP - Stanford Large Network Dataset Collection
- UC Irvine Machine Learning Repository
- USA Data Gov
(Order by Name ↑)
- Excalidraw - online Diagram tool
- LLM Visualization
- Tiktokenizer
| Benchmark | Status in 2026 | Strengths | Weaknesses | Still useful? |
|---|---|---|---|---|
| MMLU | Saturated | Historical reference for broad knowledge | Frontier models cluster ~90%+; little differentiation; contamination risk | Mostly no (use MMLU-Pro if needed) |
| GSM8K | Saturated | Simple math word problems | Top models near ceiling (~95%+); many invalid questions noted | No for frontier comparison |
| GPQA (esp. Diamond) | Near-saturating | Graduate-level “Google-proof” science reasoning | Leaders pushing 90%+; small set → noise | Still decent for science reasoning, but headroom shrinking |
| SWE-bench (Verified/Pro variants) | Contested / partially saturating | Real GitHub issue resolution; strong signal for coding agents | Contamination concerns; scores climbing into high 70s–90s on easier versions | Yes for coding/agentic work (prefer newer/harder variants like Pro) |
| GAIA | Active but contested | Multi-step agentic tasks (browsing, tools, files) that are easy for humans | Large score gaps depending on scaffolding/tools; some saturation reports | Yes for assistant/agent evaluation |
| LiveBench | Highly regarded | Monthly-refreshed questions from recent sources; objective ground-truth scoring; covers reasoning, coding, math, data analysis, language, instruction following | Not human preference; not pure agentic long-horizon | One of the strongest current objective capability benchmarks |
| LMArena (Chatbot/Arena) | Highly regarded | Large-scale blind human preference (Elo-style); real user prompts | Style/verbosity bias; overlapping confidence intervals; not verifiable correctness | Best for “which model feels best in open-ended use” |
| BenchLM | Aggregator | Combines many sources into overall indices | Depends on underlying benchmarks; not a primary eval itself | Useful meta-view, not a standalone benchmark |
| Mensa Norway IQ Test | Niche | Attempts IQ-style scoring | Narrow, not standard in frontier AI evaluation | Rarely used for serious model comparison |
| WebDev Arena | Specialized | Web development tasks | Narrow domain | Useful only if you care specifically about web-dev agents |
- ollama - CLI
- ollama-ui - Simple HTML UI for Ollama. Available as Chrome extension.
- LM Studio - GUI
| Aspect | LM Studio | Ollama |
|---|---|---|
| Primary focus | Polished desktop GUI for discovery, chat & experimentation | CLI + background service for scripting, APIs & integrations |
| License | Proprietary (free for personal + internal business use) | Fully open-source (MIT) |
| Interface | Excellent GUI + CLI (lms) + headless daemon (llmster) |
CLI-first + local API; has added a basic desktop app |
| Model discovery | Built-in Hugging Face browser with size/VRAM estimates | Curated library + simple ollama pull model:tag |
| API | OpenAI- + Anthropic-compatible (port 1234) | OpenAI-compatible + native endpoints (port 11434) |
| Best hardware edge | Stronger MLX performance on Apple Silicon | Lower overhead, often slightly faster on NVIDIA |
| Headless/server | Yes (llmster since early 2026) | Native from the start (systemd/Docker-friendly) |
| Extras | Built-in document chat (RAG), MCP client, LM Link (remote device sharing), continuous batching | Excellent ecosystem integrations, Modelfiles, official Docker image, optional cloud models |
- You are a developer building scripts, agents, or apps (many frameworks and coding tools assume Ollama by default).
- You need Docker, Kubernetes, or long-running headless servers with minimal footprint.
- You prefer the lowest idle RAM and fastest cold-start times (especially on NVIDIA GPUs).
- You want the absolute simplest CLI workflow (ollama run model).
(Order by Name ↑)
- Allen Institute AI - Tulu 3:405B
- Anthropic - Claude
- Deepseek - R1
- GitHub Models
- Google AI Studio
- Granite
- Grok
- Nova
- Qwen
- Together AI
(Order by Name ↑)
(Sorted by Surname ↑)
- Dario Amodei - Anthropic
- AmandA Askell - Anthropic
- Lex Fridman - MIT
- Salim Ismail - Anthropic
- Nathan Lambert - Allen Institute for AI
- Emad Mostaque - stability.ai
- Christopher Olah - Anthropic
(Order by Name ↑)
(Sorted by Publication Date ↓)
- 2025-04-25 BitNet b1.58 2B4T Technical Report
- 2024-04-09 RULER: What's the Real Context Size of Your Long-Context Language Models?
- 2023-07-27 Universal and Transferable Adversarial Attacks on Aligned Language Models - LLM Attacks
- 2019-06-17 Superposition of many models into one
- 2017-06-12 Attention Is All You Need
(Sorted by Publication Date ↓)
- 2025-02-12 Unlocking the Effective Context Length: Benchmarking the Granite-3.1-8b Model
- 2025-02-02 A Detailed Analysis of Fine-Tuning, Direct Preference Optimization (DPO), and Reinforcement Learning with Verifiable Rewards (RLVR) on the LLama3.1 405B Model
- 2025-01-31 DeepSeek-V3 Explained 1: Multi-head Latent Attention
- 2025-01-19 Top 5 Mistakes to Avoid When Learning Machine Learning
- 2024-10-02 How to Get Started with Machine Learning: A Beginner’s Step-by-Step Guide
- 2024-07-13 MHA vs MQA vs GQA vs MLA
- 2020-00-00 Over 200 of the Best Machine Learning, NLP, and Python Tutorials — 2018 Edition
(Sorted by Publication Date ↓)
- 2024 Machine Learning and Deep Learning with R
- 2024 Natural Language Processing (NLP) Course
- 2023 State of Open Source AI Book - 2023 Edition
- 2023 Understanding Deep Learning
- 2019 Neural Networks and Deep Learning
- 2018 Deep Learning
- 2016 Deep Learning
(Sorted by Publication Date ↓)
- 2026-05-13 The Coding Gopher: Uni-1 AI image generation model
- 2026-03-26 The Diary of a CEO: AI Whistleblower: We Are Being Gaslit By AI Companies, They’re Hiding The Truth! - Karen Hao
- 2025-04-28 Logically Answered: AMD's $243 Billion AI Disaster...What Happened?
- 2025-04022 Computerphile: What is Cuda?
- 2025-03-05 ByteByteGo: What Is the Most Popular Open-Source AI Stack?
- 2025-02-05 Andrej Karpathy: Deep Dive into LLMs like ChatGPT
- 2025-02-02 Lex Fridman: DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters | Lex Fridman Podcast #459
- 2025-01-29 Peter H. Diamandis: DeepSeek vs. Open AI - The State of AI w/ Emad Mostaque & Salim Ismail | EP #146
- 2024-11-11 Lex Fridman: Dario Amodei: Anthropic CEO on Claude, AGI & the Future of AI & Humanity | Lex Fridman Podcast #452
- 2024-09-09 IBM Technology: RAG vs. Fine Tuning
- 2024-04-01 3Blue1Brown: Neural networks Playlist
- 2023-08-23 IBM Technology: What is Retrieval-Augmented Generation (RAG)?
- 2018-07-30 Lex Fridman: Deep Learning State of the Art (2020)
- Channel: StatQuest with Josh Starmer