AI Engineer — retrieval, agents, computer vision and applied ML.
Based in Pakistan. Every repository below was built, run and measured: the numbers in each README are what the code produced on a machine, and each one says how to reproduce them.
71 public repositories · 277 projects · Python.
Sorted A–Z. The Projects column is filled in where a repository holds a set of related projects rather than one.
| # | Repository | Area | Projects | What it is |
|---|---|---|---|---|
| 1 | 3d-computer-vision | Computer vision | 10 | Reconstructing geometry from images and point clouds, measured against known or planted ground truth |
| 2 | agent-memory | Agents | A cross-session memory layer for agents, and what "remember everything and search" actually finds | |
| 3 | agent-sandbox | Agents | Runs AI-generated code under named hardening profiles and measures what each one stops | |
| 4 | agentic-ai-llm | Agents | 41 | Agent infrastructure, measurement apps and business agents, running on local models |
| 5 | assay-drift | Computational biology | Checks whether published PCR diagnostic assays still match circulating sequences | |
| 6 | blast-radius | Developer tools | Reports what a dependency upgrade changes in a package's public API | |
| 7 | book-to-skill | Agents | Turns a technical book PDF into an agent skill and measures how much of the book it keeps | |
| 8 | bounded-agent-runtime | Agents | An agent runtime that enforces limits in code, with a chaos suite that tests them | |
| 9 | browser-agent | Agents | Page encodings for browser agents on MiniWoB++, and what each one throws away | |
| 10 | cartographer | Developer tools | Compares what a repository's imports declare against what its git history shows | |
| 11 | causal-inference | Applied ML | What it costs to get the causal structure wrong, on data with a planted true effect | |
| 12 | classical-computer-vision | Computer vision | 57 | Classical CV methods implemented and compared — no deep learning, no training, no GPU |
| 13 | clcuv-surveillance | Computational biology | Genomic surveillance: variant frequency, selection, phylogeny and diagnostic coverage | |
| 14 | code-eval-harness | Benchmarks | How much a code benchmark's score depends on the extraction protocol, measured | |
| 15 | code-llm-lab | LLM evaluation | 20 | Local coder models measured on MBPP, HumanEval and Devign |
| 16 | commit-history-forensics | Developer tools | Tests whether a commit history describes work that actually happened | |
| 17 | context-bench | Retrieval / RAG | RAG, CAG and MAG compared on real data with ground-truth answers | |
| 18 | contract-reader | Benchmarks | Contract question answering on CUAD, and what a short context window can reach | |
| 19 | credit-risk-engine | Applied ML | A logistic credit scorecard with WoE/IV binning, reason codes and a fairness audit | |
| 20 | deep-computer-vision | Computer vision | 3 | Recognition and understanding: segmentation, localisation and detection |
| 21 | deep-research-agent | Agents | A research agent that deduplicates sources and reports disagreement between them | |
| 22 | demand-forecast-platform | Applied ML | Hierarchical forecasting with reconciliation, Croston, ETS and a rolling-origin backtest | |
| 23 | devign-leakage | Benchmarks | A train/test leakage and duplicate-label audit of the Devign vulnerability dataset | |
| 24 | doc-intelligence-api | Documents | Document extraction with confidence routing and cross-field arithmetic validation | |
| 25 | docstring-drift | Developer tools | Static detection of docstrings that no longer match the code they document | |
| 26 | doubt | Benchmarks | An exact bound on how well a fact-verification model scores while ignoring the evidence | |
| 27 | enterprise-ops-crew | Agents | A multi-agent back office with playbooks, SLAs, approval gates and an audit trail | |
| 28 | faithful | Benchmarks | A summarisation faithfulness checker evaluated on 10,066 human judgements | |
| 29 | flake-detective | Developer tools | Finds flaky tests and reports what each one depends on — order, hash seed, clock | |
| 30 | generative-vision-lab | Computer vision | 10 | Image-to-image restoration and translation, plus generation from text |
| 31 | harness-ablation | Agents | A coding-agent harness with every component switchable, and the ablation that prices each | |
| 32 | harness-agent-scratch | Agents | A minimal coding agent from scratch: loop, tools, permissions, sandbox, compaction, subagents | |
| 33 | incident-copilot | Operations | Reduces log lines and alarms to one incident with a suspect, using Drain and robust z-scores | |
| 34 | insurance-mlops | MLOps | A point-in-time feature store, skew detection and a release gate that refuses rather than warns | |
| 35 | knowledge-tracing | Applied ML | BKT, PFA, IRT and DKT from scratch, and how much of DKT's lead was duplicated rows | |
| 36 | langchain-llm | Agents | 5 | LangChain projects: structured output, retrieval, memory, injection defence, judge bias |
| 37 | langgraph-llm | Agents | 5 | One project per LangGraph shape: revision loops, routing, parallel merge, checkpoints |
| 38 | ledger-truth | Benchmarks | An audit of FinQA questions against the filings they were written from | |
| 39 | llm-agentic-datasets | Datasets | Index of the eight published reasoning datasets, with their shortcut baselines | |
| 40 | llm-gateway | LLM platform | One entry point for every LLM call: routing, per-tenant budgets, fallbacks, redaction | |
| 41 | llm-observability-platform | LLM platform | Cost per success, latency percentiles and PSI prompt-drift detection without storing prompts | |
| 42 | machine-learning | Applied ML | 20 | Twenty applied ML projects on public data, each measured end to end |
| 43 | mbpp-false-accepts | Benchmarks | How much wrong code MBPP's three assert statements accept | |
| 44 | mcp-llm-rag | Agents | 6 | Six agentic projects on real benchmarks, running entirely on local models |
| 45 | minimal-diff | LLM evaluation | Among one-edit patches that pass the tests, what "the smallest" actually depends on | |
| 46 | model-serving-platform | MLOps | Multi-model serving with A/B, canary and shadow routing, SLO monitoring and auto-rollback | |
| 47 | nlp-llm-ml | NLP | 23 | Classic NLP techniques measured against each other on one corpus |
| 48 | notebook-to-package | Developer tools | Turns a Jupyter notebook into an installable package and checks the conversion | |
| 49 | outcome | Benchmarks | What the ECtHR outcome benchmark scores when the facts are ignored entirely | |
| 50 | pak-law-assistant | Retrieval / RAG | Question answering over Pakistani statutes that will not cite a repealed provision | |
| 51 | perf-hunter | Developer tools | Finds performance regressions and measures the gate's own false-alarm rate | |
| 52 | pr-referee | Developer tools | A code reviewer that reports only what it can prove, measured against one that guesses | |
| 53 | primer-designer | Computational biology | Diagnostic PCR primer design with nearest-neighbour thermodynamics and drift alerts | |
| 54 | qlora-finetune-suite | Fine-tuning | Loss masking, VRAM budgeting and leak-free splits — the pre-GPU parts of fine-tuning | |
| 55 | rag-forge | Retrieval / RAG | Production RAG that locates every quote in the source before it becomes a citation | |
| 56 | rag-llm-eval | Retrieval / RAG | 9 | Nine RAG techniques measured as retrieval on one corpus, with almost no LLM in the loop |
| 57 | repo-surgeon | Developer tools | Mutation-gated codebase migration that refuses the changes it cannot prove are safe | |
| 58 | router-14b | LLM evaluation | When a smaller model is sufficient, and what predicts it | |
| 59 | sql-analyst-agent | Agents | English questions to SQL, with three safety layers and the generated SQL always shown | |
| 60 | suite-auditor | Developer tools | Mutation testing that reports a coverage gap only with the input that demonstrates it | |
| 61 | super-resolution | Computer vision | Single-image 4× super-resolution: several architectures fine-tuned and compared | |
| 62 | swebench-localization | Benchmarks | How often a SWE-bench issue names the file that has to be changed | |
| 63 | terminal-agent | Agents | A terminal coding agent whose harness replays every gold patch before a model sees a task | |
| 64 | test-impact-oracle | Developer tools | Runs only the tests a change could affect, and measures what that skips | |
| 65 | trace-to-patch | Developer tools | A failing test to a verified patch, or a reported "could not reproduce" | |
| 66 | trial-match | Benchmarks | Parsing eligibility criteria from clinical trial registry text, and where it breaks | |
| 67 | urdu-desk | NLP | An encoding audit of Urdu news text, where Arabic letters stand in for Urdu ones | |
| 68 | urdunlp | NLP | Urdu and Roman Urdu text processing. Pure Python, zero dependencies, no model downloads | |
| 69 | visual-analytics | Computer vision | 2 | Making models and datasets legible: training behaviour, embeddings, attribution |
| 70 | web-to-markdown | Developer tools | Web pages to agent-ready markdown with the standard library alone, measured on 3,975 pages | |
| 71 | worktree-fleet | Agents | N coding agents in parallel git worktrees behind a tested merge queue, measured on real history |
Retrieval that cites its sources and declines when the evidence is thin. Agents whose limits are enforced by the runtime rather than requested in a prompt. Generative AI as an engineering problem — fine-tuning, routing, cost and evaluation. Computer vision, classical and deep — restoration, translation, segmentation, detection. Applied ML and statistics on public data, measured end to end. Urdu NLP, because tooling for 240 million speakers should not be rewritten by every project that needs it.
Generative vision — one architecture family per project rather than several wrappers around the same checkpoint. Video understanding — detection on video, temporal action localization, activity recognition. Video segmentation. Concept bottleneck models. Vision-language models and the evaluation that goes with them.
Python · FastAPI · PostgreSQL + pgvector · Redis · Docker · uv · PyTorch ·
Transformers / PEFT / bitsandbytes · scikit-learn · pandas / NumPy / SciPy /
statsmodels · sentence-transformers · OpenCV / scikit-image ·
Ollama / Claude / HuggingFace behind one interface · Model Context Protocol ·
LangChain / LangGraph · GitHub Actions
📫 Open to AI/ML engineering roles.