Skip to content
View hammasbuilds's full-sized avatar

Block or report hammasbuilds

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hammasbuilds/README.md

Muhammad Hammas

AI Engineer — retrieval, agents, computer vision and applied ML.

Based in Pakistan. Every repository below was built, run and measured: the numbers in each README are what the code produced on a machine, and each one says how to reproduce them.

71 public repositories · 277 projects · Python.


All repositories

Sorted A–Z. The Projects column is filled in where a repository holds a set of related projects rather than one.

# Repository Area Projects What it is
1 3d-computer-vision Computer vision 10 Reconstructing geometry from images and point clouds, measured against known or planted ground truth
2 agent-memory Agents A cross-session memory layer for agents, and what "remember everything and search" actually finds
3 agent-sandbox Agents Runs AI-generated code under named hardening profiles and measures what each one stops
4 agentic-ai-llm Agents 41 Agent infrastructure, measurement apps and business agents, running on local models
5 assay-drift Computational biology Checks whether published PCR diagnostic assays still match circulating sequences
6 blast-radius Developer tools Reports what a dependency upgrade changes in a package's public API
7 book-to-skill Agents Turns a technical book PDF into an agent skill and measures how much of the book it keeps
8 bounded-agent-runtime Agents An agent runtime that enforces limits in code, with a chaos suite that tests them
9 browser-agent Agents Page encodings for browser agents on MiniWoB++, and what each one throws away
10 cartographer Developer tools Compares what a repository's imports declare against what its git history shows
11 causal-inference Applied ML What it costs to get the causal structure wrong, on data with a planted true effect
12 classical-computer-vision Computer vision 57 Classical CV methods implemented and compared — no deep learning, no training, no GPU
13 clcuv-surveillance Computational biology Genomic surveillance: variant frequency, selection, phylogeny and diagnostic coverage
14 code-eval-harness Benchmarks How much a code benchmark's score depends on the extraction protocol, measured
15 code-llm-lab LLM evaluation 20 Local coder models measured on MBPP, HumanEval and Devign
16 commit-history-forensics Developer tools Tests whether a commit history describes work that actually happened
17 context-bench Retrieval / RAG RAG, CAG and MAG compared on real data with ground-truth answers
18 contract-reader Benchmarks Contract question answering on CUAD, and what a short context window can reach
19 credit-risk-engine Applied ML A logistic credit scorecard with WoE/IV binning, reason codes and a fairness audit
20 deep-computer-vision Computer vision 3 Recognition and understanding: segmentation, localisation and detection
21 deep-research-agent Agents A research agent that deduplicates sources and reports disagreement between them
22 demand-forecast-platform Applied ML Hierarchical forecasting with reconciliation, Croston, ETS and a rolling-origin backtest
23 devign-leakage Benchmarks A train/test leakage and duplicate-label audit of the Devign vulnerability dataset
24 doc-intelligence-api Documents Document extraction with confidence routing and cross-field arithmetic validation
25 docstring-drift Developer tools Static detection of docstrings that no longer match the code they document
26 doubt Benchmarks An exact bound on how well a fact-verification model scores while ignoring the evidence
27 enterprise-ops-crew Agents A multi-agent back office with playbooks, SLAs, approval gates and an audit trail
28 faithful Benchmarks A summarisation faithfulness checker evaluated on 10,066 human judgements
29 flake-detective Developer tools Finds flaky tests and reports what each one depends on — order, hash seed, clock
30 generative-vision-lab Computer vision 10 Image-to-image restoration and translation, plus generation from text
31 harness-ablation Agents A coding-agent harness with every component switchable, and the ablation that prices each
32 harness-agent-scratch Agents A minimal coding agent from scratch: loop, tools, permissions, sandbox, compaction, subagents
33 incident-copilot Operations Reduces log lines and alarms to one incident with a suspect, using Drain and robust z-scores
34 insurance-mlops MLOps A point-in-time feature store, skew detection and a release gate that refuses rather than warns
35 knowledge-tracing Applied ML BKT, PFA, IRT and DKT from scratch, and how much of DKT's lead was duplicated rows
36 langchain-llm Agents 5 LangChain projects: structured output, retrieval, memory, injection defence, judge bias
37 langgraph-llm Agents 5 One project per LangGraph shape: revision loops, routing, parallel merge, checkpoints
38 ledger-truth Benchmarks An audit of FinQA questions against the filings they were written from
39 llm-agentic-datasets Datasets Index of the eight published reasoning datasets, with their shortcut baselines
40 llm-gateway LLM platform One entry point for every LLM call: routing, per-tenant budgets, fallbacks, redaction
41 llm-observability-platform LLM platform Cost per success, latency percentiles and PSI prompt-drift detection without storing prompts
42 machine-learning Applied ML 20 Twenty applied ML projects on public data, each measured end to end
43 mbpp-false-accepts Benchmarks How much wrong code MBPP's three assert statements accept
44 mcp-llm-rag Agents 6 Six agentic projects on real benchmarks, running entirely on local models
45 minimal-diff LLM evaluation Among one-edit patches that pass the tests, what "the smallest" actually depends on
46 model-serving-platform MLOps Multi-model serving with A/B, canary and shadow routing, SLO monitoring and auto-rollback
47 nlp-llm-ml NLP 23 Classic NLP techniques measured against each other on one corpus
48 notebook-to-package Developer tools Turns a Jupyter notebook into an installable package and checks the conversion
49 outcome Benchmarks What the ECtHR outcome benchmark scores when the facts are ignored entirely
50 pak-law-assistant Retrieval / RAG Question answering over Pakistani statutes that will not cite a repealed provision
51 perf-hunter Developer tools Finds performance regressions and measures the gate's own false-alarm rate
52 pr-referee Developer tools A code reviewer that reports only what it can prove, measured against one that guesses
53 primer-designer Computational biology Diagnostic PCR primer design with nearest-neighbour thermodynamics and drift alerts
54 qlora-finetune-suite Fine-tuning Loss masking, VRAM budgeting and leak-free splits — the pre-GPU parts of fine-tuning
55 rag-forge Retrieval / RAG Production RAG that locates every quote in the source before it becomes a citation
56 rag-llm-eval Retrieval / RAG 9 Nine RAG techniques measured as retrieval on one corpus, with almost no LLM in the loop
57 repo-surgeon Developer tools Mutation-gated codebase migration that refuses the changes it cannot prove are safe
58 router-14b LLM evaluation When a smaller model is sufficient, and what predicts it
59 sql-analyst-agent Agents English questions to SQL, with three safety layers and the generated SQL always shown
60 suite-auditor Developer tools Mutation testing that reports a coverage gap only with the input that demonstrates it
61 super-resolution Computer vision Single-image 4× super-resolution: several architectures fine-tuned and compared
62 swebench-localization Benchmarks How often a SWE-bench issue names the file that has to be changed
63 terminal-agent Agents A terminal coding agent whose harness replays every gold patch before a model sees a task
64 test-impact-oracle Developer tools Runs only the tests a change could affect, and measures what that skips
65 trace-to-patch Developer tools A failing test to a verified patch, or a reported "could not reproduce"
66 trial-match Benchmarks Parsing eligibility criteria from clinical trial registry text, and where it breaks
67 urdu-desk NLP An encoding audit of Urdu news text, where Arabic letters stand in for Urdu ones
68 urdunlp NLP Urdu and Roman Urdu text processing. Pure Python, zero dependencies, no model downloads
69 visual-analytics Computer vision 2 Making models and datasets legible: training behaviour, embeddings, attribution
70 web-to-markdown Developer tools Web pages to agent-ready markdown with the standard library alone, measured on 3,975 pages
71 worktree-fleet Agents N coding agents in parallel git worktrees behind a tested merge queue, measured on real history

By area

Area Repositories
Agents agentic-ai-llm · bounded-agent-runtime · deep-research-agent · enterprise-ops-crew · langchain-llm · langgraph-llm · mcp-llm-rag · sql-analyst-agent · agent-memory · agent-sandbox · book-to-skill · browser-agent · harness-ablation · harness-agent-scratch · terminal-agent · worktree-fleet
Retrieval / RAG rag-forge · rag-llm-eval · context-bench · pak-law-assistant
Computer vision classical-computer-vision · generative-vision-lab · deep-computer-vision · super-resolution · 3d-computer-vision · visual-analytics
Benchmarks code-eval-harness · contract-reader · devign-leakage · doubt · faithful · ledger-truth · mbpp-false-accepts · outcome · swebench-localization · trial-match
Developer tools blast-radius · cartographer · commit-history-forensics · docstring-drift · flake-detective · notebook-to-package · perf-hunter · pr-referee · repo-surgeon · suite-auditor · test-impact-oracle · trace-to-patch · web-to-markdown
Applied ML and MLOps machine-learning · credit-risk-engine · demand-forecast-platform · insurance-mlops · model-serving-platform · causal-inference · knowledge-tracing
LLM platform and evaluation llm-gateway · llm-observability-platform · qlora-finetune-suite · code-llm-lab · router-14b · minimal-diff
Datasets llm-agentic-datasets - eight code-verified reasoning datasets (402,000 examples), each with a real input and expected output, data on Kaggle
NLP nlp-llm-ml · urdunlp · urdu-desk
Computational biology assay-drift · clcuv-surveillance · primer-designer
Operations and documents incident-copilot · doc-intelligence-api

What I work on

Retrieval that cites its sources and declines when the evidence is thin. Agents whose limits are enforced by the runtime rather than requested in a prompt. Generative AI as an engineering problem — fine-tuning, routing, cost and evaluation. Computer vision, classical and deep — restoration, translation, segmentation, detection. Applied ML and statistics on public data, measured end to end. Urdu NLP, because tooling for 240 million speakers should not be rewritten by every project that needs it.

In progress

Generative vision — one architecture family per project rather than several wrappers around the same checkpoint. Video understanding — detection on video, temporal action localization, activity recognition. Video segmentation. Concept bottleneck models. Vision-language models and the evaluation that goes with them.


Stack

Python · FastAPI · PostgreSQL + pgvector · Redis · Docker · uv · PyTorch · Transformers / PEFT / bitsandbytes · scikit-learn · pandas / NumPy / SciPy / statsmodels · sentence-transformers · OpenCV / scikit-image · Ollama / Claude / HuggingFace behind one interface · Model Context Protocol · LangChain / LangGraph · GitHub Actions


📫 Open to AI/ML engineering roles.

Popular repositories Loading

  1. urdunlp urdunlp Public

    Python · Unicode normalisation · transliteration — Urdu and Roman Urdu NLP: normalisation, tokenisation, transliteration, stopwords. Zero dependencies, 40 tests.

    Python

  2. bounded-agent-runtime bounded-agent-runtime Public

    FastAPI · Pydantic · Anthropic API · Typer — An agent runtime that enforces its own limits - cost, steps, time, loop detection, approval gates - and a chaos suite that proves it. 21 tests, no model…

    Python

  3. llm-gateway llm-gateway Public

    Python · httpx · Anthropic API — One entry point for every LLM call: model routing by difficulty, per-tenant budgets and rate limits, fallback chains, caching, and guardrails that redact secrets be…

    Python

  4. model-serving-platform model-serving-platform Public

    Python · A/B + canary + shadow · SLO monitor — Multi-model serving: A/B testing with sticky assignment, canary rollout, shadow traffic, per-version SLOs and auto-rollback that knows the difference …

    Python

  5. llm-observability-platform llm-observability-platform Public

    Python · PSI drift detection · cost telemetry — Monitoring built for LLM applications: cost per feature and per success, latency percentiles that exclude cache hits, grounding rates, and prompt dri…

    Python

  6. incident-copilot incident-copilot Public

    Python · Drain templates · robust z-score (MAD) — AIOps: Drain log-template extraction, anomaly detection robust to the outliers it is looking for, and correlation that turns forty alarms into one …

    Python