Profiling-driven micro-architectural case study & custom Triton GPU operators for batch-1 edge vision inference on NVIDIA Ada Lovelace.
-
Updated
Aug 28, 2026 - Python
Profiling-driven micro-architectural case study & custom Triton GPU operators for batch-1 edge vision inference on NVIDIA Ada Lovelace.
Field-artillery simulation supplement to the Coordination Tax working paper/manuscript.
3D visualization toolkit for parallel computing performance analysis: speedup, efficiency, and iso-efficiency metrics with Amdahl's and Gustafson's Law implementations.
Parallelizing Federated Learning client simulation with ProcessPoolExecutor: 1.27x stable speedup, empirical Amdahl analysis (~24% parallel fraction), NumPy IPC to avoid PyTorch pickling deadlocks
Multi-threaded matrix multiplication written in java to test Amdahl's Law.
Repository of the lab2 assignment for the Parallel Programming course.
High-performance $\pi$ digit calculation service designed for the Software Architecture (ARSW) course at Escuela Colombiana de Ingeniería Julio Garavito. Features a transition from sequential to parallel execution using Java 21 platform threads, layered architectural patterns, and performance benchmarking across diverse CPU architectures.
Interactive dashboard for learning parallel computing speedup using Amdahl’s and Gustafson’s Laws with real-time simulations
Distributed prime search in C with Open MPI and hybrid MPI+OpenMP; SLURM jobs and Amdahl analysis
Sequential vs fork()+mmap vs Pthreads on 1M CSV records in C: 2.8x speedup, Amdahl's Law fit (22.4% serial) and a load-imbalance diagnosis.
Pipeline de imagens (cinza → blur gaussiano → Sobel) em versões sequencial e paralela: Python com multiprocessing (PC e AWS EC2) e JavaScript com Web Workers no navegador. Mede speedup, compara com a Lei de Amdahl e verifica os resultados por SHA-256.
Data-parallel pipeline for detecting DGA-based malware using linguistic feature extraction on multi-core CPUs. Achieves 5.75× speedup on 8 cores with 93.18% classification accuracy. Built with Python multiprocessing and Random Forest. [AMLCCZG516 — BITS Pilani]
A parameterized multi-cycle Harvard architecture processor designed from the ground up in Verilog. Developed using a complete RTL design workflow including simulation, linting, synthesis, Sky130 technology mapping, area estimation, static timing analysis, cocotb self-checking testbenches and post synthesis GLS 🖥️.
To associate your repository with the amdahls-law topic, visit your repo's landing page and select "manage topics."