Muhammad Ahmad Waseem-image

Muhammad Ahmad Waseem

I'm a ML Engineer completing my M.S. in Computer Science at University at Buffalo, specializing in ASR & Speech AI, GPU kernel optimization, and efficient model deployment.

3+ years shipping production ML systems across research and government settings. 7 peer-reviewed publications. Available for full-time roles from July 2026.

about-me-image

About me

ML Engineer with deep expertise in automatic speech recognition, CUDA kernel development, and model quantization. I build end-to-end ML systems — from training efficient ASR models for children's speech to implementing custom attention kernels on A100 GPUs. My work spans research (7 publications across IEEE IGARSS, BMVC, and SDSC) and production deployments serving government agencies at city scale.

  • Location:Buffalo, NY (Open to relocation)
  • Availability:July 2026 (STEM OPT)
  • Nationality:Pakistani
  • Interests:ASR, GPU Systems, Quantization, Diffusion Models
  • Study:M.S. CS, University at Buffalo
  • Employment:Graduate Research Assistant, UB

Education

Master of Science in Computer Science

University at Buffalo (SUNY), Buffalo, NYAugust 2024 – July 2026

GPA: 3.945 / 4.0. Relevant coursework: Algorithm Analysis & Design, GPU Computing and its Applications to AI. Research focus: ASR, model quantization, GPU systems, diffusion model interpretability.

Master of Science in Electrical Engineering

Information Technology University (ITU), LahoreSeptember 2020 – March 2023

GPA: 3.83 / 4.0. Thesis: Unsupervised Detection of Facial Landmarks using Computer Vision (GCN-based consistency-guided bottleneck; published BMVC 2023).

Bachelor of Science in Electrical Engineering

University of Engineering & Technology (UET), LahoreOctober 2016 – June 2020

GPA: 3.52 / 4.0. Major courses: Intro to Machine Learning, Applied Probability & Random Processes, Linear Algebra, Digital Signal Processing.

Publications

Under Review — IEEE SMC 2026
IEEE IGARSS 2025
IEEE IGARSS 2023
IEEE IGARSS 2023
SDSC 2022

Work

Graduate Research Assistant | ASR · GPU Systems · Quantization

University at Buffalo (SUNY), Buffalo, NYAugust 2024 – July 2026
  • Reduced Phoneme Error Rate from 80% to 15% on children's speech via Wav2Vec 2.0 fine-tuning and AutoPhon auto-annotation pipeline (CSLU Kids, MyST, PhonBank)
  • Developed outlier-aware QAT for Conformer ASR: 7.17 WER vs 93.7 naive INT8, enabling standard INT8 deployment without custom kernels
  • Implemented CUDA kernels for scaled dot-product attention and FFN on A100; profiled compute- vs memory-bound regimes against cuBLAS and PyTorch
  • Benchmarked AWQ, GPTQ, SmoothQuant across A100, Jetson, Raspberry Pi; derived device-specific quantization strategies
  • Proposed reverseDAAM: post-hoc diffusion model interpretability framework (under review IEEE SMC 2026)

ML Engineer | Computer Vision · GIS · Production Systems

CITY at LUMS, Lahore, PakistanAugust 2021 – July 2024
  • TensorRT optimization on NVIDIA AGX Xavier: 5× throughput (3→14 FPS); deployed city-wide with Punjab Safe Cities Authority
  • Designed TFNet segmentation (Dilated ResNet + dual decoders): 94% F1 on SpaceNet2 and WHU; published IEEE IGARSS 2025
  • Built city-scale waste routing platform (Clarke-Wright, PostGIS, FastAPI, Vue.js): ~100K liters/month fuel savings, production deployment across Lahore
  • Multi-sensor flood mapping on Google Earth Engine (Sentinel-1/2, Landsat-9): identified 1,410 km road damage during 2022 Pakistan floods
  • Population disaggregation pipeline: 30m resolution population density, outperforming WorldPop and Meta baselines

Graduate Research Assistant | Computer Vision · 3D Reconstruction

Information Technology University (ITU), LahoreSeptember 2020 – July 2021
  • Developed consistency-guided bottleneck for unsupervised facial landmark detection via GCN clustering; published BMVC 2023
  • Implemented DSAC* (Differentiable RANSAC) for 6-DoF camera pose estimation; COLMAP-based ground truth generation
  • Designed FPGA hardware accelerator for LeNet on NEXYS4: custom compute engines for convolutional and dense layers

Skills

Speech & ASR
Wav2Vec 2.0
Conformer / CTC
NVIDIA NeMo
Kaldi / HMM-ASR
Phoneme Modeling
Whisper
GPU & Systems
CUDA / nvcc
TensorRT
Triton
NVIDIA A100 / Jetson
INT8 / FP16 Inference
Quantization & Compression
QAT (Quantization-Aware Training)
AWQ / GPTQ
SmoothQuant / LLM.int8
Knowledge Distillation
ML Frameworks & Languages
PyTorch
HuggingFace Transformers
Python
C++ / C
TensorFlow
OpenCV

I have been really proud of having you on board in our team at CITY at LUMS. You bring a lot of experience to our team and your presence allows other team members to grow as well.

-- Dr. Momin Uppal — Associate Professor, LUMS

Your exceptional skills in computer vision give us a lot of confidence in picking new projects. The way you have handled many problems, even at short notice, shows your dedication.

-- Dr. Muhammad Tahir — Associate Professor, LUMS

It has been a wonderful experience supervising your Master's Thesis. You have been a very hard-working student and your never-give-up attitude is not found in many people.

-- Dr. Arif Mahmood — Thesis Supervisor, ITU

Get in touch.

How can I help you? Select a topic to get started.

Open to ML Engineer, Applied Scientist, and Research Engineer roles from July 2026. Available on STEM OPT. Feel free to reach out for collaboration, research discussions, or opportunities.

© 2026 Muhammad Ahmad Waseem