Research

Quantisation, inference economics and the evidence behind the engineering decisions.

LLM quantisation for edge GPUs

MSc, Machine Learning - American University of Sharjah · 2024 – expected 2027

Post-training quantisation of open-weight language models, benchmarked across both NVIDIA and AMD accelerators: throughput, memory footprint, accuracy retention per method, and the cost-per-token implication of each. Cross-vendor comparisons are scarce because few people have both estates, which is the whole reason it is worth doing.

Code and benchmarks: llm-quant-lab

Publications