Skip to content

Publications

2026Under reviewarXiv preprint

Stable FP4 Training via Transposition-Invariant Block Quantization

Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi, Xing Huang, Yao Wang, Zhijun Tu, Yufei Cui, Yunke Peng, Hongliang Li

Identifies tensor-transposition-induced scale inconsistency as a key cause of FP4 training instability and proposes a 2D block quantization scheme with transposition-invariant scaling. Achieves stable end-to-end FP4 training within 1.3% of BF16, validated on LLMs up to 7B parameters and a 30B MoE model.

  • FP4
  • Low-precision training
  • LLMs

Selected projects

Research paper · HuaweiUnder review

Boundary-Aware Neighborhood Shrinkage for Test-Time Adaptation

Proposes GTA, a test-time adaptation method that pairs predictive entropy with a feature-space margin radius and a neighborhood-shrinkage operator to identify and repair unreliable pseudo-labels, improving sample selection and adaptation across multiple datasets and architectures.

Research project · Huawei2025

FP8 Training Pipeline for Transformers

Designed and implemented an end-to-end FP8 mixed-precision training pipeline for large transformers, with custom kernels for Flash Attention, matrix multiplication, and quantization scaling. Achieves state-of-the-art activation-memory reduction and throughput gains among low-bit approaches while matching BF16 training stability; validated by pre-training models up to 7B parameters on multi-node, multi-GPU clusters.

Course project · UdeM2024

Self-Supervised ResNet

Implemented SimCLR from scratch to train a ResNet on various datasets and demonstrated its advantages over training the model directly on the datasets.