Publications

You can also find my articles on my Google Scholar profile.

Conference Papers


Preprint


MiniMax Sparse Attention

Published in arxiv, 2026

This paper introduces MiniMax Sparse Attention, a blockwise sparse attention mechanism built on GQA that preserves model quality while enabling efficient 1M-context inference with large reductions in attention compute and wall-clock latency.

Download Paper

Model Merging in Pre-training of Large Language Models

Published in arxiv, 2025

This paper comprehensively investigates model merging in pre-training, showing that merging constant-learning-rate checkpoints on dense/MoE architectures (millions to 100B+ params) improves performance, predicts annealing, boosts efficiency, reduces costs, and provides ablation-driven insights.

Download Paper