木叶吟
木叶吟
Home
Experience
Posts
Publications
Services
CV
Light
Dark
Automatic
1
OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling
We propose OctoPipe, a novel pipeline parallelism system that reduces pipeline bubbles on heterogeneous models by co-optimizing partitioning, placement, and scheduling, achieving 1.22-2.14x throughput improvement over Megatron-LM.
Jihu Guo
,
Tenghui Ma
,
Wei Gao
,
Peng Sun
,
Xun Chen
,
Jiaxing Li
,
Zhisheng YE
,
Yuyang Jin
,
Dahua Lin
Preprint
PDF
Cite
UniTG: A Unified System for Efficient and Seamless Textual Graph Learning
We propose UniTG, the first unified system that fuses the LM and GNN phases of textual graph learning into a single end-to-end procedure, reducing learning makespan by up to 17.3x without compromising model quality.
Meng Zhang
,
Zhisheng YE
,
Qiyu Liu
,
Jingshu Peng
,
Tianwei Zhang
PDF
Cite
Code
DOI
LEMUR: Large Scale End-to-End Multimodal Recommendation
Traditional ID-based recommender systems often struggle with cold-start and generalization challenges. Multimodal recommendation …
Xintian Han
,
Honggang Chen
,
Quan Lin
,
Jingyue Gao
,
Xiangyuan Ren
,
Lifei Zhu
,
Zhisheng YE
,
Shikang Wu
,
XiongHang Xie
,
Xiaochu Gan
,
Bingzheng Wei
,
Peng Xu
,
Zhe Wang
,
Yuchao Zheng
,
Jingjian Lin
,
Di Wu
,
Junfeng Ge
Preprint
PDF
Cite
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
Batch inference for agentic workloads stresses the GPU key-value (KV) cache in a sustained and cumulative manner, often causing severe …
Qiaoling Chen
,
Zhisheng YE
,
Tian Tang
,
Peng Sun
,
Boyu Tian
,
Guoteng Wang
,
Shenggui Li
,
Yonggang Wen
,
Zhenhua Han
,
Tianwei Zhang
Preprint
PDF
Cite
ICML 2026
Blog
FlowGPU: Transparent and Efficient GPU Checkpointing and Restore
GPU checkpointing and restore promises to enable emerging tasks, such as deep learning, to benefit from functionalities like task …
Zehua Yang
,
Xiao Zheng
,
Yonghao Zou
,
Junyang Zhang
,
Zhisheng YE
,
Feng Xie
,
Xiaolin Wang
,
Yingwei Luo
,
Zhenlin Wang
,
Diyu Zhou
PDF
Cite
DOI
Latency-SLO-Aware Memory Offloading for Large Language Model Inference
Offloading large language models (LLMs) state to host memory during inference promises to reduce operational costs by supporting larger …
Chenxiang Ma
,
Hanyu Zhao
,
Zhisheng YE
,
Zehua Yang
,
Tianhao Fu
,
Jiaxun Han
,
Jie Zhang
,
Yingwei Luo
,
Xiaolin Wang
,
Zhenlin Wang
,
Yong Li
,
Diyu Zhou
Preprint
PDF
Cite
DOI
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
Hybrid parallelism underpins large-scale LLM training across tens of thousands of GPUs. At such scale, hardware failures on individual …
Tenghui Ma
,
Jihu Guo
,
Wei Gao
,
Sitian Lu
,
Zhisheng YE
,
Dahua Lin
,
Hanjing Wang
Preprint
Cite
DOI
Characterization of Large Language Model Development in the Datacenter
Large Language Models (LLMs) have presented impressive performance across several transformative tasks. However, it is non-trivial to …
Qinghao Hu
,
Zhisheng YE
,
Zerui Wang
,
Guoteng Wang
,
Meng Zhang
,
Qiaoling Chen
,
Peng Sun
,
Dahua Lin
,
Xiaolin Wang
,
Yingwei Luo
,
Yonggang Wen
,
Tianwei Zhang
Preprint
Cite
Hydro: Surrogate-Based Hyperparameter Tuning Service in Datacenters
Hyperparameter tuning is an essential step in deep learning model development that provides better model performance at the cost of …
Qinghao Hu
,
Zhisheng YE
,
Meng Zhang
,
Qiaoling Chen
,
Peng Sun
,
Yonggang Wen
,
Tianwei Zhang
PDF
Cite
Code
Slides
Video
Tear Up the Bubble Boom: Lessons Learned From a Deep Learning Research and Development Cluster
With the proliferation of deep learning, there exists a strong need to efficiently operate GPU clusters for deep learning production in …
Zehua Yang
,
Zhisheng YE
,
Tianhao Fu
,
Jing Luo
,
Xiong Wei
,
Yingwei Luo
,
Xiaolin Wang
,
Zhenlin Wang
,
Tianwei Zhang
PDF
Cite
Dataset
DOI
»
Cite
×