Xu Yuzhuang

CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis

Proceedings of the AAAI Conference on Artificial Intelligence, 40(32), 27395--27404, 2026.

Xu, Yuzhuang and Han, Xu and Zhang, Yuanchi and Wang, Yixuan and Liu, Yijun and Ji, Shiyu and Zhu, Qingfu and Che, Wanxiang

CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis

Proceedings of the AAAI Conference on Artificial Intelligence, 40(32), 27395--27404, 2026.

Xu, Yuzhuang and Han, Xu and Zhang, Yuanchi and Wang, Yixuan and Liu, Yijun and Ji, Shiyu and Zhu, Qingfu and Che, Wanxiang

Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction

Proceedings of the AAAI Conference on Artificial Intelligence, 40(38), 32240--32248, 2026.

Liu, Yijun and Wang, Yixuan and Xu, Yuzhuang and Ji, Shiyu and Xu, Yang and Zhu, Qingfu and Che, Wanxiang

Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction

Proceedings of the AAAI Conference on Artificial Intelligence, 40(38), 32240--32248, 2026.

Liu, Yijun and Wang, Yixuan and Xu, Yuzhuang and Ji, Shiyu and Xu, Yang and Zhu, Qingfu and Che, Wanxiang

CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs

Transactions of the Association for Computational Linguistics, 13, 1488--1506, 2025.

Xu, Yuzhuang and Ji, Shiyu and Zhu, Qingfu and Che, Wanxiang

CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs

Transactions of the Association for Computational Linguistics, 13, 1488--1506, 2025.

Xu, Yuzhuang and Ji, Shiyu and Zhu, Qingfu and Che, Wanxiang

Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query

Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 34158--34174, 2025.

Wang, Yixuan and Ji, Shiyu and Liu, Yijun and Xu, Yuzhuang and Xu, Yang and Zhu, Qingfu and Che, Wanxiang

Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query

Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 34158--34174, 2025.

Wang, Yixuan and Ji, Shiyu and Liu, Yijun and Xu, Yuzhuang and Xu, Yang and Zhu, Qingfu and Che, Wanxiang

OneBit: Towards Extremely Low-bit Large Language Models

Advances in Neural Information Processing Systems, 66357--66382, 2024.

Xu, Yuzhuang and Han, Xu and Yang, Zonghan and Wang, Shuo and Zhu, Qingfu and Liu, Zhiyuan and Liu, Weidong and Che, Wanxiang

OneBit: Towards Extremely Low-bit Large Language Models

Advances in Neural Information Processing Systems, 66357--66382, 2024.

Xu, Yuzhuang and Han, Xu and Yang, Zonghan and Wang, Shuo and Zhu, Qingfu and Liu, Zhiyuan and Liu, Weidong and Che, Wanxiang