📝 Publications
-
X³-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin EMNLP 2026 Submission (First Author)
We propose X³-OPD, a cross-modal on-policy distillation framework that transfers reasoning from a text teacher to an audio-language student by scoring the student’s own audio-conditioned trajectories with matched textual inputs. Built on a 71.8K three-tier symmetric corpus spanning logical, audio-event, and prosody-aware dialogue reasoning, X³-OPD improves BIG Bench Audio from 87.9 to 93.6 (+5.7) while largely preserving out-of-domain capabilities.
-
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning Dongjie Fu, Fangming Feng, Xize Cheng*, Linjun Li, Zhou Zhao, Tao Jin ACM MM 2026 (First Author)
We propose RoleJudge, a multimodal and multidimensional evaluator for voice role-playing agents, and build RoleChat, the first reasoning-enhanced voice role-playing evaluation dataset. A multi-stage SFT-RL framework with Standard Alignment improves character-alignment evaluation while mitigating reward misalignment.

-
PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue Dongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin EMNLP 2025 (First Author)
We build Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset with user profiles, and propose PAChat, which jointly models semantic and speaker representations to enable persona-aware personalized responses.

-
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs Di Cao*, Dongjie Fu*, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin Interspeech 2026 (Co-first Author)
We propose X-OPD, a cross-modal on-policy distillation framework that aligns the capabilities of Speech LLMs to their text-based counterparts via token-level feedback from a text teacher, closing the performance gap while preserving speech abilities.

-
Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts Dongjie Fu, Xize Cheng, Xiaoda Yang, Hanting Wang, Zhou Zhao, Tao Jin ACM MM 2024 (Oral, First Author)
We propose PCD, a contrast-augmented prompt framework that enhances the robustness of audio-visual speech recognition to modality-distortion (asynchrony, visual noise and audio noise), achieving SOTA under distorted resources on LRS2.

-
Emphasizing Domain Differences through Interactive-Augmented Prompts in Continual Audio-Visual Speech Recognition Dongjie Fu, Xize Cheng, Jingyuan Chen, Tao Jin, Zhongfei Zhang IEEE TIP (First Author)
We introduce the Continual Audio-Visual Speech Recognition (CL-AVSR) problem and propose IMP, an interaction-enhanced multimodal prompt learning framework that transfers cross-domain knowledge with minimal parameter overhead, mitigating forgetting via cross-modal alignment and cross-task contrast.
Full Publication List
[*] denotes co-first authors, [#] denotes co-supervised.
I. Spoken Dialogue Systems & Audio LLMs
-
Under ReviewX³-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment. Dongjie Fu*, Di Cao*, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin. (EMNLP 2026 submission, First Author) -
ACMMM2026Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning. Dongjie Fu*, Fangming Feng*, Xize Cheng*, Linjun Li, Zhou Zhao, Tao Jin. (First Author) -
ACMMM2026VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference. Wenxu Jia*, Dongjie Fu*, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin. (Co-first Author) -
EMNLP2025PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue. Dongjie Fu*, Xize Cheng*, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin. (First Author) -
Interspeech2026X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs. Di Cao*, Dongjie Fu*, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin. (Co-first Author) -
Under ReviewVox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models. Xize Cheng, Dongjie Fu, et al. (NeurIPS 2026 submission, Co-first Author) -
PreprintOmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios. Xize Cheng*, Dongjie Fu*, Xiaoda Yang, Minghui Fang, Ruofan Hu, et al. (Co-first Author) -
ACL2026Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech. Fangming Feng, Dongjie Fu, Zequn Xie, Yu Zhang, Yangyang Wu, Zhou Zhao, Tao Jin. (Co-first Author) -
ICLR2025VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words? Xize Cheng, Dongjie Fu, et al. (Major Author) -
Interspeech2026Audio-NSP: Data-Centric Semi-Autoregressive Generation for Large Audio-Language Models. Liang Cao, Xize Cheng, Dongjie Fu, et al. (Co-first Author)
II. Audio-Visual Understanding
-
ACMMM2024 OralBoosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts. Dongjie Fu*, Xize Cheng*, Xiaoda Yang, Hanting Wang, Zhou Zhao, Tao Jin. (First Author) -
IEEE TIPEmphasizing Domain Differences through Interactive-Augmented Prompts in Continual Audio-Visual Speech Recognition. Dongjie Fu, Xize Cheng, Jingyuan Chen, Tao Jin, Zhongfei Zhang. (First Author) -
NeurIPS2025AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models. Xize Cheng, Dongjie Fu, et al. (Co-first Author) -
ACMMM2024SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning. Xiaoda Yang*, Xize Cheng*, Dongjie Fu, Minghui Fang, Jialong Zuo, Shengpeng Ji, Zhou Zhao, Tao Jin. (Co-first Author) -
ICLR2026MARS-Sep: Multimodal-Aligned Reinforced Sound Separation. Zihan Zhang, Xize Cheng, Dongjie Fu, et al. (Major Author)
III. Others
-
TOISDiversifying Sequential Recommendation with Retrospective and Prospective Transformers. Chaoyu Shi, Dongjie Fu, et al. (Co-first Author) -
KDD2025Multimodal Conditional Retrieval with High Controllability. Xiaoda Yang, Xize Cheng, Dongjie Fu, et al. (Major Author) -
CVPR2024MPOD123: One Image to 3D Content Generation Using Mask-enhanced Progressive Outline-to-Detail Optimization. Jimin Xu, Dongjie Fu, et al. (Major Author)