About Me
I am Dongjie Fu (付栋杰), a second-year master’s student at the School of Software Technology, Zhejiang University. My advisor is Professor Tao Jin (金涛), in Professor Zhou Zhao (赵洲)’s lab. Before that, I received my B.E. degree from Shandong University.
My research focuses on Speech / Audio Large Language Models and Audio-Visual Understanding. I have published 11 papers as first author / co-first author in top-tier conferences and journals such as NeurIPS, EMNLP, ACM MM and TIP.
My current research topics include:
- Spoken Dialogue Systems / Audio Large Language Models
- Audio-Visual Speech Understanding
I am actively looking for job opportunities, feel free to drop me an email.
🔥 News
- 2026.07: 🎉🎉 2 papers (1 first-author paper and 1 co-first-author paper) are accepted by ACM MM 2026!
- 2026.05: 🎉🎉 2 papers (co-first author) are accepted by Interspeech 2026!
- 2026.04: 🎉🎉 1 paper (first author) is accepted by IEEE TIP!
- 2026.01: 🎉🎉 1 paper (co-first author) is accepted by ACL 2026!
- 2026.01: 🎉🎉 1 paper (major author) is accepted by ICLR 2026!
- 2025.12: I start my internship at Tencent, Hunyuan Multimodal Model Department, Speech Algorithm Center.
- 2025.09: 🎉🎉 1 paper (first author) is accepted by EMNLP 2025!
- 2025.09: 🎉🎉 1 paper (co-first author) is accepted by NeurIPS 2025!
- 2025.05: 🎉🎉 1 paper (major author) is accepted by KDD 2025!
- 2025.01: 🎉🎉 1 paper (major author) is accepted by ICLR 2025!
- 2024.07: 🎉🎉 1 paper (first author, Oral) and 1 paper (co-first author) are accepted by ACM MM 2024!
- 2024.07: I start my internship at Meituan, Financial Services Platform.
- 2024.06: 🎓 I receive my B.E. degree from Shandong University and join Zhejiang University as a master's student.
📝 Publications
-
X³-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin EMNLP 2026 Submission (First Author)
We propose X³-OPD, a cross-modal on-policy distillation framework that transfers reasoning from a text teacher to an audio-language student by scoring the student’s own audio-conditioned trajectories with matched textual inputs. Built on a 71.8K three-tier symmetric corpus spanning logical, audio-event, and prosody-aware dialogue reasoning, X³-OPD improves BIG Bench Audio from 87.9 to 93.6 (+5.7) while largely preserving out-of-domain capabilities.
-
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning Dongjie Fu, Fangming Feng, Xize Cheng*, Linjun Li, Zhou Zhao, Tao Jin ACM MM 2026 (First Author)
We propose RoleJudge, a multimodal and multidimensional evaluator for voice role-playing agents, and build RoleChat, the first reasoning-enhanced voice role-playing evaluation dataset. A multi-stage SFT-RL framework with Standard Alignment improves character-alignment evaluation while mitigating reward misalignment.

-
PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue Dongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin EMNLP 2025 (First Author)
We build Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset with user profiles, and propose PAChat, which jointly models semantic and speaker representations to enable persona-aware personalized responses.

-
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs Di Cao*, Dongjie Fu*, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin Interspeech 2026 (Co-first Author)
We propose X-OPD, a cross-modal on-policy distillation framework that aligns the capabilities of Speech LLMs to their text-based counterparts via token-level feedback from a text teacher, closing the performance gap while preserving speech abilities.

-
Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts Dongjie Fu, Xize Cheng, Xiaoda Yang, Hanting Wang, Zhou Zhao, Tao Jin ACM MM 2024 (Oral, First Author)
We propose PCD, a contrast-augmented prompt framework that enhances the robustness of audio-visual speech recognition to modality-distortion (asynchrony, visual noise and audio noise), achieving SOTA under distorted resources on LRS2.

-
Emphasizing Domain Differences through Interactive-Augmented Prompts in Continual Audio-Visual Speech Recognition Dongjie Fu, Xize Cheng, Jingyuan Chen, Tao Jin, Zhongfei Zhang IEEE TIP (First Author)
We introduce the Continual Audio-Visual Speech Recognition (CL-AVSR) problem and propose IMP, an interaction-enhanced multimodal prompt learning framework that transfers cross-domain knowledge with minimal parameter overhead, mitigating forgetting via cross-modal alignment and cross-task contrast.
Full Publication List
[*] denotes co-first authors, [#] denotes co-supervised.
I. Spoken Dialogue Systems & Audio LLMs
-
Under ReviewX³-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment. Dongjie Fu*, Di Cao*, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin. (EMNLP 2026 submission, First Author) -
ACMMM2026Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning. Dongjie Fu*, Fangming Feng*, Xize Cheng*, Linjun Li, Zhou Zhao, Tao Jin. (First Author) -
ACMMM2026VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference. Wenxu Jia*, Dongjie Fu*, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin. (Co-first Author) -
EMNLP2025PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue. Dongjie Fu*, Xize Cheng*, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin. (First Author) -
Interspeech2026X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs. Di Cao*, Dongjie Fu*, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin. (Co-first Author) -
Under ReviewVox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models. Xize Cheng, Dongjie Fu, et al. (NeurIPS 2026 submission, Co-first Author) -
PreprintOmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios. Xize Cheng*, Dongjie Fu*, Xiaoda Yang, Minghui Fang, Ruofan Hu, et al. (Co-first Author) -
ACL2026Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech. Fangming Feng, Dongjie Fu, Zequn Xie, Yu Zhang, Yangyang Wu, Zhou Zhao, Tao Jin. (Co-first Author) -
ICLR2025VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words? Xize Cheng, Dongjie Fu, et al. (Major Author) -
Interspeech2026Audio-NSP: Data-Centric Semi-Autoregressive Generation for Large Audio-Language Models. Liang Cao, Xize Cheng, Dongjie Fu, et al. (Co-first Author)
II. Audio-Visual Understanding
-
ACMMM2024 OralBoosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts. Dongjie Fu*, Xize Cheng*, Xiaoda Yang, Hanting Wang, Zhou Zhao, Tao Jin. (First Author) -
IEEE TIPEmphasizing Domain Differences through Interactive-Augmented Prompts in Continual Audio-Visual Speech Recognition. Dongjie Fu, Xize Cheng, Jingyuan Chen, Tao Jin, Zhongfei Zhang. (First Author) -
NeurIPS2025AHa-Bench: Benchmarking Audio Hallucinations in Large Audio-Language Models. Xize Cheng, Dongjie Fu, et al. (Co-first Author) -
ACMMM2024SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning. Xiaoda Yang*, Xize Cheng*, Dongjie Fu, Minghui Fang, Jialong Zuo, Shengpeng Ji, Zhou Zhao, Tao Jin. (Co-first Author) -
ICLR2026MARS-Sep: Multimodal-Aligned Reinforced Sound Separation. Zihan Zhang, Xize Cheng, Dongjie Fu, et al. (Major Author)
III. Others
-
TOISDiversifying Sequential Recommendation with Retrospective and Prospective Transformers. Chaoyu Shi, Dongjie Fu, et al. (Co-first Author) -
KDD2025Multimodal Conditional Retrieval with High Controllability. Xiaoda Yang, Xize Cheng, Dongjie Fu, et al. (Major Author) -
CVPR2024MPOD123: One Image to 3D Content Generation Using Mask-enhanced Progressive Outline-to-Detail Optimization. Jimin Xu, Dongjie Fu, et al. (Major Author)
📖 Educations
-
2024.09 - 2027.06, Master, School of Software Technology, Zhejiang University, Ningbo.
-
2020.09 - 2024.06, Undergraduate, School of Computer Science and Technology, Shandong University, Qingdao.
🎖 Honors and Awards
- Top 5% in major ranking, Zhejiang University. (2025-2026)
- Huawei Intelligent Base Scholarship (华为智能基座奖学金). (2024)
- Top 10% in major ranking, Shandong University. (2020-2024)
- Academic Scholarship, Shandong University. (2020-2023)
💬 Professional Services
- Conference Reviewer: ICLR, NeurIPS, EMNLP, ACM MM, Interspeech, ICASSP.
💻 Internships & Projects
- 2025.12 - Now: Algorithm Intern, Tencent, Hunyuan Multimodal Model Department, Speech Algorithm Center.
- On-Policy Distillation (OPD) stage training for the speech model.
- Building agent capabilities for the speech model.
- Optimizing the multi-turn interaction of the speech model.
- 2024.07 - 2025.12: Algorithm Intern, Meituan, Financial Services Platform, LLM Applications.
- SFT training of the interactive model.
- RL optimization of the model on verifiable answers.