About Me

I am Dongjie Fu (付栋杰), a second-year master’s student at the School of Software Technology, Zhejiang University. My advisor is Professor Tao Jin (金涛), in Professor Zhou Zhao (赵洲)’s lab. Before that, I received my B.E. degree from Shandong University.

My research focuses on Speech / Audio Large Language Models and Audio-Visual Understanding. I have published 11 papers as first author / co-first author in top-tier conferences and journals such as NeurIPS, EMNLP, ACM MM and TIP.

My current research topics include:

  • Spoken Dialogue Systems / Audio Large Language Models
  • Audio-Visual Speech Understanding

I am actively looking for job opportunities, feel free to drop me an email.

🔥 News

  • 2026.07: 🎉🎉 2 papers (1 first-author paper and 1 co-first-author paper) are accepted by ACM MM 2026!
  • 2026.05: 🎉🎉 2 papers (co-first author) are accepted by Interspeech 2026!
  • 2026.04: 🎉🎉 1 paper (first author) is accepted by IEEE TIP!
  • 2026.01: 🎉🎉 1 paper (co-first author) is accepted by ACL 2026!
  • 2026.01: 🎉🎉 1 paper (major author) is accepted by ICLR 2026!
  • 2025.12: I start my internship at Tencent, Hunyuan Multimodal Model Department, Speech Algorithm Center.
  • 2025.09: 🎉🎉 1 paper (first author) is accepted by EMNLP 2025!
  • 2025.09: 🎉🎉 1 paper (co-first author) is accepted by NeurIPS 2025!
  • 2025.05: 🎉🎉 1 paper (major author) is accepted by KDD 2025!
  • 2025.01: 🎉🎉 1 paper (major author) is accepted by ICLR 2025!
  • 2024.07: 🎉🎉 1 paper (first author, Oral) and 1 paper (co-first author) are accepted by ACM MM 2024!
  • 2024.07: I start my internship at Meituan, Financial Services Platform.
  • 2024.06: 🎓 I receive my B.E. degree from Shandong University and join Zhejiang University as a master's student.

📝 Publications

Arxiv
X3-OPD framework
  • X³-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin EMNLP 2026 Submission (First Author)

    Paper

    We propose X³-OPD, a cross-modal on-policy distillation framework that transfers reasoning from a text teacher to an audio-language student by scoring the student’s own audio-conditioned trajectories with matched textual inputs. Built on a 71.8K three-tier symmetric corpus spanning logical, audio-event, and prosody-aware dialogue reasoning, X³-OPD improves BIG Bench Audio from 87.9 to 93.6 (+5.7) while largely preserving out-of-domain capabilities.

ACM MM 2026
RoleJudge framework
EMNLP 2025
sym
  • PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue Dongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin EMNLP 2025 (First Author)

    Demo | Paper

    We build Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset with user profiles, and propose PAChat, which jointly models semantic and speaker representations to enable persona-aware personalized responses.

Interspeech 2026
sym
  • X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs Di Cao*, Dongjie Fu*, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin Interspeech 2026 (Co-first Author)

    Paper

    We propose X-OPD, a cross-modal on-policy distillation framework that aligns the capabilities of Speech LLMs to their text-based counterparts via token-level feedback from a text teacher, closing the performance gap while preserving speech abilities.

ACM MM 2024 Oral
sym
IEEE TIP
sym

Full Publication List

[*] denotes co-first authors, [#] denotes co-supervised.

I. Spoken Dialogue Systems & Audio LLMs

II. Audio-Visual Understanding

III. Others

📖 Educations

  • 2024.09 - 2027.06, Master, School of Software Technology, Zhejiang University, Ningbo.

  • 2020.09 - 2024.06, Undergraduate, School of Computer Science and Technology, Shandong University, Qingdao.

🎖 Honors and Awards

  • Top 5% in major ranking, Zhejiang University. (2025-2026)
  • Huawei Intelligent Base Scholarship (华为智能基座奖学金). (2024)
  • Top 10% in major ranking, Shandong University. (2020-2024)
  • Academic Scholarship, Shandong University. (2020-2023)

💬 Professional Services

  • Conference Reviewer: ICLR, NeurIPS, EMNLP, ACM MM, Interspeech, ICASSP.

💻 Internships & Projects

  • 2025.12 - Now: Algorithm Intern, Tencent, Hunyuan Multimodal Model Department, Speech Algorithm Center.
    • On-Policy Distillation (OPD) stage training for the speech model.
    • Building agent capabilities for the speech model.
    • Optimizing the multi-turn interaction of the speech model.
  • 2024.07 - 2025.12: Algorithm Intern, Meituan, Financial Services Platform, LLM Applications.
    • SFT training of the interactive model.
    • RL optimization of the model on verifiable answers.