📝 Publications

Arxiv
X3-OPD framework
  • X³-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin EMNLP 2026 Submission (First Author)

    Paper

    We propose X³-OPD, a cross-modal on-policy distillation framework that transfers reasoning from a text teacher to an audio-language student by scoring the student’s own audio-conditioned trajectories with matched textual inputs. Built on a 71.8K three-tier symmetric corpus spanning logical, audio-event, and prosody-aware dialogue reasoning, X³-OPD improves BIG Bench Audio from 87.9 to 93.6 (+5.7) while largely preserving out-of-domain capabilities.

ACM MM 2026
RoleJudge framework
EMNLP 2025
sym
  • PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue Dongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin EMNLP 2025 (First Author)

    Demo | Paper

    We build Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset with user profiles, and propose PAChat, which jointly models semantic and speaker representations to enable persona-aware personalized responses.

Interspeech 2026
sym
  • X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs Di Cao*, Dongjie Fu*, Hai Yu, Siqi Zheng, Xu Tan, Tao Jin Interspeech 2026 (Co-first Author)

    Paper

    We propose X-OPD, a cross-modal on-policy distillation framework that aligns the capabilities of Speech LLMs to their text-based counterparts via token-level feedback from a text teacher, closing the performance gap while preserving speech abilities.

ACM MM 2024 Oral
sym
IEEE TIP
sym

Full Publication List

[*] denotes co-first authors, [#] denotes co-supervised.

I. Spoken Dialogue Systems & Audio LLMs

II. Audio-Visual Understanding

III. Others