arxiv:2409.18042

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Published on Sep 26

· Submitted by

akhaliq on Sep 27

#2 Paper of the day

Upvote

Authors:

Kai Chen ,

Yunhao Gou ,

Runhui Huang ,

Zhili Liu ,

Daxin Tan ,

Chunwei Wang ,

Yihan Zeng ,

Kuo Yang ,

Abstract

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging in the open-source community. Existing vision-language models rely on external tools for the speech processing, while speech-language models still suffer from limited or even without vision-understanding abilities. To address this gap, we propose EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech capabilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we notice surprisingly that omni-modal alignment can further enhance vision-language and speech abilities compared with the corresponding bi-modal aligned counterparts. Moreover, a lightweight style module is proposed for flexible speech style controls (e.g., emotions and pitches). For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions.

View arXiv page View PDF Add to collection

Community

akhaliq

Paper submitter Sep 27

https://emova-ollm.github.io/

MichaelKarpe

Oct 15

I have a connection not secure warning when trying to access the URL from several devices, would you be able to fix this? Do you have any update on when a model demo would be available on HuggingFace?

Lyte

Sep 27

will the model weights be released?

KaiChen1998

Paper author Sep 27

Will soon release the checkpoint after we get back from ECCV (the main authors are all catching flights for Milano today 😂). We are busy preparing an HF demo for temporal usage. Stay tuned!