Hi CosyVoice team,
I am Shijun Luo from NetEase Cloud Music, where I work on AI news briefing / AI news podcast generation. In our workflow, we use and evaluate major TTS systems to generate spoken news and information podcast content.
During this work, we found that many current TTS systems still struggle with Chinese news-style text, especially compact expressions that frequently appear in real news. These errors are not just voice-quality issues; they can change the information heard by listeners.
For example:
苏-27 may be read as "苏负二十七" instead of the intended aircraft model name.
96-91 may be read as a numeric range instead of a sports score.
620N·m may be read letter by letter or as symbol fragments instead of a torque unit.
3.5% may be read as "三点五百分号" or confused with percentage points.
AI / CEO may be expanded into "人工智能" / "首席执行官" when the original abbreviation should be preserved.
This motivated us to release CN-NewsTTS Bench, a raw-input Chinese news TTS benchmark focused on real-world news reading cases such as dates, numbers, units, named entities, mixed-script text, and text normalization.
The current public leaderboard includes an Alibaba Cloud TTS entry:
- model_id:
aliyun_tts
- model name:
cosyvoice-v3-plus
- voice:
default Chinese voice
Repository:
https://github.com/Jayden-X-L/cn-news-tts-bench
We would like to invite the Alibaba Cloud / DashScope team to:
- Confirm or correct the public model metadata.
- Submit an official result if the current configuration is not representative.
- Provide a system/model card if available.
Submission guide:
https://github.com/Jayden-X-L/cn-news-tts-bench/blob/main/SUBMIT.md
For questions or corrections, feel free to contact me:
xiaobiluo@gmail.com
Thanks!
Best,
Shijun Luo
Hi CosyVoice team,
I am Shijun Luo from NetEase Cloud Music, where I work on AI news briefing / AI news podcast generation. In our workflow, we use and evaluate major TTS systems to generate spoken news and information podcast content.
During this work, we found that many current TTS systems still struggle with Chinese news-style text, especially compact expressions that frequently appear in real news. These errors are not just voice-quality issues; they can change the information heard by listeners.
For example:
苏-27may be read as "苏负二十七" instead of the intended aircraft model name.96-91may be read as a numeric range instead of a sports score.620N·mmay be read letter by letter or as symbol fragments instead of a torque unit.3.5%may be read as "三点五百分号" or confused with percentage points.AI/CEOmay be expanded into "人工智能" / "首席执行官" when the original abbreviation should be preserved.This motivated us to release CN-NewsTTS Bench, a raw-input Chinese news TTS benchmark focused on real-world news reading cases such as dates, numbers, units, named entities, mixed-script text, and text normalization.
The current public leaderboard includes an Alibaba Cloud TTS entry:
aliyun_ttscosyvoice-v3-plusdefault Chinese voiceRepository:
https://github.com/Jayden-X-L/cn-news-tts-bench
We would like to invite the Alibaba Cloud / DashScope team to:
Submission guide:
https://github.com/Jayden-X-L/cn-news-tts-bench/blob/main/SUBMIT.md
For questions or corrections, feel free to contact me:
xiaobiluo@gmail.com
Thanks!
Best,
Shijun Luo