Realtime AI News
ByteDance Releases Seed Audio 1.0 with Fine-Grained Temporal Control for Cinema-Grade Sound Creation
ByteDance has officially launched Seed Audio 1.0, an audio generation model that jointly models dialogue, sound effects, background music, and ambience within a unified framework from a single prompt. The model is now available on the Volcano Fangzhou Experience Center for creator testing.
ByteDance released Seed Audio 1.0 on July 20, marking a new stage in AI audio generation technology that moves beyond single speech synthesis to complete sound scene creation. The model is now available for testing at the Volcano Fangzhou Experience Center.
Traditional cinema-grade audio production requires separate models for vocals, sound effects, and background music, followed by manual editing and mixing — a lengthy process that struggles to maintain narrative coherence. Seed Audio 1.0's core breakthrough is not simply concatenating materials but jointly modeling various audio elements within a unified framework, generating complete sound works that serve storytelling end-to-end.
The model offers three core capabilities: precise temporal-spatial arrangement, supporting 100-millisecond accuracy for dialogue and sound effect timing along the timeline, ideal for video dubbing and ad production; stable voice performance with zero-shot generation and long audio extension, maintaining character voice consistency while naturally expressing different emotions including anger and joy, even allowing the same voice to perform multiple characters; and fluent multilingual support covering more than 20 languages including Chinese, English, and Japanese.
Evaluation data shows the model's audio availability exceeds 90% across nine common creative scenarios, with MOS scores for multilingual generation consistently above 4.0 (excellent level).
On the BytePlus platform, Seed Audio 1.0 supports multiple input modes including text prompts (up to 3,000 characters), reference audio, and reference images, generating up to 120 seconds of audio per segment with extension capabilities for longer formats. BytePlus positions it as the first reference-based audio generation model in its Dola product line.
From Seed-TTS to Seed Audio 1.0, ByteDance's audio AI roadmap is progressively unfolding. The team plans to further integrate multimodal inputs such as video references and explore controllable translation technology, continuously optimizing long-form audio and track generation capabilities.
Why it matters
Seed Audio 1.0 compresses cinema-grade audio production from a multi-model pipeline into a single generation pass, dramatically lowering the barrier to high-quality audio content creation across podcasts, film, advertising, and gaming.
Nearby Updates
All07/20, 16:54
Alibaba Qwen Releases Qwen-Audio-3.0-TTS, New Speech Synthesis Large Model
Alibaba Cloud's Qwen team has released Qwen-Audio-3.0-TTS, the latest iteration of its large speech synthesis model. The model enhances naturalness, rhythm, and emotional expressiveness in text-to-speech generation, expanding Qwen's multimodal product ecosystem into the audio domain.
07/20, 14:57
ZhiXiang Future Launches vivago R1, Claimed World's First Unlimited-Duration AI Video Agent
ZhiXiang Future has launched vivago R1, which it claims is the world's first AI video creation agent with unlimited duration, breaking the previous 15-second generation limit. The company reports an 85% commercial usability rate, marking a significant step toward professional-grade AI video production.
07/20, 17:41
China unveils world-first AI for ADANES technical roadmap, integrating AI with advanced nuclear energy
At the 2026 World AI Conference, Chinese scientists unveiled the world's first AI for ADANES technical roadmap, using physics-intrinsic world models to intelligently control advanced nuclear energy systems. A new industry alliance was also launched to drive AI-nuclear fusion across data, modeling, LLMs, simulation, and safety frameworks.
07/20, 17:54
AI enters the most human-dependent industry: a fourth-tier city rehab clinic sees 40% profit growth
RICE AI, developed by Chinese special education firm Damihexiaomi, has been deployed in rehabilitation centers across third- and fourth-tier cities for six months, cutting assessment time by over 10x. One clinic in Henan reported a 40% profit increase in the first half of 2026, validating the business case for AI-assisted therapy for children with special needs.