SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, ...
每天自动聚合 AI 领域最新动态
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, ...
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains dif...
Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work ha...
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and mo...
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them i...
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledg...
Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially...
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient m...
Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still d...
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text...
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool us...
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the tea...