CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets m...
每天自动聚合 AI 领域最新动态
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets m...
High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models...
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieva...
Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structur...
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves p...
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different...
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-S...
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although...
This first release of Prior Labs in relational learning shows our continued commitment to open science. We open-sourc...
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whet...
Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fie...
Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Exist...