FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage,...
每天自动聚合 AI 领域最新动态
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage,...
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animat...
Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-determinis...
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-pu...
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models...
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterior...
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the curren...
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue re...
Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support...
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains...
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, ...
The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in stru...