OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by c...
每天自动聚合 AI 领域最新动态
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by c...
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interac...
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching ...
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but th...
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregres...
Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly ...
Scientific literature search often requires more than retrieving papers from a single query: users' intents are under...
Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remai...
We study Generated Contents Enrichment (GCE), a conditional image-generation task in which a sparse scene description...
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by no...
Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractiv...
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deploym...