Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encour...
每天自动聚合 AI 领域最新动态
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encour...
The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy reg...
We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minim...
Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly ...
Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases...
For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here...
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy image...
We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. ...
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute ...
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative think...
With large pretrained models, existing methods have effectively improved instruction-based video editing. However, mo...
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safet...