arXiv

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encour...

AI 聚合
2026-08-24
arXiv

From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry

The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy reg...

AI 聚合
2026-08-24
arXiv

Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets

We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minim...

AI 聚合
2026-08-24
arXiv

Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation

Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly ...

AI 聚合
2026-08-24
arXiv

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases...

AI 聚合
2026-08-24
arXiv

AI with Authority, from Application to Silicon

For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here...

AI 聚合
2026-08-24
arXiv

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy image...

AI 聚合
2026-08-24
arXiv

Primal Acceleration of Newton's Method

We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. ...

AI 聚合
2026-08-24
HuggingFace

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute ...

AI 聚合
2026-08-24
HuggingFace

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative think...

AI 聚合
2026-08-24
HuggingFace

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

With large pretrained models, existing methods have effectively improved instruction-based video editing. However, mo...

AI 聚合
2026-08-24
HuggingFace

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safet...

AI 聚合
2026-08-24
首页 上一页 第 34 / 212 页 下一页 末页