An AI4AI Framework for Visual Token Pruning
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet exi...
每天自动聚合 AI 领域最新动态
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet exi...
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, ...
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recent...
Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disrup...
Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffu...
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evi...
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packag...
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely f...
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primar...
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inhere...
Large language models often fail when answer options require combining atomic judgments under explicit logical operat...
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-...