Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability...
每天自动聚合 AI 领域最新动态
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability...
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically...
Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future querie...
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension wi...
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large...
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While ...
Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but prac...
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can f...
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach aba...
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames se...
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equ...
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that...