How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention
Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a ta...
每天自动聚合 AI 领域最新动态
Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a ta...
Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requir...
Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer ...
Intent misinterpretation during vehicle interactions causes recurring planning failures. We study a decision layer in...
Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions t...
Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning...
Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning d...
Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic v...
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying...
Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, o...
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed...
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and vi...