Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting ...
每天自动聚合 AI 领域最新动态
We study reference-free post-training for multilingual machine translation with open large language models. Starting ...
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but cont...
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, compl...
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus ...
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how ...
Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained ...
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of cap...
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a ph...
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attenti...
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone...
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs pres...
Gradient descent on a factored model W = UV^top is implicitly biased toward low-rank solutions, while Adam, starting ...