When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents
Long-horizon LLM agents can fail quietly: they settle on one reading of the evidence early, then spend the rest of th...
每天自动聚合 AI 领域最新动态
Long-horizon LLM agents can fail quietly: they settle on one reading of the evidence early, then spend the rest of th...
We introduce ShotcreteDepth, a bi-modal dataset from the construction domain that captures both an active shotcreting...
Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis c...
Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observatio...
Discrete text-trigger optimization -- searching for text sequences that, when ingested by a model, steer it toward a ...
Long-context reasoning is an essential capability for large language models, particularly when they are deployed as a...
Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately...
The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, m...
Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demo...
Can representations learned for image generation also support the evaluation of generated images? We study text-to-im...
Remote patient monitoring depends on patient-reported data to capture the subjective dimension of recovery that devic...
A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produc...