An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it fr...
每天自动聚合 AI 领域最新动态
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it fr...
Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven i...
With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metr...
As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have f...
Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and...
We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly ...
Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene ...
Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet t...
AI-assisted software development has moved from line-level autocomplete to agents that can plan changes, edit files, ...
Wearable devices and smartphones generate rich behavioural time series that can support proactive health intervention...
Survival prediction plays a central role for healthcare providers and clinical researchers. Accurate risk stratificat...
We show that the three movements of Beethoven's "Moonlight Sonata" (Op. 27 No. 2) instantiate three distinct machine ...