Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user...
每天自动聚合 AI 领域最新动态
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user...
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged e...
Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games w...
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a ...
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling too...
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scal...
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they ra...
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone...
In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-she...
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a ...
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand syst...
Recent foundation image and video generation models offer strong generalization and controllability, but their direct...