PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a c...
每天自动聚合 AI 领域最新动态
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a c...
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applicati...
Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet ex...
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic...
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors g...
Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical ...
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (...
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local vi...
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them tow...
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult t...
Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smalle...