arXiv

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a c...

AI 聚合
2026-07-31
arXiv

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...

AI 聚合
2026-07-31
arXiv

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applicati...

AI 聚合
2026-07-31
arXiv

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet ex...

AI 聚合
2026-07-31
arXiv

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic...

AI 聚合
2026-07-31
arXiv

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors g...

AI 聚合
2026-07-31
arXiv

Learning to Trace Seiberg Dualities

Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical ...

AI 聚合
2026-07-31
HuggingFace

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (...

AI 聚合
2026-07-31
HuggingFace

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local vi...

AI 聚合
2026-07-31
HuggingFace

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them tow...

AI 聚合
2026-07-31
HuggingFace

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult t...

AI 聚合
2026-07-31
HuggingFace

Can Large Language Models Execute Parent Orders?

Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smalle...

AI 聚合
2026-07-31
首页 上一页 第 93 / 215 页 下一页 末页