HuggingFace

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track o...

AI 聚合
2026-08-19
HuggingFace

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, tempo...

AI 聚合
2026-08-19
HuggingFace

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with ...

AI 聚合
2026-08-19
HuggingFace

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome r...

AI 聚合
2026-08-19
arXiv

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding

We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report...

AI 聚合
2026-08-18
arXiv

Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment

Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition reso...

AI 聚合
2026-08-18
arXiv

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to t...

AI 聚合
2026-08-18
arXiv

When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step ...

AI 聚合
2026-08-18
arXiv

Quipu: A Governed Bitemporal Knowledge Graph Store

Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: acce...

AI 聚合
2026-08-18
arXiv

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but exis...

AI 聚合
2026-08-18
arXiv

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Generative pretraining established reusable task representations; later work on language-based task conditioning and ...

AI 聚合
2026-08-18
arXiv

Model Hypnosis: Strong control of AI via additive subliminal effects

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually w...

AI 聚合
2026-08-18
首页 上一页 第 46 / 212 页 下一页 末页