HuggingFace

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model...

AI 聚合
2026-09-02
HuggingFace

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we ...

AI 聚合
2026-09-02
HuggingFace

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-ca...

AI 聚合
2026-09-02
HuggingFace

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation

Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and count...

AI 聚合
2026-09-02
HuggingFace

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose...

AI 聚合
2026-09-02
HuggingFace

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model mu...

AI 聚合
2026-09-02
arXiv

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an e...

AI 聚合
2026-09-01
arXiv

Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents

Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated sta...

AI 聚合
2026-09-01
arXiv

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, t...

AI 聚合
2026-09-01
arXiv

Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring

We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n...

AI 聚合
2026-09-01
arXiv

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR...

AI 聚合
2026-09-01
arXiv

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literatu...

AI 聚合
2026-09-01
首页 上一页 第 12 / 212 页 下一页 末页