HuggingFace

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-st...

AI 聚合
2026-06-18
HuggingFace

Native Active Perception as Reasoning for Omni-Modal Understanding

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly ...

AI 聚合
2026-06-18
HuggingFace

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest co...

AI 聚合
2026-06-18
HuggingFace

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are ...

AI 聚合
2026-06-18
arXiv

ReAge3D: Re-Aging 3D Faces with View Consistency

We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity...

AI 聚合
2026-06-17
arXiv

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act

Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate wheth...

AI 聚合
2026-06-17
arXiv

All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code

Software practitioners increasingly use AI coding agents that generate test code alongside production code in open so...

AI 聚合
2026-06-17
arXiv

IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction

Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable...

AI 聚合
2026-06-17
arXiv

Kolmogorov Regression for Robust Diffusion Policies

Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long...

AI 聚合
2026-06-17
arXiv

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly foc...

AI 聚合
2026-06-17
arXiv

The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data

As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce a...

AI 聚合
2026-06-17
arXiv

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 a...

AI 聚合
2026-06-17
首页 上一页 第 194 / 212 页 下一页 末页