HuggingFace

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it fr...

AI 聚合
2026-06-16
HuggingFace

Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven i...

AI 聚合
2026-06-16
HuggingFace

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metr...

AI 聚合
2026-06-16
HuggingFace

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions

As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have f...

AI 聚合
2026-06-16
HuggingFace

AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties

Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and...

AI 聚合
2026-06-16
HuggingFace

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly ...

AI 聚合
2026-06-16
HuggingFace

Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion

Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene ...

AI 聚合
2026-06-16
arXiv

AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet t...

AI 聚合
2026-06-15
arXiv

Regulating the Machine Contributor: Governance and Policy Alignment in Open Source

AI-assisted software development has moved from line-level autocomplete to agents that can plan changes, edit files, ...

AI 聚合
2026-06-15
arXiv

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

Wearable devices and smartphones generate rich behavioural time series that can support proactive health intervention...

AI 聚合
2026-06-15
arXiv

Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts

Survival prediction plays a central role for healthcare providers and clinical researchers. Accurate risk stratificat...

AI 聚合
2026-06-15
arXiv

Moonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven's Op. 27 No. 2 and Machine Learning Mechanisms

We show that the three movements of Beethoven's "Moonlight Sonata" (Op. 27 No. 2) instantiate three distinct machine ...

AI 聚合
2026-06-15
首页 上一页 第 201 / 212 页 下一页 末页