HuggingFace

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame r...

AI 聚合
2026-07-02
HuggingFace

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of action...

AI 聚合
2026-07-02
HuggingFace

Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature

The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains...

AI 聚合
2026-07-02
HuggingFace

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet t...

AI 聚合
2026-07-02
HuggingFace

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities...

AI 聚合
2026-07-02
HuggingFace

Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents

Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such ev...

AI 聚合
2026-07-02
HuggingFace

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on roboti...

AI 聚合
2026-07-02
HuggingFace

Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surf...

AI 聚合
2026-07-02
HuggingFace

Hierarchical Experimentalist Agents

Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-makin...

AI 聚合
2026-07-02
HuggingFace

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assemb...

AI 聚合
2026-07-02
HuggingFace

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven softwar...

AI 聚合
2026-07-02
HuggingFace

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edit...

AI 聚合
2026-07-02
首页 上一页 第 161 / 215 页 下一页 末页