HuggingFace

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability...

AI 聚合
2026-08-27
HuggingFace

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring m...

AI 聚合
2026-08-27
HuggingFace

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unf...

AI 聚合
2026-08-27
HuggingFace

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbo...

AI 聚合
2026-08-27
HuggingFace

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchr...

AI 聚合
2026-08-27
HuggingFace

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval

Document retrieval increasingly supports high-stakes information access in finance, healthcare, and law. Modern retri...

AI 聚合
2026-08-27
HuggingFace

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to su...

AI 聚合
2026-08-27
HuggingFace

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and la...

AI 聚合
2026-08-27
HuggingFace

DREAM Technical Report

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient...

AI 聚合
2026-08-27
HuggingFace

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recog...

AI 聚合
2026-08-27
HuggingFace

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating str...

AI 聚合
2026-08-27
HuggingFace

Automata from Agent Traces: Failure and Next-Step Prediction

LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces re...

AI 聚合
2026-08-27
首页 上一页 第 25 / 212 页 下一页 末页