Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability...
每天自动聚合 AI 领域最新动态
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability...
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring m...
Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unf...
Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbo...
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchr...
Document retrieval increasingly supports high-stakes information access in finance, healthcare, and law. Modern retri...
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to su...
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and la...
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient...
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recog...
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating str...
LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces re...