MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, b...
每天自动聚合 AI 领域最新动态
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, b...
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context...
Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be un...
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously updates ...
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they cl...
Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in pen...
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirr...
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals...
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating la...
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs)....
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond ...
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic stru...