HuggingFace

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, b...

AI 聚合
2026-08-12
HuggingFace

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context...

AI 聚合
2026-08-12
HuggingFace

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be un...

AI 聚合
2026-08-12
HuggingFace

A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization

In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously updates ...

AI 聚合
2026-08-12
HuggingFace

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Benchmarks for systems that are optimized against the evaluation signal measure something different from what they cl...

AI 聚合
2026-08-12
HuggingFace

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in pen...

AI 聚合
2026-08-12
HuggingFace

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirr...

AI 聚合
2026-08-12
HuggingFace

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals...

AI 聚合
2026-08-12
HuggingFace

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating la...

AI 聚合
2026-08-12
HuggingFace

On-Policy Self-Distillation without Any Supervision

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs)....

AI 聚合
2026-08-12
HuggingFace

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond ...

AI 聚合
2026-08-12
HuggingFace

Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic stru...

AI 聚合
2026-08-12
首页 上一页 第 63 / 212 页 下一页 末页