HuggingFace

EnvHarness: Awakening Static Worlds for Agent Learning

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agen...

AI 聚合
2026-08-21
HuggingFace

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people....

AI 聚合
2026-08-21
HuggingFace

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native con...

AI 聚合
2026-08-21
HuggingFace

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Memory has become a key component of large language models, enabling them to retain information and learn from long-t...

AI 聚合
2026-08-21
HuggingFace

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reco...

AI 聚合
2026-08-21
HuggingFace

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Large language models often fail to answer questions about a bounded document collection when the source documents ar...

AI 聚合
2026-08-21
HuggingFace

Repo0: Design-Driven Zero-to-All Code Generation

Large language model agents have made substantial progress in code generation, yet most existing systems assume a pre...

AI 聚合
2026-08-21
HuggingFace

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no ...

AI 聚合
2026-08-21
HuggingFace

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention re...

AI 聚合
2026-08-21
HuggingFace

Chain-of-Experience for Continual LLM Improvement

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the mod...

AI 聚合
2026-08-21
HuggingFace

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capab...

AI 聚合
2026-08-21
HuggingFace

Towards Quantifying Benchmark Optimization in ASR Models

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature...

AI 聚合
2026-08-21
首页 上一页 第 38 / 212 页 下一页 末页