HuggingFace

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. Howev...

AI 聚合
2026-08-28
HuggingFace

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory s...

AI 聚合
2026-08-28
HuggingFace

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. R...

AI 聚合
2026-08-28
HuggingFace

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remai...

AI 聚合
2026-08-28
HuggingFace

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local ...

AI 聚合
2026-08-28
HuggingFace

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more th...

AI 聚合
2026-08-28
HuggingFace

CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension

Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal ...

AI 聚合
2026-08-28
HuggingFace

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies i...

AI 聚合
2026-08-28
HuggingFace

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they...

AI 聚合
2026-08-28
HuggingFace

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and upda...

AI 聚合
2026-08-28
HuggingFace

TTPO: Test-Time Policy Optimization

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), h...

AI 聚合
2026-08-28
HuggingFace

GameWAM: A World Action Model for Video Games

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous n...

AI 聚合
2026-08-28
首页 上一页 第 21 / 212 页 下一页 末页