HuggingFace

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...

AI 聚合
2026-06-19
HuggingFace

LLM-Enabled NWDAF: A Step Toward AI-Native 6G Network Intelligence

The Network Data Analytics Function (NWDAF) is central to enabling zero-touch network management in fifth-generation ...

AI 聚合
2026-06-19
HuggingFace

A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets

Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more ...

AI 聚合
2026-06-19
HuggingFace

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe t...

AI 聚合
2026-06-19
HuggingFace

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between eval...

AI 聚合
2026-06-19
HuggingFace

iOSWorld: A Benchmark for Personally Intelligent Phone Agents

A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and prefer...

AI 聚合
2026-06-19
HuggingFace

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns

We show the standard basis of transformer hidden states already provides a training-free, architecture-general featur...

AI 聚合
2026-06-19
HuggingFace

ViT-Up: Faithful Feature Upsampling for Vision Transformers

Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptio...

AI 聚合
2026-06-19
HuggingFace

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction

Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan a...

AI 聚合
2026-06-19
HuggingFace

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-l...

AI 聚合
2026-06-19
HuggingFace

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

As an increasing majority of global video content is consumed on social platforms for interactive social purposes, vi...

AI 聚合
2026-06-19
HuggingFace

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with...

AI 聚合
2026-06-19
首页 上一页 第 190 / 212 页 下一页 末页