arXiv

Learning Process Rewards via Success Visitation Matching for Efficient RL

In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently s...

AI 聚合
2026-06-23
arXiv

TailorMind: Towards Preference-Aligned Multimodal Content Generation

Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly...

AI 聚合
2026-06-23
arXiv

Tapered Language Models

Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack ...

AI 聚合
2026-06-23
arXiv

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Ma...

AI 聚合
2026-06-23
arXiv

PsyBridge: A Hybrid Intelligent Framework for Multi-Dimensional Mental Health Assessment and Decision Support

Mental health assessment commonly relies on isolated screening instruments or data-driven models that often lack inte...

AI 聚合
2026-06-23
arXiv

Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostl...

AI 聚合
2026-06-23
arXiv

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large lang...

AI 聚合
2026-06-23
arXiv

Semantic Browsing: Controllable Diversity for Image Generation

Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at th...

AI 聚合
2026-06-23
arXiv

CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation

Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipula...

AI 聚合
2026-06-23
HuggingFace

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

Multi-view 3D Visual Question Answering (MV3D-VQA) requires integrating partial observations into a coherent 3D scene...

AI 聚合
2026-06-23
HuggingFace

Self-Compacting Language Model Agents

Long agent traces composed of chains of thought and tool calls accumulate stale content that anchor subsequent genera...

AI 聚合
2026-06-23
HuggingFace

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

Latent action pretraining learns representations of visual change from pairs of observations, but existing methods ty...

AI 聚合
2026-06-23
首页 上一页 第 183 / 212 页 下一页 末页