FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model
Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame r...
每天自动聚合 AI 领域最新动态
Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame r...
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of action...
The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains...
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet t...
Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities...
Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such ev...
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on roboti...
Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surf...
Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-makin...
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assemb...
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven softwar...
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edit...