arXiv

SceneBind: Binding What and Where Across Vision, Audio and Language

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understandi...

AI 聚合
2026-07-17
arXiv

Pretraining Data Can Be Poisoned through Computational Propaganda

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior wo...

AI 聚合
2026-07-17
arXiv

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors r...

AI 聚合
2026-07-17
arXiv

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-T...

AI 聚合
2026-07-17
HuggingFace

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-T...

AI 聚合
2026-07-17
HuggingFace

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn inte...

AI 聚合
2026-07-17
HuggingFace

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without cha...

AI 聚合
2026-07-17
HuggingFace

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them ...

AI 聚合
2026-07-17
HuggingFace

BadWAM: When World-Action Models Dream Right but Act Wrong

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting action...

AI 聚合
2026-07-17
HuggingFace

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...

AI 聚合
2026-07-17
HuggingFace

WanSong v1.0 Technical Report

Music generation foundation models have recently attracted significant industry attention. However, achieving efficie...

AI 聚合
2026-07-17
HuggingFace

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, ...

AI 聚合
2026-07-17
首页 上一页 第 127 / 216 页 下一页 末页