RoPE-Aware Bit Allocation for KV-Cache Quantization
Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contr...
每天自动聚合 AI 领域最新动态
Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contr...
The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. ...
Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Op...
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality train...
Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LL...
Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a da...
While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynam...
We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single succ...
Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evid...
Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and ...
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safe...
"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whe...