Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its co...
每天自动聚合 AI 领域最新动态
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its co...
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactio...
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is ...
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural ...
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot e...
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation...
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural represen...
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the p...
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks t...
When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differ...
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming exp...
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answ...