One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation...
每天自动聚合 AI 领域最新动态
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation...
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation whi...
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an...
Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundat...
With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingl...
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, ...
User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavio...
Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devi...
Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organ...
Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, ...
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle e...
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution path...