ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, produ...
每天自动聚合 AI 领域最新动态
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, produ...
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurem...
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate,...
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become m...
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-t...
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing...
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thoug...
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected sce...
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera t...