arXiv

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carr...

AI 聚合
2026-07-30
HuggingFace

Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

Most video editing systems still lack explicit layered video representations, limiting their ability to perform reali...

AI 聚合
2026-07-30
HuggingFace

GPT-Red: Automated Red Teaming via Self-Play at Scale

We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks again...

AI 聚合
2026-07-30
HuggingFace

HumanCLAW: Can Vision-Language Models Act Through a Body?

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an ac...

AI 聚合
2026-07-30
HuggingFace

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are ...

AI 聚合
2026-07-30
HuggingFace

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carr...

AI 聚合
2026-07-30
HuggingFace

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing be...

AI 聚合
2026-07-30
HuggingFace

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelli...

AI 聚合
2026-07-30
HuggingFace

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-mak...

AI 聚合
2026-07-30
HuggingFace

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather tha...

AI 聚合
2026-07-30
HuggingFace

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet sta...

AI 聚合
2026-07-30
HuggingFace

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit cri...

AI 聚合
2026-07-30
首页 上一页 第 97 / 215 页 下一页 末页