Harmonizing AI Safety Thresholds
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third p...
每天自动聚合 AI 领域最新动态
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third p...
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next mo...
AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, wh...
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded e...
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantag...
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapsh...
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-traini...
Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, ...
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover...
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components w...
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory...
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily r...