Constitutional Midtraining: Content Presence Drives Alignment Gains
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isola...
每天自动聚合 AI 领域最新动态
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isola...
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language mo...
Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying thei...
Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, pro...
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. Howe...
Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in sc...
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites ...
The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-...
Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for infe...
The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable ...
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications...
Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical...