HuggingFace

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-languag...

AI 聚合
2026-08-20
HuggingFace

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the ge...

AI 聚合
2026-08-20
HuggingFace

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM...

AI 聚合
2026-08-20
HuggingFace

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, whi...

AI 聚合
2026-08-20
HuggingFace

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the...

AI 聚合
2026-08-20
HuggingFace

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-speci...

AI 聚合
2026-08-20
HuggingFace

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integrati...

AI 聚合
2026-08-20
HuggingFace

The Problem Is the Problem: Towards Scalable Mathematical Discovery

AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model re...

AI 聚合
2026-08-20
HuggingFace

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remai...

AI 聚合
2026-08-20
arXiv

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and ...

AI 聚合
2026-08-19
arXiv

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement b...

AI 聚合
2026-08-19
arXiv

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although r...

AI 聚合
2026-08-19
首页 上一页 第 42 / 212 页 下一页 末页