arXiv

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-lev...

AI 聚合
2026-09-01
arXiv

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...

AI 聚合
2026-09-01
arXiv

Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations

Accurate daily predictions of cold hardiness in woody plants are critical in regions where freezing temperatures can ...

AI 聚合
2026-09-01
arXiv

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

Industrial post-training is a brownfield regime. Teams inherit a deployed checkpoint and must land targeted improveme...

AI 聚合
2026-09-01
arXiv

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deploymen...

AI 聚合
2026-09-01
arXiv

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently character...

AI 聚合
2026-09-01
arXiv

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural ali...

AI 聚合
2026-09-01
arXiv

Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platf...

AI 聚合
2026-09-01
arXiv

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies

Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: co...

AI 聚合
2026-09-01
HuggingFace

Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation

Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop b...

AI 聚合
2026-09-01
HuggingFace

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observe...

AI 聚合
2026-09-01
HuggingFace

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B paramet...

AI 聚合
2026-09-01
首页 上一页 第 13 / 212 页 下一页 末页