HuggingFace

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...

AI 聚合
2026-09-03
HuggingFace

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decom...

AI 聚合
2026-09-03
HuggingFace

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step ...

AI 聚合
2026-09-03
HuggingFace

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video...

AI 聚合
2026-09-03
HuggingFace

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You

Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption ...

AI 聚合
2026-09-03
HuggingFace

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...

AI 聚合
2026-09-03
HuggingFace

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-...

AI 聚合
2026-09-03
HuggingFace

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger ...

AI 聚合
2026-09-03
HuggingFace

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, o...

AI 聚合
2026-09-03
HuggingFace

Agent Memory Is a Surface for Endogenous Authorization Laundering

Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrict...

AI 聚合
2026-09-03
HuggingFace

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated Gi...

AI 聚合
2026-09-03
arXiv

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due...

AI 聚合
2026-09-02
首页 上一页 第 8 / 212 页 下一页 末页