HuggingFace

Self-Evolving Coding Agents

Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect r...

AI 聚合
2026-08-06
HuggingFace

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matchin...

AI 聚合
2026-08-06
HuggingFace

Multi-Task Multi-Frame Visual Piano Transcription

Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist...

AI 聚合
2026-08-06
HuggingFace

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows flo...

AI 聚合
2026-08-06
HuggingFace

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability...

AI 聚合
2026-08-06
HuggingFace

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically...

AI 聚合
2026-08-06
HuggingFace

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future querie...

AI 聚合
2026-08-06
HuggingFace

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension wi...

AI 聚合
2026-08-06
HuggingFace

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large...

AI 聚合
2026-08-06
HuggingFace

LegalPincite: Multi-level Legal Information Retrieval Dataset

A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While ...

AI 聚合
2026-08-06
HuggingFace

ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but prac...

AI 聚合
2026-08-06
HuggingFace

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can f...

AI 聚合
2026-08-06
首页 上一页 第 78 / 212 页 下一页 末页