HuggingFace

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...

AI 聚合
2026-09-03
HuggingFace

Aspire: Can Models Self-Evolve from Vague Goals?

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at re...

AI 聚合
2026-09-03
HuggingFace

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external ex...

AI 聚合
2026-09-03
HuggingFace

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts m...

AI 聚合
2026-09-03
HuggingFace

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe com...

AI 聚合
2026-09-03
HuggingFace

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language ...

AI 聚合
2026-09-03
HuggingFace

VibeVoice-ASR-Streaming Technical Report

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-t...

AI 聚合
2026-09-03
HuggingFace

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-...

AI 聚合
2026-09-03
HuggingFace

ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize...

AI 聚合
2026-09-03
HuggingFace

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable ima...

AI 聚合
2026-09-03
HuggingFace

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involv...

AI 聚合
2026-09-03
HuggingFace

Exploring Collaboration between a language and a non-language agent

LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through ...

AI 聚合
2026-09-03
首页 上一页 第 7 / 212 页 下一页 末页