HuggingFace

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovi...

AI 聚合
2026-07-09
HuggingFace

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematicall...

AI 聚合
2026-07-09
HuggingFace

Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo

Large Language Models (LLMs) unlocked new possibilities in automated code writing, becoming the backbone of most code...

AI 聚合
2026-07-09
HuggingFace

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reacti...

AI 聚合
2026-07-09
HuggingFace

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

Complex image creation and editing often require more than a single generation or editing model. A user request may i...

AI 聚合
2026-07-09
HuggingFace

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding ...

AI 聚合
2026-07-09
HuggingFace

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies documen...

AI 聚合
2026-07-09
HuggingFace

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer visio...

AI 聚合
2026-07-09
HuggingFace

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without ...

AI 聚合
2026-07-09
HuggingFace

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation re...

AI 聚合
2026-07-09
HuggingFace

Attending to Multimodal Generation One Token at a Time

Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic infor...

AI 聚合
2026-07-09
HuggingFace

RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules

We present RuleChef, a framework that uses large language models (LLMs) to generate executable rules for NLP tasks su...

AI 聚合
2026-07-09
首页 上一页 第 145 / 216 页 下一页 末页