HuggingFace

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained kno...

AI 聚合
2026-07-30
HuggingFace

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations r...

AI 聚合
2026-07-30
HuggingFace

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host arti...

AI 聚合
2026-07-30
HuggingFace

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

Recent game world models can generate visually realistic and interactive environments conditioned on player actions. ...

AI 聚合
2026-07-30
HuggingFace

CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents

Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, l...

AI 聚合
2026-07-30
HuggingFace

GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels

Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definit...

AI 聚合
2026-07-30
HuggingFace

Edge-Aware Thermal Infrared UAV Swarm Tracking

Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, trac...

AI 聚合
2026-07-30
HuggingFace

Projection Pursuit CPCANet for Domain Generalization

Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment me...

AI 聚合
2026-07-30
HuggingFace

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation d...

AI 聚合
2026-07-30
HuggingFace

OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distrib...

AI 聚合
2026-07-30
HuggingFace

Reinforcement Learning for Code Optimization

RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and ...

AI 聚合
2026-07-30
HuggingFace

Uncovering Latent Reasoning Strategies in Language Models

A language model p_θ(y mid x) trained on reasoning tasks learns to solve problems via multiple distinct strategies, y...

AI 聚合
2026-07-30
首页 上一页 第 98 / 215 页 下一页 末页