返回
arXiv

Efficient Test-Time Adaptation through Human-AI Interaction

AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet...

AI 聚合 09/04
arXiv

A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle

This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniatu...

AI 聚合 09/04
arXiv

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...

AI 聚合 09/04
arXiv

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them...

AI 聚合 09/04
arXiv

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to la...

AI 聚合 09/04
arXiv

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but exi...

AI 聚合 09/04
arXiv

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and bui...

AI 聚合 09/04
arXiv

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...

AI 聚合 09/04
arXiv

A Computationally Feasible Framework for Causal Probabilistic Explanation

Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical...

AI 聚合 09/04
arXiv

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit ...

AI 聚合 09/04
arXiv

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only...

AI 聚合 09/04
arXiv

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editi...

AI 聚合 09/04
arXiv

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, produ...

AI 聚合 09/04
arXiv

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurem...

AI 聚合 09/04
arXiv

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...

AI 聚合 09/04
HuggingFace

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...

AI 聚合 09/04
HuggingFace

Principia: Relational Physics Tests for Video Models

Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate,...

AI 聚合 09/04
HuggingFace

Environment Evolution for Terminal Agents

Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become m...

AI 聚合 09/04
HuggingFace

Editable Visual Design

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-t...

AI 聚合 09/04
HuggingFace

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing...

AI 聚合 09/04
HuggingFace

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...

AI 聚合 09/04
HuggingFace

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thoug...

AI 聚合 09/04
HuggingFace

WorldReward: Reward Modeling for Camera-Conditioned World Models

Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected sce...

AI 聚合 09/04
HuggingFace

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera t...

AI 聚合 09/04
HuggingFace

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...

AI 聚合 09/04
HuggingFace

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs a...

AI 聚合 09/04
HuggingFace

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation...

AI 聚合 09/04
HuggingFace

Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training

Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often ...

AI 聚合 09/04
HuggingFace

PACE: Towards Surfacing Hidden Conflicts in User Requests

Personalized assistants should not only comply with user requests but also assess whether those requests are appropri...

AI 聚合 09/04
HuggingFace

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invar...

AI 聚合 09/04
HuggingFace

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...

AI 聚合 09/04
HuggingFace

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 p...

AI 聚合 09/04
HuggingFace

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fi...

AI 聚合 09/04
HuggingFace

Using Grounded Theory for Agent Behavior Analysis at Scale

Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in lon...

AI 聚合 09/04
HuggingFace

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origam...

AI 聚合 09/04
HuggingFace

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

Multimodal models often build on architectures designed for generative vision-language modeling, typically combining ...

AI 聚合 09/04
HuggingFace

Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations

Spatial return models take the interaction matrix as given and leave feedback uninterpreted. We construct a bandwidth...

AI 聚合 09/04
HuggingFace

Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations

Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are diffic...

AI 聚合 09/04
HuggingFace

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of...

AI 聚合 09/04
HuggingFace

Small Language Models as Judges for Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring res...

AI 聚合 09/04
HuggingFace

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

Agent performance depends jointly on the model parameters and the executable harness code that manages context and co...

AI 聚合 09/04
HuggingFace

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

A language model's prediction of its next token develops across layers, and lens methods track this process by decodi...

AI 聚合 09/04
HuggingFace

An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems

Conversational Recommender Systems (CRS) typically require domain-specific dialogue data, which is costly, scarce, an...

AI 聚合 09/04
arXiv

Language Models Can Control Their Own Attention

Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...

AI 聚合 09/03
arXiv

HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design

Polymeric materials are central to modern technologies, with applications ranging from energy to health and transport...

AI 聚合 09/03
arXiv

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...

AI 聚合 09/03
arXiv

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve throu...

AI 聚合 09/03
arXiv

Untangling the Mechanisms of Misleading Context in Medical Question Answering

Large language models now answer medical questions with expert-level performance. However, the context these systems ...

AI 聚合 09/03
arXiv

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

On-premise assistants can give factory workers conversational access to machine documentation, but models capable of ...

AI 聚合 09/03
arXiv

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention va...

AI 聚合 09/03
arXiv

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with th...

AI 聚合 09/03
arXiv

Dutch Books for Language Models

People increasingly use language models to support life decisions. Many such decisions involve a probabilistic foreca...

AI 聚合 09/03
arXiv

frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study

For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record ha...

AI 聚合 09/03
arXiv

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations i...

AI 聚合 09/03
arXiv

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupationa...

AI 聚合 09/03
arXiv

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Competitive programming has become a key test of large language model reasoning, with international competitions such...

AI 聚合 09/03
arXiv

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investiga...

AI 聚合 09/03
arXiv

Discriminative World Models for Web Agents

Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resul...

AI 聚合 09/03
HuggingFace

Cliff: Learning Process Rewards from the First Mistake

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LL...

AI 聚合 09/03
HuggingFace

Kirin: Animal Motion Generation from In-the-Wild Video

Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area la...

AI 聚合 09/03
HuggingFace

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Competitive programming has become a key test of large language model reasoning, with international competitions such...

AI 聚合 09/03
HuggingFace

On the Design Fundamentals of Pixel Text Representation Learning

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet ex...

AI 聚合 09/03
HuggingFace

Language Models Can Control Their Own Attention

Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...

AI 聚合 09/03
HuggingFace

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single ...

AI 聚合 09/03
HuggingFace

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...

AI 聚合 09/03
HuggingFace

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...

AI 聚合 09/03
HuggingFace

Aspire: Can Models Self-Evolve from Vague Goals?

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at re...

AI 聚合 09/03
HuggingFace

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external ex...

AI 聚合 09/03
HuggingFace

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts m...

AI 聚合 09/03
HuggingFace

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe com...

AI 聚合 09/03
HuggingFace

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language ...

AI 聚合 09/03
HuggingFace

VibeVoice-ASR-Streaming Technical Report

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-t...

AI 聚合 09/03
HuggingFace

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-...

AI 聚合 09/03
HuggingFace

ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize...

AI 聚合 09/03
HuggingFace

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable ima...

AI 聚合 09/03
HuggingFace

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involv...

AI 聚合 09/03
HuggingFace

Exploring Collaboration between a language and a non-language agent

LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through ...

AI 聚合 09/03
HuggingFace

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...

AI 聚合 09/03
HuggingFace

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decom...

AI 聚合 09/03
HuggingFace

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step ...

AI 聚合 09/03
HuggingFace

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video...

AI 聚合 09/03
HuggingFace

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You

Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption ...

AI 聚合 09/03
HuggingFace

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...

AI 聚合 09/03
HuggingFace

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-...

AI 聚合 09/03
HuggingFace

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger ...

AI 聚合 09/03
HuggingFace

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, o...

AI 聚合 09/03
HuggingFace

Agent Memory Is a Surface for Endogenous Authorization Laundering

Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrict...

AI 聚合 09/03
HuggingFace

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated Gi...

AI 聚合 09/03
arXiv

Can LLMs Design Video Coding Tools? A Case Study on Planar Mode

This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due...

AI 聚合 09/02
arXiv

A Mathematical Theory of Reusable Neural Bases for Network Compression

As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critic...

AI 聚合 09/02
arXiv

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hy...

AI 聚合 09/02
arXiv

BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net

Automated lesion segmentation in whole-body PET/CT is complicated by the variety of physiological tracer uptake patte...

AI 聚合 09/02
arXiv

Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories

We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that sha...

AI 聚合 09/02
arXiv

H3-World: Turning Language Understanding into World Control

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world m...

AI 聚合 09/02
arXiv

From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification

Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the ...

AI 聚合 09/02
arXiv

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as polic...

AI 聚合 09/02
arXiv

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during L...

AI 聚合 09/02
arXiv

Designing Proactive Thought Partners for Writing

Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals ...

AI 聚合 09/02
arXiv

Mechanism Design for Alignment and Control

We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible a...

AI 聚合 09/02
arXiv

The Rise of Verbal Reinforcement Learning

Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent...

AI 聚合 09/02
arXiv

CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?

Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility bri...

AI 聚合 09/02
arXiv

Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation

The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining...

AI 聚合 09/02
arXiv

Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation

Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code...

AI 聚合 09/02
HuggingFace

The Mechanics of Democratic Dominance: A System Dynamics Paradigm for Dynamic Consent Engineering

Traditional frameworks of political communication operate under linear, event-driven assumptions that treat voter per...

AI 聚合 09/02
HuggingFace

InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between ...

AI 聚合 09/02
HuggingFace

ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation

SAR-to-EO image translation aims to generate electro-optical (EO) imagery from synthetic aperture radar (SAR) observa...

AI 聚合 09/02
HuggingFace

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve ...

AI 聚合 09/02
HuggingFace

Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering

A central bottleneck in multi-hop Question Answering (QA) is that the granularity at which a question is expressed of...

AI 聚合 09/02
HuggingFace

Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System

While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, fun...

AI 聚合 09/02
HuggingFace

Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled ro...

AI 聚合 09/02
HuggingFace

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends...

AI 聚合 09/02
HuggingFace

StudentSim: Training LLM-based Student Simulators

AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but eviden...

AI 聚合 09/02
HuggingFace

H3-World: Turning Language Understanding into World Control

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world m...

AI 聚合 09/02
HuggingFace

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environ...

AI 聚合 09/02
HuggingFace

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, ...

AI 聚合 09/02
HuggingFace

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Dr...

AI 聚合 09/02
HuggingFace

UI-Venus-2 Technical Report

Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchm...

AI 聚合 09/02
HuggingFace

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirement...

AI 聚合 09/02
HuggingFace

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at f...

AI 聚合 09/02
HuggingFace

Safin-1: Safety from Within through Memory-Native State Evolution

Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt o...

AI 聚合 09/02
HuggingFace

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can pl...

AI 聚合 09/02
HuggingFace

Agents in the Large: Perception-Centered Architecture for Persistent Agents

Cognitive language agents have achieved substantial progress by equipping language models with memory, tools, and dec...

AI 聚合 09/02
HuggingFace

Recursive Criticality of AI Self-Improvement

AI is increasingly used in the R\&D process that produces future AI systems. We study the conditions under which this...

AI 聚合 09/02
HuggingFace

Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions

End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth ...

AI 聚合 09/02
HuggingFace

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an exe...

AI 聚合 09/02
HuggingFace

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model...

AI 聚合 09/02
HuggingFace

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we ...

AI 聚合 09/02
HuggingFace

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-ca...

AI 聚合 09/02
HuggingFace

MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation

Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and count...

AI 聚合 09/02
HuggingFace

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose...

AI 聚合 09/02
HuggingFace

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model mu...

AI 聚合 09/02
arXiv

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an e...

AI 聚合 09/01
arXiv

Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents

Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated sta...

AI 聚合 09/01
arXiv

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, t...

AI 聚合 09/01
arXiv

Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring

We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n...

AI 聚合 09/01
arXiv

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR...

AI 聚合 09/01
arXiv

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literatu...

AI 聚合 09/01
arXiv

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-lev...

AI 聚合 09/01
arXiv

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...

AI 聚合 09/01
arXiv

Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations

Accurate daily predictions of cold hardiness in woody plants are critical in regions where freezing temperatures can ...

AI 聚合 09/01
arXiv

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

Industrial post-training is a brownfield regime. Teams inherit a deployed checkpoint and must land targeted improveme...

AI 聚合 09/01
arXiv

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deploymen...

AI 聚合 09/01
arXiv

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently character...

AI 聚合 09/01
arXiv

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural ali...

AI 聚合 09/01
arXiv

Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platf...

AI 聚合 09/01
arXiv

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies

Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: co...

AI 聚合 09/01
HuggingFace

Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation

Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop b...

AI 聚合 09/01
HuggingFace

Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling

Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observe...

AI 聚合 09/01
HuggingFace

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B paramet...

AI 聚合 09/01
HuggingFace

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al.,...

AI 聚合 09/01
HuggingFace

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interact...

AI 聚合 09/01
HuggingFace

Normalized Low-Rank Adaptation

While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its train...

AI 聚合 09/01
HuggingFace

Dynamic Important Example Mining for Reinforcement Finetuning

Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its e...

AI 聚合 09/01
HuggingFace

SHAPE of Chain-of-Thought in Math Reasoning

Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically ...

AI 聚合 09/01
HuggingFace

Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching

Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is sco...

AI 聚合 09/01
HuggingFace

Verification-Aware Training for Speculative Decoding

Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, ...

AI 聚合 09/01
HuggingFace

Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions

Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction ...

AI 聚合 09/01
HuggingFace

DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabiliti...

AI 聚合 09/01
HuggingFace

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating i...

AI 聚合 09/01
HuggingFace

WebWorld: The Browser as a World Model for Self-Improving Web Code

VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that j...

AI 聚合 09/01
HuggingFace

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across task...

AI 聚合 09/01
HuggingFace

Evaluating the Hidden Costs of Personalization in Large Language Models

While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, the...

AI 聚合 09/01
HuggingFace

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant be...

AI 聚合 09/01
HuggingFace

Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase

Organizations often develop and maintain portfolios of related applications: independently deployable codebases that ...

AI 聚合 09/01
HuggingFace

Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended ...

AI 聚合 09/01
HuggingFace

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual...

AI 聚合 09/01
HuggingFace

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing e...

AI 聚合 09/01
HuggingFace

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are sti...

AI 聚合 09/01
HuggingFace

Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned com...

AI 聚合 09/01
HuggingFace

Sliding-window beats linear attention

Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new ...

AI 聚合 09/01
HuggingFace

Generative Semantic Scene Completion

Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the ta...

AI 聚合 09/01
HuggingFace

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. S...

AI 聚合 09/01
HuggingFace

Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents

Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state ...

AI 聚合 09/01
HuggingFace

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what th...

AI 聚合 09/01
arXiv

How Proper Scoring Rules Shape LLM Forecasting

This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare fi...

AI 聚合 08/31
arXiv

LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment

Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypo...

AI 聚合 08/31
arXiv

AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet mon...

AI 聚合 08/31
arXiv

On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces

AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly ext...

AI 聚合 08/31
arXiv

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration

When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine...

AI 聚合 08/31
arXiv

Conformal Uncertainty Quantification Guarantees for Neural Operators

Neural operators provide fast surrogate models for approximating operators between function spaces, but their predict...

AI 聚合 08/31
arXiv

When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI

We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodi...

AI 聚合 08/31
arXiv

Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks

Texture image classification plays a significant role in computer vision applications, including industrial inspectio...

AI 聚合 08/31
arXiv

InstructMesh: Selective Refinement of Generative 3D Models for Fabrication

Recent advances in generative AI allow users to create 3D models from text or images. However, these models prioritiz...

AI 聚合 08/31
arXiv

An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models

A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wr...

AI 聚合 08/31
arXiv

Video Generative Models as Geometry Learner

Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as ima...

AI 聚合 08/31
arXiv

Logos: An Agent Harness on a Cross-Process Bus

Modern agent systems assemble capabilities at runtime, and this dynamic composition has recently received a complete ...

AI 聚合 08/31
arXiv

Blog: Survey of Optimizers

Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design...

AI 聚合 08/31
arXiv

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as...

AI 聚合 08/31
arXiv

Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning

Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capabil...

AI 聚合 08/31
HuggingFace

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and...

AI 聚合 08/31
HuggingFace

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are...

AI 聚合 08/31
HuggingFace

Training, learning and inference: unified dynamics of neural systems

We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete oc...

AI 聚合 08/31
HuggingFace

LMSM: LLM Security Framework Inspired by Linux Security Modules

Large language models (LLMs) are increasingly deployed with layered defenses, yet malicious prompts can still bypass ...

AI 聚合 08/31
HuggingFace

Video Generative Models as Geometry Learner

Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as ima...

AI 聚合 08/31
HuggingFace

Rubric-to-Code Credit Assignment for Reinforcement Learning

Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from ...

AI 聚合 08/31
HuggingFace

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces...

AI 聚合 08/31
HuggingFace

Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online un...

AI 聚合 08/31
HuggingFace

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage ...

AI 聚合 08/31
HuggingFace

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fix...

AI 聚合 08/31
HuggingFace

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a ...

AI 聚合 08/31
HuggingFace

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of th...

AI 聚合 08/31
HuggingFace

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent conte...

AI 聚合 08/31
HuggingFace

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory...

AI 聚合 08/31
HuggingFace

Fast Weight Attention for Continual Learning

Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recur...

AI 聚合 08/31
HuggingFace

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the targ...

AI 聚合 08/31
HuggingFace

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous...

AI 聚合 08/31
HuggingFace

Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities

Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of ...

AI 聚合 08/31
HuggingFace

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often le...

AI 聚合 08/31
HuggingFace

GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demog...

AI 聚合 08/31
HuggingFace

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis o...

AI 聚合 08/29
HuggingFace

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily lice...

AI 聚合 08/29
HuggingFace

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon....

AI 聚合 08/29
HuggingFace

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Recent advances in inference-time scaling have significantly improved the reasoning performance of large language mod...

AI 聚合 08/29
HuggingFace

Luce: Relightable Gaussians for 3D Asset Generation

High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To supp...

AI 聚合 08/29
GitHub

[GitHub] koala73/worldmonitor

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tra...

AI 聚合 08/28
arXiv

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer en...

AI 聚合 08/28
arXiv

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is...

AI 聚合 08/28
arXiv

Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions

Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint perfo...

AI 聚合 08/28
arXiv

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

Video carries the temporal structure of the physical world, yet learning representations from it has remained computa...

AI 聚合 08/28
arXiv

Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction

Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploi...

AI 聚合 08/28
arXiv

How Language Models Organize and Structure Moral Knowledge

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a...

AI 聚合 08/28
arXiv

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing th...

AI 聚合 08/28
arXiv

Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study

Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsel...

AI 聚合 08/28
arXiv

Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners

Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifac...

AI 聚合 08/28
arXiv

Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentati...

AI 聚合 08/28
arXiv

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model th...

AI 聚合 08/28
arXiv

RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution

LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful...

AI 聚合 08/28
arXiv

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

In real-world software development, code review typically involves iterative interactions between developers and revi...

AI 聚合 08/28
arXiv

SWE-Prime: Fewer Trajectories, Better Performance

To improve large language models' ability to resolve real-world software issues, prior work has focused on constructi...

AI 聚合 08/28
arXiv

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. R...

AI 聚合 08/28
HuggingFace

What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents

LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agent...

AI 聚合 08/28
HuggingFace

Magpie: Real-Time World Renderer for Interactive Games

Modern game development relies heavily on conventional graphics pipelines. High-quality visual content requires model...

AI 聚合 08/28
HuggingFace

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. Howev...

AI 聚合 08/28
HuggingFace

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory s...

AI 聚合 08/28
HuggingFace

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. R...

AI 聚合 08/28
HuggingFace

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remai...

AI 聚合 08/28
HuggingFace

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local ...

AI 聚合 08/28
HuggingFace

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more th...

AI 聚合 08/28
HuggingFace

CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension

Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal ...

AI 聚合 08/28
HuggingFace

Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report

AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies i...

AI 聚合 08/28
HuggingFace

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they...

AI 聚合 08/28
HuggingFace

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and upda...

AI 聚合 08/28
HuggingFace

TTPO: Test-Time Policy Optimization

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), h...

AI 聚合 08/28
HuggingFace

GameWAM: A World Action Model for Video Games

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous n...

AI 聚合 08/28
HuggingFace

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or shor...

AI 聚合 08/28
HuggingFace

Procedura: Agentic 3D Modeling with Procedural Control

Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where...

AI 聚合 08/28
HuggingFace

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this st...

AI 聚合 08/28
HuggingFace

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvem...

AI 聚合 08/28
HuggingFace

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducibl...

AI 聚合 08/28
HuggingFace

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the ...

AI 聚合 08/28
HuggingFace

Prefix Sliding for efficient test-time scaling

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer ...

AI 聚合 08/28
HuggingFace

Skill Issue: Are Skills Language-Invariant in LLMs?

Large language models access knowledge inconsistently across languages, but to what extent do they differ in their sk...

AI 聚合 08/28
HuggingFace

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing be...

AI 聚合 08/28
arXiv

DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without a...

AI 聚合 08/27
arXiv

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-corr...

AI 聚合 08/27
arXiv

The Value of Human Expertise

We consider optimization applications with unknown parameters where the decision maker believes that the optimal valu...

AI 聚合 08/27
arXiv

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a ta...

AI 聚合 08/27
arXiv

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requir...

AI 聚合 08/27
arXiv

Prefix Sliding for efficient test-time scaling

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer ...

AI 聚合 08/27
arXiv

Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving

Intent misinterpretation during vehicle interactions causes recurring planning failures. We study a decision layer in...

AI 聚合 08/27
arXiv

SwarmWorld: Stigmergic technological evolution in societies of language-model agents

Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions t...

AI 聚合 08/27
arXiv

ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning...

AI 聚合 08/27
arXiv

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning d...

AI 聚合 08/27
arXiv

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic v...

AI 聚合 08/27
arXiv

Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying...

AI 聚合 08/27
arXiv

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, o...

AI 聚合 08/27
arXiv

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed...

AI 聚合 08/27
arXiv

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and vi...

AI 聚合 08/27
HuggingFace

VGI-BENCH: Probing Visual Intelligence in Video Generation Models

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through g...

AI 聚合 08/27
HuggingFace

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning st...

AI 聚合 08/27
HuggingFace

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single u...

AI 聚合 08/27
HuggingFace

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...

AI 聚合 08/27
HuggingFace

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art model...

AI 聚合 08/27
HuggingFace

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) m...

AI 聚合 08/27
HuggingFace

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, cha...

AI 聚合 08/27
HuggingFace

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While uni...

AI 聚合 08/27
HuggingFace

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empatheti...

AI 聚合 08/27
HuggingFace

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objecti...

AI 聚合 08/27
HuggingFace

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution ...

AI 聚合 08/27
HuggingFace

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evalu...

AI 聚合 08/27
HuggingFace

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability...

AI 聚合 08/27
HuggingFace

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring m...

AI 聚合 08/27
HuggingFace

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unf...

AI 聚合 08/27
HuggingFace

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbo...

AI 聚合 08/27
HuggingFace

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchr...

AI 聚合 08/27
HuggingFace

RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval

Document retrieval increasingly supports high-stakes information access in finance, healthcare, and law. Modern retri...

AI 聚合 08/27
HuggingFace

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to su...

AI 聚合 08/27
HuggingFace

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and la...

AI 聚合 08/27
HuggingFace

DREAM Technical Report

Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient...

AI 聚合 08/27
HuggingFace

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recog...

AI 聚合 08/27
HuggingFace

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating str...

AI 聚合 08/27
HuggingFace

Automata from Agent Traces: Failure and Next-Step Prediction

LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces re...

AI 聚合 08/27
HuggingFace

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel ...

AI 聚合 08/27
HuggingFace

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files,...

AI 聚合 08/27
HuggingFace

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents f...

AI 聚合 08/27
HuggingFace

When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows

Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream sta...

AI 聚合 08/27
HuggingFace

MARS: Multi-Specialist LLM Relay System for Competitive Programming

Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existi...

AI 聚合 08/27
arXiv

Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks

The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tas...

AI 聚合 08/26
arXiv

Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy

Brain stroke, known for its high mortality and incidence rates, poses significant health risks and requires rapid int...

AI 聚合 08/26
arXiv

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces...

AI 聚合 08/26
arXiv

Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought

Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain pla...

AI 聚合 08/26
arXiv

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize inter...

AI 聚合 08/26
arXiv

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fix...

AI 聚合 08/26
arXiv

Automatic Model Card Generation Using an LLM

Model cards are structured documents that summarize key information about machine learning models to improve transpar...

AI 聚合 08/26
arXiv

Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core

Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existing approaches sti...

AI 聚合 08/26
arXiv

Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA

Large language models are increasingly used for knowledge graph question answering (KGQA), but can fail to correctly ...

AI 聚合 08/26
arXiv

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified...

AI 聚合 08/26
arXiv

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI...

AI 聚合 08/26
arXiv

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific...

AI 聚合 08/26
arXiv

FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs

Real-world data for knowledge graph question answering is often distributed across different organizations due to gov...

AI 聚合 08/26
arXiv

SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL

Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and var...

AI 聚合 08/26
arXiv

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state a...

AI 聚合 08/26
HuggingFace

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the...

AI 聚合 08/26
HuggingFace

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize inter...

AI 聚合 08/26
HuggingFace

Meta^n: Recursive Self-Improvement through Emergent Depth

Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level ...

AI 聚合 08/26
HuggingFace

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state a...

AI 聚合 08/26
HuggingFace

On-Policy Self-Distillation in Diffusion Models

Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint r...

AI 聚合 08/26
HuggingFace

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interaction...

AI 聚合 08/26
HuggingFace

On-policy Distillation with Verifiable Reward

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted...

AI 聚合 08/26
HuggingFace

Best Practice Critic Optimization

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...

AI 聚合 08/26
HuggingFace

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithf...

AI 聚合 08/26
HuggingFace

TorchMorph: CUDA-accelerated Morphological Transforms

Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementa...

AI 聚合 08/26
HuggingFace

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to...

AI 聚合 08/26
HuggingFace

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific...

AI 聚合 08/26
HuggingFace

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect...

AI 聚合 08/26
HuggingFace

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, pos...

AI 聚合 08/26
HuggingFace

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

Video games provide a scalable source of training data for video world models, offering diverse environments, complex...

AI 聚合 08/26
HuggingFace

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild

As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has draw...

AI 聚合 08/26
HuggingFace

Latent Action as Intention Enables Efficient Future Imagination for World Action Models

World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observati...

AI 聚合 08/26
HuggingFace

Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection

Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the...

AI 聚合 08/26
HuggingFace

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which se...

AI 聚合 08/26
HuggingFace

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a f...

AI 聚合 08/26
HuggingFace

AutoResearch: Insight In, Hallucination Out

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does ...

AI 聚合 08/26
HuggingFace

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

Large language models are increasingly expected to execute complex workflows whose success depends on maintaining int...

AI 聚合 08/26
HuggingFace

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bot...

AI 聚合 08/26
HuggingFace

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-spe...

AI 聚合 08/26
HuggingFace

What AstroPT knows about galaxies, and what that can teach us about LLMs

Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover re...

AI 聚合 08/26
HuggingFace

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

We formalize prefix invariance: representations at position t must not depend on future inputs. We give a lightweight...

AI 聚合 08/26
arXiv

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathemat...

AI 聚合 08/25
arXiv

When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World

Historical people may appear under different languages, scripts, and transcription traditions, while distinct individ...

AI 聚合 08/25
arXiv

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as t...

AI 聚合 08/25
arXiv

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing...

AI 聚合 08/25
arXiv

Correcting a learned physical invariant improves world-model rollouts

World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen Dreame...

AI 聚合 08/25
arXiv

Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition

The continual evolution of malware variants necessitates detection systems that can adapt to new threats without retr...

AI 聚合 08/25
arXiv

The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams

Does multi-agent LLM interaction help or hurt? Some work reports gains from debate (Du et al., 2024), critique loops ...

AI 聚合 08/25
arXiv

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

While AI assistance can improve human task performance in the short term, it may also undermine the development of sk...

AI 聚合 08/25
arXiv

ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings

Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive wi...

AI 聚合 08/25
arXiv

Prime Agent: A Self-Improving RLM Harness

Language models are sequential processors, but long-horizon agency requires external information and computation beyo...

AI 聚合 08/25
arXiv

Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography

Ballistocardiography (BCG) is promising for unobtrusive long-term blood pressure (BP) monitoring in laboratory settin...

AI 聚合 08/25
arXiv

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety audit...

AI 聚合 08/25
arXiv

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and la...

AI 聚合 08/25
arXiv

ReWorld: An Interactive World Model with Long-Horizon Memory

An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time....

AI 聚合 08/25
arXiv

How to Train a Critic Stably and Efficiently

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...

AI 聚合 08/25
HuggingFace

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing stude...

AI 聚合 08/25
HuggingFace

Prime Agent: A Self-Improving RLM Harness

Language models are sequential processors, but long-horizon agency requires external information and computation beyo...

AI 聚合 08/25
HuggingFace

ReWorld: An Interactive World Model with Long-Horizon Memory

An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time....

AI 聚合 08/25
HuggingFace

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. Howev...

AI 聚合 08/25
HuggingFace

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requ...

AI 聚合 08/25
HuggingFace

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarificatio...

AI 聚合 08/25
HuggingFace

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness...

AI 聚合 08/25
HuggingFace

Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module,...

AI 聚合 08/25
HuggingFace

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this ...

AI 聚合 08/25
HuggingFace

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains cha...

AI 聚合 08/25
HuggingFace

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...

AI 聚合 08/25
HuggingFace

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation...

AI 聚合 08/25
HuggingFace

EchoWM: Open and Enterable Omnimodal World Models

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation whi...

AI 聚合 08/25
HuggingFace

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an...

AI 聚合 08/25
HuggingFace

RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundat...

AI 聚合 08/25
HuggingFace

From Generation to Simulation: How Far Are World Models from Being True Simulators?

With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingl...

AI 聚合 08/25
HuggingFace

Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports

Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, ...

AI 聚合 08/25
HuggingFace

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavio...

AI 聚合 08/25
HuggingFace

Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning

Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devi...

AI 聚合 08/25
HuggingFace

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organ...

AI 聚合 08/25
HuggingFace

Human-Centric Intelligence in the Era of Foundation Models: A Survey

Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, ...

AI 聚合 08/25
HuggingFace

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle e...

AI 聚合 08/25
HuggingFace

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution path...

AI 聚合 08/25
HuggingFace

Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not...

AI 聚合 08/25
HuggingFace

Hydra-0: Action Flow for Generalist World Modeling and Control

We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel mo...

AI 聚合 08/25
HuggingFace

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We...

AI 聚合 08/25
HuggingFace

PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration

We present PhysCaP, a Physics-Informed Code-as-Policy agent for active perception in robotic manipulation. While visi...

AI 聚合 08/25
HuggingFace

WorldMind: Decoupled Game World Model for State-Aware NPC Behavior

Game world models have recently demonstrated promising capabilities in generating visually coherent and action-contro...

AI 聚合 08/25
arXiv

Ontology-supported AI Model and Dataset Management

Recently, there has been a great deal of research into improving AI methods and their application. The main focus is ...

AI 聚合 08/24
arXiv

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future ...

AI 聚合 08/24
arXiv

Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems

The Traveling Salesman Problem (TSP) is one of the most extensively studied NP-hard optimization problems. Genetic Al...

AI 聚合 08/24
arXiv

Adapting Knowledge Graphs for Behavior Denoising in Sequential Recommendation

Sequential recommendation predicts the next item from a user's interaction history, but not every interaction is equa...

AI 聚合 08/24
arXiv

EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering

Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multipl...

AI 聚合 08/24
arXiv

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safet...

AI 聚合 08/24
arXiv

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support...

AI 聚合 08/24
arXiv

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encour...

AI 聚合 08/24
arXiv

From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry

The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy reg...

AI 聚合 08/24
arXiv

Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets

We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minim...

AI 聚合 08/24
arXiv

Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation

Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly ...

AI 聚合 08/24
arXiv

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases...

AI 聚合 08/24
arXiv

AI with Authority, from Application to Silicon

For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here...

AI 聚合 08/24
arXiv

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy image...

AI 聚合 08/24
arXiv

Primal Acceleration of Newton's Method

We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. ...

AI 聚合 08/24
HuggingFace

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute ...

AI 聚合 08/24
HuggingFace

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative think...

AI 聚合 08/24
HuggingFace

InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter

With large pretrained models, existing methods have effectively improved instruction-based video editing. However, mo...

AI 聚合 08/24
HuggingFace

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safet...

AI 聚合 08/24
HuggingFace

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolutio...

AI 聚合 08/24
HuggingFace

Towards Faithful Simulation of Human Shopping Behavior

Simulating realistic user shopping behavior underpins offline evaluation and reinforcement learning in e-commerce sce...

AI 聚合 08/24
HuggingFace

Hadith computational science in the age of large language models: a critical narrative review

We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, an...

AI 聚合 08/24
HuggingFace

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

Agents learn to act through interaction with environments, yet the environments used for training are often manually ...

AI 聚合 08/24
HuggingFace

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continu...

AI 聚合 08/24
HuggingFace

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computa...

AI 聚合 08/24
HuggingFace

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's ow...

AI 聚合 08/24
HuggingFace

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to ...

AI 聚合 08/24
HuggingFace

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite:...

AI 聚合 08/24
HuggingFace

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditionin...

AI 聚合 08/24
HuggingFace

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tun...

AI 聚合 08/22
HuggingFace

CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be gras...

AI 聚合 08/22
HuggingFace

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to general...

AI 聚合 08/22
HuggingFace

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffol...

AI 聚合 08/22
HuggingFace

QuoteBench: How Matched Scores Can Hide Command-Path Failures

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched ...

AI 聚合 08/22
HuggingFace

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coheren...

AI 聚合 08/22
HuggingFace

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-a...

AI 聚合 08/22
HuggingFace

TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity

We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 param...

AI 聚合 08/22
HuggingFace

FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills

Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures di...

AI 聚合 08/22
arXiv

Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation

Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-s...

AI 聚合 08/21
arXiv

Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI

The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses ...

AI 聚合 08/21
arXiv

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to pe...

AI 聚合 08/21
arXiv

QUASAR: A Quantum-Classical Neural Network for SAR Satellite Physical-Layer Authentication

X-band SAR satellites (8-12 GHz) play a critical role in disaster response, environmental monitoring, and military in...

AI 聚合 08/21
arXiv

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather tha...

AI 聚合 08/21
arXiv

Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

The rapid proliferation of memecoins on blockchain platforms has increased the risk of fraudulent activities, particu...

AI 聚合 08/21
arXiv

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable wi...

AI 聚合 08/21
arXiv

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Large language models often fail to answer questions about a bounded document collection when the source documents ar...

AI 聚合 08/21
arXiv

Phantom Gains: Auditing Self-Improvement Against a Measured Null

Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual prob...

AI 聚合 08/21
arXiv

MidTool: Mid-training Data Synthesis for Agentic Tool Use

Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Re...

AI 聚合 08/21
arXiv

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improv...

AI 聚合 08/21
arXiv

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that ...

AI 聚合 08/21
arXiv

Inducing Task Models from Computer-Use Traces

Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resour...

AI 聚合 08/21
arXiv

An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are...

AI 聚合 08/21
arXiv

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressi...

AI 聚合 08/21
HuggingFace

PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures ari...

AI 聚合 08/21
HuggingFace

EXIMO: VLM Guided Exploration of VLA Policies

How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation polic...

AI 聚合 08/21
HuggingFace

EnvHarness: Awakening Static Worlds for Agent Learning

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agen...

AI 聚合 08/21
HuggingFace

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people....

AI 聚合 08/21
HuggingFace

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native con...

AI 聚合 08/21
HuggingFace

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Memory has become a key component of large language models, enabling them to retain information and learn from long-t...

AI 聚合 08/21
HuggingFace

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reco...

AI 聚合 08/21
HuggingFace

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

Large language models often fail to answer questions about a bounded document collection when the source documents ar...

AI 聚合 08/21
HuggingFace

Repo0: Design-Driven Zero-to-All Code Generation

Large language model agents have made substantial progress in code generation, yet most existing systems assume a pre...

AI 聚合 08/21
HuggingFace

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no ...

AI 聚合 08/21
HuggingFace

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention re...

AI 聚合 08/21
HuggingFace

Chain-of-Experience for Continual LLM Improvement

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the mod...

AI 聚合 08/21
HuggingFace

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capab...

AI 聚合 08/21
HuggingFace

Towards Quantifying Benchmark Optimization in ASR Models

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature...

AI 聚合 08/21
HuggingFace

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing meth...

AI 聚合 08/21
HuggingFace

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evol...

AI 聚合 08/21
HuggingFace

Bounded Agents: Delegation Security for Multi-Agent AI Systems

LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start...

AI 聚合 08/21
HuggingFace

LLMs Get Smarter from Targeted Synthetic Multilingual Data

Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the ...

AI 聚合 08/21
HuggingFace

SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection

Object detectors often produce over-confident predictions for objects outside their training categories, leading to s...

AI 聚合 08/21
HuggingFace

Towards Real-Time and Adaptable LiDAR Scene Completion

LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed i...

AI 聚合 08/21
HuggingFace

Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems

Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game develo...

AI 聚合 08/21
HuggingFace

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing met...

AI 聚合 08/21
arXiv

DA-WAM: Decision-Aligned Future Latents for Driving World Models

Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of...

AI 聚合 08/20
arXiv

Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift

Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause tar...

AI 聚合 08/20
arXiv

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...

AI 聚合 08/20
arXiv

Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the com...

AI 聚合 08/20
arXiv

PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints

Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug di...

AI 聚合 08/20
arXiv

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class r...

AI 聚合 08/20
arXiv

Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles

A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than...

AI 聚合 08/20
arXiv

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output c...

AI 聚合 08/20
arXiv

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable ...

AI 聚合 08/20
arXiv

Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication

Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, crea...

AI 聚合 08/20
arXiv

Interpretable AI predicts a 2026 summer dry anomaly in central China

Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict wit...

AI 聚合 08/20
arXiv

Finetuning Strategies for Querying Sounds by Vocal Imitation

This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by v...

AI 聚合 08/20
arXiv

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger t...

AI 聚合 08/20
arXiv

ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning

We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for...

AI 聚合 08/20
arXiv

SPADE: Self-Play in Adaptive Synthetic Executable Environments

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language ...

AI 聚合 08/20
HuggingFace

SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution

Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resoluti...

AI 聚合 08/20
HuggingFace

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over lon...

AI 聚合 08/20
HuggingFace

Looped Language Models Improve Compositional Tool Calling

Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use...

AI 聚合 08/20
HuggingFace

SPADE: Self-Play in Adaptive Synthetic Executable Environments

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language ...

AI 聚合 08/20
HuggingFace

SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation

Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independe...

AI 聚合 08/20
HuggingFace

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-...

AI 聚合 08/20
HuggingFace

Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification

Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented...

AI 聚合 08/20
HuggingFace

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formula...

AI 聚合 08/20
HuggingFace

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success ...

AI 聚合 08/20
HuggingFace

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (...

AI 聚合 08/20
HuggingFace

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not...

AI 聚合 08/20
HuggingFace

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows,...

AI 聚合 08/20
HuggingFace

Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion

High-quality creative writing data for large language models (LLMs) remains dominated by story-centric data, limiting...

AI 聚合 08/20
HuggingFace

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-languag...

AI 聚合 08/20
HuggingFace

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the ge...

AI 聚合 08/20
HuggingFace

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM...

AI 聚合 08/20
HuggingFace

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, whi...

AI 聚合 08/20
HuggingFace

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the...

AI 聚合 08/20
HuggingFace

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-speci...

AI 聚合 08/20
HuggingFace

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integrati...

AI 聚合 08/20
HuggingFace

The Problem Is the Problem: Towards Scalable Mathematical Discovery

AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model re...

AI 聚合 08/20
HuggingFace

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remai...

AI 聚合 08/20
arXiv

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and ...

AI 聚合 08/19
arXiv

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement b...

AI 聚合 08/19
arXiv

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although r...

AI 聚合 08/19
arXiv

Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity

Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotation...

AI 聚合 08/19
arXiv

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand indi...

AI 聚合 08/19
arXiv

Traceable Trust for action-ready artificial intelligence in bioscience

Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict...

AI 聚合 08/19
arXiv

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of L...

AI 聚合 08/19
arXiv

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, c...

AI 聚合 08/19
arXiv

Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete to...

AI 聚合 08/19
arXiv

Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors

MRI reconstruction methods for undersampled k-space data naturally utilize complex-valued measurements. Parallel deve...

AI 聚合 08/19
arXiv

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code rep...

AI 聚合 08/19
arXiv

HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance

Urban traffic congestion reduces productivity and increases travel cost and emissions. Network-wide live travel-time ...

AI 聚合 08/19
arXiv

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet thei...

AI 聚合 08/19
arXiv

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintain...

AI 聚合 08/19
arXiv

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet c...

AI 聚合 08/19
HuggingFace

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliv...

AI 聚合 08/19
HuggingFace

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) ...

AI 聚合 08/19
HuggingFace

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet c...

AI 聚合 08/19
HuggingFace

Energy-Guided Flow Matching

Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and...

AI 聚合 08/19
HuggingFace

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance o...

AI 聚合 08/19
HuggingFace

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment...

AI 聚合 08/19
HuggingFace

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastruct...

AI 聚合 08/19
HuggingFace

Personalized Auto-Research: Towards a True AI Co-Scientist

AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full pa...

AI 聚合 08/19
HuggingFace

Unifying Graph Neural Networks Through a Common Layer Equation

Graph neural networks are commonly described through family-specific equations whose notation obscures shared computa...

AI 聚合 08/19
HuggingFace

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent sta...

AI 聚合 08/19
HuggingFace

Agent Lightning v1.0: Towards Harnessed Agentic RL

Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a criti...

AI 聚合 08/19
HuggingFace

MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable ...

AI 聚合 08/19
HuggingFace

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets m...

AI 聚合 08/19
HuggingFace

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models...

AI 聚合 08/19
HuggingFace

Cross-Model Memory Transfer via Target-Side Reader Adaptation

Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieva...

AI 聚合 08/19
HuggingFace

Demystifying Agent Skills: Why They Work-Until They Don't

Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structur...

AI 聚合 08/19
HuggingFace

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves p...

AI 聚合 08/19
HuggingFace

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different...

AI 聚合 08/19
HuggingFace

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-S...

AI 聚合 08/19
HuggingFace

V-RAE: Rethinking Video Latent Spaces for Generation

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although...

AI 聚合 08/19
HuggingFace

Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI

This first release of Prior Labs in relational learning shows our continued commitment to open science. We open-sourc...

AI 聚合 08/19
HuggingFace

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whet...

AI 聚合 08/19
HuggingFace

Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays

Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fie...

AI 聚合 08/19
HuggingFace

StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Exist...

AI 聚合 08/19
HuggingFace

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track o...

AI 聚合 08/19
HuggingFace

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, tempo...

AI 聚合 08/19
HuggingFace

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with ...

AI 聚合 08/19
HuggingFace

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome r...

AI 聚合 08/19
arXiv

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding

We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report...

AI 聚合 08/18
arXiv

Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment

Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition reso...

AI 聚合 08/18
arXiv

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to t...

AI 聚合 08/18
arXiv

When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step ...

AI 聚合 08/18
arXiv

Quipu: A Governed Bitemporal Knowledge Graph Store

Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: acce...

AI 聚合 08/18
arXiv

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but exis...

AI 聚合 08/18
arXiv

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Generative pretraining established reusable task representations; later work on language-based task conditioning and ...

AI 聚合 08/18
arXiv

Model Hypnosis: Strong control of AI via additive subliminal effects

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually w...

AI 聚合 08/18
arXiv

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-la...

AI 聚合 08/18
arXiv

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on m...

AI 聚合 08/18
arXiv

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control...

AI 聚合 08/18
arXiv

Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

A language model's output does not by itself provide verifiable evidence about the internal computation that produced...

AI 聚合 08/18
arXiv

AutoSR: Automatic Symbolic Regression by Searching Research States

We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbol...

AI 聚合 08/18
arXiv

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve

The current best bounds on the matrix multiplication exponent $ω$ are obtained through a refinement of the laser meth...

AI 聚合 08/18
arXiv

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VL...

AI 聚合 08/18
HuggingFace

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs)...

AI 聚合 08/18
HuggingFace

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize...

AI 聚合 08/18
HuggingFace

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve

The current best bounds on the matrix multiplication exponent ω are obtained through a refinement of the laser method...

AI 聚合 08/18
HuggingFace

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Althoug...

AI 聚合 08/18
HuggingFace

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justi...

AI 聚合 08/18
HuggingFace

Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form

Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form w...

AI 聚合 08/18
HuggingFace

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as th...

AI 聚合 08/18
HuggingFace

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any...

AI 聚合 08/18
HuggingFace

Drive, Pack, Fly: The Travelling Thief Problem with Drone

In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on rout...

AI 聚合 08/18
HuggingFace

HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation

Boundary representation (B-Rep) generation is a fundamental task in computer-aided design, yet the direct synthesis o...

AI 聚合 08/18
HuggingFace

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This...

AI 聚合 08/18
HuggingFace

Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training languag...

AI 聚合 08/18
HuggingFace

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strate...

AI 聚合 08/18
HuggingFace

WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations

Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene ...

AI 聚合 08/18
HuggingFace

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain...

AI 聚合 08/18
HuggingFace

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended...

AI 聚合 08/18
HuggingFace

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landsc...

AI 聚合 08/18
HuggingFace

A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models

Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resultin...

AI 聚合 08/18
HuggingFace

MOSS-VL Technical Report

We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it spe...

AI 聚合 08/18
HuggingFace

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interfac...

AI 聚合 08/18
HuggingFace

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We...

AI 聚合 08/18
HuggingFace

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier f...

AI 聚合 08/18
HuggingFace

Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings

Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make mu...

AI 聚合 08/18
HuggingFace

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a...

AI 聚合 08/18
HuggingFace

Modular Cognitive Architecture Emerges in Large Language Models

The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, ...

AI 聚合 08/18
HuggingFace

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented traini...

AI 聚合 08/18
HuggingFace

Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead

Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers...

AI 聚合 08/18
HuggingFace

Is this Citation on Point?

In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations ...

AI 聚合 08/18
arXiv

SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning

Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet rea...

AI 聚合 08/17
arXiv

Shift Aware Transfer Learning with Adaptive Dual-Encoder Fusion for PM Forecasting in Data-Limited Environments

Short-horizon forecasting of fine particulate matter (PM2.5) remains difficult when observations from the target doma...

AI 聚合 08/17
arXiv

LP-NAS: Linear Programming-based Neural Architecture Search

Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expe...

AI 聚合 08/17
arXiv

Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes

As heterogeneous robotic systems deploy across diverse urban zones, maintaining safety amid complex human-robot inter...

AI 聚合 08/17
arXiv

Twin: Playing an Unknown Game with a Test-Time Digital Twin

We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable wor...

AI 聚合 08/17
arXiv

Optimal Scheduling of Road Maintenance Jobs Considering Impact on Traffic Flows

Network-level maintenance planning requires repeated evaluations of equilibrium traffic flows under road capacity red...

AI 聚合 08/17
arXiv

Generating Benchmark Health Data Using a Tabular Diffusion Transformer

Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce...

AI 聚合 08/17
arXiv

Universal Thermodynamic Interatomic Potentials for Crystalline Materials

Free energies govern solid-state phase stability, yet computational materials discovery still relies largely on groun...

AI 聚合 08/17
arXiv

RecipeNet: A Hierarchical Transformer for Recipe Data

Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, ...

AI 聚合 08/17
arXiv

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. T...

AI 聚合 08/17
arXiv

Learning-to-Transition for Large-scale and High-Order MIMO Detection

High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol spa...

AI 聚合 08/17
arXiv

Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. ...

AI 聚合 08/17
arXiv

Handover of In-Context Learning State Across Session Boundaries

This study investigates the methodological and theoretical properties of session handover in applications that use la...

AI 聚合 08/17
arXiv

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing st...

AI 聚合 08/17
arXiv

Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils

Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging p...

AI 聚合 08/17
HuggingFace

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing corres...

AI 聚合 08/17
HuggingFace

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reli...

AI 聚合 08/17
HuggingFace

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-wor...

AI 聚合 08/17
HuggingFace

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with wh...

AI 聚合 08/17
HuggingFace

Scaling Domain Data Repetition in LLM Pretraining

As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per...

AI 聚合 08/17
HuggingFace

Verifier-Induced Support Reshaping in On-Policy Optimization

We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while ...

AI 聚合 08/17
HuggingFace

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation...

AI 聚合 08/17
HuggingFace

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

With the rapid advancement of image editing models and their widespread application across various domains, there is ...

AI 聚合 08/17
HuggingFace

Multimodal Model Diffing for Feature Discovery and Control

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause th...

AI 聚合 08/17
HuggingFace

MobileMem: Learning from a Year of Mobile Experiences

The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persiste...

AI 聚合 08/17
HuggingFace

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule...

AI 聚合 08/17
HuggingFace

Latent On-Policy Self-Distillation

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-ev...

AI 聚合 08/17
HuggingFace

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is conf...

AI 聚合 08/17
HuggingFace

Forecast Collapse in Time-Series Foundation Models

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly...

AI 聚合 08/17
HuggingFace

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently ha...

AI 聚合 08/17
HuggingFace

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skil...

AI 聚合 08/17
HuggingFace

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, ...

AI 聚合 08/17
HuggingFace

Self-Supervised Visual On-Policy Distillation

Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, st...

AI 聚合 08/17
HuggingFace

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to tar...

AI 聚合 08/17
HuggingFace

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and ...

AI 聚合 08/17
GitHub

[GitHub] rasbt/LLMs-from-scratch

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step(⭐102694)

AI 聚合 08/15
HuggingFace

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding...

AI 聚合 08/15
HuggingFace

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on st...

AI 聚合 08/15
HuggingFace

Mitigating Gender Bias in English to Romanian Machine Translation

Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-n...

AI 聚合 08/15
HuggingFace

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This...

AI 聚合 08/15
HuggingFace

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step di...

AI 聚合 08/15
HuggingFace

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

Rib fractures are common and time-consuming to localize on computed tomography (CT). We ask whether fractures detecte...

AI 聚合 08/15
HuggingFace

Thought-Level Beam Search for Reasoning

Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficien...

AI 聚合 08/15
HuggingFace

Maglev: Sliding Recurrent Memory

We introduce , a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention ...

AI 聚合 08/15
arXiv

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Modern image classification models excel when trained on single task-specific datasets but often struggle to generali...

AI 聚合 08/14
arXiv

Concept Drift Detection and Adaptive Retraining of Malware Classification Models

Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was use...

AI 聚合 08/14
arXiv

AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models

Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that re...

AI 聚合 08/14
arXiv

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM promp...

AI 聚合 08/14
arXiv

Synthetic Persona Pretraining: Alignment from Token Zero

As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those ...

AI 聚合 08/14
arXiv

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rathe...

AI 聚合 08/14
arXiv

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive ge...

AI 聚合 08/14
arXiv

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier f...

AI 聚合 08/14
arXiv

The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity

We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \e...

AI 聚合 08/14
arXiv

Vero: Can AI Agents Build Formally Verified Software Repositories?

AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated cod...

AI 聚合 08/14
arXiv

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skil...

AI 聚合 08/14
arXiv

QuoteBench: How Matched Scores Can Hide Command-Path Failures

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched ...

AI 聚合 08/14
arXiv

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with wh...

AI 聚合 08/14
arXiv

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows,...

AI 聚合 08/14
arXiv

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a ...

AI 聚合 08/14
HuggingFace

Full-bandwidth transformer

Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through mode...

AI 聚合 08/14
HuggingFace

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Video world models simulate future states conditioned on current observations and user actions. Recent systems have d...

AI 聚合 08/14
HuggingFace

LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation

Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory sys...

AI 聚合 08/14
HuggingFace

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a ...

AI 聚合 08/14
HuggingFace

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two arch...

AI 聚合 08/14
HuggingFace

DarwinX: Evolving Agent Harnesses Through Natural Selection

An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control f...

AI 聚合 08/14
HuggingFace

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To im...

AI 聚合 08/14
HuggingFace

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed f...

AI 聚合 08/14
HuggingFace

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essen...

AI 聚合 08/14
HuggingFace

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive ...

AI 聚合 08/14
HuggingFace

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation...

AI 聚合 08/14
HuggingFace

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose s...

AI 聚合 08/14
HuggingFace

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source...

AI 聚合 08/14
HuggingFace

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within ...

AI 聚合 08/14
HuggingFace

An AI4AI Framework for Visual Token Pruning

Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet exi...

AI 聚合 08/14
HuggingFace

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, ...

AI 聚合 08/14
HuggingFace

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recent...

AI 聚合 08/14
HuggingFace

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disrup...

AI 聚合 08/14
HuggingFace

PixSDS: Why Latent SDS Makes Noisy Pixels

Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffu...

AI 聚合 08/14
HuggingFace

Parameter Exploration for RLVR via Variational Learning

Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evi...

AI 聚合 08/14
HuggingFace

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packag...

AI 聚合 08/14
HuggingFace

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely f...

AI 聚合 08/14
HuggingFace

Gaze Target Estimation Anywhere with Concepts

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primar...

AI 聚合 08/14
HuggingFace

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inhere...

AI 聚合 08/14
HuggingFace

From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

Large language models often fail when answer options require combining atomic judgments under explicit logical operat...

AI 聚合 08/14
arXiv

HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression

Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-...

AI 聚合 08/13
arXiv

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VI...

AI 聚合 08/13
arXiv

An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS

Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases ca...

AI 聚合 08/13
arXiv

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simu...

AI 聚合 08/13
arXiv

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. F...

AI 聚合 08/13
arXiv

Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents

LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction...

AI 聚合 08/13
arXiv

A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery

Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT ...

AI 聚合 08/13
arXiv

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to acc...

AI 聚合 08/13
arXiv

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benc...

AI 聚合 08/13
arXiv

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lac...

AI 聚合 08/13
arXiv

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial i...

AI 聚合 08/13
arXiv

Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models

Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional o...

AI 聚合 08/13
arXiv

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only...

AI 聚合 08/13
arXiv

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...

AI 聚合 08/13
arXiv

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan futur...

AI 聚合 08/13
HuggingFace

Agent Safety Should Be a Runtime Contract

The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitu...

AI 聚合 08/13
HuggingFace

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design...

AI 聚合 08/13
HuggingFace

AVA-Encoder: Towards Agent-Native Video Representation Learning

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce...

AI 聚合 08/13
HuggingFace

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

World modeling is an unsettled field: architectures, training objectives, and state representations interact in compl...

AI 聚合 08/13
HuggingFace

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet exis...

AI 聚合 08/13
HuggingFace

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vis...

AI 聚合 08/13
HuggingFace

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...

AI 聚合 08/13
HuggingFace

NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs

Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the la...

AI 聚合 08/13
HuggingFace

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. Howev...

AI 聚合 08/13
HuggingFace

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and flui...

AI 聚合 08/13
HuggingFace

Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models

Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-...

AI 聚合 08/13
HuggingFace

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely l...

AI 聚合 08/13
HuggingFace

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

AI agents operate in persistent environments where early state changes can influence decisions far into the future. U...

AI 聚合 08/13
HuggingFace

Persistent Recursive Worlds Enable Autonomous Software Evolution

Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agenti...

AI 聚合 08/13
HuggingFace

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands

Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these ...

AI 聚合 08/13
HuggingFace

Simplex Relaxation for Discrete Diffusion

Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the interme...

AI 聚合 08/13
HuggingFace

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulati...

AI 聚合 08/13
HuggingFace

Self-Evolving Embodied Agents via Skill-Harness Evolution

Embodied agents are increasingly built as systems around foundation models, where performance depends not only on mod...

AI 聚合 08/13
HuggingFace

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome,...

AI 聚合 08/13
HuggingFace

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities a...

AI 聚合 08/13
arXiv

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

The application of computer vision in agriculture has shown significant potential for improving crop monitoring and p...

AI 聚合 08/13
arXiv

Entropy-Centric Explainable AI for Remote Sensing Image Segmentation

Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many con...

AI 聚合 08/13
arXiv

Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory

We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses se...

AI 聚合 08/13
arXiv

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the ...

AI 聚合 08/13
arXiv

RTSKG: Building a Rail Transit Station Knowledge Graph Dataset

Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems,...

AI 聚合 08/13
arXiv

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository reti...

AI 聚合 08/13
arXiv

Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting

Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon set...

AI 聚合 08/13
arXiv

sLTN: Structural Logic Tensor Networks

Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tenso...

AI 聚合 08/13
arXiv

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution b...

AI 聚合 08/13
arXiv

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021,...

AI 聚合 08/13
arXiv

How to Verify Consistency of Probabilistic Claims

When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can...

AI 聚合 08/13
arXiv

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters af...

AI 聚合 08/13
arXiv

Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration

AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards...

AI 聚合 08/13
arXiv

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are ...

AI 聚合 08/13
arXiv

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: te...

AI 聚合 08/13
HuggingFace

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoi...

AI 聚合 08/13
HuggingFace

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mo...

AI 聚合 08/12
HuggingFace

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging pers...

AI 聚合 08/12
HuggingFace

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualizatio...

AI 聚合 08/12
HuggingFace

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often boun...

AI 聚合 08/12
HuggingFace

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the ...

AI 聚合 08/12
HuggingFace

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain...

AI 聚合 08/12
HuggingFace

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchma...

AI 聚合 08/12
HuggingFace

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity

The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis,...

AI 聚合 08/12
HuggingFace

iFAN: Inference-Aware Learning for Plain Mask Transformers

Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of...

AI 聚合 08/12
HuggingFace

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can b...

AI 聚合 08/12
HuggingFace

Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

We study reference-free post-training for multilingual machine translation with open large language models. Starting ...

AI 聚合 08/12
HuggingFace

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but cont...

AI 聚合 08/12
HuggingFace

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, compl...

AI 聚合 08/12
HuggingFace

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus ...

AI 聚合 08/12
HuggingFace

Articulated Object Reconstruction from Rest-State Observation

Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how ...

AI 聚合 08/12
HuggingFace

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained ...

AI 聚合 08/12
HuggingFace

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of cap...

AI 聚合 08/12
HuggingFace

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a ph...

AI 聚合 08/12
HuggingFace

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attenti...

AI 聚合 08/12
HuggingFace

InSight-doc: Agentic Visual Perception for Long-Document Understanding

Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone...

AI 聚合 08/12
HuggingFace

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs pres...

AI 聚合 08/12
HuggingFace

The Loss Does Not See the Basis, but Adam Does

Gradient descent on a factored model W = UV^top is implicitly biased toward low-rank solutions, while Adam, starting ...

AI 聚合 08/12
HuggingFace

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, b...

AI 聚合 08/12
HuggingFace

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context...

AI 聚合 08/12
HuggingFace

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be un...

AI 聚合 08/12
HuggingFace

A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization

In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously updates ...

AI 聚合 08/12
HuggingFace

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Benchmarks for systems that are optimized against the evaluation signal measure something different from what they cl...

AI 聚合 08/12
HuggingFace

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in pen...

AI 聚合 08/12
HuggingFace

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirr...

AI 聚合 08/12
HuggingFace

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals...

AI 聚合 08/12
HuggingFace

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating la...

AI 聚合 08/12
HuggingFace

On-Policy Self-Distillation without Any Supervision

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs)....

AI 聚合 08/12
HuggingFace

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond ...

AI 聚合 08/12
HuggingFace

Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic stru...

AI 聚合 08/12
HuggingFace

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each t...

AI 聚合 08/12
arXiv

Agentic Auto-Research is Fuzz Testing

Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have respo...

AI 聚合 08/11
arXiv

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of su...

AI 聚合 08/11
arXiv

Towards Expert-level Medical AI for Real-time Video Consultations

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effe...

AI 聚合 08/11
arXiv

Stealing Reasoning Traces from Proprietary LLM APIs

Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to prot...

AI 聚合 08/11
arXiv

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation...

AI 聚合 08/11
arXiv

ArchAgent v2: A Case Study with the Data Prefetching Championship

Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar technique...

AI 聚合 08/11
arXiv

Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning

Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must ...

AI 聚合 08/11
arXiv

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that ...

AI 聚合 08/11
arXiv

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs pres...

AI 聚合 08/11
arXiv

Fusion Training for Mathematical Generalization in Large Language Models

Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by...

AI 聚合 08/11
arXiv

DSLE: A Learning Environment for Dark Souls Boss Encounters

We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounter...

AI 聚合 08/11
arXiv

GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from ...

AI 聚合 08/11
arXiv

From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch

Large language models are increasingly being deployed in governmental settings, yet few existing evaluation framework...

AI 聚合 08/11
arXiv

Multimodal Model Diffing for Feature Discovery and Control

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause th...

AI 聚合 08/11
arXiv

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Model...

AI 聚合 08/11
HuggingFace

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-ind...

AI 聚合 08/11
HuggingFace

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation...

AI 聚合 08/11
HuggingFace

Motif 3: Technical Report

We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 bil...

AI 聚合 08/11
HuggingFace

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are ra...

AI 聚合 08/11
HuggingFace

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments...

AI 聚合 08/11
HuggingFace

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generati...

AI 聚合 08/11
HuggingFace

Evo-Bench: Can Language Models Improve Agent Harness?

Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confine...

AI 聚合 08/11
HuggingFace

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation...

AI 聚合 08/11
HuggingFace

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however...

AI 聚合 08/11
HuggingFace

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning...

AI 聚合 08/11
HuggingFace

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for...

AI 聚合 08/11
HuggingFace

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-conte...

AI 聚合 08/11
HuggingFace

Scaling Inherently Interpretable Language Models

Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explain...

AI 聚合 08/11
HuggingFace

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems prim...

AI 聚合 08/11
HuggingFace

A^2E : An End-to-End Agent Auditing Engine

With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploy...

AI 聚合 08/11
HuggingFace

Stealing Reasoning Traces from Proprietary LLM APIs

Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to prot...

AI 聚合 08/11
HuggingFace

CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems

The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interacti...

AI 聚合 08/11
HuggingFace

Vision-Language Grounding as Bidirectional Concept Correspondence

Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a u...

AI 聚合 08/11
HuggingFace

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System

We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data col...

AI 聚合 08/11
HuggingFace

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment...

AI 聚合 08/11
HuggingFace

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints...

AI 聚合 08/11
HuggingFace

FATE: Frame-Level Audio-Visual Temporal Embedding

When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio...

AI 聚合 08/11
HuggingFace

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's inter...

AI 聚合 08/11
HuggingFace

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Lega...

AI 聚合 08/11
HuggingFace

Small Foundation Models of Human Cognition and Behaviour

Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the...

AI 聚合 08/11
HuggingFace

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and...

AI 聚合 08/11
HuggingFace

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious eth...

AI 聚合 08/11
HuggingFace

DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds

CLI-based software-engineering agents have matured rapidly, yet the open ecosystem has converged on a single training...

AI 聚合 08/11
HuggingFace

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are ...

AI 聚合 08/11
HuggingFace

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting

Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence le...

AI 聚合 08/11
HuggingFace

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies wit...

AI 聚合 08/11
HuggingFace

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open wheth...

AI 聚合 08/11
HuggingFace

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

World generative models are typically used through what they produce: a rendered future, a video-conditioned action, ...

AI 聚合 08/11
arXiv

CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing

Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a ...

AI 聚合 08/10
arXiv

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical in...

AI 聚合 08/10
arXiv

TEPA: Revoking Stale Memories for Conflict-Robust Language Agents

Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also crea...

AI 聚合 08/10
arXiv

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoisin...

AI 聚合 08/10
arXiv

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to...

AI 聚合 08/10
arXiv

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition fas...

AI 聚合 08/10
arXiv

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents ...

AI 聚合 08/10
arXiv

PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents

Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolv...

AI 聚合 08/10
arXiv

Blast Radius

Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive mem...

AI 聚合 08/10
arXiv

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and gover...

AI 聚合 08/10
arXiv

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lig...

AI 聚合 08/10
arXiv

Strategy-first synthesis planning for complex natural products

The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistr...

AI 聚合 08/10
arXiv

Interaction Creates Dynamical AI Behavior Absent in Isolation

What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a cou...

AI 聚合 08/10
arXiv

CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG

Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoi...

AI 聚合 08/10
arXiv

CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity

While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diver...

AI 聚合 08/10
HuggingFace

Modular TTT: Rethinking Test-Time Training as Composable Modules

Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by a...

AI 聚合 08/10
HuggingFace

Characterizing the Quality Profile of AI-Generated C++ in Production

The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent stud...

AI 聚合 08/10
HuggingFace

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers s...

AI 聚合 08/10
HuggingFace

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio...

AI 聚合 08/10
HuggingFace

Douyin Multimodal Embedding Model Technical Report

Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vec...

AI 聚合 08/10
HuggingFace

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal...

AI 聚合 08/10
HuggingFace

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action pred...

AI 聚合 08/10
HuggingFace

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher...

AI 聚合 08/10
HuggingFace

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-tim...

AI 聚合 08/10
HuggingFace

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors

Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it a...

AI 聚合 08/10
HuggingFace

PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. ...

AI 聚合 08/10
HuggingFace

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable ...

AI 聚合 08/10
HuggingFace

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and rol...

AI 聚合 08/10
HuggingFace

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain di...

AI 聚合 08/10
HuggingFace

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing m...

AI 聚合 08/10
HuggingFace

Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning

Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply incr...

AI 聚合 08/10
HuggingFace

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorize...

AI 聚合 08/10
HuggingFace

Towards Interpretable Foundation Models for Retinal Fundus Images

Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically vi...

AI 聚合 08/10
HuggingFace

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherent...

AI 聚合 08/10
HuggingFace

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However...

AI 聚合 08/10
HuggingFace

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yid...

AI 聚合 08/08
HuggingFace

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by dee...

AI 聚合 08/08
HuggingFace

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattere...

AI 聚合 08/08
HuggingFace

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an ag...

AI 聚合 08/08
HuggingFace

KVAE: Family of Tokenizers for Multimodal Generative Models

Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed represen...

AI 聚合 08/08
HuggingFace

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

World models have attracted significant attention for their ability to capture and predict the structure and dynamics...

AI 聚合 08/08
HuggingFace

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action,...

AI 聚合 08/08
HuggingFace

GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization

Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene e...

AI 聚合 08/08
arXiv

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) ...

AI 聚合 08/07
arXiv

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but...

AI 聚合 08/07
arXiv

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and ...

AI 聚合 08/07
arXiv

Does FLAIR super-resolution erase or hallucinate small white-matter lesions?

White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associat...

AI 聚合 08/07
arXiv

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark ...

AI 聚合 08/07
arXiv

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

From natural-language query interfaces to automated report generation, data analysis tools need a description of the ...

AI 聚合 08/07
arXiv

TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories

LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading error...

AI 聚合 08/07
arXiv

Challenges in Evaluating Explanation Methods for Static and Evolving Data

This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evalua...

AI 聚合 08/07
arXiv

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mecha...

AI 聚合 08/07
arXiv

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and v...

AI 聚合 08/07
arXiv

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, ...

AI 聚合 08/07
arXiv

An Optimal Agnostic PAC Algorithm

Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_...

AI 聚合 08/07
arXiv

Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including f...

AI 聚合 08/07
arXiv

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for ...

AI 聚合 08/07
arXiv

Learning When to Trust via Selective Context Preference Optimization

Language models increasingly condition their answers on external signals, and a single misleading one can turn a corr...

AI 聚合 08/07
HuggingFace

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiri...

AI 聚合 08/07
HuggingFace

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogen...

AI 聚合 08/07
HuggingFace

On-Policy Delta Distillation for Multilingual Math Reasoning

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, ...

AI 聚合 08/07
HuggingFace

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception fro...

AI 聚合 08/07
HuggingFace

WorldClaw: Agentic 3D Open-World Generation at Scale

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jo...

AI 聚合 08/07
HuggingFace

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fai...

AI 聚合 08/07
HuggingFace

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately cha...

AI 聚合 08/07
HuggingFace

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous informa...

AI 聚合 08/07
HuggingFace

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesi...

AI 聚合 08/07
HuggingFace

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, over...

AI 聚合 08/07
HuggingFace

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one au...

AI 聚合 08/07
HuggingFace

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...

AI 聚合 08/07
HuggingFace

Invisible Shortcuts: Why Vision Encoders Know Your Camera

Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused...

AI 聚合 08/07
HuggingFace

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fi...

AI 聚合 08/07
HuggingFace

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet t...

AI 聚合 08/07
HuggingFace

MASS: Multiplayer World Models with Authoritative Shared State

Current video world models struggle in multiplayer environments because they entangle world state with view-dependent...

AI 聚合 08/07
HuggingFace

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modelin...

AI 聚合 08/07
HuggingFace

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, desp...

AI 聚合 08/07
HuggingFace

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despi...

AI 聚合 08/07
HuggingFace

Continual Learning in Transition

Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through par...

AI 聚合 08/07
HuggingFace

DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoisi...

AI 聚合 08/07
HuggingFace

Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what...

AI 聚合 08/07
HuggingFace

Lossless Tensor Compression as Program Synthesis

Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly co...

AI 聚合 08/07
HuggingFace

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textu...

AI 聚合 08/07
HuggingFace

What AI Red-Team Evaluations Can and Cannot Prove

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable ...

AI 聚合 08/07
HuggingFace

FinanceHarness: Autonomous Financial Deep Research Framework

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic pr...

AI 聚合 08/07
HuggingFace

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Collaborative music agents need internal representations rich enough to support both understanding and generation, ye...

AI 聚合 08/07
arXiv

VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection

Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variabi...

AI 聚合 08/06
arXiv

MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation

Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks s...

AI 聚合 08/06
arXiv

Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection

Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, ...

AI 聚合 08/06
arXiv

Item Response Theory for AI Safety

Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggrega...

AI 聚合 08/06
arXiv

Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite

Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new fac...

AI 聚合 08/06
arXiv

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate e...

AI 聚合 08/06
arXiv

CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs

AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions acr...

AI 聚合 08/06
arXiv

Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth

Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum genera...

AI 聚合 08/06
arXiv

Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition

Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally...

AI 聚合 08/06
arXiv

Chained Recursive Language Models for Multi-Iteration Reasoning

Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference tra...

AI 聚合 08/06
arXiv

SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant

Achieving local differential privacy in distributed optimization while maintaining low communication cost remains cha...

AI 聚合 08/06
arXiv

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in mul...

AI 聚合 08/06
arXiv

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, desp...

AI 聚合 08/06
arXiv

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, ...

AI 聚合 08/06
arXiv

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and p...

AI 聚合 08/06
HuggingFace

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

The abundance of casually captured monocular videos and images on social media provides a valuable source for immersi...

AI 聚合 08/06
HuggingFace

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing ...

AI 聚合 08/06
HuggingFace

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean perfor...

AI 聚合 08/06
HuggingFace

When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents

Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the enviro...

AI 聚合 08/06
HuggingFace

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain,...

AI 聚合 08/06
HuggingFace

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains...

AI 聚合 08/06
HuggingFace

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense to...

AI 聚合 08/06
HuggingFace

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are ...

AI 聚合 08/06
HuggingFace

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compound...

AI 聚合 08/06
HuggingFace

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as ...

AI 聚合 08/06
HuggingFace

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language...

AI 聚合 08/06
HuggingFace

K-EXAONE 2.0 Technical Report

This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research...

AI 聚合 08/06
HuggingFace

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric hu...

AI 聚合 08/06
HuggingFace

Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning

As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross...

AI 聚合 08/06
HuggingFace

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore cri...

AI 聚合 08/06
HuggingFace

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visua...

AI 聚合 08/06
HuggingFace

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks m...

AI 聚合 08/06
HuggingFace

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction...

AI 聚合 08/06
HuggingFace

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate e...

AI 聚合 08/06
HuggingFace

Self-Evolving Coding Agents

Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect r...

AI 聚合 08/06
HuggingFace

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matchin...

AI 聚合 08/06
HuggingFace

Multi-Task Multi-Frame Visual Piano Transcription

Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist...

AI 聚合 08/06
HuggingFace

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows flo...

AI 聚合 08/06
HuggingFace

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability...

AI 聚合 08/06
HuggingFace

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically...

AI 聚合 08/06
HuggingFace

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future querie...

AI 聚合 08/06
HuggingFace

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension wi...

AI 聚合 08/06
HuggingFace

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large...

AI 聚合 08/06
HuggingFace

LegalPincite: Multi-level Legal Information Retrieval Dataset

A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While ...

AI 聚合 08/06
HuggingFace

ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but prac...

AI 聚合 08/06
HuggingFace

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can f...

AI 聚合 08/06
HuggingFace

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach aba...

AI 聚合 08/06
arXiv

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames se...

AI 聚合 08/05
arXiv

Equivariant Music Transformer

Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equ...

AI 聚合 08/05
arXiv

The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections

This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that...

AI 聚合 08/05
arXiv

PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection

Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, ...

AI 聚合 08/05
arXiv

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. Howev...

AI 聚合 08/05
arXiv

TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring

Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-languag...

AI 聚合 08/05
arXiv

A game theory for foundation models shows new paths to rational cooperation through similarity inference

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, under...

AI 聚合 08/05
arXiv

Interpretable Adaptive Sampling for LLM Test-Time Scaling

Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines...

AI 聚合 08/05
arXiv

Separating quantum circuits from classical LLMs

Modern large language models - transformers and diffusion language models - are built around two canonical algorithmi...

AI 聚合 08/05
arXiv

Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations

Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic...

AI 聚合 08/05
arXiv

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large...

AI 聚合 08/05
arXiv

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video stream...

AI 聚合 08/05
arXiv

Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?

Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed progr...

AI 聚合 08/05
arXiv

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "t...

AI 聚合 08/05
arXiv

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, exi...

AI 聚合 08/05
HuggingFace

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate under...

AI 聚合 08/05
HuggingFace

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Real-time video editing requires low-latency causal generation with bounded computational resources while preserving ...

AI 聚合 08/05
HuggingFace

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, exi...

AI 聚合 08/05
HuggingFace

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. Howeve...

AI 聚合 08/05
HuggingFace

AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling

Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in contin...

AI 聚合 08/05
HuggingFace

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks...

AI 聚合 08/05
HuggingFace

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learnin...

AI 聚合 08/05
HuggingFace

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI ag...

AI 聚合 08/05
HuggingFace

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift rai...

AI 聚合 08/05
HuggingFace

ExplainBench: Evaluating Code Explanations from Agents

Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in ...

AI 聚合 08/05
HuggingFace

Quo Vadis, World Modeling?

Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environm...

AI 聚合 08/05
HuggingFace

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual d...

AI 聚合 08/05
HuggingFace

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories

Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robo...

AI 聚合 08/05
HuggingFace

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video stream...

AI 聚合 08/05
HuggingFace

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. W...

AI 聚合 08/05
HuggingFace

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling beha...

AI 聚合 08/05
HuggingFace

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems ...

AI 聚合 08/05
HuggingFace

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it...

AI 聚合 08/05
HuggingFace

MiniWorld: Democratizing the Training of Video World Models from Scratch

Video world models predict future observations conditioned on historical observations and control signals, enabling l...

AI 聚合 08/05
HuggingFace

Decoding Children's Gait Behavior

We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behavior...

AI 聚合 08/05
HuggingFace

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and w...

AI 聚合 08/05
HuggingFace

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging becau...

AI 聚合 08/05
HuggingFace

GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation

Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating th...

AI 聚合 08/05
HuggingFace

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-...

AI 聚合 08/05
HuggingFace

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries an...

AI 聚合 08/05
HuggingFace

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors...

AI 聚合 08/05
HuggingFace

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalit...

AI 聚合 08/05
HuggingFace

Zero-Mem: Zero-Token Memory Operations for LLM Agents

LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to opera...

AI 聚合 08/05
HuggingFace

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on ...

AI 聚合 08/05
HuggingFace

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing...

AI 聚合 08/05
HuggingFace

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilitie...

AI 聚合 08/05
HuggingFace

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-...

AI 聚合 08/05
HuggingFace

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. How...

AI 聚合 08/05
arXiv

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid in...

AI 聚合 08/04
arXiv

DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation

Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging becau...

AI 聚合 08/04
arXiv

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and mo...

AI 聚合 08/04
arXiv

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet th...

AI 聚合 08/04
arXiv

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of s...

AI 聚合 08/04
arXiv

Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies

Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncer...

AI 聚合 08/04
arXiv

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in co...

AI 聚合 08/04
arXiv

Analytic Planning under Uncertainty with Moment Closure

Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictiv...

AI 聚合 08/04
arXiv

Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation

Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. ...

AI 聚合 08/04
arXiv

A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI

Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustain...

AI 聚合 08/04
arXiv

Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection

Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Tra...

AI 聚合 08/04
arXiv

AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: th...

AI 聚合 08/04
arXiv

CoWAM: Coordination Contracts for Selective Policy Intervention with WAMs

World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future a...

AI 聚合 08/04
arXiv

UEmbed: Unified Sparse and Dense Multimodal Embeddings

Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work ha...

AI 聚合 08/04
arXiv

Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework

Artificial intelligence (AI) is increasingly central to power and energy systems, supporting modeling, forecasting, o...

AI 聚合 08/04
HuggingFace

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, ...

AI 聚合 08/04
HuggingFace

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains dif...

AI 聚合 08/04
HuggingFace

UEmbed: Unified Sparse and Dense Multimodal Embeddings

Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work ha...

AI 聚合 08/04
HuggingFace

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and mo...

AI 聚合 08/04
HuggingFace

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them i...

AI 聚合 08/04
HuggingFace

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledg...

AI 聚合 08/04
HuggingFace

Progressive Agent Skill Generation via Reinforcement Learning

Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially...

AI 聚合 08/04
HuggingFace

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient m...

AI 聚合 08/04
HuggingFace

CADENA: Stepwise CAD Reverse Engineering

Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still d...

AI 聚合 08/04
HuggingFace

DiffusionGemma Technical Report

We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text...

AI 聚合 08/04
HuggingFace

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool us...

AI 聚合 08/04
HuggingFace

DAPD: Dual-Anchored Policy Distillation

On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the tea...

AI 聚合 08/04
HuggingFace

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into worl...

AI 聚合 08/04
HuggingFace

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations

Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations la...

AI 聚合 08/04
HuggingFace

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametri...

AI 聚合 08/04
HuggingFace

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling...

AI 聚合 08/04
HuggingFace

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integ...

AI 聚合 08/04
HuggingFace

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous ...

AI 聚合 08/04
HuggingFace

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. P...

AI 聚合 08/04
HuggingFace

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, object...

AI 聚合 08/04
HuggingFace

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance ofte...

AI 聚合 08/04
HuggingFace

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while...

AI 聚合 08/04
HuggingFace

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understandi...

AI 聚合 08/04
HuggingFace

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life...

AI 聚合 08/04
HuggingFace

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isola...

AI 聚合 08/04
arXiv

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Standard AI-text detection benchmarks compare human-written text against text generated directly by large language mo...

AI 聚合 08/03
arXiv

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying thei...

AI 聚合 08/03
arXiv

COntExt: Towards Context-Aware Ontology Extension from Operational Metrics

Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, pro...

AI 聚合 08/03
arXiv

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. Howe...

AI 聚合 08/03
arXiv

MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in sc...

AI 聚合 08/03
arXiv

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites ...

AI 聚合 08/03
arXiv

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-...

AI 聚合 08/03
arXiv

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for infe...

AI 聚合 08/03
arXiv

A Human-Centered Validation of the Explainability-Performance Coefficient

The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable ...

AI 聚合 08/03
arXiv

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications...

AI 聚合 08/03
arXiv

CENDRe: Concept Extraction with Natural Domain Representations

Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical...

AI 聚合 08/03
arXiv

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and o...

AI 聚合 08/03
arXiv

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct ex...

AI 聚合 08/03
arXiv

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency,...

AI 聚合 08/03
arXiv

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-def...

AI 聚合 08/03
HuggingFace

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language...

AI 聚合 08/03
HuggingFace

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-ric...

AI 聚合 08/03
HuggingFace

QQWorld: Quantile-Quantile Matching for World Model Regularization

Latent world models enable efficient planning by predicting future states in a compact representation space, but thei...

AI 聚合 08/03
HuggingFace

Scaling Properties of Text Conditioning in Visual Generation

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been me...

AI 聚合 08/03
HuggingFace

Meshy T2: Fast Native Mesh Generation with Flow Matching

Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes w...

AI 聚合 08/03
HuggingFace

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined sc...

AI 聚合 08/03
HuggingFace

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly...

AI 聚合 08/03
HuggingFace

Enhancing Rubric-based RL via Self-Distillation

Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of r...

AI 聚合 08/03
HuggingFace

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a bas...

AI 聚合 08/03
HuggingFace

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surfac...

AI 聚合 08/03
HuggingFace

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future visi...

AI 聚合 08/03
HuggingFace

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioriti...

AI 聚合 08/03
HuggingFace

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning par...

AI 聚合 08/03
HuggingFace

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limite...

AI 聚合 08/03
HuggingFace

Mental World Modeling

World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physica...

AI 聚合 08/03
HuggingFace

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often c...

AI 聚合 08/03
HuggingFace

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NI...

AI 聚合 08/03
HuggingFace

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing

Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This cre...

AI 聚合 08/03
HuggingFace

SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing

Autonomous multi-vehicle racing requires real-time planning of diverse competitive behaviors in intense interactions....

AI 聚合 08/03
HuggingFace

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language mo...

AI 聚合 08/03
HuggingFace

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and fut...

AI 聚合 08/01
HuggingFace

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account ...

AI 聚合 08/01
HuggingFace

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains britt...

AI 聚合 08/01
HuggingFace

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

Existing token compression methods for omnimodal large language models typically rely on one modality to determine wh...

AI 聚合 08/01
arXiv

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonl...

AI 聚合 07/31
arXiv

MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems

Large language model-based multi-agent systems improve complex problem solving through task decomposition, agent spec...

AI 聚合 07/31
arXiv

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something diffe...

AI 聚合 07/31
arXiv

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While...

AI 聚合 07/31
arXiv

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs

Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and p...

AI 聚合 07/31
arXiv

Algorithms for Structured Elections under Thiele Voting Rules

We study the computational complexity of winner determination problems in approval-based committee elections under Th...

AI 聚合 07/31
arXiv

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B

Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of ...

AI 聚合 07/31
arXiv

DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation

While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex...

AI 聚合 07/31
arXiv

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a c...

AI 聚合 07/31
arXiv

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...

AI 聚合 07/31
arXiv

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applicati...

AI 聚合 07/31
arXiv

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet ex...

AI 聚合 07/31
arXiv

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic...

AI 聚合 07/31
arXiv

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors g...

AI 聚合 07/31
arXiv

Learning to Trace Seiberg Dualities

Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical ...

AI 聚合 07/31
HuggingFace

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (...

AI 聚合 07/31
HuggingFace

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local vi...

AI 聚合 07/31
HuggingFace

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them tow...

AI 聚合 07/31
HuggingFace

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult t...

AI 聚合 07/31
HuggingFace

Can Large Language Models Execute Parent Orders?

Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smalle...

AI 聚合 07/31
HuggingFace

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placi...

AI 聚合 07/31
HuggingFace

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. User...

AI 聚合 07/31
HuggingFace

AI Tour Meeting: Group Travel Planning by LLM Agents

This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)...

AI 聚合 07/31
HuggingFace

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose...

AI 聚合 07/31
HuggingFace

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem i...

AI 聚合 07/31
HuggingFace

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perce...

AI 聚合 07/31
HuggingFace

INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

Forward latent world models predict how actions change a scene, but recover actions for a desired change only through...

AI 聚合 07/31
HuggingFace

Multi-Head Attention Residuals

Transformers propagate information across depth through a single additive residual stream: every sublayer reads only ...

AI 聚合 07/31
HuggingFace

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models t...

AI 聚合 07/31
HuggingFace

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The o...

AI 聚合 07/31
HuggingFace

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence ...

AI 聚合 07/31
HuggingFace

Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing

In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely ...

AI 聚合 07/31
HuggingFace

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that...

AI 聚合 07/31
HuggingFace

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reason...

AI 聚合 07/31
HuggingFace

Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems

Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rath...

AI 聚合 07/31
HuggingFace

CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches...

AI 聚合 07/31
HuggingFace

Memory for Large Language Models

Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an impl...

AI 聚合 07/31
HuggingFace

πR^2: Reactive Real-time Flow Policies

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretraine...

AI 聚合 07/31
HuggingFace

Voice Memory for Agentic Speech Recognition

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector r...

AI 聚合 07/31
HuggingFace

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based mu...

AI 聚合 07/31
HuggingFace

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including ...

AI 聚合 07/31
HuggingFace

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a ...

AI 聚合 07/31
arXiv

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnera...

AI 聚合 07/30
arXiv

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, age...

AI 聚合 07/30
arXiv

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and ...

AI 聚合 07/30
arXiv

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptio...

AI 聚合 07/30
arXiv

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The...

AI 聚合 07/30
arXiv

Linguistic Monoculture in LLM-Assisted Language Use

Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, rev...

AI 聚合 07/30
arXiv

DLAM: Distributional Latent Actions with Temporal Constraints

Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free video...

AI 聚合 07/30
arXiv

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable u...

AI 聚合 07/30
arXiv

Anatomy Contextualized Adaption of CT Foundation Models

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typical...

AI 聚合 07/30
arXiv

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing be...

AI 聚合 07/30
arXiv

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision ...

AI 聚合 07/30
arXiv

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc...

AI 聚合 07/30
arXiv

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its prese...

AI 聚合 07/30
arXiv

APEX-Accounting

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models...

AI 聚合 07/30
arXiv

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carr...

AI 聚合 07/30
HuggingFace

Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

Most video editing systems still lack explicit layered video representations, limiting their ability to perform reali...

AI 聚合 07/30
HuggingFace

GPT-Red: Automated Red Teaming via Self-Play at Scale

We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks again...

AI 聚合 07/30
HuggingFace

HumanCLAW: Can Vision-Language Models Act Through a Body?

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an ac...

AI 聚合 07/30
HuggingFace

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are ...

AI 聚合 07/30
HuggingFace

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carr...

AI 聚合 07/30
HuggingFace

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing be...

AI 聚合 07/30
HuggingFace

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelli...

AI 聚合 07/30
HuggingFace

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-mak...

AI 聚合 07/30
HuggingFace

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather tha...

AI 聚合 07/30
HuggingFace

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet sta...

AI 聚合 07/30
HuggingFace

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit cri...

AI 聚合 07/30
HuggingFace

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained kno...

AI 聚合 07/30
HuggingFace

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations r...

AI 聚合 07/30
HuggingFace

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host arti...

AI 聚合 07/30
HuggingFace

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

Recent game world models can generate visually realistic and interactive environments conditioned on player actions. ...

AI 聚合 07/30
HuggingFace

CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents

Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, l...

AI 聚合 07/30
HuggingFace

GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels

Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definit...

AI 聚合 07/30
HuggingFace

Edge-Aware Thermal Infrared UAV Swarm Tracking

Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, trac...

AI 聚合 07/30
HuggingFace

Projection Pursuit CPCANet for Domain Generalization

Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment me...

AI 聚合 07/30
HuggingFace

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation d...

AI 聚合 07/30
HuggingFace

OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distrib...

AI 聚合 07/30
HuggingFace

Reinforcement Learning for Code Optimization

RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and ...

AI 聚合 07/30
HuggingFace

Uncovering Latent Reasoning Strategies in Language Models

A language model p_θ(y mid x) trained on reasoning tasks learns to solve problems via multiple distinct strategies, y...

AI 聚合 07/30
HuggingFace

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update ru...

AI 聚合 07/30
HuggingFace

Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection

Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how...

AI 聚合 07/30
arXiv

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in ...

AI 聚合 07/29
arXiv

Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no sin...

AI 聚合 07/29
arXiv

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. ...

AI 聚合 07/29
arXiv

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health beha...

AI 聚合 07/29
arXiv

Reinforcement Learning for Code Optimization

RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and ...

AI 聚合 07/29
arXiv

MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents

Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon...

AI 聚合 07/29
arXiv

Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?

Kubernetes is central to the cloud-native ecosystem, orchestrating containerised workloads. Recent work suggests that...

AI 聚合 07/29
arXiv

Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models

Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competi...

AI 聚合 07/29
arXiv

Pictura: Perspective-View Self-Play at Scale for Driving

Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made usi...

AI 聚合 07/29
arXiv

MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar

Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficien...

AI 聚合 07/29
arXiv

CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer

Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains a...

AI 聚合 07/29
arXiv

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even ...

AI 聚合 07/29
arXiv

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks p...

AI 聚合 07/29
arXiv

$π\mathbf{R}^2$: Reactive Real-time Flow Policies

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretraine...

AI 聚合 07/29
arXiv

Pass the Baton: Trajectory-Relayed On-Policy Distillation

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix...

AI 聚合 07/29
HuggingFace

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often c...

AI 聚合 07/29
HuggingFace

Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms mod...

AI 聚合 07/29
HuggingFace

Wonder: Video World Model Done Better

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an...

AI 聚合 07/29
HuggingFace

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but thei...

AI 聚合 07/29
HuggingFace

Visual prompt engineering for video models

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has becom...

AI 聚合 07/29
HuggingFace

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. Th...

AI 聚合 07/29
HuggingFace

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities...

AI 聚合 07/29
HuggingFace

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scal...

AI 聚合 07/29
HuggingFace

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval...

AI 聚合 07/29
HuggingFace

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning ...

AI 聚合 07/29
HuggingFace

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, ...

AI 聚合 07/29
HuggingFace

Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion

We present a reproducible pipeline for mapping Common Vulnerabilities and Exposures (CVEs) to MITRE ATT&CK Enterprise...

AI 聚合 07/29
HuggingFace

Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However...

AI 聚合 07/29
HuggingFace

Pass the Baton: Trajectory-Relayed On-Policy Distillation

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix...

AI 聚合 07/29
HuggingFace

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in ...

AI 聚合 07/29
HuggingFace

Parallel Decoding Distillation for Fast Image and Video Generation

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling proc...

AI 聚合 07/29
HuggingFace

Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than...

AI 聚合 07/29
HuggingFace

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

Different research lines use the term world model in different ways, yet they share a common aim: to capture how the ...

AI 聚合 07/29
HuggingFace

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

Compressed short-text generators can fail in two different places: the codec may discard information before generatio...

AI 聚合 07/29
HuggingFace

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and train...

AI 聚合 07/29
HuggingFace

A Vocabulary for Multi-Agent Automated Research Systems

We introduce a vocabulary for automated research systems built from one or more agents to make their design choices e...

AI 聚合 07/29
HuggingFace

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the act...

AI 聚合 07/29
HuggingFace

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deploy...

AI 聚合 07/29
HuggingFace

TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time ...

AI 聚合 07/29
HuggingFace

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features

Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fa...

AI 聚合 07/29
arXiv

Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents

Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks ...

AI 聚合 07/28
arXiv

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between S...

AI 聚合 07/28
arXiv

Efficiency Matters in Autonomous Research

AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their p...

AI 聚合 07/28
arXiv

Reason-Mediated Behavioral Models for Auditing LLM Social Simulators

Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most eva...

AI 聚合 07/28
arXiv

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic...

AI 聚合 07/28
arXiv

Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating

A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed meth...

AI 聚合 07/28
arXiv

Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

Multi-modal classification leverages complementary information across diverse data sources to enhance predictive perf...

AI 聚合 07/28
arXiv

Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines

Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline...

AI 聚合 07/28
arXiv

ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only ...

AI 聚合 07/28
arXiv

Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures

Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit...

AI 聚合 07/28
arXiv

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many ...

AI 聚合 07/28
arXiv

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains...

AI 聚合 07/28
arXiv

KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability

Computer vision models have become highly effective for medical applications, yet their black-box nature continues to...

AI 聚合 07/28
arXiv

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the curren...

AI 聚合 07/28
arXiv

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying the...

AI 聚合 07/28
HuggingFace

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of ...

AI 聚合 07/28
HuggingFace

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning ...

AI 聚合 07/28
HuggingFace

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretra...

AI 聚合 07/28
HuggingFace

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent ge...

AI 聚合 07/28
HuggingFace

Kimi K3: Open Frontier Intelligence

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision ...

AI 聚合 07/28
HuggingFace

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage,...

AI 聚合 07/28
HuggingFace

GNM Head: A Generative aNthropometric Model of the human head

Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animat...

AI 聚合 07/28
HuggingFace

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-determinis...

AI 聚合 07/28
HuggingFace

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-pu...

AI 聚合 07/28
HuggingFace

dRAE: Representation Autoencoder with Hyper-Spherical Codes

In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models...

AI 聚合 07/28
HuggingFace

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterior...

AI 聚合 07/28
HuggingFace

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the curren...

AI 聚合 07/28
HuggingFace

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue re...

AI 聚合 07/28
HuggingFace

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support...

AI 聚合 07/28
HuggingFace

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains...

AI 聚合 07/28
HuggingFace

Codifying the Judge: Scalable Evaluation via Program Distillation

LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, ...

AI 聚合 07/28
HuggingFace

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling

The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in stru...

AI 聚合 07/28
HuggingFace

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a do...

AI 聚合 07/28
HuggingFace

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models

Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may c...

AI 聚合 07/28
HuggingFace

Characterizing Warp Divergence from Pascal to Blackwell

Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp diver...

AI 聚合 07/28
HuggingFace

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, whi...

AI 聚合 07/28
HuggingFace

Interactive Training 2: Auditable Control Plane for Live Model Training

Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific...

AI 聚合 07/28
arXiv

Robot Learning to Communicate through Projected Visual Abstractions

Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Y...

AI 聚合 07/27
arXiv

Hyperball May Not Be a Free Lunch

For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training b...

AI 聚合 07/27
arXiv

Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture

Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role ...

AI 聚合 07/27
arXiv

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio mo...

AI 聚合 07/27
arXiv

Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming Education

Generative AI is reshaping programming education, yet educators often infer students' AI-supported learning from clas...

AI 聚合 07/27
arXiv

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deploy...

AI 聚合 07/27
arXiv

Learning to Prepare Molecular Ground States with Transformer Models

Quantum state preparation is a key component of many quantum algorithms. Performing this step efficiently is essentia...

AI 聚合 07/27
arXiv

MineValiCoder: Reliable Code Generation with Test Case Quality Mining and Bipartite Graph-Based Mutual Validation

Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, exist...

AI 聚合 07/27
arXiv

\k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth Updating

Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposin...

AI 聚合 07/27
arXiv

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

Automating theoretical research is constrained not only by the generation of candidate results, but also by their rel...

AI 聚合 07/27
arXiv

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

Commercial large language models are increasingly used as knowledge references, yet their stance on contested scienti...

AI 聚合 07/27
arXiv

Quantum Spectral Model: Data Reuploading with Input-Conditioned Frequency Support

A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bia...

AI 聚合 07/27
arXiv

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents

Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this...

AI 聚合 07/27
arXiv

Explainable Reinforcement Learning for assisting Air Traffic Controllers

To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviat...

AI 聚合 07/27
arXiv

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capabil...

AI 聚合 07/27
HuggingFace

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing se...

AI 聚合 07/27
HuggingFace

SceneActBench: Can Agents Act on the 3D Scenes They See?

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existin...

AI 聚合 07/27
HuggingFace

Scaling Native Multimodal Pre-Training From Scratch

Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-trai...

AI 聚合 07/27
HuggingFace

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as...

AI 聚合 07/27
HuggingFace

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new ...

AI 聚合 07/27
HuggingFace

LAMAR: An Open Language-Aware Multilingual Alignment Reranker

In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple langu...

AI 聚合 07/27
HuggingFace

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and ...

AI 聚合 07/27
HuggingFace

IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation

Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years...

AI 聚合 07/27
HuggingFace

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot m...

AI 聚合 07/27
HuggingFace

Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making

Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-toke...

AI 聚合 07/27
HuggingFace

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unifie...

AI 聚合 07/27
HuggingFace

Three-Body Scattering for Generative Modeling

Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressi...

AI 聚合 07/27
HuggingFace

Spectral Prior for Reducing Exposure Bias in Diffusion Models

Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure...

AI 聚合 07/27
HuggingFace

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions t...

AI 聚合 07/27
HuggingFace

Self-Supervised Learning of Structured Dynamics from Videos

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two ...

AI 聚合 07/25
HuggingFace

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale...

AI 聚合 07/25
HuggingFace

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified a...

AI 聚合 07/25
HuggingFace

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multip...

AI 聚合 07/25
HuggingFace

Dataset Distillation by Influence Matching

We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-ste...

AI 聚合 07/25
HuggingFace

OpenForgeRL: Train Harness-native Agents in Any Environment

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn r...

AI 聚合 07/25
arXiv

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents ...

AI 聚合 07/24
arXiv

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challe...

AI 聚合 07/24
arXiv

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, ...

AI 聚合 07/24
arXiv

From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs

Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing...

AI 聚合 07/24
arXiv

The Boundaries of Automation: A Theory of Persistent Human Participation

The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with ...

AI 聚合 07/24
arXiv

MIRROR: Learning from the Other View for Multi-Modal Reasoning

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggl...

AI 聚合 07/24
arXiv

Visual Contrastive Self-Distillation

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation...

AI 聚合 07/24
arXiv

OpenForgeRL: Train Harness-native Agents in Any Environment

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn r...

AI 聚合 07/24
arXiv

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

Building socially calibrated large language models, which can learn from others without simply yielding to them, requ...

AI 聚合 07/24
arXiv

Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana

A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to id...

AI 聚合 07/24
arXiv

Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity

Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserv...

AI 聚合 07/24
arXiv

Synthetic data generation framework for quality control automation in gravure printing

Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manu...

AI 聚合 07/24
arXiv

Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$

Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence ...

AI 聚合 07/24
arXiv

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactio...

AI 聚合 07/24
arXiv

3D-Aware VLMs with Implicit and Explicit Geometries

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when h...

AI 聚合 07/24
HuggingFace

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely co...

AI 聚合 07/24
HuggingFace

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states...

AI 聚合 07/24
HuggingFace

TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation

The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-sca...

AI 聚合 07/24
HuggingFace

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We ...

AI 聚合 07/24
HuggingFace

Predictive Divergence Masks for LLM RL

Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-polic...

AI 聚合 07/24
HuggingFace

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its co...

AI 聚合 07/24
HuggingFace

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactio...

AI 聚合 07/24
HuggingFace

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is ...

AI 聚合 07/24
HuggingFace

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural ...

AI 聚合 07/24
HuggingFace

Robostral Navigate

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot e...

AI 聚合 07/24
HuggingFace

Visual Contrastive Self-Distillation

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation...

AI 聚合 07/24
HuggingFace

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural represen...

AI 聚合 07/24
HuggingFace

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the p...

AI 聚合 07/24
HuggingFace

LLMs Get Lost in Evolving User Intent

As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks t...

AI 聚合 07/24
HuggingFace

Color Pass-Through via Camera-Display Coupling

When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differ...

AI 聚合 07/24
HuggingFace

Sample-Efficient Learning from Agent Experience

Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming exp...

AI 聚合 07/24
HuggingFace

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answ...

AI 聚合 07/24
HuggingFace

SLAM in Low-Light Environments: Project Report

Simultaneous localization and mapping (SLAM) is one of the fundamental problems in robotics, as it enables autonomous...

AI 聚合 07/24
HuggingFace

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and...

AI 聚合 07/24
HuggingFace

Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices

Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural ...

AI 聚合 07/24
HuggingFace

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independentl...

AI 聚合 07/24
arXiv

The Ethics of Autonomous AI Agents for Offensive Security

LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- dete...

AI 聚合 07/23
arXiv

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating struc...

AI 聚合 07/23
arXiv

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires...

AI 聚合 07/23
arXiv

Self-supervision drives representational convergence in medical foundation models more than clinical supervision

Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scal...

AI 聚合 07/23
arXiv

Sound Probabilistic Safety Bounds for Large Language Models

We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) gener...

AI 聚合 07/23
arXiv

Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments

We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a tim...

AI 聚合 07/23
arXiv

Don't Trust the Label: License Laundering in AI Supply Chains

AI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and application...

AI 聚合 07/23
arXiv

Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition Dropout

RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always av...

AI 聚合 07/23
arXiv

Understanding Generative AI-mediated User Engagement with Academic Library Resources

This study empirically analyzed generative AI as an emerging discovery pathway to academic library resources. Utilizi...

AI 聚合 07/23
arXiv

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Visio...

AI 聚合 07/23
arXiv

Generative AI floods and dilutes the market for books

Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-qual...

AI 聚合 07/23
arXiv

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed fa...

AI 聚合 07/23
arXiv

FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization

Clinical biomarker workflows in translational research settings often rely on spreadsheet-driven tracking, manual qua...

AI 聚合 07/23
arXiv

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages des...

AI 聚合 07/23
arXiv

SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data

In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensi...

AI 聚合 07/23
HuggingFace

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represe...

AI 聚合 07/23
HuggingFace

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D ...

AI 聚合 07/23
HuggingFace

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requ...

AI 聚合 07/23
HuggingFace

G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data f...

AI 聚合 07/23
HuggingFace

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in e...

AI 聚合 07/23
HuggingFace

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges fo...

AI 聚合 07/23
HuggingFace

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

As large language models and AI agents become the primary consumers of search results, document set quality determine...

AI 聚合 07/23
HuggingFace

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the ...

AI 聚合 07/23
HuggingFace

Self Gradient Forcing: Native Long Video Extrapolation

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained ...

AI 聚合 07/23
HuggingFace

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization

Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL al...

AI 聚合 07/23
HuggingFace

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its ex...

AI 聚合 07/23
HuggingFace

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in hig...

AI 聚合 07/23
HuggingFace

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed fa...

AI 聚合 07/23
HuggingFace

An Exam for Active Observers

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapsh...

AI 聚合 07/23
HuggingFace

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypern...

AI 聚合 07/23
HuggingFace

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become cri...

AI 聚合 07/23
HuggingFace

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an u...

AI 聚合 07/23
HuggingFace

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike st...

AI 聚合 07/23
HuggingFace

AutoIndex: Learning Representation Programs for Retrieval

We present AutoIndex, a framework for learning representation programs: executable transformations that map raw docum...

AI 聚合 07/23
arXiv

GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models

The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. W...

AI 聚合 07/22
arXiv

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic m...

AI 聚合 07/22
arXiv

Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes

This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative...

AI 聚合 07/22
arXiv

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example ...

AI 聚合 07/22
arXiv

Riemannian Deep Learning:Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components re...

AI 聚合 07/22
arXiv

From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs

Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan sa...

AI 聚合 07/22
arXiv

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR ...

AI 聚合 07/22
arXiv

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the...

AI 聚合 07/22
arXiv

Associative Emotional Learning in Convolutional Neural Networks

Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence o...

AI 聚合 07/22
arXiv

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language mod...

AI 聚合 07/22
arXiv

Provable diffusion-based posterior sampling for linear inverse problems via DDIM

Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existin...

AI 聚合 07/22
arXiv

Agents in the Wild: Where Research Meets Deployment

Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinati...

AI 聚合 07/22
arXiv

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rat...

AI 聚合 07/22
arXiv

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Controllable image generation remains challenging for creative professionals, who often require precise regional cont...

AI 聚合 07/22
arXiv

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, ...

AI 聚合 07/22
HuggingFace

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Controllable image generation remains challenging for creative professionals, who often require precise regional cont...

AI 聚合 07/22
HuggingFace

Masked Visual Actions for Unified World Modeling

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them prom...

AI 聚合 07/22
HuggingFace

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where mode...

AI 聚合 07/22
HuggingFace

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-...

AI 聚合 07/22
HuggingFace

SciForma: Structure-Faithful Generation of Scientific Diagrams

Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams mu...

AI 聚合 07/22
HuggingFace

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but stale...

AI 聚合 07/22
HuggingFace

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically...

AI 聚合 07/22
HuggingFace

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation...

AI 聚合 07/22
HuggingFace

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the clai...

AI 聚合 07/22
HuggingFace

HPD-Parsing: Hierarchical Parallel Document Parsing

Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflecte...

AI 聚合 07/22
HuggingFace

Trajectory-aware Cross-view Geo-localization with Sequential Observations

Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods sh...

AI 聚合 07/22
HuggingFace

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction,...

AI 聚合 07/22
HuggingFace

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language mod...

AI 聚合 07/22
HuggingFace

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation dur...

AI 聚合 07/22
HuggingFace

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused i...

AI 聚合 07/22
HuggingFace

Generative World Renderer at the Speed of Play

Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesize...

AI 聚合 07/22
HuggingFace

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their p...

AI 聚合 07/22
HuggingFace

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B...

AI 聚合 07/22
HuggingFace

Delineate Anything v2: A Global Foundation Model for Field Delineation

Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chai...

AI 聚合 07/22
HuggingFace

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on ...

AI 聚合 07/22
HuggingFace

Can Multimodal Large Language Models Understand OCT?

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although...

AI 聚合 07/22
HuggingFace

ShotPlan: Cinematic Video Generation with Learnable Planning Token

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic...

AI 聚合 07/22
HuggingFace

Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy Distillation

Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. T...

AI 聚合 07/22
HuggingFace

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resource...

AI 聚合 07/22
HuggingFace

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial...

AI 聚合 07/22
HuggingFace

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous ...

AI 聚合 07/22
HuggingFace

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driv...

AI 聚合 07/22
HuggingFace

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand...

AI 聚合 07/22
HuggingFace

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the ma...

AI 聚合 07/22
HuggingFace

Nonuniformity Principle in Human-AI Coworking

As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvem...

AI 聚合 07/22
arXiv

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models p...

AI 聚合 07/21
arXiv

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an ...

AI 聚合 07/21
arXiv

SGA: Plug&Play Geometric Verification for Educational Video Synthesis

Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libra...

AI 聚合 07/21
arXiv

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, whi...

AI 聚合 07/21
arXiv

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial...

AI 聚合 07/21
arXiv

LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications

Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools ...

AI 聚合 07/21
arXiv

Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices

Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural ...

AI 聚合 07/21
arXiv

TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization

Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, buil...

AI 聚合 07/21
arXiv

OR Else: A Differentiable Trust Region for Policy Optimization

PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduc...

AI 聚合 07/21
arXiv

A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive Manufacturing

Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as opera...

AI 聚合 07/21
arXiv

Learning Adaptive Safety Margins for Visual Navigation

Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fix...

AI 聚合 07/21
arXiv

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer ...

AI 聚合 07/21
arXiv

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syll...

AI 聚合 07/21
arXiv

Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pi...

AI 聚合 07/21
arXiv

Automated Discovery Has No Universally Superior Harness

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However...

AI 聚合 07/21
HuggingFace

Group Entropy-Controlled Policy Optimization

Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping ...

AI 聚合 07/21
HuggingFace

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with ...

AI 聚合 07/21
HuggingFace

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. Whil...

AI 聚合 07/21
HuggingFace

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Pruning long context for coding agents has been a vital technology for efficient context management. While existing c...

AI 聚合 07/21
HuggingFace

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous mod...

AI 聚合 07/21
HuggingFace

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

Building assistants that can continually watch the world, remember what they see, and reason over their accumulated e...

AI 聚合 07/21
HuggingFace

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, ex...

AI 聚合 07/21
HuggingFace

Distilled Reinforcement Learning for LLM Post-training

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing me...

AI 聚合 07/21
HuggingFace

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a t...

AI 聚合 07/21
HuggingFace

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive li...

AI 聚合 07/21
HuggingFace

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the sup...

AI 聚合 07/21
HuggingFace

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physica...

AI 聚合 07/21
HuggingFace

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

In line with the prevailing direction of vision research, we explore the integration of both generation and editing c...

AI 聚合 07/21
HuggingFace

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?

Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromis...

AI 聚合 07/21
HuggingFace

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong...

AI 聚合 07/21
HuggingFace

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Predicting a football match before kickoff requires more than knowing past results: a model must use changing informa...

AI 聚合 07/21
HuggingFace

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies o...

AI 聚合 07/21
HuggingFace

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffer...

AI 聚合 07/21
HuggingFace

GigaChat Audio: Time-aware Large Audio Language Model

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio L...

AI 聚合 07/21
HuggingFace

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architec...

AI 聚合 07/21
arXiv

When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis

Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting th...

AI 聚合 07/21
arXiv

Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography

Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative bi...

AI 聚合 07/21
arXiv

LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization

Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation towa...

AI 聚合 07/21
arXiv

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embod...

AI 聚合 07/21
arXiv

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from inc...

AI 聚合 07/21
arXiv

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch ...

AI 聚合 07/21
arXiv

Understanding Reasoning from Pretraining to Post-Training

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, ...

AI 聚合 07/21
arXiv

Harmonizing AI Safety Thresholds

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third p...

AI 聚合 07/21
arXiv

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

Evaluations should do more than measure a models current performance. They should tell us what to fix for the next mo...

AI 聚合 07/21
arXiv

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance

AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, wh...

AI 聚合 07/21
arXiv

ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded e...

AI 聚合 07/21
arXiv

When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantag...

AI 聚合 07/21
arXiv

An Exam for Active Observers

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapsh...

AI 聚合 07/21
arXiv

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-traini...

AI 聚合 07/21
arXiv

Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities

Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, ...

AI 聚合 07/21
HuggingFace

Cura 1T: Specialized Model for Agentic Healthcare

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover...

AI 聚合 07/21
HuggingFace

Recursive Harness Self-Improvement

Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components w...

AI 聚合 07/21
HuggingFace

xHC: Expanded Hyper-Connections

Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory...

AI 聚合 07/21
HuggingFace

DSWorld: A Data Science World Model for Efficient Autonomous Agents

Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily r...

AI 聚合 07/21
HuggingFace

On-Policy Delta Distillation

On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constrain...

AI 聚合 07/21
HuggingFace

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse lang...

AI 聚合 07/21
HuggingFace

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, con...

AI 聚合 07/21
HuggingFace

Qwen-Music Technical Report

In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and hi...

AI 聚合 07/21
HuggingFace

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existin...

AI 聚合 07/21
HuggingFace

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI...

AI 聚合 07/21
HuggingFace

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-traini...

AI 聚合 07/21
HuggingFace

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy c...

AI 聚合 07/21
HuggingFace

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural know...

AI 聚合 07/21
HuggingFace

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grad...

AI 聚合 07/21
HuggingFace

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These a...

AI 聚合 07/21
HuggingFace

Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies

Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. Wh...

AI 聚合 07/21
HuggingFace

REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation

Training-free in-context segmentation enables new object categories to be introduced at inference time from a single ...

AI 聚合 07/21
HuggingFace

Benchmarking Sensor Robustness in Plasma Diagnostic Models: A Systematic Evaluation on TokaMark

Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. ...

AI 聚合 07/21
HuggingFace

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a ...

AI 聚合 07/21
HuggingFace

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing v...

AI 聚合 07/21
GitHub

[GitHub] OpenHands/OpenHands

🙌 OpenHands: AI-Driven Development(⭐81227)

AI 聚合 07/19
HuggingFace

On Locality and Length Generalization in Visual Reasoning

A striking feature of the human visual system is that it ingests visual information through a series of local foveate...

AI 聚合 07/18
HuggingFace

AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling

Recent generalizable 3D Gaussian Splatting models have advanced long-sequence novel view synthesis (NVS), but at the ...

AI 聚合 07/18
HuggingFace

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a sin...

AI 聚合 07/18
HuggingFace

Token Time Continuous Diffusion for Language Modeling

In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates ...

AI 聚合 07/18
HuggingFace

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We i...

AI 聚合 07/18
HuggingFace

Hierarchical Denoising For Multi-Step Visual Reasoning

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streami...

AI 聚合 07/18
HuggingFace

Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving

Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based test...

AI 聚合 07/18
HuggingFace

Rethinking the Evaluation of Harness Evolution for Agents

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit ...

AI 聚合 07/18
arXiv

Subjective Risk Decomposition: A New View for Uncertainty Quantification

We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axio...

AI 聚合 07/17
arXiv

Mask-Aware Policy Gradients for Diffusion Language Models

Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Mas...

AI 聚合 07/17
arXiv

Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation

Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems f...

AI 聚合 07/17
arXiv

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and ...

AI 聚合 07/17
arXiv

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misalign...

AI 聚合 07/17
arXiv

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically beni...

AI 聚合 07/17
arXiv

In-Place Tokenizer Expansion for Pre-trained LLMs

A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflect...

AI 聚合 07/17
arXiv

AutoSynthesis: An agentic system for automated meta-analysis

Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, ...

AI 聚合 07/17
arXiv

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of t...

AI 聚合 07/17
arXiv

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...

AI 聚合 07/17
arXiv

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing v...

AI 聚合 07/17
arXiv

SceneBind: Binding What and Where Across Vision, Audio and Language

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understandi...

AI 聚合 07/17
arXiv

Pretraining Data Can Be Poisoned through Computational Propaganda

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior wo...

AI 聚合 07/17
arXiv

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors r...

AI 聚合 07/17
arXiv

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-T...

AI 聚合 07/17
HuggingFace

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-T...

AI 聚合 07/17
HuggingFace

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn inte...

AI 聚合 07/17
HuggingFace

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without cha...

AI 聚合 07/17
HuggingFace

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them ...

AI 聚合 07/17
HuggingFace

BadWAM: When World-Action Models Dream Right but Act Wrong

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting action...

AI 聚合 07/17
HuggingFace

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...

AI 聚合 07/17
HuggingFace

WanSong v1.0 Technical Report

Music generation foundation models have recently attracted significant industry attention. However, achieving efficie...

AI 聚合 07/17
HuggingFace

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, ...

AI 聚合 07/17
HuggingFace

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly ...

AI 聚合 07/17
HuggingFace

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference imag...

AI 聚合 07/17
HuggingFace

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multip...

AI 聚合 07/17
HuggingFace

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field...

AI 聚合 07/17
HuggingFace

DeepLoop: Depth Scaling for Looped Transformers

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, ...

AI 聚合 07/17
HuggingFace

UniVR: Thinking in Visual Space for Unified Visual Reasoning

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduc...

AI 聚合 07/17
HuggingFace

From Pixels to States: Rethinking Interactive World Models as Game Engines

Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphic...

AI 聚合 07/17
HuggingFace

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a conte...

AI 聚合 07/17
HuggingFace

Spectral Rewiring for Exploration, Purification, and Model Merging

Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter...

AI 聚合 07/17
HuggingFace

GRASP: GRanularity-Aware Search Policy for Agentic RAG

Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, ge...

AI 聚合 07/17
HuggingFace

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual informati...

AI 聚合 07/17
HuggingFace

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-to...

AI 聚合 07/17
HuggingFace

Self-Improvements in Modern Agentic Systems: A Survey

Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is control...

AI 聚合 07/17
HuggingFace

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide li...

AI 聚合 07/17
HuggingFace

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard determin...

AI 聚合 07/17
HuggingFace

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) m...

AI 聚合 07/17
HuggingFace

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world...

AI 聚合 07/17
HuggingFace

SPEAR: A Simulator for Photorealistic Embodied AI Research

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, ...

AI 聚合 07/17
HuggingFace

Length Penalties Make Chain-of-Thought Less Monitorable

Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives ...

AI 聚合 07/17
HuggingFace

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their stri...

AI 聚合 07/17
arXiv

AIMO Interpretability Challenge

We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in fro...

AI 聚合 07/16
arXiv

Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings

Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $...

AI 聚合 07/16
arXiv

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their stri...

AI 聚合 07/16
arXiv

A Self-Evolving Agent for Longitudinal Personal Health Management

Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolati...

AI 聚合 07/16
arXiv

Music-to-Dance Generation via Atomic Movements

Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically co...

AI 聚合 07/16
arXiv

The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce

The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradi...

AI 聚合 07/16
arXiv

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and...

AI 聚合 07/16
arXiv

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, c...

AI 聚合 07/16
arXiv

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank s...

AI 聚合 07/16
arXiv

Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study

With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable ener...

AI 聚合 07/16
arXiv

Early Adoption of Agentic Coding Tools by GitHub Projects

Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, ...

AI 聚合 07/16
arXiv

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study

Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running sc...

AI 聚合 07/16
arXiv

AI-accelerated End-to-End Framework for Rapid Professional Upskilling

By 2030, 59 of every 100 workers will need reskilling or upskilling, yet the average time to close an enterprise skil...

AI 聚合 07/16
arXiv

Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education

This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational...

AI 聚合 07/16
arXiv

Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LL...

AI 聚合 07/16
HuggingFace

OvisOCR2 Technical Report

We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document ...

AI 聚合 07/16
HuggingFace

Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

The capability of a modern AI agent depends not only on its foundation model but also on its harness, which construct...

AI 聚合 07/16
HuggingFace

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platfo...

AI 聚合 07/16
HuggingFace

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerge...

AI 聚合 07/16
HuggingFace

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recogni...

AI 聚合 07/16
HuggingFace

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an in...

AI 聚合 07/16
HuggingFace

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, ...

AI 聚合 07/16
HuggingFace

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception o...

AI 聚合 07/16
HuggingFace

Tracing Agentic Failure from the Flow of Success

Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the t...

AI 聚合 07/16
HuggingFace

From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization

The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language mo...

AI 聚合 07/16
HuggingFace

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving ...

AI 聚合 07/16
HuggingFace

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D di...

AI 聚合 07/16
HuggingFace

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising...

AI 聚合 07/16
HuggingFace

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes...

AI 聚合 07/16
HuggingFace

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling too...

AI 聚合 07/16
HuggingFace

Registers Matter for Pixel-Space Diffusion Transformers

Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a pr...

AI 聚合 07/16
HuggingFace

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans ...

AI 聚合 07/16
HuggingFace

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

Starting from the utilization of deep neural networks to approximate the state-action value function that led to winn...

AI 聚合 07/16
HuggingFace

Towards Autonomous and Auditable Medical Imaging Model Development

Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, ...

AI 聚合 07/16
HuggingFace

MuScriptor: An Open Model for Multi-Instrument Music Transcription

Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on compl...

AI 聚合 07/16
HuggingFace

Let RGB Be the Language of Vision

This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natura...

AI 聚合 07/16
HuggingFace

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images with...

AI 聚合 07/16
HuggingFace

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbound...

AI 聚合 07/16
HuggingFace

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as Do...

AI 聚合 07/16
HuggingFace

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (C...

AI 聚合 07/16
HuggingFace

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right ...

AI 聚合 07/16
HuggingFace

Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering

Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their d...

AI 聚合 07/16
HuggingFace

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, ov...

AI 聚合 07/16
HuggingFace

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a def...

AI 聚合 07/16
arXiv

A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study

Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reache...

AI 聚合 07/15
arXiv

UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies

Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, g...

AI 聚合 07/15
arXiv

MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations

Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, mult...

AI 聚合 07/15
arXiv

Real-time fall detection based on vision for low-power edge platforms

Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches...

AI 聚合 07/15
arXiv

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each d...

AI 聚合 07/15
arXiv

ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark

Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interact...

AI 聚合 07/15
arXiv

Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models

Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still meas...

AI 聚合 07/15
arXiv

FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation

Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLM...

AI 聚合 07/15
arXiv

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user...

AI 聚合 07/15
arXiv

Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation

Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged e...

AI 聚合 07/15
arXiv

Dynamic Resource Allocation for Ensemble Determinization MCTS

Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games w...

AI 聚合 07/15
arXiv

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a ...

AI 聚合 07/15
arXiv

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling too...

AI 聚合 07/15
arXiv

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scal...

AI 聚合 07/15
arXiv

Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution

Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they ra...

AI 聚合 07/15
HuggingFace

Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution

LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone...

AI 聚合 07/15
HuggingFace

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-she...

AI 聚合 07/15
HuggingFace

Latent-Identity Tuning in Text-to-Image Personalization Models

Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a ...

AI 聚合 07/15
HuggingFace

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand syst...

AI 聚合 07/15
HuggingFace

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct...

AI 聚合 07/15
HuggingFace

A Theory of Contrastive Learning with Natural Images

Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks?...

AI 聚合 07/15
HuggingFace

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet ...

AI 聚合 07/15
HuggingFace

Evidence-Backed Video Question Answering

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes,...

AI 聚合 07/15
HuggingFace

Multi-Agent LLMs Fail to Explore Each Other

Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language ...

AI 聚合 07/15
arXiv

Active Offline-to-Online Reinforcement Learning

Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously colle...

AI 聚合 07/14
arXiv

Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories

Prefabricated prefinished volumetric construction moves most building work into module factories, whose production fl...

AI 聚合 07/14
arXiv

Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement

Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes h...

AI 聚合 07/14
arXiv

Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal

Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regula...

AI 聚合 07/14
arXiv

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, ...

AI 聚合 07/14
arXiv

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a ...

AI 聚合 07/14
arXiv

Introducing Human-Centeredness in AI-Assisted Lexicography

This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While gen...

AI 聚合 07/14
arXiv

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The fram...

AI 聚合 07/14
arXiv

Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search

Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive ...

AI 聚合 07/14
arXiv

LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments

This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity r...

AI 聚合 07/14
arXiv

Evidence-Backed Video Question Answering

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes,...

AI 聚合 07/14
arXiv

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, meas...

AI 聚合 07/14
arXiv

A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kin...

AI 聚合 07/14
arXiv

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language ...

AI 聚合 07/14
arXiv

Metacognition in LLMs: Foundations, Progress, and Opportunities

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-m...

AI 聚合 07/14
HuggingFace

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad ve...

AI 聚合 07/14
HuggingFace

NeuroCogMap Reveals Cognitive Organization of Large Language Models

Understanding how complex cognitive functions are organized within artificial systems is central to interpreting larg...

AI 聚合 07/14
HuggingFace

LightMem-Ego: Your AI Memory for Everyday Life

Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and aud...

AI 聚合 07/14
HuggingFace

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet ...

AI 聚合 07/14
HuggingFace

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents s...

AI 聚合 07/14
HuggingFace

Weak-to-Strong Generalization via Direct On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, bu...

AI 聚合 07/14
HuggingFace

LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow

Flow matching over carefully designed latent representations has recently emerged as a powerful paradigm for topology...

AI 聚合 07/14
HuggingFace

Motion4Motion: Motion Transfer Across Subjects at Inference

This work explores the motion transfer from one video to another, which is crucial in animation for diverse character...

AI 聚合 07/14
HuggingFace

4D Human-Scene Reconstruction from Low-Overlap Captures

Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in...

AI 聚合 07/14
HuggingFace

Metacognition in LLMs: Foundations, Progress, and Opportunities

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-m...

AI 聚合 07/14
HuggingFace

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet m...

AI 聚合 07/14
HuggingFace

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existin...

AI 聚合 07/14
HuggingFace

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet th...

AI 聚合 07/14
arXiv

Large-Scale Portfolio Optimization Problem Under Cardinality Constraint With Enhanced Multi-Objective Evolutionary Algorithms

Decision-making is posing an increasingly formidable challenge to investors because of the growing number of alternat...

AI 聚合 07/13
arXiv

Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph Approach

We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight ty...

AI 聚合 07/13
arXiv

Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining

Pre-demolition assessment, the regulated audit process at the heart of urban mining, is an information process in whi...

AI 聚合 07/13
arXiv

TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems

The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of gen...

AI 聚合 07/13
arXiv

PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers

Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force...

AI 聚合 07/13
arXiv

Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse...

AI 聚合 07/13
arXiv

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Questi...

AI 聚合 07/13
arXiv

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout...

AI 聚合 07/13
arXiv

Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory

Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information process...

AI 聚合 07/13
arXiv

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to ...

AI 聚合 07/13
arXiv

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concep...

AI 聚合 07/13
arXiv

VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecu...

AI 聚合 07/13
arXiv

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluati...

AI 聚合 07/13
arXiv

Scalable Visual Pretraining for Language Intelligence

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpor...

AI 聚合 07/13
arXiv

PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis

Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical mom...

AI 聚合 07/13
HuggingFace

Scalable Visual Pretraining for Language Intelligence

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpor...

AI 聚合 07/13
HuggingFace

A Sovereign, Open-Source Foundation Model for German and English

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation mod...

AI 聚合 07/13
HuggingFace

Video Generation Models are General-Purpose Vision Learners

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. Wh...

AI 聚合 07/13
HuggingFace

Self-Guided Test-Time Training for Long-Context LLMs

Long-context processing has become increasingly important for large language models (LLMs), but simply extending the ...

AI 聚合 07/13
HuggingFace

Trust Region Policy Distillation

Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Di...

AI 聚合 07/13
HuggingFace

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to...

AI 聚合 07/13
HuggingFace

PanoWorld: Real-World Panoramic Generation

In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotati...

AI 聚合 07/13
HuggingFace

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining lea...

AI 聚合 07/13
HuggingFace

Phone Segmentation and Recognition through Phonological Activation Mapping

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separatel...

AI 聚合 07/13
HuggingFace

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benc...

AI 聚合 07/13
HuggingFace

KronQ: LLM Quantization via Kronecker-Factored Hessian

Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without r...

AI 聚合 07/13
HuggingFace

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitizati...

AI 聚合 07/13
HuggingFace

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation

Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks pr...

AI 聚合 07/13
HuggingFace

A Sparse and Truncated State Vector Simulator for Peaked Circuits

In a class of quantum circuits known as peaked circuits, the goal is to predict the most probable bit string at the o...

AI 聚合 07/11
HuggingFace

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approach...

AI 聚合 07/11
HuggingFace

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

We introduce PAST-TIDE, our stance detection system addressing both subtasks of the StanceNakba Shared Task at NakbaN...

AI 聚合 07/11
HuggingFace

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action ag...

AI 聚合 07/11
arXiv

WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question ...

AI 聚合 07/10
arXiv

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment o...

AI 聚合 07/10
arXiv

A Practical Investigation of Training-free Relaxed Speculative Decoding

Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft token...

AI 聚合 07/10
arXiv

ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation

Repository-level code generation requires implementing target functions while accounting for complex cross-file depen...

AI 聚合 07/10
arXiv

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action ag...

AI 聚合 07/10
arXiv

Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction

Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and...

AI 聚合 07/10
arXiv

Validity of LLMs as data annotators: AMALIA on authority

A national language model offers a linguistic community its own instrument for measuring what its citizens say and va...

AI 聚合 07/10
arXiv

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its e...

AI 聚合 07/10
arXiv

Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows

Large language model (LLM) applications increasingly use explicit workflows for tool use, retrieval, branching, check...

AI 聚合 07/10
arXiv

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved ...

AI 聚合 07/10
arXiv

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embed...

AI 聚合 07/10
arXiv

Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis

In this study, we present a large-scale descriptive analysis of the use of an AI-based learning assistant (Syntea) in...

AI 聚合 07/10
arXiv

SLORR: Simple and Efficient In-Training Low-Rank Regularization

Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable...

AI 聚合 07/10
arXiv

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pie...

AI 聚合 07/10
arXiv

OpenCoF: Learning to Reason Through Video Generation

Reasoning has become a core capability for large models, especially when reliable decisions require understanding log...

AI 聚合 07/10
HuggingFace

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffu...

AI 聚合 07/10
HuggingFace

OpenCoF: Learning to Reason Through Video Generation

Reasoning has become a core capability for large models, especially when reliable decisions require understanding log...

AI 聚合 07/10
HuggingFace

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pie...

AI 聚合 07/10
HuggingFace

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE

Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository...

AI 聚合 07/10
HuggingFace

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, ...

AI 聚合 07/10
HuggingFace

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

Inference-time scaling for text-to-image generation has progressed from simple Best-of-N (BoN) sampling to guided sea...

AI 聚合 07/10
HuggingFace

Vidu S1: A Real-Time Interactive Video Generation Model

We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. ...

AI 聚合 07/10
HuggingFace

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of larg...

AI 聚合 07/10
HuggingFace

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with ad...

AI 聚合 07/10
HuggingFace

PhyMRI-SR: Toward Physics-Aware MRI Image Super-Resolution

Magnetic resonance imaging (MRI) super-resolution is vital for improving diagnostic accessibility, yet most methods t...

AI 聚合 07/10
HuggingFace

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects...

AI 聚合 07/10
HuggingFace

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLM...

AI 聚合 07/10
HuggingFace

Video-Oasis: Rethinking Evaluation of Video Understanding

The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performanc...

AI 聚合 07/10
HuggingFace

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

The rapid development of large language models and multimodal large language models has accelerated the emergence of ...

AI 聚合 07/10
HuggingFace

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of p...

AI 聚合 07/10
HuggingFace

A Quantized Native Runtime for On-Device Semantic Audio Generation

Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather th...

AI 聚合 07/10
HuggingFace

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-...

AI 聚合 07/10
HuggingFace

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures...

AI 聚合 07/10
HuggingFace

DrugGen 2: A disease-aware language model for enhancing drug discovery

Current computational approaches for drug design typically focus on generating molecules conditioned on specific targ...

AI 聚合 07/10
HuggingFace

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length ...

AI 聚合 07/10
HuggingFace

Automating the Design of Embodied Agent Architectures

Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules...

AI 聚合 07/10
HuggingFace

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does n...

AI 聚合 07/10
HuggingFace

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limit...

AI 聚合 07/10
HuggingFace

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previo...

AI 聚合 07/10
HuggingFace

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and complian...

AI 聚合 07/10
HuggingFace

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce ...

AI 聚合 07/10
HuggingFace

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

Accurate breast cancer classification from mammography requires effective integration of complementary information fr...

AI 聚合 07/10
HuggingFace

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but ...

AI 聚合 07/10
HuggingFace

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spat...

AI 聚合 07/10
HuggingFace

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often d...

AI 聚合 07/10
arXiv

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators th...

AI 聚合 07/09
arXiv

Towards Agentic AI Governance: A Preliminary Assessment

Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning an...

AI 聚合 07/09
arXiv

Future Confidence Distillation in Large Language Models

Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, w...

AI 聚合 07/09
arXiv

QCNN with Rough Path Signature Kernels

Time series analysis plays a vital role across a wide range of scientific and engineering domains but poses substanti...

AI 聚合 07/09
arXiv

ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation

Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on lo...

AI 聚合 07/09
arXiv

RL Post-Training Builds Compositional Reasoning Strategies

Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive ski...

AI 聚合 07/09
arXiv

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses du...

AI 聚合 07/09
arXiv

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-le...

AI 聚合 07/09
arXiv

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operationa...

AI 聚合 07/09
arXiv

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grad...

AI 聚合 07/09
arXiv

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models w...

AI 聚合 07/09
arXiv

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hol...

AI 聚合 07/09
arXiv

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access...

AI 聚合 07/09
arXiv

Co-LMLM: Continuous-Query Limited Memory Language Models

Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rat...

AI 聚合 07/09
arXiv

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reacti...

AI 聚合 07/09
HuggingFace

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their prima...

AI 聚合 07/09
HuggingFace

Infinite Worlds with Versatile Interactions

We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring...

AI 聚合 07/09
HuggingFace

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilom...

AI 聚合 07/09
HuggingFace

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovi...

AI 聚合 07/09
HuggingFace

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematicall...

AI 聚合 07/09
HuggingFace

Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo

Large Language Models (LLMs) unlocked new possibilities in automated code writing, becoming the backbone of most code...

AI 聚合 07/09
HuggingFace

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reacti...

AI 聚合 07/09
HuggingFace

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

Complex image creation and editing often require more than a single generation or editing model. A user request may i...

AI 聚合 07/09
HuggingFace

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding ...

AI 聚合 07/09
HuggingFace

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies documen...

AI 聚合 07/09
HuggingFace

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer visio...

AI 聚合 07/09
HuggingFace

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without ...

AI 聚合 07/09
HuggingFace

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation re...

AI 聚合 07/09
HuggingFace

Attending to Multimodal Generation One Token at a Time

Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic infor...

AI 聚合 07/09
HuggingFace

RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules

We present RuleChef, a framework that uses large language models (LLMs) to generate executable rules for NLP tasks su...

AI 聚合 07/09
HuggingFace

SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

Geometry-conditioned 3D scene generation enables the creation of 3D environments from user-provided geometry, offerin...

AI 聚合 07/09
HuggingFace

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases ...

AI 聚合 07/09
HuggingFace

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little ...

AI 聚合 07/09
HuggingFace

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-sp...

AI 聚合 07/09
HuggingFace

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, th...

AI 聚合 07/09
HuggingFace

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

JD.com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchan...

AI 聚合 07/09
arXiv

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical ...

AI 聚合 07/08
arXiv

Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms

The dairy industry in Ireland has a large potential for the integration of renewable energy and the reduction of carb...

AI 聚合 07/08
arXiv

Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine

Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the groundin...

AI 聚合 07/08
arXiv

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail,...

AI 聚合 07/08
arXiv

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which...

AI 聚合 07/08
arXiv

Industry Classification of GitHub Repositories Using the North American Industry Classification System (NAICS)

GitHub hosts hundreds of millions of public repositories, but the platform exposes no native mapping from repositorie...

AI 聚合 07/08
arXiv

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

We present FootsiesGym, an open-source environment for learning in a non-trivial two-player, zero-sum, imperfect-info...

AI 聚合 07/08
arXiv

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive comp...

AI 聚合 07/08
arXiv

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment

Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tas...

AI 聚合 07/08
arXiv

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store ...

AI 聚合 07/08
arXiv

RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation

Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelin...

AI 聚合 07/08
arXiv

The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology

- Objective: Multimodal deep learning models in oncology are currently limited by monolithic designs that rigidly cou...

AI 聚合 07/08
arXiv

Rethinking Indic AI from a Lens of Cultural Heritage Preservation

As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant i...

AI 聚合 07/08
arXiv

Graph Convolutional Attention: A Spectral Perspective on Graph Denoising and Diffusion

Denoising graphs is a fundamental problem in graph learning and the core operation of graph diffusion models. Attenti...

AI 聚合 07/08
arXiv

ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone,...

AI 聚合 07/08
HuggingFace

TREK: Distill to Explore, Reinforce to Refine

Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning traje...

AI 聚合 07/08
HuggingFace

Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models

Late-interaction retrieval models that use the MaxSim similarity function have shown strong empirical performance, of...

AI 聚合 07/08
HuggingFace

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level...

AI 聚合 07/08
HuggingFace

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods re...

AI 聚合 07/08
HuggingFace

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game d...

AI 聚合 07/08
HuggingFace

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target veri...

AI 聚合 07/08
HuggingFace

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor l...

AI 聚合 07/08
HuggingFace

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic po...

AI 聚合 07/08
HuggingFace

CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation

Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of indivi...

AI 聚合 07/08
HuggingFace

Vision as Unified Multimodal Generation

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the ...

AI 聚合 07/08
HuggingFace

MentalThink: Shaping Thoughts in Mental SVG World

We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable...

AI 聚合 07/08
HuggingFace

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve gener...

AI 聚合 07/08
HuggingFace

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories...

AI 聚合 07/08
HuggingFace

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions...

AI 聚合 07/08
HuggingFace

When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers

LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management polici...

AI 聚合 07/08
HuggingFace

Gemma 4 Technical Report

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family....

AI 聚合 07/08
HuggingFace

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), ye...

AI 聚合 07/08
HuggingFace

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destin...

AI 聚合 07/08
HuggingFace

From Foundation to Application: Improving VLA Models in Practice

Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applicat...

AI 聚合 07/08
HuggingFace

Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing

Academic output is produced across a fragmented toolchain: literature discovery in one application, reference managem...

AI 聚合 07/08
HuggingFace

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes

We present the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes. The task is compose...

AI 聚合 07/08
HuggingFace

MANCE: Manifold Aware Concept Erasure

Concept erasure aims to remove a target concept from a representation while preserving the other information encoded ...

AI 聚合 07/08
HuggingFace

Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports

In longitudinal clinical practice, every chest X-ray is read in the context of the patients prior exam, and much of w...

AI 聚合 07/08
HuggingFace

Taste-aware music retrieval from audio embeddings

Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely a...

AI 聚合 07/08
HuggingFace

Unified Audio Intelligence Without Regressing on Text Intelligence

Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we in...

AI 聚合 07/08
HuggingFace

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained ...

AI 聚合 07/08
HuggingFace

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a st...

AI 聚合 07/08
HuggingFace

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively...

AI 聚合 07/08
HuggingFace

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models

Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their ge...

AI 聚合 07/08
HuggingFace

Learning to Trigger: Reinforcement Learning at the Large Hadron Collider

High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (triggeri...

AI 聚合 07/08
HuggingFace

SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottlenec...

AI 聚合 07/08
HuggingFace

ACID: Action Consistency via Inverse Dynamics for Planning with World Models

Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. Howev...

AI 聚合 07/08
HuggingFace

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems ...

AI 聚合 07/08
arXiv

OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement

We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research pr...

AI 聚合 07/07
arXiv

Multiplayer Interactive World Models with Representation Autoencoders

We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interacti...

AI 聚合 07/07
arXiv

Selective Disclosure Watermarking for Large Language Models

Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). ...

AI 聚合 07/07
arXiv

Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning

Planning under uncertainty in continuous domains is essential for autonomous systems, yet computationally demanding. ...

AI 聚合 07/07
arXiv

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediate...

AI 聚合 07/07
arXiv

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without f...

AI 聚合 07/07
arXiv

SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models

Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, stan...

AI 聚合 07/07
arXiv

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems ...

AI 聚合 07/07
arXiv

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle ...

AI 聚合 07/07
arXiv

What Does a Discrete Diffusion Model Learn?

What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of...

AI 聚合 07/07
arXiv

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbound...

AI 聚合 07/07
arXiv

LLM-as-a-Verifier: A General-Purpose Verification Framework

Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabi...

AI 聚合 07/07
arXiv

Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification

Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated...

AI 聚合 07/07
arXiv

Weak-to-Strong Generalization via Direct On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, bu...

AI 聚合 07/07
arXiv

From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model

Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositi...

AI 聚合 07/07
HuggingFace

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-rela...

AI 聚合 07/07
HuggingFace

Vision Pretraining for Dense Spatial Perception

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structu...

AI 聚合 07/07
HuggingFace

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task ex...

AI 聚合 07/07
HuggingFace

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstructio...

AI 聚合 07/07
HuggingFace

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling d...

AI 聚合 07/07
HuggingFace

PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation

Semi-supervised semantic segmentation (SSSS) has long turned on one question, which pseudo-labels to trust, and answe...

AI 聚合 07/07
HuggingFace

Multiplayer Interactive World Models with Representation Autoencoders

We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interacti...

AI 聚合 07/07
HuggingFace

LLM-as-a-Verifier: A General-Purpose Verification Framework

Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabi...

AI 聚合 07/07
HuggingFace

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulati...

AI 聚合 07/07
HuggingFace

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from r...

AI 聚合 07/07
HuggingFace

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and th...

AI 聚合 07/07
HuggingFace

GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving

Increasingly, LLM inference services proxy client requests to engine replicas distributed globally. Load-balancing po...

AI 聚合 07/07
HuggingFace

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel...

AI 聚合 07/07
HuggingFace

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by c...

AI 聚合 07/07
HuggingFace

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interac...

AI 聚合 07/07
HuggingFace

Perceptual Flow Matching for Few-Step Generative Modeling

We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching ...

AI 聚合 07/07
HuggingFace

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but th...

AI 聚合 07/07
HuggingFace

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregres...

AI 聚合 07/07
HuggingFace

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly ...

AI 聚合 07/07
HuggingFace

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

Scientific literature search often requires more than retrieving papers from a single query: users' intents are under...

AI 聚合 07/07
HuggingFace

AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer

Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remai...

AI 聚合 07/07
HuggingFace

Generated Contents Enrichment

We study Generated Contents Enrichment (GCE), a conditional image-generation task in which a sparse scene description...

AI 聚合 07/07
HuggingFace

Measuring the Gap Between Human and LLM Research Ideas

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by no...

AI 聚合 07/07
HuggingFace

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care

Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractiv...

AI 聚合 07/07
HuggingFace

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deploym...

AI 聚合 07/06
HuggingFace

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training...

AI 聚合 07/06
HuggingFace

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation

GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referr...

AI 聚合 07/06
HuggingFace

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context...

AI 聚合 07/06
HuggingFace

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To re...

AI 聚合 07/06
HuggingFace

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to eviden...

AI 聚合 07/06
HuggingFace

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and ...

AI 聚合 07/06
HuggingFace

Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment

Cloud removal (CR) is essential for optical remote sensing, serving as a prerequisite for reliable downstream interpr...

AI 聚合 07/06
HuggingFace

DataComp-VLM: Improved Open Datasets for Vision-Language Models

Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the ...

AI 聚合 07/06
HuggingFace

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations....

AI 聚合 07/04
HuggingFace

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span ...

AI 聚合 07/04
HuggingFace

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over mult...

AI 聚合 07/04
HuggingFace

AutoMem: Automated Learning of Memory as a Cognitive Skill

Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capac...

AI 聚合 07/04
HuggingFace

WARP: Weight-Space Analysis for Recovering Training Data Portfolios

Foundation models are routinely released to the public, yet the data recipes used to train them -- such as domain mix...

AI 聚合 07/04
HuggingFace

Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting

Traffic matrices (TMs) capture network-wide origin-destination demand and are central to traffic engineering, yet acc...

AI 聚合 07/04
HuggingFace

Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions

Grid-based approaches to approximate nearest neighbor (ANN) search have been absent from modern scaling analyses. We ...

AI 聚合 07/04
arXiv

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and ...

AI 聚合 07/03
arXiv

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triple...

AI 聚合 07/03
arXiv

Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting

Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money pred...

AI 聚合 07/03
arXiv

TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution

Software tests and code evolve together: a code change should be followed by new or updated tests that record the new...

AI 聚合 07/03
arXiv

Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning

Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing...

AI 聚合 07/03
arXiv

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to...

AI 聚合 07/03
arXiv

Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials

Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts...

AI 聚合 07/03
arXiv

DemoPSD: Disagreement-Modulated Policy Self-Distillation

On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to rea...

AI 聚合 07/03
arXiv

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex s...

AI 聚合 07/03
arXiv

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

LLM agents will increasingly act in socially structured settings where role, audience, and relational context can sha...

AI 聚合 07/03
arXiv

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs...

AI 聚合 07/03
arXiv

Online Safety Monitoring for LLMs

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs onl...

AI 聚合 07/03
arXiv

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repa...

AI 聚合 07/03
arXiv

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need ...

AI 聚合 07/03
arXiv

Distributed Attacks in Persistent-State AI Control

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting acr...

AI 聚合 07/03
HuggingFace

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amoun...

AI 聚合 07/03
HuggingFace

Optimizing Visual Generative Models via Distribution-wise Rewards

Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, y...

AI 聚合 07/03
HuggingFace

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repa...

AI 聚合 07/03
HuggingFace

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, c...

AI 聚合 07/03
HuggingFace

Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Re...

AI 聚合 07/03
HuggingFace

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object me...

AI 聚合 07/03
HuggingFace

Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling

Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature cach...

AI 聚合 07/03
HuggingFace

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triple...

AI 聚合 07/03
HuggingFace

Morphing into Hybrid Attention Models

Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and repla...

AI 聚合 07/03
HuggingFace

Representation Distribution Matching for One-Step Visual Generation

We elucidate the design space of Representation Distribution Matching (RDM), our name for the paradigm that trains a ...

AI 聚合 07/03
HuggingFace

PACE: A Proxy for Agentic Capability Evaluation

Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex in...

AI 聚合 07/03
HuggingFace

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts...

AI 聚合 07/03
HuggingFace

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking task...

AI 聚合 07/03
HuggingFace

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens...

AI 聚合 07/03
HuggingFace

From SRA to Self-Flow: Data Augmentation or Self-Supervision?

Representation alignment has become an effective way to accelerate diffusion transformer training and improve generat...

AI 聚合 07/03
HuggingFace

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest con...

AI 聚合 07/03
HuggingFace

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations...

AI 聚合 07/03
HuggingFace

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-train...

AI 聚合 07/03
HuggingFace

InstanceControl: Controllable Complex Image Generation without Instance Labeling

Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visua...

AI 聚合 07/03
HuggingFace

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain r...

AI 聚合 07/03
HuggingFace

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leadi...

AI 聚合 07/03
HuggingFace

CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion

Blind image deblurring demands the recovery of high-fidelity details and coherent structures from complex, unknown de...

AI 聚合 07/03
HuggingFace

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they stil...

AI 聚合 07/03
HuggingFace

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of syntheti...

AI 聚合 07/03
HuggingFace

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents b...

AI 聚合 07/03
HuggingFace

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is esse...

AI 聚合 07/03
HuggingFace

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumul...

AI 聚合 07/03
HuggingFace

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning syst...

AI 聚合 07/03
HuggingFace

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity

Three of the most popular methods for training language models to reason look like three different tricks. They are n...

AI 聚合 07/03
HuggingFace

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be e...

AI 聚合 07/03
arXiv

Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search

While generative models have enabled training-free reward alignment, current methods typically excel in local explora...

AI 聚合 07/02
arXiv

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: w...

AI 聚合 07/02
arXiv

Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-order...

AI 聚合 07/02
arXiv

Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined succ...

AI 聚合 07/02
arXiv

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing th...

AI 聚合 07/02
arXiv

World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from mono...

AI 聚合 07/02
arXiv

GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics

This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) ...

AI 聚合 07/02
arXiv

Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering...

AI 聚合 07/02
arXiv

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents b...

AI 聚合 07/02
arXiv

FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model

Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduc...

AI 聚合 07/02
arXiv

The State-Prediction Separation Hypothesis

Transformers use the same forward computation stream to both predict the next token and store useful state for future...

AI 聚合 07/02
arXiv

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the pr...

AI 聚合 07/02
arXiv

AutoMem: Automated Learning of Memory as a Cognitive Skill

Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capac...

AI 聚合 07/02
arXiv

Language-Critique Imitation Learning from Suboptimal Demonstrations

Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals su...

AI 聚合 07/02
arXiv

Measuring the Gap Between Human and LLM Research Ideas

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by no...

AI 聚合 07/02
HuggingFace

Valdi: Value Diffusion World Models

World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enoug...

AI 聚合 07/02
HuggingFace

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied lear...

AI 聚合 07/02
HuggingFace

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as ...

AI 聚合 07/02
HuggingFace

CausalMix: Data Mixture as Causal Inference for Language Model Training

In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent met...

AI 聚合 07/02
HuggingFace

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents becom...

AI 聚合 07/02
HuggingFace

TurboServe: Serving Streaming Video Generation Efficiently and Economically

Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions tha...

AI 聚合 07/02
HuggingFace

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visu...

AI 聚合 07/02
HuggingFace

The State-Prediction Separation Hypothesis

Transformers use the same forward computation stream to both predict the next token and store useful state for future...

AI 聚合 07/02
HuggingFace

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e.,...

AI 聚合 07/02
HuggingFace

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistant...

AI 聚合 07/02
HuggingFace

ASPIRE: Agentic /Skills Discovery for Robotics

Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical cont...

AI 聚合 07/02
HuggingFace

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approac...

AI 聚合 07/02
HuggingFace

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing...

AI 聚合 07/02
HuggingFace

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual...

AI 聚合 07/02
HuggingFace

NoPA: Non-Parametric Online 3D Scene Graph Generation

Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environ...

AI 聚合 07/02
HuggingFace

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmar...

AI 聚合 07/02
HuggingFace

Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle w...

AI 聚合 07/02
HuggingFace

Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks

Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Thing...

AI 聚合 07/02
HuggingFace

AI translation of literary texts is "fine", but readers still prefer human translations

AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know...

AI 聚合 07/02
HuggingFace

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-...

AI 聚合 07/02
HuggingFace

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce...

AI 聚合 07/02
HuggingFace

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame r...

AI 聚合 07/02
HuggingFace

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of action...

AI 聚合 07/02
HuggingFace

Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature

The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains...

AI 聚合 07/02
HuggingFace

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet t...

AI 聚合 07/02
HuggingFace

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities...

AI 聚合 07/02
HuggingFace

Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents

Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such ev...

AI 聚合 07/02
HuggingFace

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models

Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on roboti...

AI 聚合 07/02
HuggingFace

Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surf...

AI 聚合 07/02
HuggingFace

Hierarchical Experimentalist Agents

Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-makin...

AI 聚合 07/02
HuggingFace

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assemb...

AI 聚合 07/02
HuggingFace

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven softwar...

AI 聚合 07/02
HuggingFace

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edit...

AI 聚合 07/02
HuggingFace

SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE

We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos ...

AI 聚合 07/02
arXiv

LUNA: Learning Universal 3D Human Animation Beyond Skinning

Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skin...

AI 聚合 07/01
arXiv

GR2 Technical Report

Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage rank...

AI 聚合 07/01
arXiv

Amplifying Membership Signal Through Chained Regeneration

The tendency of large generative models to memorize training data makes sample verification critical for privacy audi...

AI 聚合 07/01
arXiv

Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization

Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case st...

AI 聚合 07/01
arXiv

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions...

AI 聚合 07/01
arXiv

PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines

Policy-grounded document review requires determining whether a target document complies with organization-specific po...

AI 聚合 07/01
arXiv

AxDafny: Agentic Verified Code Generation in Dafny

We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts f...

AI 聚合 07/01
arXiv

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edit...

AI 聚合 07/01
arXiv

FLORA: A deep learning approach to predict forest attributes from heterogeneous LiDAR data

Forest attributes are essential for national-scale resource monitoring. Airborne LiDAR metrics are among the auxiliar...

AI 聚合 07/01
arXiv

AdaJEPA: An Adaptive Latent World Model

Latent world models enable planning from high-dimensional observations by predicting future states in a compact laten...

AI 聚合 07/01
arXiv

Freeform Preference Learning for Robotic Manipulation

Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipu...

AI 聚合 07/01
arXiv

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e.,...

AI 聚合 07/01
arXiv

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own co...

AI 聚合 07/01
arXiv

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of action...

AI 聚合 07/01
arXiv

Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision

When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, ...

AI 聚合 07/01
HuggingFace

DOPD: Dual On-policy Distillation

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...

AI 聚合 07/01
HuggingFace

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Lar...

AI 聚合 07/01
HuggingFace

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality soundi...

AI 聚合 07/01
HuggingFace

Orca: The World is in Your Mind

We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent s...

AI 聚合 07/01
HuggingFace

MemLearner: Learning to Query Context memory for Video World Models

Video World Models are interactive video generation models that predict future world states based on user actions and...

AI 聚合 07/01
HuggingFace

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their in...

AI 聚合 07/01
HuggingFace

Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real appli...

AI 聚合 07/01
HuggingFace

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in paralle...

AI 聚合 07/01
HuggingFace

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged a...

AI 聚合 07/01
HuggingFace

PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising

Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms...

AI 聚合 07/01
HuggingFace

TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive

Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unli...

AI 聚合 07/01
HuggingFace

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own co...

AI 聚合 07/01
HuggingFace

Multi-Block Diffusion Language Models

Block Diffusion Language Models (BD-LMs) improve diffusion-based text generation with KV caching and flexible-length ...

AI 聚合 07/01
HuggingFace

Little Brains, Big Feats: Exploring Compact Language Models

While large language models have been dominating the research landscape recently, small language models remain highly...

AI 聚合 07/01
HuggingFace

GEAR: Guided End-to-End AutoRegression for Image Synthesis

Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and the...

AI 聚合 07/01
HuggingFace

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fai...

AI 聚合 07/01
HuggingFace

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction m...

AI 聚合 07/01
HuggingFace

MuSViT: A Foundation Vision Model for Sheet Music Representation

Foundation models have transformed vision and language processing by providing rich, reusable representations that tr...

AI 聚合 07/01
HuggingFace

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation

Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneou...

AI 聚合 07/01
HuggingFace

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

Agent skills extend language-model agents with task-specific procedures, scripts, and references, but the tasks and e...

AI 聚合 07/01
HuggingFace

Beyond IID: How General Are Tabular Foundation Models, Really?

Foundation models for predictive machine learning on tabular data have recently gained significant traction in academ...

AI 聚合 07/01
HuggingFace

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

The advancement of generative AI models capable of producing text and image marks a critical step forward in the real...

AI 聚合 07/01
HuggingFace

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful seman...

AI 聚合 07/01
HuggingFace

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multipl...

AI 聚合 07/01
HuggingFace

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations...

AI 聚合 07/01
HuggingFace

SAM2Matting: Generalized Image and Video Matting

Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-...

AI 聚合 07/01
HuggingFace

One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications

Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhanceme...

AI 聚合 07/01
HuggingFace

A Gravitational Interpretation of Fine-Tuning Reversion

Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign...

AI 聚合 07/01
HuggingFace

RocketSmith: Agentic Additive Manufacturing of High-Powered Rockets

RocketSmith is an agentic system which intelligently automates the DFAM process for the development of high powered r...

AI 聚合 07/01
HuggingFace

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by...

AI 聚合 07/01
HuggingFace

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by ...

AI 聚合 07/01
HuggingFace

LLM Program Optimization via Retrieval Augmented Search

Recent work has demonstrated the potential of large language models (LLMs) for program optimization, a key challenge ...

AI 聚合 07/01
HuggingFace

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?

Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized langua...

AI 聚合 07/01
HuggingFace

Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement

Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, b...

AI 聚合 07/01
HuggingFace

MirrorPPR: Exemplar-Based Portrait Photo Retouching

While text-guided image editing has made remarkable progress, it remains limited in structural portrait retouching. T...

AI 聚合 07/01
arXiv

Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where sp...

AI 聚合 06/30
arXiv

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challeng...

AI 聚合 06/30
arXiv

The Human Creativity Benchmark

Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, profession...

AI 聚合 06/30
arXiv

A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution

Malware classification remains a challenging problem due to its inherent heterogeneity, the presence of packed binari...

AI 聚合 06/30
arXiv

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) w...

AI 聚合 06/30
arXiv

Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection

Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection. ...

AI 聚合 06/30
arXiv

MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems

Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agen...

AI 聚合 06/30
arXiv

C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders

Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse,...

AI 聚合 06/30
arXiv

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms

Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosi...

AI 聚合 06/30
arXiv

DOPD: Dual On-policy Distillation

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...

AI 聚合 06/30
arXiv

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy ...

AI 聚合 06/30
arXiv

GROW$^2$: Grounding Which and Where for Robot Tool Use

Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to...

AI 聚合 06/30
arXiv

Self-Evolving World Models for LLM Agent Planning

World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequenc...

AI 聚合 06/30
arXiv

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics...

AI 聚合 06/30
arXiv

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whol...

AI 聚合 06/30
HuggingFace

PoseShield: Neural Collision Fields for Human Self-Collision Resolution

Self-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extrem...

AI 聚合 06/30
HuggingFace

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enab...

AI 聚合 06/30
HuggingFace

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to app...

AI 聚合 06/30
HuggingFace

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent vid...

AI 聚合 06/30
HuggingFace

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world comp...

AI 聚合 06/30
HuggingFace

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly c...

AI 聚合 06/30
HuggingFace

Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit Dependence

Current models of representational reliability in neural populations focus on temporal stability: whether population ...

AI 聚合 06/30
HuggingFace

DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model

We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulatio...

AI 聚合 06/30
HuggingFace

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

Streaming video editing has made rapid progress, yet practical deployment is still limited by two core issues: mainta...

AI 聚合 06/30
HuggingFace

Trimming the Long-Tail of Visual World Modeling Evaluation

Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human exp...

AI 聚合 06/30
HuggingFace

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic v...

AI 聚合 06/30
HuggingFace

Interleaved Speech Language Models Latently Work In Text

Speech language models (SLMs) have been extensively studied, with the common paradigm incorporating text data and pre...

AI 聚合 06/30
HuggingFace

The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction

4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-b...

AI 聚合 06/30
HuggingFace

Learning Transferable Dynamics Priors from Action to World Modeling

We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning...

AI 聚合 06/30
HuggingFace

TheoremGraph: Bridging Formal and Informal Mathematics

Mathematical knowledge is organized around statements and their dependencies, but this structure is exposed unevenly:...

AI 聚合 06/30
HuggingFace

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on la...

AI 聚合 06/30
HuggingFace

Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path Planner

Hydropower tunnel inspection is critical for infrastructure integrity yet remains inefficient and hazardous using man...

AI 聚合 06/30
HuggingFace

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

Adapting a foundation vision-language encoder to a specialized retrieval task creates a fundamental tradeoff: gains o...

AI 聚合 06/30
HuggingFace

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete...

AI 聚合 06/30
HuggingFace

How Good Can Linear Models Be for Time-Series Forecasting?

Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers ...

AI 聚合 06/30
HuggingFace

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring...

AI 聚合 06/30
HuggingFace

MultiHashFormer: Hash-based Generative Language Models

Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To cons...

AI 聚合 06/30
HuggingFace

Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving

Efficient deployment of large language models (LLMs) in production forces a trade-off between accuracy and cost. Oper...

AI 聚合 06/30
HuggingFace

Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation

Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches,...

AI 聚合 06/30
HuggingFace

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consisten...

AI 聚合 06/30
HuggingFace

How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring

LLM-based code agents navigate repositories through keyword search but miss the structural relationships, such as cal...

AI 聚合 06/30
HuggingFace

To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair

LLM-based agents for program repair are increasingly built on a "generate-run-revise" paradigm, iteratively executing...

AI 聚合 06/30
HuggingFace

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models

Tokenization is central to adapting scientific data for transformer-based foundation models, yet its impact on learne...

AI 聚合 06/30
HuggingFace

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents

For agents to learn continuously from interaction with the world at test time, they must be able to explore effective...

AI 聚合 06/30
HuggingFace

Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured ...

AI 聚合 06/30
HuggingFace

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a ...

AI 聚合 06/30
HuggingFace

Simplified Sparse Attention via Gist Tokens

Sparse attention can reduce the cost of long-context inference, but most variants introduce new architectural compone...

AI 聚合 06/30
HuggingFace

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee tha...

AI 聚合 06/30
HuggingFace

Vesta: A Generalist Embodied Reasoning Model

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, an...

AI 聚合 06/30
arXiv

LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

Embodied agents operating in decentralized and partially observable environments have attracted growing attention in ...

AI 聚合 06/29
arXiv

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and ...

AI 聚合 06/29
arXiv

The Remittance Blueprint: Data-driven Intelligence for Sri Lanka

This study analyzes Sri Lankan migration and remittances over 32 years (1994-2025). Using a 384-month harmonized data...

AI 聚合 06/29
arXiv

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data c...

AI 聚合 06/29
arXiv

Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives

We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agent...

AI 聚合 06/29
arXiv

Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching

Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepan...

AI 聚合 06/29
arXiv

Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software

Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates th...

AI 聚合 06/29
arXiv

How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks

Understanding how performance scales jointly with model size and data is a central problem in modern machine learning...

AI 聚合 06/29
arXiv

Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly Segmentation

Test-time adaptation (TTA) has emerged as a promising paradigm for mitigating distribution shifts in deep models. How...

AI 聚合 06/29
arXiv

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and m...

AI 聚合 06/29
arXiv

Parameter Efficient Hybrid Transformer (PEHT) for Network Traffic Prediction via Dynamic Urban Congestion Integration

Accurate network traffic prediction is a critical element for efficient resource allocation in dynamic urban cellular...

AI 聚合 06/29
arXiv

Towards Automating Scientific Review with Google's Paper Assistant Tool

Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis gene...

AI 聚合 06/29
arXiv

Agentic Hardware Design as Repository-Level Code Evolution

We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A...

AI 聚合 06/29
arXiv

Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes

Many two-player zero-sum games admit not a unique Nash equilibrium but a convex set of them: a polytope of profiles t...

AI 聚合 06/29
arXiv

DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand

Dexterous manipulation policies can solve individual skills, but composing them to perform multiple tasks with a sing...

AI 聚合 06/29
HuggingFace

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-dom...

AI 聚合 06/29
HuggingFace

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a...

AI 聚合 06/29
HuggingFace

Towards Automating Scientific Review with Google's Paper Assistant Tool

Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis gene...

AI 聚合 06/29
HuggingFace

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

We introduce an axiomatic evaluation framework for latent thought representations in LLMs, comprising metrics that ar...

AI 聚合 06/29
HuggingFace

Qwen-Image-2.0-RL Technical Report

We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF)...

AI 聚合 06/29
HuggingFace

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications....

AI 聚合 06/29
HuggingFace

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades ...

AI 聚合 06/29
HuggingFace

Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents

Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- whil...

AI 聚合 06/29
HuggingFace

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowl...

AI 聚合 06/29
HuggingFace

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel gripper...

AI 聚合 06/29
HuggingFace

Boundary-Aware Context Grounding for A Low-Channel EEG Agent

Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatica...

AI 聚合 06/29
HuggingFace

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system...

AI 聚合 06/29
HuggingFace

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-ba...

AI 聚合 06/29
HuggingFace

GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems

Multi-agent systems (MAS) built on large language models (LLMs) provide a promising framework for solving complex tas...

AI 聚合 06/29
HuggingFace

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

Voice agents face a fundamental tension: the reasoning, retrieval, and tool use that make foundation models capable a...

AI 聚合 06/29
HuggingFace

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-mode...

AI 聚合 06/27
HuggingFace

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to f...

AI 聚合 06/27
HuggingFace

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermed...

AI 聚合 06/27
HuggingFace

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under ch...

AI 聚合 06/27
HuggingFace

Information-Aware KV Cache Compression for Long Reasoning

Reasoning capability has advanced rapidly in large language models (LLMs), leading to an increasing size of key-value...

AI 聚合 06/27
HuggingFace

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings rema...

AI 聚合 06/27
HuggingFace

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting...

AI 聚合 06/27
arXiv

Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts

We present a conceptual framework for analyzing dialogue in collaborative problem-solving contexts, with an emphasis ...

AI 聚合 06/26
arXiv

From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan

AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real ...

AI 聚合 06/26
arXiv

Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy

Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools...

AI 聚合 06/26
arXiv

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major chal...

AI 聚合 06/26
arXiv

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under ch...

AI 聚合 06/26
arXiv

Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC

Mechanistic epidemiological models are widely used to support infectious disease forecasting and public-health decisi...

AI 聚合 06/26
arXiv

Prompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection Settings

Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candida...

AI 聚合 06/26
arXiv

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-mode...

AI 聚合 06/26
arXiv

AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns

AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their per...

AI 聚合 06/26
arXiv

Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation model...

AI 聚合 06/26
arXiv

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning

Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential...

AI 聚合 06/26
arXiv

Language-Based Digital Twins for Elderly Cognitive Assistance

Digital twins have emerged as a promising paradigm for personalized healthcare, enabling modeling of individual behav...

AI 聚合 06/26
arXiv

Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching

Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are c...

AI 聚合 06/26
arXiv

Error-Conditioned Neural Solvers

Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat so...

AI 聚合 06/26
arXiv

Autoregressive Boltzmann Generators

Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. ...

AI 聚合 06/26
HuggingFace

GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods

With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily e...

AI 聚合 06/26
HuggingFace

PhysiFormer: Learning to Simulate Mechanics in World Space

We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models ...

AI 聚合 06/26
HuggingFace

DanceOPD: On-Policy Generative Field Distillation

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), loca...

AI 聚合 06/26
HuggingFace

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are o...

AI 聚合 06/26
HuggingFace

COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physica...

AI 聚合 06/26
HuggingFace

Fast LeWorldModel

Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising found...

AI 聚合 06/26
HuggingFace

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse traj...

AI 聚合 06/26
HuggingFace

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more...

AI 聚合 06/26
HuggingFace

Confidence-Aware Tool Orchestration for Robust Video Understanding

Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we t...

AI 聚合 06/26
HuggingFace

In-Context World Modeling for Robotic Control

Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoint...

AI 聚合 06/26
HuggingFace

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the c...

AI 聚合 06/26
HuggingFace

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL...

AI 聚合 06/26
HuggingFace

GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents

Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces...

AI 聚合 06/26
HuggingFace

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this i...

AI 聚合 06/26
HuggingFace

Hallucination in World Models is Predictable and Preventable

Modern generative world models render increasingly realistic action-controllable futures, yet they frequently halluci...

AI 聚合 06/26
HuggingFace

How Post-Training Shapes Biological Reasoning Models

Scientific reasoning models for biology combine language models with foundation models trained on multimodal biologic...

AI 聚合 06/26
HuggingFace

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and ver...

AI 聚合 06/26
HuggingFace

Discretizing Reward Models

Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Rewar...

AI 聚合 06/26
HuggingFace

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies

As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is ...

AI 聚合 06/26
HuggingFace

ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation

On-policy distillation (OPD) improves LLM reasoning by training a student model on its own generated outputs, but sta...

AI 聚合 06/26
HuggingFace

Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation

Continual Test-Time Adaptation (CTTA) aims to maintain model performance under evolving target domains by adapting on...

AI 聚合 06/26
HuggingFace

What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit...

AI 聚合 06/26
HuggingFace

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decis...

AI 聚合 06/26
HuggingFace

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computati...

AI 聚合 06/26
HuggingFace

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (...

AI 聚合 06/26
HuggingFace

Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can cont...

AI 聚合 06/26
HuggingFace

Forecasting Future Behavior as a Learning Task

Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior...

AI 聚合 06/26
HuggingFace

Do Thinking Tokens Help with Safety?

Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tune...

AI 聚合 06/26
HuggingFace

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate v...

AI 聚合 06/26
arXiv

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality train...

AI 聚合 06/25
arXiv

Variable Bound Tightening for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games

There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player ze...

AI 聚合 06/25
arXiv

Hierarchical Reinforcement Learning for Neural Network Compression (HiReLC): Pruning and Quantization

We present HiReLC, a hierarchical ensemble-reinforcement learning framework for automated joint quantization and stru...

AI 聚合 06/25
arXiv

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation

Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While...

AI 聚合 06/25
arXiv

Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection of High-Risk Queries

Tabular foundation models are commonly assumed to present limited privacy concerns as they are often pre-trained on l...

AI 聚合 06/25
arXiv

Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem

As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges...

AI 聚合 06/25
arXiv

TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchma...

AI 聚合 06/25
arXiv

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's...

AI 聚合 06/25
arXiv

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems...

AI 聚合 06/25
arXiv

A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks

The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in ach...

AI 聚合 06/25
arXiv

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on det...

AI 聚合 06/25
arXiv

A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding

Supervised deep learning has been widely used for weld penetration state classification; however, its performance oft...

AI 聚合 06/25
arXiv

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings rema...

AI 聚合 06/25
arXiv

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity

On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with...

AI 聚合 06/25
arXiv

Learning Action Priors for Cross-embodiment Robot Manipulation

Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action modul...

AI 聚合 06/25
HuggingFace

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still st...

AI 聚合 06/25
HuggingFace

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

While Large Language Models (LLMs) have substantially advanced text-to-code synthesis, many real programming tasks sp...

AI 聚合 06/25
HuggingFace

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time strea...

AI 聚合 06/25
HuggingFace

Improved Large Language Diffusion Models

Modern large language models are predominantly trained with autoregressive factorization and causal attention. We pre...

AI 聚合 06/25
HuggingFace

RoPE-Aware Bit Allocation for KV-Cache Quantization

Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contr...

AI 聚合 06/25
HuggingFace

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. ...

AI 聚合 06/25
HuggingFace

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Op...

AI 聚合 06/25
HuggingFace

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality train...

AI 聚合 06/25
HuggingFace

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LL...

AI 聚合 06/25
HuggingFace

Are We Ready For An Agent-Native Memory System?

Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a da...

AI 聚合 06/25
HuggingFace

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynam...

AI 聚合 06/25
HuggingFace

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single succ...

AI 聚合 06/25
HuggingFace

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evid...

AI 聚合 06/25
HuggingFace

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and ...

AI 聚合 06/25
HuggingFace

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safe...

AI 聚合 06/25
HuggingFace

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whe...

AI 聚合 06/25
HuggingFace

RL-Index: Reinforcement Learning for Retrieval Index Reasoning

Retrieving external knowledge is essential for solving real-world tasks, yet it remains challenging when the relation...

AI 聚合 06/25
HuggingFace

ShutterMuse: Capture-Time Photography Guidance with MLLMs

Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aestheti...

AI 聚合 06/25
HuggingFace

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geome...

AI 聚合 06/25
HuggingFace

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joi...

AI 聚合 06/25
HuggingFace

ChartWalker: Benchmarking the Cross-Chart RAG Task

Cross-Chart Retrieval-Augmented Generation (RAG) is critical for complex multi-modal analytical tasks in scientific, ...

AI 聚合 06/25
HuggingFace

AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

Large language models are increasingly deployed as agents that reason over documents rather than answer from parametr...

AI 聚合 06/25
HuggingFace

Semantic Browsing: Controllable Diversity for Image Generation

Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at th...

AI 聚合 06/25
HuggingFace

Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation

Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation...

AI 聚合 06/25
HuggingFace

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bou...

AI 聚合 06/25
HuggingFace

Critique of Agent Model

What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding a...

AI 聚合 06/25
HuggingFace

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as d...

AI 聚合 06/25
HuggingFace

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery

Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving unders...

AI 聚合 06/25
arXiv

Grad Detect: Gradient-Based Hallucination Detection in LLMs

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to...

AI 聚合 06/24
arXiv

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answ...

AI 聚合 06/24
arXiv

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generat...

AI 聚合 06/24
arXiv

Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution

Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quan...

AI 聚合 06/24
arXiv

Solving Inverse Problems of Chaotic Systems with Bidirectional Conditional Flow Matching

Modeling chaotic systems is crucial yet challenging. Inverse problems in chaotic dynamics, namely inferring initial c...

AI 聚合 06/24
arXiv

Difference-Making without Making a Difference

Over a series of seven papers, Andreas & Günther have introduced seven definitions of actual causation and have class...

AI 聚合 06/24
arXiv

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks...

AI 聚合 06/24
arXiv

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This m...

AI 聚合 06/24
arXiv

Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models

Prompt-based learning has emerged as a dominant paradigm in natural language processing. This study explores the impa...

AI 聚合 06/24
arXiv

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a wo...

AI 聚合 06/24
arXiv

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still st...

AI 聚合 06/24
arXiv

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are abl...

AI 聚合 06/24
arXiv

OpenThoughts-Agent: Data Recipes for Agentic Models

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate t...

AI 聚合 06/24
arXiv

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation...

AI 聚合 06/24
arXiv

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bou...

AI 聚合 06/24
HuggingFace

MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization

MLLM-based mobile GUI agents have made substantial progress in UI understanding and action execution, but adapting th...

AI 聚合 06/24
HuggingFace

FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation

Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation...

AI 聚合 06/24
HuggingFace

OpenThoughts-Agent: Data Recipes for Agentic Models

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate t...

AI 聚合 06/24
HuggingFace

MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management

MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet remain unreliable on long-hor...

AI 聚合 06/24
HuggingFace

Qwen-AgentWorld: Language World Models for General Agents

A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive me...

AI 聚合 06/24
HuggingFace

World Value Models for Robotic Manipulation

Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. ...

AI 聚合 06/24
HuggingFace

Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from n...

AI 聚合 06/24
HuggingFace

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of...

AI 聚合 06/24
HuggingFace

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narra...

AI 聚合 06/24
HuggingFace

FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs

Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ...

AI 聚合 06/24
HuggingFace

DREAM: Dense Retrieval Embeddings via Autoregressive Modeling

Dense retrieval embedding models are a fundamental component of modern retrieval-based AI systems. Most dense retriev...

AI 聚合 06/24
HuggingFace

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publica...

AI 聚合 06/24
HuggingFace

An Efficient Method for the Optimal Control of Microgrids Under Uncertainties using Local Reduction

The problem of optimal sizing and power scheduling in microgrids subject to uncertainties is well known to the contro...

AI 聚合 06/24
HuggingFace

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from d...

AI 聚合 06/24
HuggingFace

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning

Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world inte...

AI 聚合 06/24
HuggingFace

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction

AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, mana...

AI 聚合 06/24
HuggingFace

FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation

Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric represen...

AI 聚合 06/24
HuggingFace

LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis

Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of i...

AI 聚合 06/24
HuggingFace

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) ...

AI 聚合 06/24
HuggingFace

QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging

Attention-based Multiple Instance Learning aggregators in medical imaging are prone to attention concentration, produ...

AI 聚合 06/24
HuggingFace

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention

Self-attention is central to Transformer performance and is often the most expensive part of the Transformer at long ...

AI 聚合 06/24
HuggingFace

Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation

Text and image conditioned 3D models now generate convincing assets, but they still offer little direct control over ...

AI 聚合 06/24
HuggingFace

Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs

Open-weight Large Language Models (LLMs) enable scientific progress and broad deployment. However, they make it diffi...

AI 聚合 06/24
HuggingFace

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, lang...

AI 聚合 06/24
HuggingFace

Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?

Computer-use agents (CUAs) now act on a user's behalf across personal applications such as email, calendars, and to-d...

AI 聚合 06/24
HuggingFace

Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach

As urban areas expand, automatic monitoring of parking lots becomes essential for efficient and sustainable cities. T...

AI 聚合 06/24
HuggingFace

AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining

Optimizing pretraining data composition is pivotal for LLM generalization. While dynamic mixing outperforms static st...

AI 聚合 06/24
HuggingFace

An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game

Large language models (LLMs) are increasingly used to support software development, but their practical usefulness in...

AI 聚合 06/24
HuggingFace

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City

As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backb...

AI 聚合 06/24
HuggingFace

Vera: A Layered Diffusion Model for Content-Preserving Video Editing

Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservatio...

AI 聚合 06/24
HuggingFace

A Verifiable Search Is Not a Learnable Chain-of-Thought

It is tempting to assume any task solvable by a short program can be taught to a model as its chain-of-thought: write...

AI 聚合 06/24
HuggingFace

Libretto: Giving LLM Agents a Sense of Musical Structure

Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspe...

AI 聚合 06/24
HuggingFace

Go-with-the-Track: Video Compositing and Motion Control with Point Tracking

Filmmaking demands precise motion control and reference image compositing -- capabilities that existing methods treat...

AI 聚合 06/24
HuggingFace

When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents

Long-horizon LLM agents can fail quietly: they settle on one reading of the evidence early, then spend the rest of th...

AI 聚合 06/24
HuggingFace

ShotcreteDepth: A Bi-modal Dataset for Robust Robotic Depth Perception in Shotcrete Construction Environments

We introduce ShotcreteDepth, a bi-modal dataset from the construction domain that captures both an active shotcreting...

AI 聚合 06/24
HuggingFace

Comparing Linear Probes with Mahalanobis Cosine Similarity

Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis c...

AI 聚合 06/24
HuggingFace

Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild

Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observatio...

AI 聚合 06/24
HuggingFace

TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization

Discrete text-trigger optimization -- searching for text sequences that, when ingested by a model, steer it toward a ...

AI 聚合 06/24
HuggingFace

Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning

Long-context reasoning is an essential capability for large language models, particularly when they are deployed as a...

AI 聚合 06/24
arXiv

Discovering Latent Groups for Robust Classification

Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately...

AI 聚合 06/23
arXiv

Data Selection Through Iterative Self-Filtering for Vision-Language Settings

The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, m...

AI 聚合 06/23
arXiv

RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models

Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demo...

AI 聚合 06/23
arXiv

DiT-Reward: Generative Representations for Text-to-Image Reward Modeling

Can representations learned for image generation also support the evaluation of generated images? We study text-to-im...

AI 聚合 06/23
arXiv

AI-driven Optimisation of Quality of Recovery (QoR) in Remote Patient Monitoring

Remote patient monitoring depends on patient-reported data to capture the subjective dimension of recovery that devic...

AI 聚合 06/23
arXiv

AI Exposure Scores: what they measure, what they miss, and what comes next

A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produc...

AI 聚合 06/23
arXiv

Learning Process Rewards via Success Visitation Matching for Efficient RL

In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently s...

AI 聚合 06/23
arXiv

TailorMind: Towards Preference-Aligned Multimodal Content Generation

Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly...

AI 聚合 06/23
arXiv

Tapered Language Models

Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack ...

AI 聚合 06/23
arXiv

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Ma...

AI 聚合 06/23
arXiv

PsyBridge: A Hybrid Intelligent Framework for Multi-Dimensional Mental Health Assessment and Decision Support

Mental health assessment commonly relies on isolated screening instruments or data-driven models that often lack inte...

AI 聚合 06/23
arXiv

Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostl...

AI 聚合 06/23
arXiv

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large lang...

AI 聚合 06/23
arXiv

Semantic Browsing: Controllable Diversity for Image Generation

Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at th...

AI 聚合 06/23
arXiv

CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation

Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipula...

AI 聚合 06/23
HuggingFace

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

Multi-view 3D Visual Question Answering (MV3D-VQA) requires integrating partial observations into a coherent 3D scene...

AI 聚合 06/23
HuggingFace

Self-Compacting Language Model Agents

Long agent traces composed of chains of thought and tool calls accumulate stale content that anchor subsequent genera...

AI 聚合 06/23
HuggingFace

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

Latent action pretraining learns representations of visual change from pairs of observations, but existing methods ty...

AI 聚合 06/23
HuggingFace

Training Open Models for Agentic Phone Use

Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable ...

AI 聚合 06/23
HuggingFace

CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents

While recent LLM-based terminal agents have demonstrated promising capabilities, the scarcity of high-quality, execut...

AI 聚合 06/23
HuggingFace

Unlimited OCR Works

Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely...

AI 聚合 06/23
HuggingFace

FastMix: Fast Data Mixture Optimization via Gradient Descent

While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture fo...

AI 聚合 06/23
HuggingFace

Tapered Language Models

Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack ...

AI 聚合 06/23
HuggingFace

UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfe...

AI 聚合 06/23
HuggingFace

Tmax: A simple recipe for terminal agents

Terminal-using agents have quickly become the most popular downstream application of language models (LMs). Despite t...

AI 聚合 06/23
HuggingFace

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

Massive unstructured multimodal streams suffer from high "data entropy," impeding both efficient human knowledge acqu...

AI 聚合 06/23
HuggingFace

HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions

With the rapid spread of retrieval-augmented generation and semantic search, choosing the right embedding and retriev...

AI 聚合 06/23
HuggingFace

MeshFlow: Mesh Generation with Equivariant Flow Matching

Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the ...

AI 聚合 06/23
HuggingFace

Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents

Long-horizon tasks are common in real-world robotic deployments, yet failure detection for such tasks remains underex...

AI 聚合 06/23
HuggingFace

Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding

Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that ...

AI 聚合 06/23
HuggingFace

Improving Text-to-Music Generation with Human Preference Rewards

We describe our entry to the efficiency track of the Academic Text-to-Music (ATTM) Grand Challenge at ICME 2026. Beyo...

AI 聚合 06/23
HuggingFace

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for...

AI 聚合 06/23
HuggingFace

BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natur...

AI 聚合 06/23
HuggingFace

Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills

Scientific discovery workflows usually contain and rely heavily on lab notes, where researchers record observations, ...

AI 聚合 06/23
HuggingFace

Counsel: A Meta-Evaluation Dataset for Agentic Tasks

As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottl...

AI 聚合 06/23
HuggingFace

StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs

Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential setting...

AI 聚合 06/23
HuggingFace

MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval

Retrieval-augmented generation (RAG) systems depend critically on how documents are chunked and searched. Fine-graine...

AI 聚合 06/23
HuggingFace

SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG

Retrieval-augmented generation (RAG) systems must balance retrieval granularity with contextual coherence, a challeng...

AI 聚合 06/23
HuggingFace

Characterizing Narrative Content in Web-scale LLM Pretraining Data

The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a ...

AI 聚合 06/23
HuggingFace

When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning

Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because ...

AI 聚合 06/23
HuggingFace

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, mo...

AI 聚合 06/22
HuggingFace

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and ...

AI 聚合 06/22
HuggingFace

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could a...

AI 聚合 06/22
HuggingFace

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

Memory benchmarks for LLM agents largely assume single-user settings, leaving shared assistants for hospitals, workpl...

AI 聚合 06/22
HuggingFace

Distilling Examples into Task Instructions: Enhanced In-Context Learning for Real-World B2B Conversations

In-context learning (ICL) is the standard method for low-resource classification, yet its efficacy in specialized dom...

AI 聚合 06/22
HuggingFace

SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction

High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human...

AI 聚合 06/22
HuggingFace

GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to pla...

AI 聚合 06/22
HuggingFace

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

While reasoning on autoregressive (AR) models is often performed by chain-of-thought reasoning and reflection, their ...

AI 聚合 06/22
HuggingFace

MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision

Personalized presentation generation requires more than conditioning on a current prompt or template: agents must pre...

AI 聚合 06/22
GitHub

[GitHub] lobehub/lobehub

🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and report...

AI 聚合 06/21
GitHub

[GitHub] obra/superpowers

An agentic skills framework & software development methodology that works.(⭐234547)

AI 聚合 06/21
arXiv

Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems

When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate t...

AI 聚合 06/21
arXiv

Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software

Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on cont...

AI 聚合 06/21
arXiv

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a c...

AI 聚合 06/21
arXiv

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models ...

AI 聚合 06/21
arXiv

Efficient and Sound Probabilistic Verification for AI Agents

Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring ap...

AI 聚合 06/21
arXiv

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on cod...

AI 聚合 06/21
arXiv

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunc...

AI 聚合 06/21
arXiv

Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes

Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutatio...

AI 聚合 06/21
arXiv

SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for s...

AI 聚合 06/21
arXiv

DeepSWIP: Quotient-WMC Counterfactuals for Neural Probabilistic Logic Programs

Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference ...

AI 聚合 06/21
arXiv

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling ...

AI 聚合 06/21
arXiv

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual word...

AI 聚合 06/21
arXiv

Toward Calibrated Mixture-of-Experts Under Distribution Shift

Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important f...

AI 聚合 06/21
arXiv

Structuring and Tokenizing Distributed User Interest Context for Generative Recommendation

Generative recommendation is an emerging paradigm that has shown promise in industrial recommendation systems, aiming...

AI 聚合 06/21
arXiv

How Transparent is DiffusionGemma?

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalign...

AI 聚合 06/21
HuggingFace

JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

Creating 3D visual illusions, a single 3D mesh that reveals entirely different semantics from various viewing angles,...

AI 聚合 06/21
HuggingFace

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and too...

AI 聚合 06/21
HuggingFace

Playful Agentic Robot Learning

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior acr...

AI 聚合 06/21
HuggingFace

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-wo...

AI 聚合 06/21
HuggingFace

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tigh...

AI 聚合 06/21
HuggingFace

DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects

Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, wher...

AI 聚合 06/21
HuggingFace

FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows

Conditional diffusion and flow models routinely fail to satisfy the very constraints that define their task. For inst...

AI 聚合 06/21
HuggingFace

Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation

Hybrid linear attention models offer an appealing path to faster long-context inference: they reduce the quadratic co...

AI 聚合 06/21
HuggingFace

No Resource, No Benchmarks, No Problem? Evaluating and Improving LLMs for Code Generation in No-Resource Languages

Large Language Models (LLMs) have significantly advanced the automation of software engineering tasks. One prominent ...

AI 聚合 06/21
HuggingFace

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on cod...

AI 聚合 06/21
HuggingFace

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hard...

AI 聚合 06/21
HuggingFace

Duration Aware Scheduling for ASR Serving Under Workload Drift

Scheduling policies in large-scale Automatic Speech Recognition (ASR) serving pipelines play a key role in determinin...

AI 聚合 06/21
HuggingFace

The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation

The Frechet Inception Distance (FID) is the de facto arbiter of image generation, yet most papers report just a singl...

AI 聚合 06/21
HuggingFace

The Data Manifold under the Microscope

A significant gap exists between theory and practice in deep learning. Generalization and approximation error bounds ...

AI 聚合 06/21
HuggingFace

Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why

Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-l...

AI 聚合 06/21
HuggingFace

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average c...

AI 聚合 06/21
HuggingFace

ReSyn: A Generalized Recursive Regular Expression Synthesis Framework

Existing Programming-By-Example (PBE) systems often rely on simplified benchmarks that fail to capture the high struc...

AI 聚合 06/21
HuggingFace

Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequ...

AI 聚合 06/21
HuggingFace

Context-Aware RL for Agentic and Multimodal LLMs

Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence wi...

AI 聚合 06/21
HuggingFace

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling ...

AI 聚合 06/21
GitHub

[GitHub] pytorch/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration(⭐100863)

AI 聚合 06/19
HuggingFace

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning an...

AI 聚合 06/19
HuggingFace

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning

Reinforcement learning pipelines for Large Language Model (LLM) training often rely on manually redesigned environmen...

AI 聚合 06/19
HuggingFace

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language mode...

AI 聚合 06/19
HuggingFace

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...

AI 聚合 06/19
HuggingFace

LLM-Enabled NWDAF: A Step Toward AI-Native 6G Network Intelligence

The Network Data Analytics Function (NWDAF) is central to enabling zero-touch network management in fifth-generation ...

AI 聚合 06/19
HuggingFace

A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets

Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more ...

AI 聚合 06/19
HuggingFace

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe t...

AI 聚合 06/19
HuggingFace

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between eval...

AI 聚合 06/19
HuggingFace

iOSWorld: A Benchmark for Personally Intelligent Phone Agents

A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and prefer...

AI 聚合 06/19
HuggingFace

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns

We show the standard basis of transformer hidden states already provides a training-free, architecture-general featur...

AI 聚合 06/19
HuggingFace

ViT-Up: Faithful Feature Upsampling for Vision Transformers

Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptio...

AI 聚合 06/19
HuggingFace

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction

Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan a...

AI 聚合 06/19
HuggingFace

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-l...

AI 聚合 06/19
HuggingFace

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

As an increasing majority of global video content is consumed on social platforms for interactive social purposes, vi...

AI 聚合 06/19
HuggingFace

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with...

AI 聚合 06/19
HuggingFace

HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing

Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer...

AI 聚合 06/19
HuggingFace

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

Offline reinforcement learning is typically analyzed under process-level reward supervision, yet many sequential deci...

AI 聚合 06/19
HuggingFace

Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities

Robotic systems perceive the world through multiple input modalities -- including visual camera streams and natural l...

AI 聚合 06/19
HuggingFace

Re-Centering Humans in LLM Personalization

Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on...

AI 聚合 06/19
HuggingFace

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM)...

AI 聚合 06/19
arXiv

Mechanism-Guided Selective Unlearning for RLVR-Induced Reasoning

We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reason...

AI 聚合 06/18
arXiv

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...

AI 聚合 06/18
arXiv

TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology

Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-m...

AI 聚合 06/18
arXiv

A Taxonomy of Mental Health and Technology Needs for Alzheimer's and Dementia Caregivers

Family members caring for individuals with Alzheimer's disease and related dementias (AD/ADRD) provide the foundation...

AI 聚合 06/18
arXiv

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific...

AI 聚合 06/18
arXiv

X+Slides: Benchmarking Audience-Conditioned Slide Generation

Automatically generating slide decks from source documents is an important application of large language models (LLMs...

AI 聚合 06/18
arXiv

A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimod...

AI 聚合 06/18
arXiv

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized doma...

AI 聚合 06/18
arXiv

NeSyCat Torch: A Differentiable Tensor Implementation of Categorical Semantics for Neurosymbolic Learning

Neurosymbolic semantics is fragmented: classical, fuzzy, probabilistic and neural systems each define truth by their ...

AI 聚合 06/18
arXiv

Correct Yourself, Keep My Trust: How Self-Correction and Social Connection Shape Credibility in Social Chatbots

When social chatbots make mistakes, and they do, how they recover determines whether users trust them again. Social c...

AI 聚合 06/18
arXiv

Explaining Attention with Program Synthesis

A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-me...

AI 聚合 06/18
arXiv

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts ...

AI 聚合 06/18
arXiv

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, mu...

AI 聚合 06/18
arXiv

Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation

Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning w...

AI 聚合 06/18
arXiv

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning

Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing ...

AI 聚合 06/18
HuggingFace

Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems

Multi-agent LLM systems share state through memory stores, vector indices, and tool registries. We model such sharing...

AI 聚合 06/18
HuggingFace

Self-Evolving Visual Questioner

Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse...

AI 聚合 06/18
HuggingFace

Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings

Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a sca...

AI 聚合 06/18
HuggingFace

Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences

In this report, we present LOGOS (Language Of Generative Objects in Science), a scientific generative language model ...

AI 聚合 06/18
HuggingFace

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games

Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observat...

AI 聚合 06/18
HuggingFace

Guava: An Effective and Universal Harness for Embodied Manipulation

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. H...

AI 聚合 06/18
HuggingFace

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ide...

AI 聚合 06/18
HuggingFace

CEO-Bench: Can Agents Play the Long Game?

Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering...

AI 聚合 06/18
HuggingFace

Physics-IQ Verified

Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a m...

AI 聚合 06/18
HuggingFace

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and l...

AI 聚合 06/18
HuggingFace

Learning User Simulators with Turing Rewards

Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation o...

AI 聚合 06/18
HuggingFace

Kairos: A Native World Model Stack for Physical AI

World models are transitioning from passive visual generators to foundational, operational infrastructure for Physica...

AI 聚合 06/18
HuggingFace

IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products

Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern pro...

AI 聚合 06/18
HuggingFace

Sumi: Open Uniform Diffusion Language Model from Scratch

Diffusion models have become a promising alternative to autoregressive models. Among these, uniform diffusion languag...

AI 聚合 06/18
HuggingFace

Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding

Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in ...

AI 聚合 06/18
HuggingFace

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space def...

AI 聚合 06/18
HuggingFace

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-st...

AI 聚合 06/18
HuggingFace

Native Active Perception as Reasoning for Omni-Modal Understanding

Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly ...

AI 聚合 06/18
HuggingFace

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest co...

AI 聚合 06/18
HuggingFace

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are ...

AI 聚合 06/18
arXiv

ReAge3D: Re-Aging 3D Faces with View Consistency

We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity...

AI 聚合 06/17
arXiv

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act

Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate wheth...

AI 聚合 06/17
arXiv

All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code

Software practitioners increasingly use AI coding agents that generate test code alongside production code in open so...

AI 聚合 06/17
arXiv

IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction

Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable...

AI 聚合 06/17
arXiv

Kolmogorov Regression for Robust Diffusion Policies

Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long...

AI 聚合 06/17
arXiv

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly foc...

AI 聚合 06/17
arXiv

The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data

As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce a...

AI 聚合 06/17
arXiv

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models

We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 a...

AI 聚合 06/17
arXiv

RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to allevi...

AI 聚合 06/17
arXiv

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require composi...

AI 聚合 06/17
arXiv

Looped World Models

Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deepe...

AI 聚合 06/17
arXiv

Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents

With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cybe...

AI 聚合 06/17
arXiv

EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation

Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any p...

AI 聚合 06/17
arXiv

ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues

Reproducing research results from papers and released code is central to scientific progress. Existing works have int...

AI 聚合 06/17
arXiv

Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement

Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism...

AI 聚合 06/17
HuggingFace

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, ex...

AI 聚合 06/17
HuggingFace

LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educ...

AI 聚合 06/17
HuggingFace

TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs

Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irr...

AI 聚合 06/17
HuggingFace

Looped World Models

Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deepe...

AI 聚合 06/17
HuggingFace

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory

Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action voc...

AI 聚合 06/17
HuggingFace

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

Game generation is an emerging application of coding agents, requiring models to transform natural-language specifica...

AI 聚合 06/17
HuggingFace

A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization ...

AI 聚合 06/17
HuggingFace

Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus

Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved source...

AI 聚合 06/17
HuggingFace

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning...

AI 聚合 06/17
HuggingFace

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Knowledge distillation transfers a teacher's competence to a small student but is brittle in the small-student regime...

AI 聚合 06/17
HuggingFace

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases l...

AI 聚合 06/17
HuggingFace

ProCUA-SFT Technical Report

Training computer-use agents (CUAs) -- models that interact with graphical desktops through screenshots and keyboard/...

AI 聚合 06/17
HuggingFace

Variable-Width Transformers

Scaling model size, specifically depth and width, has driven significant progress in transformer-based language model...

AI 聚合 06/17
HuggingFace

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory ...

AI 聚合 06/17
HuggingFace

Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion

Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoi...

AI 聚合 06/17
HuggingFace

MotionVLA: Vision-Language-Action Model for Humanoid Motion

Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-f...

AI 聚合 06/17
HuggingFace

ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions

Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this abilit...

AI 聚合 06/17
HuggingFace

Rethinking the Role of Efficient Attention in Hybrid Architectures

Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention m...

AI 聚合 06/17
HuggingFace

Learning from the Self-future: On-policy Self-distillation for dLLMs

On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its appli...

AI 聚合 06/17
HuggingFace

Text-Vision Co-Instructed Image Editing

Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based on...

AI 聚合 06/17
arXiv

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and in...

AI 聚合 06/17
arXiv

Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a priv...

AI 聚合 06/17
arXiv

Probing Low Frame Rate Degradation in Neural Audio Codecs

Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost ...

AI 聚合 06/17
arXiv

How Much Do Reviews Really Contribute? A Study on Text-Enriched Matrix Factorization for Recommendations

Incorporating textual reviews into a Recommender System has become a prominent strategy for enriching collaborative s...

AI 聚合 06/17
arXiv

The embrace of open science: An analysis of a decade of AI research and 56 800 conference papers

The reproducibility crisis has directed the AI research community toward improving documentation practices. Several s...

AI 聚合 06/17
arXiv

Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification

Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, tr...

AI 聚合 06/17
arXiv

Stable Menus of Public Goods: AI-Enabled Progress

Using an open problem from the EC 2025 paper "Stable Menus of Public Goods" as a testbed, we conduct experiments to u...

AI 聚合 06/17
arXiv

When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning

Reinforcement Learning (RL) policies often degrade in unfamiliar environments because they lack explicit deliberation...

AI 聚合 06/17
arXiv

ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation

Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it...

AI 聚合 06/17
arXiv

Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series...

AI 聚合 06/17
arXiv

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preferen...

AI 聚合 06/17
arXiv

TokenPilot: Cache-Efficient Context Management for LLM Agents

As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approac...

AI 聚合 06/17
arXiv

FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains c...

AI 聚合 06/17
arXiv

HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting

Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, an...

AI 聚合 06/17
arXiv

The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim Test of Image Classifiers

Oppenheim and Lim (1981) showed that natural images stay recognizable when reconstructed from their Fourier phase alo...

AI 聚合 06/17
HuggingFace

Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs

Standard accuracy benchmarks are designed to test how closely large language models (LLMs) approach correct answers, ...

AI 聚合 06/17
HuggingFace

Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks

A content-moderation system can score well on every standard accuracy metric and still cause real harm, if its mistak...

AI 聚合 06/17
HuggingFace

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon gener...

AI 聚合 06/17
HuggingFace

MMDiff: Extending Diffusion Transformers for Multi-Modal Generation

Diffusion transformers have demonstrated remarkable generative capabilities, yet the rich perceptual representations ...

AI 聚合 06/17
HuggingFace

GD^2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization

As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultiv...

AI 聚合 06/17
HuggingFace

SP^3: Spherical Priors for Plug-and-Play Restoration

In this paper, we introduce SP^3, a novel Plug-and-Play algorithm that accelerates maximum a posteriori image restora...

AI 聚合 06/17
HuggingFace

Memento: Reconstruct to Remember for Consistent Long Video Generation

Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions...

AI 聚合 06/17
HuggingFace

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary out...

AI 聚合 06/17
HuggingFace

Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning

Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency an...

AI 聚合 06/17
HuggingFace

The Ghosts of Polymarket: When Off-Chain Matches Meet On-Chain Reverts

Polymarket has emerged as a prominent prediction market platform and one of the fastest-growing applications in DeFi....

AI 聚合 06/17
HuggingFace

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong re...

AI 聚合 06/17
HuggingFace

MVEB: Massive Video Embedding Benchmark

We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classifi...

AI 聚合 06/17
HuggingFace

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but...

AI 聚合 06/17
HuggingFace

Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders

Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on ...

AI 聚合 06/17
HuggingFace

EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video

Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformabl...

AI 聚合 06/17
HuggingFace

Human Universal Grasping

Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue ...

AI 聚合 06/17
HuggingFace

ExpRL: Exploratory RL for LLM Mid-Training

Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success dep...

AI 聚合 06/17
HuggingFace

Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera traj...

AI 聚合 06/17
HuggingFace

You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences

Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with wea...

AI 聚合 06/17
HuggingFace

Attacks on Machine-Text Detectors Retain Stylistic Fingerprints

Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be m...

AI 聚合 06/17
GitHub

[GitHub] TauricResearch/TradingAgents

TradingAgents: Multi-Agents LLM Financial Trading Framework(⭐86448)

AI 聚合 06/16
HuggingFace

ActiveMimic: Egocentric Video Pretraining with Active Perception

Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such vid...

AI 聚合 06/16
HuggingFace

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the ...

AI 聚合 06/16
HuggingFace

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

Vision-Language-Action (VLA) models that couple pretrained Vision-Language Models (VLMs) with continuous action exper...

AI 聚合 06/16
HuggingFace

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment

As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisi...

AI 聚合 06/16
HuggingFace

Squeeze-Release: Iterative Pruning with Exact Structural Minimization

Unstructured pruning produces sparse weight tensors, but the standard implementation keeps tensor shapes unchanged so...

AI 聚合 06/16
HuggingFace

LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations

Online group chats are social spaces with local conversational norms that are rarely stated explicitly. The ability a...

AI 聚合 06/16
HuggingFace

iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environme...

AI 聚合 06/16
HuggingFace

Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics

Token-level hallucination detectors are evaluated as classifiers, by AUC over all tokens, yet a streaming monitor is ...

AI 聚合 06/16
HuggingFace

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but ...

AI 聚合 06/16
HuggingFace

FVSpec: Real-World Property-Based Tests as Lean Challenges

We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first...

AI 聚合 06/16
HuggingFace

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models

Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it fr...

AI 聚合 06/16
HuggingFace

Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation

Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven i...

AI 聚合 06/16
HuggingFace

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metr...

AI 聚合 06/16
HuggingFace

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions

As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have f...

AI 聚合 06/16
HuggingFace

AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties

Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and...

AI 聚合 06/16
HuggingFace

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly ...

AI 聚合 06/16
HuggingFace

Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion

Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene ...

AI 聚合 06/16
arXiv

AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet t...

AI 聚合 06/15
arXiv

Regulating the Machine Contributor: Governance and Policy Alignment in Open Source

AI-assisted software development has moved from line-level autocomplete to agents that can plan changes, edit files, ...

AI 聚合 06/15
arXiv

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

Wearable devices and smartphones generate rich behavioural time series that can support proactive health intervention...

AI 聚合 06/15
arXiv

Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts

Survival prediction plays a central role for healthcare providers and clinical researchers. Accurate risk stratificat...

AI 聚合 06/15
arXiv

Moonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven's Op. 27 No. 2 and Machine Learning Mechanisms

We show that the three movements of Beethoven's "Moonlight Sonata" (Op. 27 No. 2) instantiate three distinct machine ...

AI 聚合 06/15
arXiv

When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks

Verifier-driven self-DPO is a common recipe for self-improving production visual-language models. In this setup, a fr...

AI 聚合 06/15
arXiv

From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing

Recent advances in speech generation have significantly improved the naturalness of synthetic speech, making spoofing...

AI 聚合 06/15
arXiv

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their prediction...

AI 聚合 06/15
arXiv

Abstracting Cross-Domain Action Sequences into Interpretable Workflows

Sequential or time-stamped interaction logs provide objective records of digital application usage, yet their granula...

AI 聚合 06/15
arXiv

Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications

Artificial Intelligence (AI) is increasingly used to automate a variety of real-world computer vision (CV) applicatio...

AI 聚合 06/15
arXiv

Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows

Large language models increasingly serve as execution engines for agentic systems, yet they still consume context thr...

AI 聚合 06/15
arXiv

CottonLeafVision: An Explainable and Robust Deep Learning Framework for Cotton Leaf Disease Classification

Globally, cotton is a highly economically beneficial crop, as the textile industry heavily depends on it. So, the pre...

AI 聚合 06/15
arXiv

Flood and Harvest: The Provable Necessity of Trivia for Generating Valuable Mathematics via the Lens of Language Generation in the Limit

AI systems coupled to proof assistants now generate formal mathematics at scale, and the gap between what a checker c...

AI 聚合 06/15
arXiv

Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning

Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, p...

AI 聚合 06/15
arXiv

ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision supp...

AI 聚合 06/15
HuggingFace

μ_0: A Scalable 3D Interaction-Trace World Model

World models that capture how actions induce physical change enable scalable robot learning without reliance on embod...

AI 聚合 06/15
HuggingFace

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control fl...

AI 聚合 06/15
HuggingFace

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full ro...

AI 聚合 06/15
HuggingFace

Avatar V: Scaling Video-Reference Avatar Video Generation

Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, f...

AI 聚合 06/15
HuggingFace

Orchestra-o1: Omnimodal Agent Orchestration

The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-ag...

AI 聚合 06/15
HuggingFace

LLM Agents Can See Code Repositories

Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Ye...

AI 聚合 06/15
HuggingFace

RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space

Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text e...

AI 聚合 06/15
HuggingFace

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and...

AI 聚合 06/15
HuggingFace

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelit...

AI 聚合 06/15
HuggingFace

RedAct: Redacting Agent Capability Traces for Procedural Skill Protection

Users rely on execution traces to observe agent behavior, diagnose failures, and ensure accountability. These traces ...

AI 聚合 06/15
HuggingFace

RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling

Video generation models based on Diffusion Transformers (DiTs) have achieved remarkable performance in video synthesi...

AI 聚合 06/15
HuggingFace

APPO: Agentic Procedural Policy Optimization

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabiliti...

AI 聚合 06/15
HuggingFace

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in re...

AI 聚合 06/15
HuggingFace

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. Wh...

AI 聚合 06/15
HuggingFace

CARVE: Certified Affordable Repair of Vetoed Maneuvers via Envelopes for Interactive Driving

Interactive driving exposes a failure mode that is easy to miss in rule-aware autonomous-driving stacks: a hard-rule ...

AI 聚合 06/15
HuggingFace

Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-rel...

AI 聚合 06/15
HuggingFace

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption tha...

AI 聚合 06/15
HuggingFace

AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization

Large reasoning models typically follow a read-then-think paradigm: they observe the complete input, reason over a st...

AI 聚合 06/15
HuggingFace

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling...

AI 聚合 06/15
HuggingFace

WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis

Large and demographically balanced datasets are essential for reliable neuroimaging biomarkers. Full-resolution 3D br...

AI 聚合 06/15
GitHub

[GitHub] firecrawl/firecrawl

The API to search, scrape, and interact with the web at scale. 🔥(⭐132792)

AI 聚合 06/15
GitHub

[GitHub] langchain-ai/langchain

The agent engineering platform.(⭐139289)

AI 聚合 06/15
GitHub

[GitHub] x1xhlol/system-prompts-and-models-of-ai-tools

FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, No...

AI 聚合 06/15
GitHub

[GitHub] open-webui/open-webui

User-friendly AI Interface (Supports Ollama, OpenAI API, ...)(⭐141524)

AI 聚合 06/15
GitHub

[GitHub] langgenius/dify

Production-ready platform for agentic workflow development.(⭐145209)

AI 聚合 06/15
GitHub

[GitHub] Snailclimb/JavaGuide

Java 面试 & 后端通用面试指南,覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发(⭐156367)

AI 聚合 06/15
GitHub

[GitHub] huggingface/transformers

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, a...

AI 聚合 06/15
GitHub

[GitHub] f/prompts.chat

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-...

AI 聚合 06/15
GitHub

[GitHub] ollama/ollama

Get up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.(⭐174170)

AI 聚合 06/15
GitHub

[GitHub] Significant-Gravitas/AutoGPT

AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so ...

AI 聚合 06/15
GitHub

[GitHub] n8n-io/n8n

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-ho...

AI 聚合 06/15
GitHub

[GitHub] NousResearch/hermes-agent

The agent that grows with you(⭐193566)

AI 聚合 06/15
GitHub

[GitHub] tensorflow/tensorflow

An Open Source Machine Learning Framework for Everyone(⭐195658)

AI 聚合 06/15
GitHub

[GitHub] affaan-m/ECC

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first developmen...

AI 聚合 06/15
GitHub

[GitHub] openclaw/openclaw

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞(⭐378719)

AI 聚合 06/15
arXiv

Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models

Chain-of-thought (CoT) reasoning is the dominant paradigm for inference-time scaling in language models, yet the caus...

自动抓取 06/15
arXiv

Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch

Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decis...

自动抓取 06/15
arXiv

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evid...

自动抓取 06/15
arXiv

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on ...

自动抓取 06/15
arXiv

One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders

Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...

自动抓取 06/15
arXiv

Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks

Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic spe...

自动抓取 06/15
arXiv

Valid Inference with Synthetic Data via Task Exchangeability

There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social sc...

自动抓取 06/15
arXiv

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slav...

自动抓取 06/15
arXiv

Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization

This paper examines three recent frameworks for understanding the cognitive and epistemic consequences of artificial ...

自动抓取 06/15
arXiv

EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery

LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and ...

自动抓取 06/15
arXiv

Agents-K1: Towards Agent-native Knowledge Orchestration

Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledg...

自动抓取 06/15
arXiv

Automated reproducibility assessments in the social and behavioral sciences using large language models

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze...

自动抓取 06/15
arXiv

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fu...

自动抓取 06/15
arXiv

Mana: Dexterous Manipulation of Articulated Tools

Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal ...

自动抓取 06/15
arXiv

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning

Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowle...

自动抓取 06/15
AI前沿

[AI News] Hoi3DGen:生成高质量的三维人-物交互场景

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: Hoi3DGen:生成高质量的三维人-物交互场景 英文名: Hoi3DGen: Generating High-Quality Human-Object-Interacti...

自动抓取 03/13
AI前沿

[AI News] 跨上下文审查:通过分离生成与审查会话提升大语言模型输出质量

来源: arXiv 爆炸度: 7/10 💥💥💥💥💥 中文名: 跨上下文审查:通过分离生成与审查会话提升大语言模型输出质量 英文名: Cross-Context Review: Improving LLM Output Quality ...

自动抓取 03/13
AI前沿

[AI News] 大语言模型应用精选集

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 大语言模型应用精选集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个收录了基于各类大语言模型(包括OpenAI、Anthropi...

自动抓取 03/13
AI前沿

[AI News] 基于机器学习与组合融合分析的NCAA锦标赛对阵预测

来源: arXiv 爆炸度: 7/10 💥💥💥💥💥 中文名: 基于机器学习与组合融合分析的NCAA锦标赛对阵预测 英文名: NCAA Bracket Prediction Using Machine Learning and Comb...

自动抓取 03/12
AI前沿

[AI News] LLM2Vec-Gen:从大语言模型生成式嵌入

来源: arXiv 爆炸度: 7/10 💥💥💥💥💥 中文名: LLM2Vec-Gen:从大语言模型生成式嵌入 英文名: LLM2Vec-Gen: Generative Embeddings from Large Language Mo...

自动抓取 03/12
AI前沿

[AI News] 超棒的大型语言模型应用集合

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超棒的大型语言模型应用集合 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个收录了基于AI智能体和检索增强生成技术的各类LLM应用...

自动抓取 03/12
AI前沿

[AI News] 全迷你语言模型L6-v2

来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 全迷你语言模型L6-v2 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 轻量高效的通用语义嵌入模型,在速...

自动抓取 03/11
AI前沿

[AI News] 大型语言模型应用精选集

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 大型语言模型应用精选集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个系统化整理当前最热门大型语言模型应用实例的开源项目库,涵盖...

自动抓取 03/11
AI前沿

[AI News] 全句嵌入MiniLM-L6-v2模型

来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: 全句嵌入MiniLM-L6-v2模型 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 一款高效轻量的通用句...

自动抓取 03/10
AI前沿

[AI News] AI工具的系统提示词与模型配置库

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: AI工具的系统提示词与模型配置库 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 逆向解析主流AI编程...

自动抓取 03/10
AI前沿

[AI News] 超赞的大语言模型应用合集

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超赞的大语言模型应用合集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个系统化收集和分类各类大语言模型实际应用案例的开源知识库,...

自动抓取 03/10
AI前沿

[AI News] Artificial Intelligence for Detecting Fetal Orofac

来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: Artificial Intelligence for Detecting Fetal Orofacial Clefts and Advancing Medical Edu...

自动抓取 03/09
AI前沿

[AI News] SG-DOR: Learning Scene Graphs with Direction-Condi

来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: SG-DOR: Learning Scene Graphs with Direction-Conditioned Occlusion Reasoning for Peppe...

自动抓取 03/09
AI前沿

[AI News] google-gemini/gemini-cli

来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: google-gemini/gemini-cli 英文名: google-gemini/gemini-cli AI 总结: ⭐ 96812 | 🍴 12008 | 👁️...

自动抓取 03/08
AI前沿

[AI News] Transformer-Based Inpainting for Real-Time 3D Stre

来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups ...

自动抓取 03/07
AI前沿

[AI News] FaceCam: Portrait Video Camera Control via Scale-A

来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning 英文名: FaceCam: Port...

自动抓取 03/07
AI前沿

[AI News] SimpliHuMoN: Simplifying Human Motion Prediction

来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: SimpliHuMoN: Simplifying Human Motion Prediction 英文名: SimpliHuMoN: Simplifying Human M...

自动抓取 03/06
AI前沿

[AI News] Accurate and Efficient Hybrid-Ensemble Atmospheric

来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: Accurate and Efficient Hybrid-Ensemble Atmospheric Data Assimilation in Latent Space w...

自动抓取 03/06
AI前沿

[AI News] openclaw/openclaw

来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: openclaw/openclaw 英文名: openclaw/openclaw AI 总结: ⭐ 260218 | 🍴 49883 | 👁️ 260218 Your ...

自动抓取 03/05
AI前沿

[AI News] x1xhlol/system-prompts-and-models-of-ai-tools

来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: x1xhlol/system-prompts-and-models-of-ai-tools 英文名: x1xhlol/system-prompts-and-models...

自动抓取 03/05
AI前沿

[AI News] Shubhamsaboo/awesome-llm-apps

来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: Shubhamsaboo/awesome-llm-apps 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: ⭐ 99599 | 🍴 ...

自动抓取 03/05
AI前沿

[AI News] HiFi-Inpaint:面向高保真参考修复的细节保持型人-物图像生成

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: HiFi-Inpaint:面向高保真参考修复的细节保持型人-物图像生成 英文名: HiFi-Inpaint: Towards High-Fidelity Reference...

自动抓取 03/04
AI前沿

[AI News] 推理核心:一个用于符号预训练与后训练的可扩展程序化数据生成套件

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: 推理核心:一个用于符号预训练与后训练的可扩展程序化数据生成套件 英文名: Reasoning Core: A Scalable Procedural Data Genera...

自动抓取 03/04
AI前沿

[AI News] AI工具的系统提示词与模型库

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: AI工具的系统提示词与模型库 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 开源逆向工程项目,收集并...

自动抓取 03/04
AI前沿

[AI News] 惊艳的LLM应用合集

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 惊艳的LLM应用合集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个系统整理大型语言模型(LLM)在AI智能体与RAG等领域实际...

自动抓取 03/04
AI前沿

[AI News] UFO-4D:从两张无位姿图像进行前馈式4D重建

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: UFO-4D:从两张无位姿图像进行前馈式4D重建 英文名: UFO-4D: Unposed Feedforward 4D Reconstruction from Two I...

自动抓取 03/03
AI前沿

[AI News] 模式寻求与均值寻求相遇:快速生成长视频的新方法

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: 模式寻求与均值寻求相遇:快速生成长视频的新方法 英文名: Mode Seeking meets Mean Seeking for Fast Long Video Gener...

自动抓取 03/03
AI前沿

[AI News] 通用句子嵌入模型

来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 通用句子嵌入模型 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 一个高效轻量的句子嵌入模型,在语义理解任...

自动抓取 03/02
AI前沿

[AI News] BERT基础未标注版

来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: BERT基础未标注版 英文名: google-bert/bert-base-uncased AI 总结: 谷歌发布的经典双向Transformer预训练模型,奠...

自动抓取 03/02
AI前沿

[AI News] OpenClaw(开源AI助手)

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: OpenClaw(开源AI助手) 英文名: openclaw/openclaw AI 总结: 一个开源、跨平台、可高度定制的个人AI助手,支持多种大模型和功能,旨在成为用...

自动抓取 03/02
AI前沿

[AI News] 超赞大语言模型应用集锦

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超赞大语言模型应用集锦 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个精心整理的、涵盖AI智能体与RAG技术的大语言模型应用实战...

自动抓取 03/02
AI前沿

[AI News] OpenClaw - 个人AI助手

来源: GitHub 爆炸度: 8/10 💥💥💥💥💥 中文名: OpenClaw - 个人AI助手 英文名: openclaw/openclaw AI 总结: 一个跨操作系统和平台的个人AI助手项目,采用龙虾(Claw)作为品牌标识,...

自动抓取 03/01
AI前沿

[AI News] AI工具系统提示词与模型集合

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: AI工具系统提示词与模型集合 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 开源逆向工程项目,系统整...

自动抓取 03/01
AI前沿

[AI News] Gemini 命令行工具

来源: GitHub 爆炸度: 8/10 💥💥💥💥💥 中文名: Gemini 命令行工具 英文名: google-gemini/gemini-cli AI 总结: Google官方开源的终端AI助手,让你在命令行中直接使用Gemini...

自动抓取 03/01
AI前沿

[AI News] 全MiniLM-L6-v2句子嵌入模型

来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 全MiniLM-L6-v2句子嵌入模型 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 轻量高效的通用句子...

自动抓取 03/01
AI前沿

[AI News] BERT基础未分词版

来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: BERT基础未分词版 英文名: google-bert/bert-base-uncased AI 总结: Google官方发布的英语基础BERT模型,是NLP领...

自动抓取 03/01
AI前沿

[AI News] MediX-R1:开放式医学强化学习框架

来源: arXiv 爆炸度: 3/10 💥💥💥 中文名: MediX-R1:开放式医学强化学习框架 英文名: MediX-R1: Open Ended Medical Reinforcement Learning AI 总结: 该研究...

自动抓取 03/01
AI前沿

[AI News] VGG-T³:大规模离线前馈式三维重建

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: VGG-T³:大规模离线前馈式三维重建 英文名: VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale AI...

自动抓取 03/01
AI前沿

[AI News] OpenClaw(开源爪)

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: OpenClaw(开源爪) 英文名: openclaw/openclaw AI 总结: OpenClaw是一个开源、跨平台、可高度定制的个人AI助手项目,旨在让用户在任何...

自动抓取 03/01
AI前沿

[AI News] AI工具系统提示词与模型库

来源: GitHub 爆炸度: 8/10 💥💥💥💥💥 中文名: AI工具系统提示词与模型库 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 开源仓库汇总了主流AI编程...

自动抓取 03/01
AI前沿

[AI News] 超赞的LLM应用集合(含AI智能体与RAG)

来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超赞的LLM应用集合(含AI智能体与RAG) 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 这是一个系统整理基于大语言模型的AI智能...

自动抓取 03/01
AI前沿

[AI News] 通用小型句子嵌入模型(MiniLM-L6-v2)

来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 通用小型句子嵌入模型(MiniLM-L6-v2) 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 这是一个...

自动抓取 03/01
AI前沿

[AI News] 谷歌BERT基础模型(不区分大小写)

来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: 谷歌BERT基础模型(不区分大小写) 英文名: google-bert/bert-base-uncased AI 总结: 这是一个由谷歌在2018年发布的基于T...

自动抓取 03/01
AI前沿

[AI News] ELECTRA基础判别器模型 - 基于判别式预训练的文本编码器

来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: ELECTRA基础判别器模型 - 基于判别式预训练的文本编码器 英文名: google/electra-base-discriminator AI 总结: EL...

自动抓取 03/01
AI前沿

[AI News] 基于锚定的模型一致性研究

来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: 基于锚定的模型一致性研究 英文名: Model Agreement via Anchoring AI 总结: 该论文研究如何通过锚定机制控制两个独立训练的机器学习模型在实值...

自动抓取 03/01
AI前沿

[AI News] sentence-transformers/all-MiniLM-L6-v2(HuggingFace

来源: HuggingFace 爆炸度: 6/10 🔴🔴🔴🔴🔴 标题: sentence-transformers/all-MiniLM-L6-v2(HuggingFace) 英文: sentence-transformers/all...

自动抓取 03/01
AI前沿

[AI News] google-bert/bert-base-uncased(HuggingFace)

来源: HuggingFace 爆炸度: 7/10 🔴🔴🔴🔴🔴 标题: google-bert/bert-base-uncased(HuggingFace) 英文: google-bert/bert-base-uncased 摘要: ...

自动抓取 03/01
AI前沿

[AI News] google/electra-base-discriminator(HuggingFace)

来源: HuggingFace 爆炸度: 6/10 🔴🔴🔴🔴🔴 标题: google/electra-base-discriminator(HuggingFace) 英文: google/electra-base-discrimina...

自动抓取 03/01
AI前沿

[AI News] Falconsai/nsfw_image_detection(HuggingFace)

来源: HuggingFace 爆炸度: 5/10 🔴🔴🔴🔴🔴 标题: Falconsai/nsfw_image_detection(HuggingFace) 英文: Falconsai/nsfw_image_detection 摘要...

自动抓取 03/01
AI前沿

[AI News] sentence-transformers/all-mpnet-base-v2(HuggingFac

来源: HuggingFace 爆炸度: 7/10 🔴🔴🔴🔴🔴 标题: sentence-transformers/all-mpnet-base-v2(HuggingFace) 英文: sentence-transformers/al...

自动抓取 03/01
AI前沿

[AI News] MediX-R1:开放式医疗强化学习(arXiv)

来源: arXiv 爆炸度: 8/10 🔴🔴🔴🔴🔴 标题: MediX-R1:开放式医疗强化学习(arXiv) 英文: MediX-R1: Open Ended Medical Reinforcement Learning 摘要: W...

自动抓取 03/01
AI前沿

[AI News] VGG-T³:大规模离线前馈3D重建(arXiv)

来源: arXiv 爆炸度: 7/10 🔴🔴🔴🔴🔴 标题: VGG-T³:大规模离线前馈3D重建(arXiv) 英文: VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scal...

自动抓取 03/01
AI前沿

[AI News] 基于锚定的模型一致性方法(arXiv)

来源: arXiv 爆炸度: 6/10 🔴🔴🔴🔴🔴 标题: 基于锚定的模型一致性方法(arXiv) 英文: Model Agreement via Anchoring 摘要: Numerous lines of aim to cont...

自动抓取 03/01
AI前沿

[AI News] SeeThrough3D:文本到图像生成中的遮挡感知3D控制(arXiv)

来源: arXiv 爆炸度: 8/10 🔴🔴🔴🔴🔴 标题: SeeThrough3D:文本到图像生成中的遮挡感知3D控制(arXiv) 英文: SeeThrough3D: Occlusion Aware 3D Control in T...

自动抓取 03/01
AI前沿

[AI News] 一个数据集仅需1MB(arXiv)

来源: arXiv 爆炸度: 9/10 🔴🔴🔴🔴🔴 标题: 一个数据集仅需1MB(arXiv) 英文: A Dataset is Worth 1 MB 摘要: A dataset server must often distribut...

自动抓取 03/01
AI前沿

[AI News] sentence-transformers/all-MiniLM-L6-v2

来源: HuggingFace 标题: sentence-transformers/all-MiniLM-L6-v2 摘要: 下载: 187393215 | 点赞: 4520 链接: https://huggingface.co/se...

自动抓取 03/01
AI前沿

[AI News] google-bert/bert-base-uncased

来源: HuggingFace 标题: google-bert/bert-base-uncased 摘要: 下载: 60873623 | 点赞: 2578 链接: https://huggingface.co/google-bert/...

自动抓取 03/01
AI前沿

[AI News] google/electra-base-discriminator

来源: HuggingFace 标题: google/electra-base-discriminator 摘要: 下载: 48227365 | 点赞: 83 链接: https://huggingface.co/google/ele...

自动抓取 03/01
AI前沿

[AI News] Falconsai/nsfw_image_detection

来源: HuggingFace 标题: Falconsai/nsfw_image_detection 摘要: 下载: 40235989 | 点赞: 1000 链接: https://huggingface.co/Falconsai/n...

自动抓取 03/01
AI前沿

[AI News] sentence-transformers/all-mpnet-base-v2

来源: HuggingFace 标题: sentence-transformers/all-mpnet-base-v2 摘要: 下载: 26331045 | 点赞: 1253 链接: https://huggingface.co/se...

自动抓取 03/01
AI前沿

[AI News] Model Agreement via Anchoring

来源: arXiv 标题: Model Agreement via Anchoring 摘要: Numerous lines of aim to control $\textit{model disagreement}$ -- the...

自动抓取 03/01
AI前沿

[AI News] SeeThrough3D: Occlusion Aware 3D Control in Text-t

来源: arXiv 标题: SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation 摘要: We identify occlusion reasonin...

自动抓取 03/01
AI前沿

[AI News] A Dataset is Worth 1 MB

来源: arXiv 标题: A Dataset is Worth 1 MB 摘要: A dataset server must often distribute the same large payload to many clien...

自动抓取 03/01
AI前沿

[AI News] SOTAlign: Semi-Supervised Alignment of Unimodal Vi

来源: arXiv 标题: SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport 摘要: Th...

自动抓取 03/01
AI前沿

[AI News] FlashOptim: Optimizers for Memory Efficient Traini

来源: arXiv 标题: FlashOptim: Optimizers for Memory Efficient Training 摘要: Standard mixed-precision training of neural ne...

自动抓取 03/01
AI前沿

[AI News] Error

来源: arXiv 标题: Error 摘要: sortOrder must be in: ascending, descending 链接: https://arxiv.org/help/api/user-manual#sort

自动抓取 03/01