Efficient Test-Time Adaptation through Human-AI Interaction
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet...
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet...
This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniatu...
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...
Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them...
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to la...
Repository-level software engineering benchmarks have significantly advanced the evaluation of coding agents, but exi...
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and bui...
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...
Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical...
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit ...
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only...
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editi...
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, produ...
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurem...
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, execut...
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate,...
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become m...
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-t...
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing...
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. E...
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thoug...
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected sce...
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera t...
Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization...
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs a...
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation...
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often ...
Personalized assistants should not only comply with user requests but also assess whether those requests are appropri...
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invar...
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remot...
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 p...
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fi...
Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in lon...
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origam...
Multimodal models often build on architectures designed for generative vision-language modeling, typically combining ...
Spatial return models take the interaction matrix as given and leave feedback uninterpreted. We construct a bandwidth...
Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are diffic...
Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of...
Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring res...
Agent performance depends jointly on the model parameters and the executable harness code that manages context and co...
A language model's prediction of its next token develops across layers, and lens methods track this process by decodi...
Conversational Recommender Systems (CRS) typically require domain-specific dialogue data, which is costly, scarce, an...
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...
Polymeric materials are central to modern technologies, with applications ranging from energy to health and transport...
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve throu...
Large language models now answer medical questions with expert-level performance. However, the context these systems ...
On-premise assistants can give factory workers conversational access to machine documentation, but models capable of ...
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention va...
The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with th...
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic foreca...
For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record ha...
Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations i...
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupationa...
Competitive programming has become a key test of large language model reasoning, with international competitions such...
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investiga...
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resul...
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LL...
Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area la...
Competitive programming has become a key test of large language model reasoning, with international competitions such...
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet ex...
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to fi...
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single ...
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model ...
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral e...
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at re...
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external ex...
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts m...
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe com...
An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language ...
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-t...
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-...
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize...
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable ima...
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involv...
LLMs are increasingly deployed as orchestrators that coordinate specialized subagents to solve complex tasks through ...
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely contr...
Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decom...
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step ...
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video...
Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption ...
Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-...
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger ...
Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, o...
Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrict...
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated Gi...
This paper explores whether large language models (LLMs) can design video coding tools, a highly challenging task due...
As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critic...
Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hy...
Automated lesion segmentation in whole-body PET/CT is complicated by the variety of physiological tracer uptake patte...
We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that sha...
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world m...
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the ...
Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as polic...
How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during L...
Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals ...
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible a...
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent...
Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility bri...
The repository-level code generation task requires synthesizing code that satisfies task requirements while remaining...
Evaluating software engineering agents on realistic benchmarks is costly, since each task may require multi-step code...
Traditional frameworks of political communication operate under linear, event-driven assumptions that treat voter per...
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between ...
SAR-to-EO image translation aims to generate electro-optical (EO) imagery from synthetic aperture radar (SAR) observa...
Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve ...
A central bottleneck in multi-hop Question Answering (QA) is that the granularity at which a question is expressed of...
While unified multimodal models (UMMs) jointly perform visual understanding and generation within a single model, fun...
Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled ro...
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends...
AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but eviden...
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world m...
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environ...
Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, ...
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Dr...
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchm...
This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirement...
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at f...
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt o...
Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can pl...
Cognitive language agents have achieved substantial progress by equipping language models with memory, tools, and dec...
AI is increasingly used in the R\&D process that produces future AI systems. We study the conditions under which this...
End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth ...
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an exe...
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model...
Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we ...
Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-ca...
Multimodal misinformation on social media is highly prevalent, potent, and harmful, yet difficult to detect and count...
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose...
Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model mu...
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an e...
Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated sta...
When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, t...
We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n...
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR...
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literatu...
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-lev...
Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and P...
Accurate daily predictions of cold hardiness in woody plants are critical in regions where freezing temperatures can ...
Industrial post-training is a brownfield regime. Teams inherit a deployed checkpoint and must land targeted improveme...
Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deploymen...
The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently character...
Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural ali...
The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platf...
Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: co...
Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop b...
Composable scene modeling aims to recover a real indoor scene as complete, editable object assets arranged as observe...
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B paramet...
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al.,...
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interact...
While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its train...
Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its e...
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically ...
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is sco...
Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, ...
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction ...
Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabiliti...
Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating i...
VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that j...
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across task...
While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, the...
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant be...
Organizations often develop and maintain portfolios of related applications: independently deployable codebases that ...
Recent work on image content manipulation based on vision-language pre-training models has been effectively extended ...
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual...
Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing e...
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are sti...
Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned com...
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy. Every new ...
Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the ta...
LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. S...
Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state ...
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what th...
This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare fi...
Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypo...
Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet mon...
AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly ext...
When training Mixture-of-Experts (MoE) language models with expert parallelism, all-to-all token dispatch and combine...
Neural operators provide fast surrogate models for approximating operators between function spaces, but their predict...
We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodi...
Texture image classification plays a significant role in computer vision applications, including industrial inspectio...
Recent advances in generative AI allow users to create 3D models from text or images. However, these models prioritiz...
A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wr...
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as ima...
Modern agent systems assemble capabilities at runtime, and this dynamic composition has recently received a complete ...
Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design...
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as...
Tendon-driven hands are anthropomorphic, and moving the actuators off the joints is what makes a hand of this capabil...
Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and...
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are...
We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete oc...
Large language models (LLMs) are increasingly deployed with layered defenses, yet malicious prompts can still bypass ...
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as ima...
Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from ...
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces...
Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online un...
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage ...
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fix...
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a ...
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of th...
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent conte...
Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory...
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recur...
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the targ...
Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous...
Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of ...
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often le...
Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demog...
Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis o...
Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily lice...
Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon....
Recent advances in inference-time scaling have significantly improved the reasoning performance of large language mod...
High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To supp...
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tra...
Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer en...
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is...
Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint perfo...
Video carries the temporal structure of the physical world, yet learning representations from it has remained computa...
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploi...
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a...
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing th...
Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsel...
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifac...
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentati...
Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model th...
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful...
In real-world software development, code review typically involves iterative interactions between developers and revi...
To improve large language models' ability to resolve real-world software issues, prior work has focused on constructi...
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. R...
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agent...
Modern game development relies heavily on conventional graphics pipelines. High-quality visual content requires model...
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. Howev...
On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory s...
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. R...
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remai...
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local ...
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more th...
Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal ...
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies i...
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they...
Explicit visual intermediates can help multimodal large language models (MLLMs) externalize spatial evidence and upda...
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), h...
Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous n...
While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or shor...
Native 3D generators now recover impressive mesh geometry from a single image. However, a dense mesh stays soft where...
A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this st...
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvem...
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducibl...
The development of 0.1^{circ} global weather forecasting models based on machine learning (ML) is constrained by the ...
Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer ...
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their sk...
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing be...
On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without a...
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-corr...
We consider optimization applications with unknown parameters where the decision maker believes that the optimal valu...
Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a ta...
Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requir...
Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer ...
Intent misinterpretation during vehicle interactions causes recurring planning failures. We study a decision layer in...
Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions t...
Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning...
Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning d...
Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic v...
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying...
Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, o...
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed...
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and vi...
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through g...
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning st...
Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single u...
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art model...
Multi-arm collaboration is becoming a core capability in embodied manipulation. Recent vision-language-action (VLA) m...
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, cha...
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While uni...
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empatheti...
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objecti...
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution ...
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evalu...
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability...
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring m...
Coding agents perform long-running tasks spanning dozens of model calls, tool uses, and code edits. As these runs unf...
Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbo...
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchr...
Document retrieval increasingly supports high-stakes information access in finance, healthcare, and law. Modern retri...
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to su...
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and la...
Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient...
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recog...
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating str...
LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long unstructured traces re...
Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel ...
Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files,...
We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents f...
Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream sta...
Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existi...
The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tas...
Brain stroke, known for its high mortality and incidence rates, poses significant health risks and requires rapid int...
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces...
Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain pla...
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize inter...
We present StarHarness, a framework for evolving environment-specific agent harnesses while keeping model weights fix...
Model cards are structured documents that summarize key information about machine learning models to improve transpar...
Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existing approaches sti...
Large language models are increasingly used for knowledge graph question answering (KGQA), but can fail to correctly ...
The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified...
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI...
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific...
Real-world data for knowledge graph question answering is often distributed across different organizations due to gov...
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and var...
Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state a...
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the...
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize inter...
Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level ...
Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state a...
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint r...
LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interaction...
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted...
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...
Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithf...
Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementa...
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to...
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific...
Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect...
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, pos...
Video games provide a scalable source of training data for video world models, offering diverse environments, complex...
As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has draw...
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observati...
Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the...
Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which se...
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a f...
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does ...
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining int...
As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bot...
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-spe...
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover re...
We formalize prefix invariance: representations at position t must not depend on future inputs. We give a lightweight...
Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathemat...
Historical people may appear under different languages, scripts, and transcription traditions, while distinct individ...
Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as t...
Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing...
World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen Dreame...
The continual evolution of malware variants necessitates detection systems that can adapt to new threats without retr...
Does multi-agent LLM interaction help or hurt? Some work reports gains from debate (Du et al., 2024), critique loops ...
While AI assistance can improve human task performance in the short term, it may also undermine the development of sk...
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive wi...
Language models are sequential processors, but long-horizon agency requires external information and computation beyo...
Ballistocardiography (BCG) is promising for unobtrusive long-term blood pressure (BP) monitoring in laboratory settin...
Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety audit...
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and la...
An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time....
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling...
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing stude...
Language models are sequential processors, but long-horizon agency requires external information and computation beyo...
An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time....
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. Howev...
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requ...
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarificatio...
We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness...
Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module,...
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this ...
While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains cha...
Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation...
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation whi...
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an...
Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundat...
With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingl...
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, ...
User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavio...
Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devi...
Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organ...
Human-centric intelligence is evolving in the foundation-model era, with growing emphasis on scale, transferability, ...
Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle e...
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution path...
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not...
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel mo...
Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We...
We present PhysCaP, a Physics-Informed Code-as-Policy agent for active perception in robotic manipulation. While visi...
Game world models have recently demonstrated promising capabilities in generating visually coherent and action-contro...
Recently, there has been a great deal of research into improving AI methods and their application. The main focus is ...
Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future ...
The Traveling Salesman Problem (TSP) is one of the most extensively studied NP-hard optimization problems. Genetic Al...
Sequential recommendation predicts the next item from a user's interaction history, but not every interaction is equa...
Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multipl...
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safet...
Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support...
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encour...
The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy reg...
We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minim...
Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly ...
Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases...
For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here...
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy image...
We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. ...
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute ...
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative think...
With large pretrained models, existing methods have effectively improved instruction-based video editing. However, mo...
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safet...
LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolutio...
Simulating realistic user shopping behavior underpins offline evaluation and reinforcement learning in e-commerce sce...
We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, an...
Agents learn to act through interaction with environments, yet the environments used for training are often manually ...
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continu...
Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computa...
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's ow...
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to ...
Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite:...
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditionin...
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tun...
Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be gras...
Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to general...
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffol...
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched ...
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coheren...
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-a...
We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 param...
Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures di...
Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-s...
The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses ...
Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to pe...
X-band SAR satellites (8-12 GHz) play a critical role in disaster response, environmental monitoring, and military in...
Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather tha...
The rapid proliferation of memecoins on blockchain platforms has increased the risk of fraudulent activities, particu...
Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable wi...
Large language models often fail to answer questions about a bounded document collection when the source documents ar...
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual prob...
Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Re...
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improv...
Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that ...
Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resour...
Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are...
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressi...
Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures ari...
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation polic...
LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agen...
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people....
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native con...
Memory has become a key component of large language models, enabling them to retain information and learn from long-t...
We present 4DAnyone, a framework for reconstructing 4D humans from an uncalibrated monocular video by generating reco...
Large language models often fail to answer questions about a bounded document collection when the source documents ar...
Large language model agents have made substantial progress in code generation, yet most existing systems assume a pre...
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no ...
Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention re...
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the mod...
Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capab...
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature...
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing meth...
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evol...
LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start...
Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the ...
Object detectors often produce over-confident predictions for objects outside their training categories, leading to s...
LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed i...
Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game develo...
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing met...
Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of...
Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause tar...
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized...
Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the com...
Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug di...
When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class r...
A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than...
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output c...
Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable ...
Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, crea...
Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict wit...
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by v...
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger t...
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for...
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language ...
Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resoluti...
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over lon...
Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use...
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language ...
Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independe...
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-...
Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented...
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formula...
Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success ...
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (...
Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not...
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows,...
High-quality creative writing data for large language models (LLMs) remains dominated by story-centric data, limiting...
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-languag...
Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the ge...
Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM...
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, whi...
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the...
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-speci...
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integrati...
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model re...
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remai...
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and ...
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement b...
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although r...
Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotation...
The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand indi...
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict...
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of L...
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, c...
GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete to...
MRI reconstruction methods for undersampled k-space data naturally utilize complex-valued measurements. Parallel deve...
AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code rep...
Urban traffic congestion reduces productivity and increases travel cost and emissions. Network-wide live travel-time ...
Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet thei...
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintain...
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet c...
We assess indirect prompt injection in DeepSeek Harness (DSH), using AI-Infra-Guard (A.I.G) to construct tests, deliv...
Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) ...
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet c...
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and...
Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance o...
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment...
Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastruct...
AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full pa...
Graph neural networks are commonly described through family-specific equations whose notation obscures shared computa...
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent sta...
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a criti...
Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable ...
The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets m...
High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models...
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieva...
Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structur...
Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves p...
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different...
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-S...
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although...
This first release of Prior Labs in relational learning shows our continued commitment to open science. We open-sourc...
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whet...
Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fie...
Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Exist...
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track o...
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, tempo...
Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with ...
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome r...
We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report...
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition reso...
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to t...
Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step ...
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: acce...
Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but exis...
Generative pretraining established reusable task representations; later work on language-based task conditioning and ...
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually w...
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-la...
The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on m...
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control...
A language model's output does not by itself provide verifiable evidence about the internal computation that produced...
We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbol...
The current best bounds on the matrix multiplication exponent $ω$ are obtained through a refinement of the laser meth...
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VL...
Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs)...
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize...
The current best bounds on the matrix multiplication exponent ω are obtained through a refinement of the laser method...
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Althoug...
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justi...
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form w...
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as th...
We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any...
In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on rout...
Boundary representation (B-Rep) generation is a fundamental task in computer-aided design, yet the direct synthesis o...
As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This...
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training languag...
Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strate...
Learning to generate or reconstruct explorable worlds requires video paired with more than RGB: camera motion, scene ...
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain...
Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended...
AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landsc...
Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resultin...
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it spe...
Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interfac...
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We...
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier f...
Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make mu...
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a...
The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, ...
Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented traini...
Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers...
In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations ...
Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet rea...
Short-horizon forecasting of fine particulate matter (PM2.5) remains difficult when observations from the target doma...
Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expe...
As heterogeneous robotic systems deploy across diverse urban zones, maintaining safety amid complex human-robot inter...
We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable wor...
Network-level maintenance planning requires repeated evaluations of equilibrium traffic flows under road capacity red...
Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce...
Free energies govern solid-state phase stability, yet computational materials discovery still relies largely on groun...
Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, ...
Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. T...
High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol spa...
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. ...
This study investigates the methodological and theoretical properties of session handover in applications that use la...
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing st...
Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging p...
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing corres...
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reli...
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-wor...
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with wh...
As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per...
We show that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while ...
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation...
With the rapid advancement of image editing models and their widespread application across various domains, there is ...
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause th...
The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persiste...
Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule...
Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-ev...
LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is conf...
When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly...
Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently ha...
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skil...
On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, ...
Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, st...
Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to tar...
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and ...
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step(⭐102694)
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding...
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on st...
Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-n...
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This...
Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step di...
Rib fractures are common and time-consuming to localize on computed tomography (CT). We ask whether fractures detecte...
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficien...
We introduce , a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention ...
Modern image classification models excel when trained on single task-specific datasets but often struggle to generali...
Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was use...
Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that re...
We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM promp...
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those ...
When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rathe...
This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive ge...
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier f...
We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \e...
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated cod...
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skil...
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched ...
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with wh...
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows,...
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a ...
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through mode...
Video world models simulate future states conditioned on current observations and user actions. Recent systems have d...
Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory sys...
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a ...
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two arch...
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control f...
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To im...
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed f...
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essen...
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive ...
Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation...
Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose s...
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source...
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within ...
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet exi...
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, ...
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recent...
Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disrup...
Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffu...
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evi...
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packag...
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely f...
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primar...
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inhere...
Large language models often fail when answer options require combining atomic judgments under explicit logical operat...
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-...
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VI...
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases ca...
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simu...
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. F...
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction...
Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT ...
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to acc...
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benc...
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lac...
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial i...
Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional o...
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only...
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan futur...
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitu...
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design...
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce...
World modeling is an unsettled field: architectures, training objectives, and state representations interact in compl...
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet exis...
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vis...
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's...
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the la...
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. Howev...
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and flui...
Recent Vision Foundation Models (VFMs) predict depth, camera pose, and pointmap in a single forward pass without per-...
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely l...
AI agents operate in persistent environments where early state changes can influence decisions far into the future. U...
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agenti...
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these ...
Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the interme...
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulati...
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on mod...
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome,...
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities a...
The application of computer vision in agriculture has shown significant potential for improving crop monitoring and p...
Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many con...
We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses se...
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the ...
Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems,...
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository reti...
Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon set...
Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tenso...
We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution b...
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021,...
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can...
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters af...
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards...
Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are ...
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: te...
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoi...
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mo...
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging pers...
Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualizatio...
Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often boun...
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the ...
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain...
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchma...
The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis,...
Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of...
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can b...
We study reference-free post-training for multilingual machine translation with open large language models. Starting ...
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but cont...
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, compl...
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus ...
Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how ...
Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained ...
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of cap...
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a ph...
The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attenti...
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone...
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs pres...
Gradient descent on a factored model W = UV^top is implicitly biased toward low-rank solutions, while Adam, starting ...
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, b...
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context...
Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be un...
In evolutionary algorithms powered by language models, the LLM acts as a single operator that simultaneously updates ...
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they cl...
Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in pen...
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirr...
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals...
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating la...
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs)....
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond ...
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to syntactic stru...
GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each t...
Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have respo...
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of su...
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effe...
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to prot...
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation...
Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar technique...
Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must ...
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that ...
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs pres...
Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by...
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounter...
Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from ...
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation framework...
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause th...
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Model...
Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-ind...
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation...
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 bil...
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are ra...
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments...
User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generati...
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confine...
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation...
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however...
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning...
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for...
Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-conte...
Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explain...
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems prim...
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploy...
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to prot...
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interacti...
Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a u...
We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data col...
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment...
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints...
When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio...
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's inter...
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Lega...
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the...
Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and...
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious eth...
CLI-based software-engineering agents have matured rapidly, yet the open ecosystem has converged on a single training...
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are ...
Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence le...
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies wit...
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open wheth...
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, ...
Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a ...
LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical in...
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also crea...
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoisin...
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to...
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition fas...
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents ...
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolv...
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive mem...
Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and gover...
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lig...
The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistr...
What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a cou...
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoi...
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diver...
Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by a...
The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent stud...
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers s...
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio...
Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vec...
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal...
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action pred...
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher...
Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-tim...
Autoregressive models accumulate error over long rollouts, yet at deployment there is no ground truth to measure it a...
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. ...
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable ...
Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and rol...
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain di...
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing m...
Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply incr...
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorize...
Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically vi...
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherent...
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However...
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yid...
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by dee...
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattere...
Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an ag...
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed represen...
World models have attracted significant attention for their ability to capture and predict the structure and dynamics...
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action,...
Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene e...
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) ...
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but...
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and ...
White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associat...
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark ...
From natural-language query interfaces to automated report generation, data analysis tools need a description of the ...
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading error...
This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evalua...
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mecha...
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and v...
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, ...
Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_...
The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including f...
Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for ...
Language models increasingly condition their answers on external signals, and a single misleading one can turn a corr...
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiri...
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogen...
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, ...
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception fro...
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jo...
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fai...
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately cha...
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous informa...
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesi...
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, over...
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one au...
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused...
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fi...
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet t...
Current video world models struggle in multiplayer environments because they entangle world state with view-dependent...
Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modelin...
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, desp...
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despi...
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through par...
Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoisi...
A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what...
Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly co...
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textu...
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable ...
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic pr...
Collaborative music agents need internal representations rich enough to support both understanding and generation, ye...
Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variabi...
Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks s...
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, ...
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggrega...
Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new fac...
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate e...
AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions acr...
Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum genera...
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally...
Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference tra...
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains cha...
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in mul...
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, desp...
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, ...
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and p...
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersi...
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing ...
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean perfor...
Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the enviro...
Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain,...
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains...
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense to...
LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are ...
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compound...
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as ...
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language...
This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research...
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric hu...
As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross...
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore cri...
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visua...
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks m...
GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction...
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate e...
Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect r...
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matchin...
Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist...
We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows flo...
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability...
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically...
Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future querie...
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension wi...
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large...
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While ...
Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but prac...
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can f...
Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach aba...
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames se...
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equ...
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that...
Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, ...
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. Howev...
Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-languag...
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, under...
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines...
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmi...
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic...
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large...
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video stream...
Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed progr...
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "t...
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, exi...
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate under...
Real-time video editing requires low-latency causal generation with bounded computational resources while preserving ...
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, exi...
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. Howeve...
Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in contin...
Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks...
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learnin...
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI ag...
Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift rai...
Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in ...
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environm...
World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual d...
Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robo...
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video stream...
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. W...
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling beha...
Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems ...
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it...
Video world models predict future observations conditioned on historical observations and control signals, enabling l...
We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behavior...
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and w...
Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging becau...
Geospatial foundation models aim to learn representations that transfer across regions and sensors, yet evaluating th...
Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-...
Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries an...
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors...
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalit...
LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to opera...
World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on ...
Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing...
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilitie...
Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-...
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. How...
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid in...
Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging becau...
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and mo...
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet th...
Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of s...
Sequential decision-making in real-world applications often involves uncertainty about the environment's model. Uncer...
The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in co...
Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictiv...
Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. ...
Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustain...
Retrieval-augmented generation (RAG) imposes a prefill cost proportional to retrieved context length, and -- with Tra...
The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: th...
World Action Models (WAMs) augment robot policies with action-conditioned predicted futures, but a plausible future a...
Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work ha...
Artificial intelligence (AI) is increasingly central to power and energy systems, supporting modeling, forecasting, o...
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, ...
Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains dif...
Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work ha...
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and mo...
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them i...
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledg...
Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially...
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient m...
Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still d...
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text...
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool us...
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the tea...
Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into worl...
Video motion transfer aims to animate a target object using dynamics from a reference video. Existing formulations la...
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametri...
Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling...
Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integ...
Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous ...
Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. P...
Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, object...
Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance ofte...
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while...
Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understandi...
The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life...
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isola...
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language mo...
Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying thei...
Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, pro...
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. Howe...
Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in sc...
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites ...
The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-...
Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for infe...
The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable ...
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications...
Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical...
Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and o...
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct ex...
Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency,...
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-def...
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language...
We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-ric...
Latent world models enable efficient planning by predicting future states in a compact representation space, but thei...
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been me...
Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes w...
Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined sc...
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly...
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of r...
Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a bas...
RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surfac...
We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future visi...
As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioriti...
In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning par...
Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limite...
World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physica...
AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often c...
Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NI...
Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This cre...
Autonomous multi-vehicle racing requires real-time planning of diverse competitive behaviors in intense interactions....
Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language mo...
This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and fut...
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account ...
On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains britt...
Existing token compression methods for omnimodal large language models typically rely on one modality to determine wh...
All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonl...
Large language model-based multi-agent systems improve complex problem solving through task decomposition, agent spec...
Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something diffe...
Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While...
Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and p...
We study the computational complexity of winner determination problems in approval-based committee elections under Th...
Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of ...
While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex...
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a c...
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's act...
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applicati...
Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet ex...
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic...
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors g...
Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical ...
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (...
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local vi...
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them tow...
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult t...
Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smalle...
Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placi...
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. User...
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)...
Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose...
The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem i...
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perce...
Forward latent world models predict how actions change a scene, but recover actions for a desired change only through...
Transformers propagate information across depth through a single additive residual stream: every sublayer reads only ...
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models t...
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The o...
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence ...
In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely ...
Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that...
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reason...
Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rath...
On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches...
Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an impl...
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretraine...
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector r...
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based mu...
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including ...
LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a ...
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnera...
As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, age...
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and ...
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptio...
Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The...
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, rev...
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free video...
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable u...
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typical...
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing be...
Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision ...
Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc...
Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its prese...
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models...
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carr...
Most video editing systems still lack explicit layered video representations, limiting their ability to perform reali...
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks again...
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an ac...
Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are ...
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carr...
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing be...
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelli...
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-mak...
Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather tha...
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet sta...
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit cri...
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained kno...
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations r...
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host arti...
Recent game world models can generate visually realistic and interactive environments conditioned on player actions. ...
Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, l...
Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definit...
Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, trac...
Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment me...
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation d...
Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distrib...
RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and ...
A language model p_θ(y mid x) trained on reasoning tasks learns to solve problems via multiple distinct strategies, y...
In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update ru...
Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how...
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in ...
Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no sin...
Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. ...
Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health beha...
RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and ...
Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon...
Kubernetes is central to the cloud-native ecosystem, orchestrating containerised workloads. Recent work suggests that...
Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competi...
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made usi...
Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficien...
Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains a...
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even ...
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks p...
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretraine...
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix...
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often c...
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms mod...
We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an...
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but thei...
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has becom...
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. Th...
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities...
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scal...
Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval...
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning ...
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, ...
We present a reproducible pipeline for mapping Common Vulnerabilities and Exposures (CVEs) to MITRE ATT&CK Enterprise...
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However...
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix...
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in ...
Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling proc...
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than...
Different research lines use the term world model in different ways, yet they share a common aim: to capture how the ...
Compressed short-text generators can fail in two different places: the codec may discard information before generatio...
Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and train...
We introduce a vocabulary for automated research systems built from one or more agents to make their design choices e...
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the act...
Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deploy...
Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time ...
Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fa...
Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks ...
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between S...
AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their p...
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most eva...
Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic...
A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed meth...
Multi-modal classification leverages complementary information across diverse data sources to enhance predictive perf...
Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline...
Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only ...
Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit...
Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many ...
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains...
Computer vision models have become highly effective for medical applications, yet their black-box nature continues to...
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the curren...
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying the...
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of ...
Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning ...
While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretra...
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent ge...
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision ...
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage,...
Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animat...
Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-determinis...
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-pu...
In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models...
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterior...
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the curren...
Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue re...
Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support...
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains...
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, ...
The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in stru...
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a do...
Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may c...
Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp diver...
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, whi...
Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific...
Humans routinely communicate through abstractions of their bodies, including shadows, silhouettes, and reflections. Y...
For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training b...
Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role ...
Do learned audio embeddings encode structure that nobody told them to encode? We probe four large pretrained audio mo...
Generative AI is reshaping programming education, yet educators often infer students' AI-supported learning from clas...
Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deploy...
Quantum state preparation is a key component of many quantum algorithms. Performing this step efficiently is essentia...
Large Language Model (LLM)-based Test-Driven Development (TDD) has advanced automated code generation. However, exist...
Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposin...
Automating theoretical research is constrained not only by the generation of candidate results, but also by their rel...
Commercial large language models are increasingly used as knowledge references, yet their stance on contested scienti...
A central design principle in modern machine learning and artificial intelligence is to align a model's inductive bia...
Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this...
To effectively integrate AI into high-stakes, critical environments such as healthcare, autonomous driving, and aviat...
Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capabil...
LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing se...
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existin...
Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-trai...
Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as...
Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new ...
In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple langu...
Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and ...
Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years...
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot m...
Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-toke...
The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unifie...
Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressi...
Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure...
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions t...
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two ...
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale...
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified a...
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multip...
We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-ste...
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn r...
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents ...
Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challe...
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, ...
Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing...
The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with ...
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggl...
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation...
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn r...
Building socially calibrated large language models, which can learn from others without simply yielding to them, requ...
A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to id...
Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserv...
Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manu...
Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence ...
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactio...
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when h...
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely co...
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states...
The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-sca...
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We ...
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-polic...
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its co...
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactio...
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is ...
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural ...
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot e...
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation...
We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural represen...
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the p...
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks t...
When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differ...
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming exp...
Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answ...
Simultaneous localization and mapping (SLAM) is one of the fundamental problems in robotics, as it enables autonomous...
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and...
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural ...
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independentl...
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- dete...
Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating struc...
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires...
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scal...
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) gener...
We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a tim...
AI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and application...
RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always av...
This study empirically analyzed generative AI as an emerging discovery pathway to academic library resources. Utilizi...
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Visio...
Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-qual...
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed fa...
Clinical biomarker workflows in translational research settings often rely on spreadsheet-driven tracking, manual qua...
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages des...
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensi...
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represe...
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D ...
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requ...
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data f...
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in e...
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges fo...
As large language models and AI agents become the primary consumers of search results, document set quality determine...
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the ...
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained ...
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL al...
Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its ex...
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in hig...
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed fa...
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapsh...
Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypern...
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become cri...
Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an u...
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike st...
We present AutoIndex, a framework for learning representation programs: executable transformations that map raw docum...
The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. W...
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic m...
This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative...
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example ...
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components re...
Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan sa...
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR ...
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the...
Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence o...
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language mod...
Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existin...
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinati...
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rat...
Controllable image generation remains challenging for creative professionals, who often require precise regional cont...
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, ...
Controllable image generation remains challenging for creative professionals, who often require precise regional cont...
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them prom...
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where mode...
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-...
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams mu...
Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but stale...
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically...
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation...
Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the clai...
Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflecte...
Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods sh...
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction,...
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language mod...
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation dur...
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused i...
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesize...
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their p...
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B...
Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chai...
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on ...
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although...
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic...
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. T...
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resource...
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial...
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous ...
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driv...
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand...
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the ma...
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvem...
Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models p...
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an ...
Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libra...
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, whi...
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial...
Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools ...
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural ...
Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, buil...
PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduc...
Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as opera...
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fix...
Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer ...
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syll...
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pi...
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However...
Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping ...
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with ...
Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. Whil...
Pruning long context for coding agents has been a vital technology for efficient context management. While existing c...
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous mod...
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated e...
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, ex...
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing me...
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a t...
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive li...
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the sup...
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physica...
In line with the prevailing direction of vision research, we explore the integration of both generation and editing c...
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromis...
Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong...
Predicting a football match before kickoff requires more than knowing past results: a model must use changing informa...
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies o...
Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffer...
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio L...
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architec...
Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting th...
Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative bi...
Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation towa...
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embod...
Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from inc...
Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch ...
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, ...
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third p...
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next mo...
AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, wh...
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded e...
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantag...
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapsh...
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-traini...
Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, ...
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover...
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components w...
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory...
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily r...
On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constrain...
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse lang...
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, con...
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and hi...
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existin...
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI...
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-traini...
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy c...
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural know...
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grad...
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These a...
Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. Wh...
Training-free in-context segmentation enables new object categories to be introduced at inference time from a single ...
Plasma diagnostic models for tokamak fusion devices are almost universally evaluated on clean, complete sensor data. ...
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a ...
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing v...
🙌 OpenHands: AI-Driven Development(⭐81227)
A striking feature of the human visual system is that it ingests visual information through a series of local foveate...
Recent generalizable 3D Gaussian Splatting models have advanced long-sequence novel view synthesis (NVS), but at the ...
CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a sin...
In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates ...
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We i...
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streami...
Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based test...
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit ...
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axio...
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Mas...
Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems f...
Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and ...
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misalign...
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically beni...
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflect...
Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, ...
Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of t...
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing v...
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understandi...
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior wo...
Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors r...
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-T...
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-T...
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn inte...
We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without cha...
MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them ...
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting action...
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seekin...
Music generation foundation models have recently attracted significant industry attention. However, achieving efficie...
Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, ...
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly ...
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference imag...
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multip...
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field...
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, ...
Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduc...
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphic...
In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a conte...
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter...
Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, ge...
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual informati...
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-to...
Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is control...
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide li...
We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard determin...
Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) m...
Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world...
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, ...
Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives ...
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their stri...
We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in fro...
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $...
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their stri...
Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolati...
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically co...
The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradi...
Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and...
Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, c...
We investigate how each component of the Transformer feedforward block architecture design determines how much rank s...
With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable ener...
Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, ...
Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running sc...
By 2030, 59 of every 100 workers will need reskilling or upskilling, yet the average time to close an enterprise skil...
This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational...
The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LL...
We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document ...
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which construct...
OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platfo...
Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerge...
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recogni...
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an in...
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, ...
Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception o...
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the t...
The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language mo...
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving ...
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D di...
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising...
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes...
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling too...
Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a pr...
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans ...
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winn...
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, ...
Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on compl...
This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natura...
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images with...
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbound...
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as Do...
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (C...
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right ...
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their d...
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, ov...
Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a def...
Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reache...
Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, g...
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, mult...
Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches...
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each d...
Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interact...
Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still meas...
Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLM...
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user...
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged e...
Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games w...
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a ...
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling too...
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scal...
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they ra...
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone...
In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-she...
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a ...
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand syst...
Recent foundation image and video generation models offer strong generalization and controllability, but their direct...
Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks?...
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet ...
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes,...
Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language ...
Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously colle...
Prefabricated prefinished volumetric construction moves most building work into module factories, whose production fl...
Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes h...
Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regula...
Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, ...
Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a ...
This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While gen...
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The fram...
Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive ...
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity r...
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes,...
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, meas...
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kin...
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language ...
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-m...
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad ve...
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting larg...
Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and aud...
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet ...
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents s...
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, bu...
Flow matching over carefully designed latent representations has recently emerged as a powerful paradigm for topology...
This work explores the motion transfer from one video to another, which is crucial in animation for diverse character...
Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in...
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-m...
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet m...
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existin...
Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet th...
Decision-making is posing an increasingly formidable challenge to investors because of the growing number of alternat...
We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight ty...
Pre-demolition assessment, the regulated audit process at the heart of urban mining, is an information process in whi...
The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of gen...
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force...
Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse...
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Questi...
Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout...
Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information process...
Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to ...
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concep...
Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecu...
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluati...
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpor...
Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical mom...
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpor...
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation mod...
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. Wh...
Long-context processing has become increasingly important for large language models (LLMs), but simply extending the ...
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Di...
Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to...
In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotati...
Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining lea...
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separatel...
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benc...
Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without r...
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitizati...
Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks pr...
In a class of quantum circuits known as peaked circuits, the goal is to predict the most probable bit string at the o...
Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approach...
We introduce PAST-TIDE, our stance detection system addressing both subtasks of the StanceNakba Shared Task at NakbaN...
In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action ag...
Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question ...
As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment o...
Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft token...
Repository-level code generation requires implementing target functions while accounting for complex cross-file depen...
In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action ag...
Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and...
A national language model offers a linguistic community its own instrument for measuring what its citizens say and va...
Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its e...
Large language model (LLM) applications increasingly use explicit workflows for tool use, retrieval, branching, check...
Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved ...
While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embed...
In this study, we present a large-scale descriptive analysis of the use of an AI-based learning assistant (Syntea) in...
Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable...
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pie...
Reasoning has become a core capability for large models, especially when reliable decisions require understanding log...
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffu...
Reasoning has become a core capability for large models, especially when reliable decisions require understanding log...
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pie...
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository...
Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, ...
Inference-time scaling for text-to-image generation has progressed from simple Best-of-N (BoN) sampling to guided sea...
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. ...
Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of larg...
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with ad...
Magnetic resonance imaging (MRI) super-resolution is vital for improving diagnostic accessibility, yet most methods t...
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects...
A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLM...
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performanc...
The rapid development of large language models and multimodal large language models has accelerated the emergence of ...
Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of p...
Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather th...
In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-...
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures...
Current computational approaches for drug design typically focus on generating molecules conditioned on specific targ...
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length ...
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules...
Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does n...
Linear attention models allow a fixed state size and a fixed amount of compute per token. However, due to their limit...
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previo...
Touch supplies the physical grounding needed to perceive intrinsic material properties, such as friction and complian...
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce ...
Accurate breast cancer classification from mammography requires effective integration of complementary information fr...
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but ...
Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spat...
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often d...
Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators th...
Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning an...
Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, w...
Time series analysis plays a vital role across a wide range of scientific and engineering domains but poses substanti...
Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on lo...
Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive ski...
AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses du...
Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-le...
Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operationa...
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grad...
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models w...
We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hol...
Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access...
Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rat...
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reacti...
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their prima...
We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring...
Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilom...
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovi...
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematicall...
Large Language Models (LLMs) unlocked new possibilities in automated code writing, becoming the backbone of most code...
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reacti...
Complex image creation and editing often require more than a single generation or editing model. A user request may i...
Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding ...
We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies documen...
Recently, Joint Embedding Predictive Architectures (JEPAs) have attracted significant attention in the computer visio...
We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without ...
Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation re...
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic infor...
We present RuleChef, a framework that uses large language models (LLMs) to generate executable rules for NLP tasks su...
Geometry-conditioned 3D scene generation enables the creation of 3D environments from user-provided geometry, offerin...
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases ...
Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little ...
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-sp...
Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, th...
JD.com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchan...
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical ...
The dairy industry in Ireland has a large potential for the integration of renewable energy and the reduction of carb...
Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the groundin...
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail,...
Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which...
GitHub hosts hundreds of millions of public repositories, but the platform exposes no native mapping from repositorie...
We present FootsiesGym, an open-source environment for learning in a non-trivial two-player, zero-sum, imperfect-info...
Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive comp...
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tas...
Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store ...
Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelin...
- Objective: Multimodal deep learning models in oncology are currently limited by monolithic designs that rigidly cou...
As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant i...
Denoising graphs is a fundamental problem in graph learning and the core operation of graph diffusion models. Attenti...
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone,...
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning traje...
Late-interaction retrieval models that use the MaxSim similarity function have shown strong empirical performance, of...
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level...
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods re...
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game d...
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target veri...
Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor l...
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic po...
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of indivi...
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the ...
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable...
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve gener...
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories...
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions...
LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management polici...
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family....
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), ye...
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destin...
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applicat...
Academic output is produced across a fragmented toolchain: literature discovery in one application, reference managem...
We present the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes. The task is compose...
Concept erasure aims to remove a target concept from a representation while preserving the other information encoded ...
In longitudinal clinical practice, every chest X-ray is read in the context of the patients prior exam, and much of w...
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely a...
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we in...
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained ...
Speech-based depression detection compresses features from short audio segments into one speaker-level decision, a st...
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively...
Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their ge...
High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (triggeri...
Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottlenec...
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. Howev...
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems ...
We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research pr...
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interacti...
Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). ...
Planning under uncertainty in continuous domains is essential for autonomous systems, yet computationally demanding. ...
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediate...
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without f...
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, stan...
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems ...
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle ...
What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of...
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbound...
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabi...
Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated...
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, bu...
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositi...
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-rela...
Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structu...
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task ex...
3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstructio...
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling d...
Semi-supervised semantic segmentation (SSSS) has long turned on one question, which pseudo-labels to trust, and answe...
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interacti...
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabi...
We present EVA-Client, an open-source framework for deployment, data collection, and evaluation of trained manipulati...
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from r...
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and th...
Increasingly, LLM inference services proxy client requests to engine replicas distributed globally. Load-balancing po...
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel...
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by c...
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interac...
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching ...
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but th...
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregres...
Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly ...
Scientific literature search often requires more than retrieving papers from a single query: users' intents are under...
Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remai...
We study Generated Contents Enrichment (GCE), a conditional image-generation task in which a sparse scene description...
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by no...
Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractiv...
Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deploym...
Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training...
GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referr...
The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context...
Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To re...
As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to eviden...
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and ...
Cloud removal (CR) is essential for optical remote sensing, serving as a prerequisite for reliable downstream interpr...
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the ...
Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations....
In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span ...
Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over mult...
Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capac...
Foundation models are routinely released to the public, yet the data recipes used to train them -- such as domain mix...
Traffic matrices (TMs) capture network-wide origin-destination demand and are central to traffic engineering, yet acc...
Grid-based approaches to approximate nearest neighbor (ANN) search have been absent from modern scaling analyses. We ...
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and ...
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triple...
Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money pred...
Software tests and code evolve together: a code change should be followed by new or updated tests that record the new...
Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing...
In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to...
Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts...
On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to rea...
Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex s...
LLM agents will increasingly act in socially structured settings where role, audience, and relational context can sha...
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs...
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs onl...
Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repa...
LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need ...
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting acr...
Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amoun...
Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, y...
Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repa...
Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, c...
Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Re...
We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object me...
Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature cach...
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triple...
Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and repla...
We elucidate the design space of Representation Distribution Matching (RDM), our name for the paradigm that trains a ...
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex in...
Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts...
Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking task...
Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens...
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generat...
Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest con...
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations...
Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-train...
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visua...
Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain r...
This paper explores multi-turn visual reasoning and observes that MLLMs repeatedly fail to localize the target, leadi...
Blind image deblurring demands the recovery of high-fidelity details and coherent structures from complex, unknown de...
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they stil...
Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of syntheti...
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents b...
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is esse...
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumul...
People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning syst...
Three of the most popular methods for training language models to reason look like three different tricks. They are n...
In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be e...
While generative models have enabled training-free reward alignment, current methods typically excel in local explora...
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: w...
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-order...
RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined succ...
In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing th...
We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from mono...
This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) ...
Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering...
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents b...
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduc...
Transformers use the same forward computation stream to both predict the next token and store useful state for future...
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the pr...
Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capac...
Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals su...
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by no...
World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enoug...
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied lear...
Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as ...
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent met...
Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents becom...
Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions tha...
Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visu...
Transformers use the same forward computation stream to both predict the next token and store useful state for future...
While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e.,...
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistant...
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical cont...
We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approac...
In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing...
Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual...
Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environ...
We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmar...
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle w...
Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Thing...
AI translation of literary works is increasingly common. While the content may be rendered adequately, we do not know...
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-...
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce...
Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame r...
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of action...
The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains...
Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet t...
Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities...
Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such ev...
Embodied Vision-Language-Action (VLA) models are typically obtained by fine-tuning powerful pretrained VLMs on roboti...
Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surf...
Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-makin...
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assemb...
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven softwar...
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edit...
We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos ...
Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skin...
Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage rank...
The tendency of large generative models to memorize training data makes sample verification critical for privacy audi...
Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case st...
Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions...
Policy-grounded document review requires determining whether a target document complies with organization-specific po...
We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts f...
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edit...
Forest attributes are essential for national-scale resource monitoring. Airborne LiDAR metrics are among the auxiliar...
Latent world models enable planning from high-dimensional observations by predicting future states in a compact laten...
Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipu...
While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e.,...
Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own co...
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of action...
When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, ...
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Lar...
Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality soundi...
We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent s...
Video World Models are interactive video generation models that predict future world states based on user actions and...
Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their in...
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real appli...
Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in paralle...
Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged a...
Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms...
Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unli...
Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own co...
Block Diffusion Language Models (BD-LMs) improve diffusion-based text generation with KV caching and flexible-length ...
While large language models have been dominating the research landscape recently, small language models remain highly...
Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and the...
Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fai...
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction m...
Foundation models have transformed vision and language processing by providing rich, reusable representations that tr...
Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneou...
Agent skills extend language-model agents with task-specific procedures, scripts, and references, but the tasks and e...
Foundation models for predictive machine learning on tabular data have recently gained significant traction in academ...
The advancement of generative AI models capable of producing text and image marks a critical step forward in the real...
Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful seman...
A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multipl...
Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations...
Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-...
Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhanceme...
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign...
RocketSmith is an agentic system which intelligently automates the DFAM process for the development of high powered r...
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by...
Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by ...
Recent work has demonstrated the potential of large language models (LLMs) for program optimization, a key challenge ...
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized langua...
Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, b...
While text-guided image editing has made remarkable progress, it remains limited in structural portrait retouching. T...
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where sp...
Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challeng...
Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, profession...
Malware classification remains a challenging problem due to its inherent heterogeneity, the presence of packed binari...
Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) w...
Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection. ...
Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agen...
Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse,...
Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosi...
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense...
Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy ...
Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to...
World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequenc...
Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics...
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whol...
Self-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extrem...
MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enab...
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to app...
Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent vid...
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world comp...
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly c...
Current models of representational reliability in neural populations focus on temporal stability: whether population ...
We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulatio...
Streaming video editing has made rapid progress, yet practical deployment is still limited by two core issues: mainta...
Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human exp...
Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic v...
Speech language models (SLMs) have been extensively studied, with the common paradigm incorporating text data and pre...
4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-b...
We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning...
Mathematical knowledge is organized around statements and their dependencies, but this structure is exposed unevenly:...
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on la...
Hydropower tunnel inspection is critical for infrastructure integrity yet remains inefficient and hazardous using man...
Adapting a foundation vision-language encoder to a specialized retrieval task creates a fundamental tradeoff: gains o...
LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete...
Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers ...
Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring...
Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To cons...
Efficient deployment of large language models (LLMs) in production forces a trade-off between accuracy and cost. Oper...
Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches,...
Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consisten...
LLM-based code agents navigate repositories through keyword search but miss the structural relationships, such as cal...
LLM-based agents for program repair are increasingly built on a "generate-run-revise" paradigm, iteratively executing...
Tokenization is central to adapting scientific data for transformer-based foundation models, yet its impact on learne...
For agents to learn continuously from interaction with the world at test time, they must be able to explore effective...
Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured ...
Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a ...
Sparse attention can reduce the cost of long-context inference, but most variants introduce new architectural compone...
Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee tha...
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, an...
Embodied agents operating in decentralized and partially observable environments have attracted growing attention in ...
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and ...
This study analyzes Sri Lankan migration and remittances over 32 years (1994-2025). Using a 384-month harmonized data...
Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data c...
We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agent...
Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepan...
Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates th...
Understanding how performance scales jointly with model size and data is a central problem in modern machine learning...
Test-time adaptation (TTA) has emerged as a promising paradigm for mitigating distribution shifts in deep models. How...
The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and m...
Accurate network traffic prediction is a critical element for efficient resource allocation in dynamic urban cellular...
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis gene...
We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A...
Many two-player zero-sum games admit not a unique Nash equilibrium but a convex set of them: a polytope of profiles t...
Dexterous manipulation policies can solve individual skills, but composing them to perform multiple tasks with a sing...
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-dom...
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a...
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis gene...
We introduce an axiomatic evaluation framework for latent thought representations in LLMs, comprising metrics that ar...
We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF)...
Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications....
Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades ...
Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- whil...
Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowl...
We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel gripper...
Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatica...
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system...
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-ba...
Multi-agent systems (MAS) built on large language models (LLMs) provide a promising framework for solving complex tas...
Voice agents face a fundamental tension: the reasoning, retrieval, and tool use that make foundation models capable a...
Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-mode...
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to f...
The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermed...
Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under ch...
Reasoning capability has advanced rapidly in large language models (LLMs), leading to an increasing size of key-value...
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings rema...
ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting...
We present a conceptual framework for analyzing dialogue in collaborative problem-solving contexts, with an emphasis ...
AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real ...
Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools...
Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major chal...
Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under ch...
Mechanistic epidemiological models are widely used to support infectious disease forecasting and public-health decisi...
Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candida...
Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-mode...
AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their per...
Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation model...
Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential...
Digital twins have emerged as a promising paradigm for personalized healthcare, enabling modeling of individual behav...
Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are c...
Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat so...
Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. ...
With the increasing development of Vision-Language Models, it becomes imperative that their predictions are readily e...
We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models ...
Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), loca...
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are o...
While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physica...
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising found...
Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse traj...
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more...
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we t...
Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoint...
A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the c...
Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL...
Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces...
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this i...
Modern generative world models render increasingly realistic action-controllable futures, yet they frequently halluci...
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biologic...
Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and ver...
Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Rewar...
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is ...
On-policy distillation (OPD) improves LLM reasoning by training a student model on its own generated outputs, but sta...
Continual Test-Time Adaptation (CTTA) aims to maintain model performance under evolving target domains by adapting on...
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit...
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decis...
Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computati...
As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (...
Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can cont...
Trust in an AI system is often anchored by explanations of how it works, which one then uses to forecast its behavior...
Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tune...
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate v...
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality train...
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player ze...
We present HiReLC, a hierarchical ensemble-reinforcement learning framework for automated joint quantization and stru...
Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While...
Tabular foundation models are commonly assumed to present limited privacy concerns as they are often pre-trained on l...
As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges...
Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchma...
Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's...
AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems...
The laser welding full-penetration is of critical importance, as it constitutes one of the fundamental factors in ach...
A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on det...
Supervised deep learning has been widely used for weld penetration state classification; however, its performance oft...
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings rema...
On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with...
Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action modul...
Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still st...
While Large Language Models (LLMs) have substantially advanced text-to-code synthesis, many real programming tasks sp...
Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time strea...
Modern large language models are predominantly trained with autoregressive factorization and causal attention. We pre...
Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contr...
The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. ...
Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Op...
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality train...
Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LL...
Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a da...
While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynam...
We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single succ...
Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evid...
Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and ...
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safe...
"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whe...
Retrieving external knowledge is essential for solving real-world tasks, yet it remains challenging when the relation...
Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aestheti...
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geome...
Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joi...
Cross-Chart Retrieval-Augmented Generation (RAG) is critical for complex multi-modal analytical tasks in scientific, ...
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametr...
Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at th...
Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation...
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bou...
What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding a...
Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as d...
Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving unders...
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to...
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answ...
Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generat...
Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quan...
Modeling chaotic systems is crucial yet challenging. Inverse problems in chaotic dynamics, namely inferring initial c...
Over a series of seven papers, Andreas & Günther have introduced seven definitions of actual causation and have class...
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks...
Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This m...
Prompt-based learning has emerged as a dominant paradigm in natural language processing. This study explores the impa...
In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a wo...
Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still st...
Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are abl...
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate t...
Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation...
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bou...
MLLM-based mobile GUI agents have made substantial progress in UI understanding and action execution, but adapting th...
Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation...
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate t...
MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet remain unreliable on long-hor...
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive me...
Generalist value models play a pivotal role in scaling robotic policy learning from large-scale, mixed-quality data. ...
Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from n...
The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of...
Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narra...
Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ...
Dense retrieval embedding models are a fundamental component of modern retrieval-based AI systems. Most dense retriev...
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publica...
The problem of optimal sizing and power scheduling in microgrids subject to uncertainties is well known to the contro...
Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from d...
Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world inte...
AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, mana...
Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric represen...
Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of i...
Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) ...
Attention-based Multiple Instance Learning aggregators in medical imaging are prone to attention concentration, produ...
Self-attention is central to Transformer performance and is often the most expensive part of the Transformer at long ...
Text and image conditioned 3D models now generate convincing assets, but they still offer little direct control over ...
Open-weight Large Language Models (LLMs) enable scientific progress and broad deployment. However, they make it diffi...
As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, lang...
Computer-use agents (CUAs) now act on a user's behalf across personal applications such as email, calendars, and to-d...
As urban areas expand, automatic monitoring of parking lots becomes essential for efficient and sustainable cities. T...
Optimizing pretraining data composition is pivotal for LLM generalization. While dynamic mixing outperforms static st...
Large language models (LLMs) are increasingly used to support software development, but their practical usefulness in...
As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backb...
Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservatio...
It is tempting to assume any task solvable by a short program can be taught to a model as its chain-of-thought: write...
Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspe...
Filmmaking demands precise motion control and reference image compositing -- capabilities that existing methods treat...
Long-horizon LLM agents can fail quietly: they settle on one reading of the evidence early, then spend the rest of th...
We introduce ShotcreteDepth, a bi-modal dataset from the construction domain that captures both an active shotcreting...
Linear probes are widely used in interpretability research and often compared by cosine similarity. The Mahalanobis c...
Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observatio...
Discrete text-trigger optimization -- searching for text sequences that, when ingested by a model, steer it toward a ...
Long-context reasoning is an essential capability for large language models, particularly when they are deployed as a...
Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately...
The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, m...
Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demo...
Can representations learned for image generation also support the evaluation of generated images? We study text-to-im...
Remote patient monitoring depends on patient-reported data to capture the subjective dimension of recovery that devic...
A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produc...
In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently s...
Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly...
Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack ...
This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Ma...
Mental health assessment commonly relies on isolated screening instruments or data-driven models that often lack inte...
AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostl...
Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large lang...
Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at th...
Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipula...
Multi-view 3D Visual Question Answering (MV3D-VQA) requires integrating partial observations into a coherent 3D scene...
Long agent traces composed of chains of thought and tool calls accumulate stale content that anchor subsequent genera...
Latent action pretraining learns representations of visual change from pairs of observations, but existing methods ty...
Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable ...
While recent LLM-based terminal agents have demonstrated promising capabilities, the scarcity of high-quality, execut...
Recently, end-to-end OCR models, exemplified by DeepSeek OCR, have once again thrust OCR into the spotlight. A widely...
While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture fo...
Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack ...
Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfe...
Terminal-using agents have quickly become the most popular downstream application of language models (LMs). Despite t...
Massive unstructured multimodal streams suffer from high "data entropy," impeding both efficient human knowledge acqu...
With the rapid spread of retrieval-augmented generation and semantic search, choosing the right embedding and retriev...
Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the ...
Long-horizon tasks are common in real-world robotic deployments, yet failure detection for such tasks remains underex...
Autoregressive generation in large language models (LLMs) conventionally decodes from the final layer, assuming that ...
We describe our entry to the efficiency track of the Academic Text-to-Music (ATTM) Grand Challenge at ICME 2026. Beyo...
Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for...
We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natur...
Scientific discovery workflows usually contain and rely heavily on lab notes, where researchers record observations, ...
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottl...
Multimodal large language models (MLLMs) are increasingly deployed in personally and societally consequential setting...
Retrieval-augmented generation (RAG) systems depend critically on how documents are chunked and searched. Fine-graine...
Retrieval-augmented generation (RAG) systems must balance retrieval granularity with contextual coherence, a challeng...
The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a ...
Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because ...
Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, mo...
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and ...
Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could a...
Memory benchmarks for LLM agents largely assume single-user settings, leaving shared assistants for hospitals, workpl...
In-context learning (ICL) is the standard method for low-resource classification, yet its efficacy in specialized dom...
High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human...
Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to pla...
While reasoning on autoregressive (AR) models is often performed by chain-of-thought reasoning and reflection, their ...
Personalized presentation generation requires more than conditioning on a current prompt or template: agents must pre...
🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and report...
An agentic skills framework & software development methodology that works.(⭐234547)
When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate t...
Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on cont...
Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a c...
Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models ...
Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring ap...
LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on cod...
Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunc...
Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutatio...
Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for s...
Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference ...
Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling ...
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual word...
Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important f...
Generative recommendation is an emerging paradigm that has shown promise in industrial recommendation systems, aiming...
LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalign...
Creating 3D visual illusions, a single 3D mesh that reveals entirely different semantics from various viewing angles,...
Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and too...
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior acr...
Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-wo...
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tigh...
Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, wher...
Conditional diffusion and flow models routinely fail to satisfy the very constraints that define their task. For inst...
Hybrid linear attention models offer an appealing path to faster long-context inference: they reduce the quadratic co...
Large Language Models (LLMs) have significantly advanced the automation of software engineering tasks. One prominent ...
LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on cod...
FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hard...
Scheduling policies in large-scale Automatic Speech Recognition (ASR) serving pipelines play a key role in determinin...
The Frechet Inception Distance (FID) is the de facto arbiter of image generation, yet most papers report just a singl...
A significant gap exists between theory and practice in deep learning. Generalization and approximation error bounds ...
Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-l...
AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average c...
Existing Programming-By-Example (PBE) systems often rely on simplified benchmarks that fail to capture the high struc...
Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequ...
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence wi...
Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling ...
Tensors and Dynamic neural networks in Python with strong GPU acceleration(⭐100863)
Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning an...
Reinforcement learning pipelines for Large Language Model (LLM) training often rely on manually redesigned environmen...
Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language mode...
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...
The Network Data Analytics Function (NWDAF) is central to enabling zero-touch network management in fifth-generation ...
Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more ...
Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe t...
Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between eval...
A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and prefer...
We show the standard basis of transformer hidden states already provides a training-free, architecture-general featur...
Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptio...
Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan a...
On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-l...
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, vi...
Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with...
Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer...
Offline reinforcement learning is typically analyzed under process-level reward supervision, yet many sequential deci...
Robotic systems perceive the world through multiple input modalities -- including visual camera streams and natural l...
Despite growing interest, most evaluations of large language models' (LLMs') personalization abilities have relied on...
Test-time scaling via sequential revision has emerged as a powerful paradigm for enhancing Large Language Model (LLM)...
We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reason...
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradi...
Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-m...
Family members caring for individuals with Alzheimer's disease and related dementias (AD/ADRD) provide the foundation...
Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific...
Automatically generating slide decks from source documents is an important application of large language models (LLMs...
Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimod...
The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized doma...
Neurosymbolic semantics is fragmented: classical, fuzzy, probabilistic and neural systems each define truth by their ...
When social chatbots make mistakes, and they do, how they recover determines whether users trust them again. Social c...
A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-me...
Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts ...
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, mu...
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning w...
Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing ...
Multi-agent LLM systems share state through memory stores, vector indices, and tool registries. We model such sharing...
Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse...
Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a sca...
In this report, we present LOGOS (Language Of Generative Objects in Science), a scientific generative language model ...
Deploying multimodal foundation models as closed-loop policies increasingly requires conditioning actions on observat...
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. H...
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ide...
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering...
Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a m...
World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and l...
Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation o...
World models are transitioning from passive visual generators to foundational, operational infrastructure for Physica...
Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern pro...
Diffusion models have become a promising alternative to autoregressive models. Among these, uniform diffusion languag...
Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in ...
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space def...
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-st...
Passive models for long video understanding typically rely on a "watch-it-all" paradigm, processing frames uniformly ...
Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest co...
Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are ...
We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity...
Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate wheth...
Software practitioners increasingly use AI coding agents that generate test code alongside production code in open so...
Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable...
Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long...
Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly foc...
As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce a...
We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 a...
The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to allevi...
Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require composi...
Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deepe...
With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cybe...
Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any p...
Reproducing research results from papers and released code is central to scientific progress. Existing works have int...
Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism...
Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, ex...
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educ...
Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irr...
Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deepe...
Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action voc...
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifica...
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization ...
Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved source...
Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning...
Knowledge distillation transfers a teacher's competence to a small student but is brittle in the small-student regime...
Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases l...
Training computer-use agents (CUAs) -- models that interact with graphical desktops through screenshots and keyboard/...
Scaling model size, specifically depth and width, has driven significant progress in transformer-based language model...
Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory ...
Pixel-space diffusion models are trained on full-bandwidth noisy images, yet the useful signal available to the denoi...
Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-f...
Large language models perform increasingly well on standardized logical reasoning benchmarks, but whether this abilit...
Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention m...
On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its appli...
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based on...
Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and in...
The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a priv...
Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost ...
Incorporating textual reviews into a Recommender System has become a prominent strategy for enriching collaborative s...
The reproducibility crisis has directed the AI research community toward improving documentation practices. Several s...
Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, tr...
Using an open problem from the EC 2025 paper "Stable Menus of Public Goods" as a testbed, we conduct experiments to u...
Reinforcement Learning (RL) policies often degrade in unfamiliar environments because they lack explicit deliberation...
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it...
Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series...
We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preferen...
As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approac...
Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains c...
Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, an...
Oppenheim and Lim (1981) showed that natural images stay recognizable when reconstructed from their Fourier phase alo...
Standard accuracy benchmarks are designed to test how closely large language models (LLMs) approach correct answers, ...
A content-moderation system can score well on every standard accuracy metric and still cause real harm, if its mistak...
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon gener...
Diffusion transformers have demonstrated remarkable generative capabilities, yet the rich perceptual representations ...
As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultiv...
In this paper, we introduce SP^3, a novel Plug-and-Play algorithm that accelerates maximum a posteriori image restora...
Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions...
When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary out...
Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency an...
Polymarket has emerged as a prominent prediction market platform and one of the fastest-growing applications in DeFi....
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong re...
We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classifi...
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but...
Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on ...
Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformabl...
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue ...
Sparse reward reinforcement learning (RL) has become a standard tool for improving LLM reasoning, but its success dep...
Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera traj...
Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with wea...
Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be m...
TradingAgents: Multi-Agents LLM Financial Trading Framework(⭐86448)
Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such vid...
Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the ...
Vision-Language-Action (VLA) models that couple pretrained Vision-Language Models (VLMs) with continuous action exper...
As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisi...
Unstructured pruning produces sparse weight tensors, but the standard implementation keeps tensor shapes unchanged so...
Online group chats are social spaces with local conversational norms that are rarely stated explicitly. The ability a...
Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environme...
Token-level hallucination detectors are evaluated as classifiers, by AUC over all tokens, yet a streaming monitor is ...
Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but ...
We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first...
Studies of human reasoning have shown that people are typically stronger at evaluating reasoning than producing it fr...
Modern Lean theorem provers achieve strong performance only with substantial training and inference compute, driven i...
With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metr...
As AI-generated reviews move from experimental tools into peer-review infrastructure, most robustness concerns have f...
Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and...
We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly ...
Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene ...
Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet t...
AI-assisted software development has moved from line-level autocomplete to agents that can plan changes, edit files, ...
Wearable devices and smartphones generate rich behavioural time series that can support proactive health intervention...
Survival prediction plays a central role for healthcare providers and clinical researchers. Accurate risk stratificat...
We show that the three movements of Beethoven's "Moonlight Sonata" (Op. 27 No. 2) instantiate three distinct machine ...
Verifier-driven self-DPO is a common recipe for self-improving production visual-language models. In this setup, a fr...
Recent advances in speech generation have significantly improved the naturalness of synthetic speech, making spoofing...
Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their prediction...
Sequential or time-stamped interaction logs provide objective records of digital application usage, yet their granula...
Artificial Intelligence (AI) is increasingly used to automate a variety of real-world computer vision (CV) applicatio...
Large language models increasingly serve as execution engines for agentic systems, yet they still consume context thr...
Globally, cotton is a highly economically beneficial crop, as the textile industry heavily depends on it. So, the pre...
AI systems coupled to proof assistants now generate formal mathematics at scale, and the gap between what a checker c...
Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, p...
Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision supp...
World models that capture how actions induce physical change enable scalable robot learning without reliance on embod...
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control fl...
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full ro...
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, f...
The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-ag...
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Ye...
Large language models (LLMs) are widely used in text-to-image (T2I) systems, but they are typically limited to text e...
Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and...
Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelit...
Users rely on execution traces to observe agent behavior, diagnose failures, and ensure accountability. These traces ...
Video generation models based on Diffusion Transformers (DiTs) have achieved remarkable performance in video synthesi...
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabiliti...
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in re...
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. Wh...
Interactive driving exposes a failure mode that is easy to miss in rule-aware autonomous-driving stacks: a hard-rule ...
Retrieval-augmented generation is moving beyond text into long, egocentric video, where systems must select query-rel...
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption tha...
Large reasoning models typically follow a read-then-think paradigm: they observe the complete input, reason over a st...
Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling...
Large and demographically balanced datasets are essential for reliable neuroimaging biomarkers. Full-resolution 3D br...
The API to search, scrape, and interact with the web at scale. 🔥(⭐132792)
The agent engineering platform.(⭐139289)
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, No...
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)(⭐141524)
Production-ready platform for agentic workflow development.(⭐145209)
Java 面试 & 后端通用面试指南,覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发(⭐156367)
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, a...
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-...
Get up and running with Kimi-K2.6, GLM-5.1, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.(⭐174170)
AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so ...
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-ho...
The agent that grows with you(⭐193566)
An Open Source Machine Learning Framework for Everyone(⭐195658)
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first developmen...
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞(⭐378719)
Chain-of-thought (CoT) reasoning is the dominant paradigm for inference-time scaling in language models, yet the caus...
Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decis...
When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evid...
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on ...
Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This cre...
Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic spe...
There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social sc...
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slav...
This paper examines three recent frameworks for understanding the cognitive and epistemic consequences of artificial ...
LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and ...
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledg...
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze...
Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fu...
Articulated tool manipulation remains a major challenge in dexterous robotics due to the need to coordinate internal ...
Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowle...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: Hoi3DGen:生成高质量的三维人-物交互场景 英文名: Hoi3DGen: Generating High-Quality Human-Object-Interacti...
来源: arXiv 爆炸度: 7/10 💥💥💥💥💥 中文名: 跨上下文审查:通过分离生成与审查会话提升大语言模型输出质量 英文名: Cross-Context Review: Improving LLM Output Quality ...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 大语言模型应用精选集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个收录了基于各类大语言模型(包括OpenAI、Anthropi...
来源: arXiv 爆炸度: 7/10 💥💥💥💥💥 中文名: 基于机器学习与组合融合分析的NCAA锦标赛对阵预测 英文名: NCAA Bracket Prediction Using Machine Learning and Comb...
来源: arXiv 爆炸度: 7/10 💥💥💥💥💥 中文名: LLM2Vec-Gen:从大语言模型生成式嵌入 英文名: LLM2Vec-Gen: Generative Embeddings from Large Language Mo...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超棒的大型语言模型应用集合 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个收录了基于AI智能体和检索增强生成技术的各类LLM应用...
来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 全迷你语言模型L6-v2 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 轻量高效的通用语义嵌入模型,在速...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 大型语言模型应用精选集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个系统化整理当前最热门大型语言模型应用实例的开源项目库,涵盖...
来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: 全句嵌入MiniLM-L6-v2模型 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 一款高效轻量的通用句...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: AI工具的系统提示词与模型配置库 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 逆向解析主流AI编程...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超赞的大语言模型应用合集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个系统化收集和分类各类大语言模型实际应用案例的开源知识库,...
来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: Artificial Intelligence for Detecting Fetal Orofacial Clefts and Advancing Medical Edu...
来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: SG-DOR: Learning Scene Graphs with Direction-Conditioned Occlusion Reasoning for Peppe...
来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: google-gemini/gemini-cli 英文名: google-gemini/gemini-cli AI 总结: ⭐ 96812 | 🍴 12008 | 👁️...
来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups ...
来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning 英文名: FaceCam: Port...
来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: SimpliHuMoN: Simplifying Human Motion Prediction 英文名: SimpliHuMoN: Simplifying Human M...
来源: arXiv 爆炸度: 5/10 💥💥💥💥💥 中文名: Accurate and Efficient Hybrid-Ensemble Atmospheric Data Assimilation in Latent Space w...
来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: openclaw/openclaw 英文名: openclaw/openclaw AI 总结: ⭐ 260218 | 🍴 49883 | 👁️ 260218 Your ...
来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: x1xhlol/system-prompts-and-models-of-ai-tools 英文名: x1xhlol/system-prompts-and-models...
来源: GitHub 爆炸度: 10/10 💥💥💥💥💥 中文名: Shubhamsaboo/awesome-llm-apps 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: ⭐ 99599 | 🍴 ...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: HiFi-Inpaint:面向高保真参考修复的细节保持型人-物图像生成 英文名: HiFi-Inpaint: Towards High-Fidelity Reference...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: 推理核心:一个用于符号预训练与后训练的可扩展程序化数据生成套件 英文名: Reasoning Core: A Scalable Procedural Data Genera...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: AI工具的系统提示词与模型库 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 开源逆向工程项目,收集并...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 惊艳的LLM应用合集 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个系统整理大型语言模型(LLM)在AI智能体与RAG等领域实际...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: UFO-4D:从两张无位姿图像进行前馈式4D重建 英文名: UFO-4D: Unposed Feedforward 4D Reconstruction from Two I...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: 模式寻求与均值寻求相遇:快速生成长视频的新方法 英文名: Mode Seeking meets Mean Seeking for Fast Long Video Gener...
来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 通用句子嵌入模型 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 一个高效轻量的句子嵌入模型,在语义理解任...
来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: BERT基础未标注版 英文名: google-bert/bert-base-uncased AI 总结: 谷歌发布的经典双向Transformer预训练模型,奠...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: OpenClaw(开源AI助手) 英文名: openclaw/openclaw AI 总结: 一个开源、跨平台、可高度定制的个人AI助手,支持多种大模型和功能,旨在成为用...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超赞大语言模型应用集锦 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 一个精心整理的、涵盖AI智能体与RAG技术的大语言模型应用实战...
来源: GitHub 爆炸度: 8/10 💥💥💥💥💥 中文名: OpenClaw - 个人AI助手 英文名: openclaw/openclaw AI 总结: 一个跨操作系统和平台的个人AI助手项目,采用龙虾(Claw)作为品牌标识,...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: AI工具系统提示词与模型集合 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 开源逆向工程项目,系统整...
来源: GitHub 爆炸度: 8/10 💥💥💥💥💥 中文名: Gemini 命令行工具 英文名: google-gemini/gemini-cli AI 总结: Google官方开源的终端AI助手,让你在命令行中直接使用Gemini...
来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 全MiniLM-L6-v2句子嵌入模型 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 轻量高效的通用句子...
来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: BERT基础未分词版 英文名: google-bert/bert-base-uncased AI 总结: Google官方发布的英语基础BERT模型,是NLP领...
来源: arXiv 爆炸度: 3/10 💥💥💥 中文名: MediX-R1:开放式医学强化学习框架 英文名: MediX-R1: Open Ended Medical Reinforcement Learning AI 总结: 该研究...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: VGG-T³:大规模离线前馈式三维重建 英文名: VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scale AI...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: OpenClaw(开源爪) 英文名: openclaw/openclaw AI 总结: OpenClaw是一个开源、跨平台、可高度定制的个人AI助手项目,旨在让用户在任何...
来源: GitHub 爆炸度: 8/10 💥💥💥💥💥 中文名: AI工具系统提示词与模型库 英文名: x1xhlol/system-prompts-and-models-of-ai-tools AI 总结: 开源仓库汇总了主流AI编程...
来源: GitHub 爆炸度: 9/10 💥💥💥💥💥 中文名: 超赞的LLM应用集合(含AI智能体与RAG) 英文名: Shubhamsaboo/awesome-llm-apps AI 总结: 这是一个系统整理基于大语言模型的AI智能...
来源: HuggingFace 爆炸度: 8/10 💥💥💥💥💥 中文名: 通用小型句子嵌入模型(MiniLM-L6-v2) 英文名: sentence-transformers/all-MiniLM-L6-v2 AI 总结: 这是一个...
来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: 谷歌BERT基础模型(不区分大小写) 英文名: google-bert/bert-base-uncased AI 总结: 这是一个由谷歌在2018年发布的基于T...
来源: HuggingFace 爆炸度: 9/10 💥💥💥💥💥 中文名: ELECTRA基础判别器模型 - 基于判别式预训练的文本编码器 英文名: google/electra-base-discriminator AI 总结: EL...
来源: arXiv 爆炸度: 8/10 💥💥💥💥💥 中文名: 基于锚定的模型一致性研究 英文名: Model Agreement via Anchoring AI 总结: 该论文研究如何通过锚定机制控制两个独立训练的机器学习模型在实值...
来源: HuggingFace 爆炸度: 6/10 🔴🔴🔴🔴🔴 标题: sentence-transformers/all-MiniLM-L6-v2(HuggingFace) 英文: sentence-transformers/all...
来源: HuggingFace 爆炸度: 7/10 🔴🔴🔴🔴🔴 标题: google-bert/bert-base-uncased(HuggingFace) 英文: google-bert/bert-base-uncased 摘要: ...
来源: HuggingFace 爆炸度: 6/10 🔴🔴🔴🔴🔴 标题: google/electra-base-discriminator(HuggingFace) 英文: google/electra-base-discrimina...
来源: HuggingFace 爆炸度: 5/10 🔴🔴🔴🔴🔴 标题: Falconsai/nsfw_image_detection(HuggingFace) 英文: Falconsai/nsfw_image_detection 摘要...
来源: HuggingFace 爆炸度: 7/10 🔴🔴🔴🔴🔴 标题: sentence-transformers/all-mpnet-base-v2(HuggingFace) 英文: sentence-transformers/al...
来源: arXiv 爆炸度: 8/10 🔴🔴🔴🔴🔴 标题: MediX-R1:开放式医疗强化学习(arXiv) 英文: MediX-R1: Open Ended Medical Reinforcement Learning 摘要: W...
来源: arXiv 爆炸度: 7/10 🔴🔴🔴🔴🔴 标题: VGG-T³:大规模离线前馈3D重建(arXiv) 英文: VGG-T$^3$: Offline Feed-Forward 3D Reconstruction at Scal...
来源: arXiv 爆炸度: 6/10 🔴🔴🔴🔴🔴 标题: 基于锚定的模型一致性方法(arXiv) 英文: Model Agreement via Anchoring 摘要: Numerous lines of aim to cont...
来源: arXiv 爆炸度: 8/10 🔴🔴🔴🔴🔴 标题: SeeThrough3D:文本到图像生成中的遮挡感知3D控制(arXiv) 英文: SeeThrough3D: Occlusion Aware 3D Control in T...
来源: arXiv 爆炸度: 9/10 🔴🔴🔴🔴🔴 标题: 一个数据集仅需1MB(arXiv) 英文: A Dataset is Worth 1 MB 摘要: A dataset server must often distribut...
来源: HuggingFace 标题: sentence-transformers/all-MiniLM-L6-v2 摘要: 下载: 187393215 | 点赞: 4520 链接: https://huggingface.co/se...
来源: HuggingFace 标题: google-bert/bert-base-uncased 摘要: 下载: 60873623 | 点赞: 2578 链接: https://huggingface.co/google-bert/...
来源: HuggingFace 标题: google/electra-base-discriminator 摘要: 下载: 48227365 | 点赞: 83 链接: https://huggingface.co/google/ele...
来源: HuggingFace 标题: Falconsai/nsfw_image_detection 摘要: 下载: 40235989 | 点赞: 1000 链接: https://huggingface.co/Falconsai/n...
来源: HuggingFace 标题: sentence-transformers/all-mpnet-base-v2 摘要: 下载: 26331045 | 点赞: 1253 链接: https://huggingface.co/se...
来源: arXiv 标题: Model Agreement via Anchoring 摘要: Numerous lines of aim to control $\textit{model disagreement}$ -- the...
来源: arXiv 标题: SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation 摘要: We identify occlusion reasonin...
来源: arXiv 标题: A Dataset is Worth 1 MB 摘要: A dataset server must often distribute the same large payload to many clien...
来源: arXiv 标题: SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport 摘要: Th...
来源: arXiv 标题: FlashOptim: Optimizers for Memory Efficient Training 摘要: Standard mixed-precision training of neural ne...
来源: arXiv 标题: Error 摘要: sortOrder must be in: ascending, descending 链接: https://arxiv.org/help/api/user-manual#sort