arXiv

Grad Detect: Gradient-Based Hallucination Detection in LLMs

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to...

AI 聚合
2026-06-24
arXiv

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answ...

AI 聚合
2026-06-24
arXiv

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generat...

AI 聚合
2026-06-24
arXiv

Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution

Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quan...

AI 聚合
2026-06-24
arXiv

Solving Inverse Problems of Chaotic Systems with Bidirectional Conditional Flow Matching

Modeling chaotic systems is crucial yet challenging. Inverse problems in chaotic dynamics, namely inferring initial c...

AI 聚合
2026-06-24
arXiv

Difference-Making without Making a Difference

Over a series of seven papers, Andreas & Günther have introduced seven definitions of actual causation and have class...

AI 聚合
2026-06-24
arXiv

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks...

AI 聚合
2026-06-24
arXiv

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This m...

AI 聚合
2026-06-24
arXiv

Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models

Prompt-based learning has emerged as a dominant paradigm in natural language processing. This study explores the impa...

AI 聚合
2026-06-24
arXiv

World Models in Pieces: Structural Certification for General Agents

In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a wo...

AI 聚合
2026-06-24
arXiv

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still st...

AI 聚合
2026-06-24
arXiv

It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces

Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are abl...

AI 聚合
2026-06-24
首页 上一页 第 178 / 212 页 下一页 末页