PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al.,...
每天自动聚合 AI 领域最新动态
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al.,...
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interact...
While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its train...
Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its e...
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically ...
Image retrieval has traditionally been formulated as a point-wise matching problem, where each candidate image is sco...
Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, ...
Strategic deception by LLM and VLM agents has emerged as a central AI alignment and safety concern. Social-deduction ...
Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabiliti...
Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating i...
VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that j...
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across task...