BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language
Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged a...
每天自动聚合 AI 领域最新动态
Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged a...
Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms...
Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unli...
Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own co...
Block Diffusion Language Models (BD-LMs) improve diffusion-based text generation with KV caching and flexible-length ...
While large language models have been dominating the research landscape recently, small language models remain highly...
Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and the...
Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fai...
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction m...
Foundation models have transformed vision and language processing by providing rich, reusable representations that tr...
Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneou...
Agent skills extend language-model agents with task-specific procedures, scripts, and references, but the tasks and e...