Persian Pixel: A large-scale synthetic OCR dataset for Persian language
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages des...
每天自动聚合 AI 领域最新动态
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages des...
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensi...
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represe...
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D ...
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requ...
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data f...
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in e...
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges fo...
As large language models and AI agents become the primary consumers of search results, document set quality determine...
Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the ...
Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained ...
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL al...