Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints...
每天自动聚合 AI 领域最新动态
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints...
When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio...
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's inter...
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Lega...
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the...
Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and...
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious eth...
CLI-based software-engineering agents have matured rapidly, yet the open ecosystem has converged on a single training...
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are ...
Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence le...
Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies wit...
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open wheth...