Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published Jul 3 • 84
Skip a Layer or Loop It? Learning Program-of-Layers in LLMs Paper • 2606.06574 • Published Jun 4 • 25
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 243
Representation Forcing for Bottleneck-Free Unified Multimodal Models Paper • 2605.31604 • Published May 29 • 63
Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Paper • 2605.02801 • Published May 4 • 10
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents Paper • 2605.05185 • Published May 6 • 108
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 239
LLM Safety From Within: Detecting Harmful Content with Internal Representations Paper • 2604.18519 • Published Apr 20 • 26
Elucidating the SNR-t Bias of Diffusion Probabilistic Models Paper • 2604.16044 • Published Apr 17 • 72