GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 2 days ago • 36
view article Article Baseten on Hugging Face Inference Providers 🔥 +6 alexker-baseten, rolandcrosby-baseten, squidarth, johan-baseten, celinah, sbrandeis, Wauplin, merve • 2 days ago • 22
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 12 days ago • 36
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Paper • 2608.05137 • Published 3 days ago • 18
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 2 days ago • 27
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published 9 days ago • 59
SQuAD: 100,000+ Questions for Machine Comprehension of Text Paper • 1606.05250 • Published Jun 16, 2016 • 5
Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation Paper • 2608.02738 • Published 5 days ago • 47
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Paper • 2608.05102 • Published 3 days ago • 62
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 5 days ago • 81
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 5 days ago • 58
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published 17 days ago • 13
🍃 MINT-1T Collection Data for "MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens" • 11 items • Updated Mar 2 • 69