view article Article I love speed. We all love speed. But what happens when the speed is slowing you down? darkc0de • 6 days ago • 3
view article Article Making LLMs Smaller Without Breaking Them: A GLU-Aware Pruning Approach oopere • Nov 24, 2024 • 22
view article Article Frontier-Assisted Single-Prompt Disposable Risk Assessment darkc0de • 20 days ago • 5
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper • 2508.06471 • Published Aug 8, 2025 • 215
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect Paper • 2403.03853 • Published Mar 6, 2024 • 66
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels Paper • 2406.17415 • Published Jun 25, 2024 • 1
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models Paper • 2608.17202 • Published Aug 17 • 1
Python Pre-compiled Binaries Collection These are python wheels that are built and saved in their completed format for specific cuda/torch combinations. Use in images or GPaaS instances. • 3 items • Updated 5 days ago • 1
Data Attribution of Emergent Misalignment with Persona Features Paper • 2608.11025 • Published Aug 11 • 1
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs Paper • 2506.17080 • Published Jun 20, 2025 • 8