Zach Mustafa PRO
Zmu
AI & ML interests
None yet
Recent Activity
upvoted an article about 4 hours ago
New in llama.cpp: Decision Models upvoted a collection about 4 hours ago
🎲 Decision 2.0 upvoted a collection about 4 hours ago
⛵ Vela 2.0Organizations
Small LLMs
-
Qwen/Qwen3.5-0.8B-Base
Image-Text-to-Text • 0.9B • Updated • 453k • 115 -
Qwen/Qwen3.5-0.8B
Image-Text-to-Text • 0.9B • Updated • 2.56M • 747 -
Cactus-Compute/needle2
Text Generation • Updated • 13.7k • 331 -
XHToken/Spark-X2.5-4B
Text Generation • 4B • Updated • 37.3k • 1.38k
Video Understanding
-
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Paper • 2310.19773 • Published • 20 -
Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models
Paper • 2310.05863 • Published • 2 -
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Paper • 2311.06242 • Published • 98 -
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
Paper • 2311.10126 • Published • 9
Multimodal
OCR
Tool Calling
LLM
-
System 2 Attention (is something you might need too)
Paper • 2311.11829 • Published • 43 -
ToolTalk: Evaluating Tool-Usage in a Conversational Setting
Paper • 2311.10775 • Published • 9 -
Adapters: A Unified Library for Parameter-Efficient and Modular Transfer Learning
Paper • 2311.11077 • Published • 28
Encoders
-
google/siglip-so400m-patch14-224
Zero-Shot Image Classification • 0.9B • Updated • 60.3k • 60 -
knowledgator/gliformer-large-v1
Token Classification • Updated • 2.07k • 165 -
convaiinnovations/laya
Text Classification • 0.4B • Updated • 20.4k • 5.3k -
AlexWortega/openjev
Text Classification • Updated • 638
Computer Vision
OCR
Small LLMs
-
Qwen/Qwen3.5-0.8B-Base
Image-Text-to-Text • 0.9B • Updated • 453k • 115 -
Qwen/Qwen3.5-0.8B
Image-Text-to-Text • 0.9B • Updated • 2.56M • 747 -
Cactus-Compute/needle2
Text Generation • Updated • 13.7k • 331 -
XHToken/Spark-X2.5-4B
Text Generation • 4B • Updated • 37.3k • 1.38k
Tool Calling
Video Understanding
-
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Paper • 2310.19773 • Published • 20 -
Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models
Paper • 2310.05863 • Published • 2 -
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
Paper • 2311.06242 • Published • 98 -
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
Paper • 2311.10126 • Published • 9
LLM
-
System 2 Attention (is something you might need too)
Paper • 2311.11829 • Published • 43 -
ToolTalk: Evaluating Tool-Usage in a Conversational Setting
Paper • 2311.10775 • Published • 9 -
Adapters: A Unified Library for Parameter-Efficient and Modular Transfer Learning
Paper • 2311.11077 • Published • 28
Multimodal
Encoders
-
google/siglip-so400m-patch14-224
Zero-Shot Image Classification • 0.9B • Updated • 60.3k • 60 -
knowledgator/gliformer-large-v1
Token Classification • Updated • 2.07k • 165 -
convaiinnovations/laya
Text Classification • 0.4B • Updated • 20.4k • 5.3k -
AlexWortega/openjev
Text Classification • Updated • 638