-
llm-jp/llm-jp-4.1-33b-thinking
Text Generation • 33B • Updated • 1.14k • 14 -
llm-jp/llm-jp-4.1-33b-thinking-gguf
Text Generation • 33B • Updated • 3.86k • 4 -
llm-jp/llm-jp-4.1-32b-a3b-thinking
Text Generation • 32B • Updated • 1.51k • 21 -
llm-jp/llm-jp-4.1-32b-a3b-thinking-gguf
Text Generation • 32B • Updated • 6.4k • 8
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
-
llm-jp/llm-jp-4-33b-thinking
Text Generation • 33B • Updated • 3.33k • 45 -
llm-jp/llm-jp-4-33b-thinking-gguf
Text Generation • 33B • Updated • 325k • 11 -
llm-jp/llm-jp-4-33b-base
Text Generation • 33B • Updated • 1.6k • 8 -
llm-jp/llm-jp-4-32b-a3b-thinking
Text Generation • 32B • Updated • 2.46k • 40
WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
-
WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
Paper • 2510.22276 • Published • 3 -
llm-jp/WAON-Bench
Viewer • Updated • 1.87k • 169 • 2 -
llm-jp/waon-siglip2-base-patch16-256
Zero-Shot Image Classification • 0.4B • Updated • 739 • 1 -
llm-jp/WAON
Updated • 144 • 8
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
-
llm-jp/optimal-sparsity-code-d512-E8-k2-320M-A170M
Text Generation • 0.3B • Updated • 190 -
llm-jp/optimal-sparsity-code-d512-E16-k2-520M-A170M
Text Generation • 0.5B • Updated • 173 -
llm-jp/optimal-sparsity-code-d512-E32-k2-920M-A170M
Text Generation • 0.9B • Updated • 165 -
llm-jp/optimal-sparsity-code-d512-E64-k2-1.7B-A170M
Text Generation • 2B • Updated • 126
Fine-tuned models in the LLM-jp-3 model series
-
llm-jp/llm-jp-3.1-8x13b-instruct4
Text Generation • 73B • Updated • 226 • 4 -
llm-jp/llm-jp-3.1-8x13b-32K-instruct4
Text Generation • 73B • Updated • 205 • 2 -
llm-jp/llm-jp-3.1-13b-instruct4
Text Generation • 14B • Updated • 1.9k • 19 -
llm-jp/llm-jp-3.1-1.8b-instruct4
Text Generation • 2B • Updated • 3.67k • 22
Pre-trained models in the LLM-jp-3.1 model series
Models in the LLM-jp ver2.0 model series
-
llm-jp/llm-jp-13b-v2.0
Text Generation • Updated • 506 • 15 -
llm-jp/llm-jp-13b-instruct-full-dolly-ichikara_004_001_single-oasst-oasst2-v2.0
Text Generation • 14B • Updated • 147 -
llm-jp/llm-jp-13b-instruct-full-ac_001-dolly-ichikara_004_001_single-oasst-oasst2-v2.0
Text Generation • 14B • Updated • 135 • 1 -
llm-jp/llm-jp-13b-instruct-full-ac_001_16x-dolly-ichikara_004_001_single-oasst-oasst2-v2.0
Text Generation • 14B • Updated • 141 • 3
Models in the LLM-jp ver1.0 model series
-
llm-jp/llm-jp-13b-v1.0
Text Generation • Updated • 416 • 41 -
llm-jp/llm-jp-13b-instruct-full-jaster-v1.0
Text Generation • Updated • 260 • 15 -
llm-jp/llm-jp-13b-instruct-full-jaster-dolly-oasst-v1.0
Text Generation • Updated • 259 • 8 -
llm-jp/llm-jp-13b-instruct-full-dolly-oasst-v1.0
Text Generation • Updated • 254 • 4
-
llm-jp/llm-jp-4.1-thinking-sft-data
Viewer • Updated • 4.87M • 2.6k • 3 -
llm-jp/llm-jp-4.1-33b-thinking-dpo-data
Viewer • Updated • 82.3k • 1.19k • 2 -
llm-jp/llm-jp-4.1-32b-a3b-thinking-dpo-data
Viewer • Updated • 255k • 1.07k • 2 -
llm-jp/llm-jp-4.1-8b-thinking-dpo-data
Viewer • Updated • 290k • 812 • 2
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision–Language Models
Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
-
llm-jp/optimal-sparsity-math-d512-E8-k2-320M-A170M
Text Generation • 0.3B • Updated • 156 -
llm-jp/optimal-sparsity-math-d512-E16-k2-520M-A170M
Text Generation • 0.5B • Updated • 121 -
llm-jp/optimal-sparsity-math-d512-E32-k2-920M-A170M
Text Generation • 0.9B • Updated • 118 -
llm-jp/optimal-sparsity-math-d512-E64-k2-1.7B-A170M
Text Generation • 2B • Updated • 146
Fine-tuned models in the LLM-jp-3 model series
-
llm-jp/llm-jp-3-8x13b-instruct3
Text Generation • 73B • Updated • 309 • 8 -
llm-jp/llm-jp-3-172b-instruct3
Text Generation • 172B • Updated • 261 • 11 -
llm-jp/llm-jp-3-13b-instruct3
Text Generation • 14B • Updated • 848 • 8 -
llm-jp/llm-jp-3-8x1.8b-instruct3
Text Generation • 9B • Updated • 144 • 4
Pre-trained models in the LLM-jp-3 model series
Models in the LLM-jp ver1.1 model series
-
llm-jp/llm-jp-13b-dpo-lora-hh_rlhf_ja-v1.1
Text Generation • Updated • 1 -
llm-jp/llm-jp-13b-instruct-full-dolly_en-dolly_ja-ichikara_003_001-oasst_en-oasst_ja-v1.1
Text Generation • 13B • Updated • 226 • 2 -
llm-jp/llm-jp-13b-instruct-lora-dolly_en-dolly_ja-ichikara_003_001-oasst_en-oasst_ja-v1.1
Text Generation • Updated • 1
-
llm-jp/llm-jp-4.1-33b-thinking
Text Generation • 33B • Updated • 1.14k • 14 -
llm-jp/llm-jp-4.1-33b-thinking-gguf
Text Generation • 33B • Updated • 3.86k • 4 -
llm-jp/llm-jp-4.1-32b-a3b-thinking
Text Generation • 32B • Updated • 1.51k • 21 -
llm-jp/llm-jp-4.1-32b-a3b-thinking-gguf
Text Generation • 32B • Updated • 6.4k • 8
-
llm-jp/llm-jp-4.1-thinking-sft-data
Viewer • Updated • 4.87M • 2.6k • 3 -
llm-jp/llm-jp-4.1-33b-thinking-dpo-data
Viewer • Updated • 82.3k • 1.19k • 2 -
llm-jp/llm-jp-4.1-32b-a3b-thinking-dpo-data
Viewer • Updated • 255k • 1.07k • 2 -
llm-jp/llm-jp-4.1-8b-thinking-dpo-data
Viewer • Updated • 290k • 812 • 2
-
llm-jp/llm-jp-4-33b-thinking
Text Generation • 33B • Updated • 3.33k • 45 -
llm-jp/llm-jp-4-33b-thinking-gguf
Text Generation • 33B • Updated • 325k • 11 -
llm-jp/llm-jp-4-33b-base
Text Generation • 33B • Updated • 1.6k • 8 -
llm-jp/llm-jp-4-32b-a3b-thinking
Text Generation • 32B • Updated • 2.46k • 40
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision–Language Models
WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
-
WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
Paper • 2510.22276 • Published • 3 -
llm-jp/WAON-Bench
Viewer • Updated • 1.87k • 169 • 2 -
llm-jp/waon-siglip2-base-patch16-256
Zero-Shot Image Classification • 0.4B • Updated • 739 • 1 -
llm-jp/WAON
Updated • 144 • 8
Llama-Mimi: Speech Language Models with Interleaved Semantic and Acoustic Tokens
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
-
llm-jp/optimal-sparsity-code-d512-E8-k2-320M-A170M
Text Generation • 0.3B • Updated • 190 -
llm-jp/optimal-sparsity-code-d512-E16-k2-520M-A170M
Text Generation • 0.5B • Updated • 173 -
llm-jp/optimal-sparsity-code-d512-E32-k2-920M-A170M
Text Generation • 0.9B • Updated • 165 -
llm-jp/optimal-sparsity-code-d512-E64-k2-1.7B-A170M
Text Generation • 2B • Updated • 126
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
-
llm-jp/optimal-sparsity-math-d512-E8-k2-320M-A170M
Text Generation • 0.3B • Updated • 156 -
llm-jp/optimal-sparsity-math-d512-E16-k2-520M-A170M
Text Generation • 0.5B • Updated • 121 -
llm-jp/optimal-sparsity-math-d512-E32-k2-920M-A170M
Text Generation • 0.9B • Updated • 118 -
llm-jp/optimal-sparsity-math-d512-E64-k2-1.7B-A170M
Text Generation • 2B • Updated • 146
Fine-tuned models in the LLM-jp-3 model series
-
llm-jp/llm-jp-3.1-8x13b-instruct4
Text Generation • 73B • Updated • 226 • 4 -
llm-jp/llm-jp-3.1-8x13b-32K-instruct4
Text Generation • 73B • Updated • 205 • 2 -
llm-jp/llm-jp-3.1-13b-instruct4
Text Generation • 14B • Updated • 1.9k • 19 -
llm-jp/llm-jp-3.1-1.8b-instruct4
Text Generation • 2B • Updated • 3.67k • 22
Fine-tuned models in the LLM-jp-3 model series
-
llm-jp/llm-jp-3-8x13b-instruct3
Text Generation • 73B • Updated • 309 • 8 -
llm-jp/llm-jp-3-172b-instruct3
Text Generation • 172B • Updated • 261 • 11 -
llm-jp/llm-jp-3-13b-instruct3
Text Generation • 14B • Updated • 848 • 8 -
llm-jp/llm-jp-3-8x1.8b-instruct3
Text Generation • 9B • Updated • 144 • 4
Pre-trained models in the LLM-jp-3.1 model series
Pre-trained models in the LLM-jp-3 model series
Models in the LLM-jp ver2.0 model series
-
llm-jp/llm-jp-13b-v2.0
Text Generation • Updated • 506 • 15 -
llm-jp/llm-jp-13b-instruct-full-dolly-ichikara_004_001_single-oasst-oasst2-v2.0
Text Generation • 14B • Updated • 147 -
llm-jp/llm-jp-13b-instruct-full-ac_001-dolly-ichikara_004_001_single-oasst-oasst2-v2.0
Text Generation • 14B • Updated • 135 • 1 -
llm-jp/llm-jp-13b-instruct-full-ac_001_16x-dolly-ichikara_004_001_single-oasst-oasst2-v2.0
Text Generation • 14B • Updated • 141 • 3
Models in the LLM-jp ver1.1 model series
-
llm-jp/llm-jp-13b-dpo-lora-hh_rlhf_ja-v1.1
Text Generation • Updated • 1 -
llm-jp/llm-jp-13b-instruct-full-dolly_en-dolly_ja-ichikara_003_001-oasst_en-oasst_ja-v1.1
Text Generation • 13B • Updated • 226 • 2 -
llm-jp/llm-jp-13b-instruct-lora-dolly_en-dolly_ja-ichikara_003_001-oasst_en-oasst_ja-v1.1
Text Generation • Updated • 1
Models in the LLM-jp ver1.0 model series
-
llm-jp/llm-jp-13b-v1.0
Text Generation • Updated • 416 • 41 -
llm-jp/llm-jp-13b-instruct-full-jaster-v1.0
Text Generation • Updated • 260 • 15 -
llm-jp/llm-jp-13b-instruct-full-jaster-dolly-oasst-v1.0
Text Generation • Updated • 259 • 8 -
llm-jp/llm-jp-13b-instruct-full-dolly-oasst-v1.0
Text Generation • Updated • 254 • 4