Instructions to use txktxkabcd/fahqgpt-1.0-nano with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use txktxkabcd/fahqgpt-1.0-nano with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M # Run inference directly in the terminal: llama cli -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M # Run inference directly in the terminal: llama cli -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
Use Docker
docker model run hf.co/txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use txktxkabcd/fahqgpt-1.0-nano with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "txktxkabcd/fahqgpt-1.0-nano" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "txktxkabcd/fahqgpt-1.0-nano", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
- Ollama
How to use txktxkabcd/fahqgpt-1.0-nano with Ollama:
ollama run hf.co/txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use txktxkabcd/fahqgpt-1.0-nano with Docker Model Runner:
docker model run hf.co/txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
- Lemonade
How to use txktxkabcd/fahqgpt-1.0-nano with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull txktxkabcd/fahqgpt-1.0-nano:Q4_K_M
Run and chat with the model
lemonade run user.fahqgpt-1.0-nano-Q4_K_M
List all available models
lemonade list
- Atomic Chat
FahQgpt 1.0 Nano
由 Bilibili @一片烂海苔 从零训练的小型语言模型。
一个 145.9M 参数的 LLaMA 式模型,在单张 RTX 5060 Ti 16GB 上从零训练完成, 没有使用任何预训练权重。
模型规格
| 项目 | 数值 |
|---|---|
| 参数量 | 145.9M |
| 架构 | LLaMA 式:RMSNorm + RoPE + GQA (14 查询头 / 2 KV 头) + SwiGLU + 权重绑定 |
| 层数 / 隐藏维 | 14 层 / 896 |
| 上下文长度 | 1024 |
| 词表 | 自训 32k BPE(32002,byte-level) |
| 训练 token | 1,899,888,640(约 19 亿) |
| 训练步数 | 14,495 步 |
| 预训练耗时 | 13.45 小时 |
| 最佳验证 loss | 3.0143 |
| 精度 | bf16 + torch.compile |
训练流程(完全从零)
数据准备 → 自训分词器 → 预训练 → 退火 → 对话继续预训练 → 混合损失 SFT → 量化
32k BPE 14,495步 lr→0 15000步 1800步
两个关键的技术决定
对话格式用纯文本,不用特殊 token。 最初用
<|user|>/<|assistant|>特殊 token 标记对话、只在 assistant 段算 loss, 全部失败(loss 卡在 10.37 = 均匀分布,一旦学习率够大就塌成':::')。 改成把对话写成普通文本(User: .../Assistant: ...)、整条序列算 loss 后, 任务退化成"续写对话"——这正是 base 模型本来就擅长的事,loss 从 10 降到 3.0。必须在预训练阶段就见过对话。 原本的预训练语料 83.4% 是 FineWeb-Edu(教育文章),完全没有对话。 直接做指令微调无效。补上"对话继续预训练"(339,389 段对话 / 101M token, 混 40% 原数据)后模型才学会对话。这与 COLING 2025 的 CLASS-IT(同样 140M 规模) 的结论一致。
SFT 必须混通用文本。 纯指令数据做梯度下降会把输出分布推向很窄的区域,train loss 掉到 ~8 以下 就必崩。每个 batch 混 60% 通用文本才能稳住。
使用方式
llama.cpp
llama-server -m fahqgpt-1.0-nano-routea-f16.gguf \
--host 0.0.0.0 --port 8080 -c 512 -ngl 99 -r "User:"
-r "User:" 很重要:模型是连续对话训练的,答完会自己接着编下一轮
"User: ...",加这个停止符输出才干净。
LM Studio / Ollama / 手机 App
直接加载 GGUF 即可,对话模板已写入文件元数据,会自动识别。
能力与局限(如实说明)
能做:
- ✅ 自我介绍(知道自己的名字和开发者)
- ✅ 简单英文闲聊
- ✅ 输出连贯的英文
做不到:
- ❌ 事实类问答很弱(
1+1、首都名等可能答错) - ❌ 不会推理、不会写代码
- ❌ 中文能力很弱(训练语料以英文为主)
这是 146M 参数量的客观限制。模型定位是"从零走通全流程"的工程成果。 GPT-2 small 是 124M,可以参考它的水平。
建议 temperature 设 0.5 左右。
文件
| 文件 | 大小 | 说明 |
|---|---|---|
fahqgpt-1.0-nano-routea-f16.gguf |
350 MB | 推荐,全精度 |
fahqgpt-1.0-nano-routea-Q4_K_M.gguf |
137 MB | 量化版 |
训练数据
| 来源 | 占比 |
|---|---|
| fineweb-edu-dedup | 83.4% |
| cosmopedia | 6.8% |
| openwebmath | 6.4% |
| wikipedia | 3.2% |
| SmolTalk(对话继续预训练 + SFT) | — |
许可
Apache-2.0
- Downloads last month
- 94
4-bit
16-bit