Text Generation
PyTorch
Safetensors
GGUF
English
byrne
spikewhale
looped-transformer
memory-cache
mla
small-language-model
conversational
Instructions to use Quazim0t0/Byrne-100M-Ultra-MC with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Quazim0t0/Byrne-100M-Ultra-MC with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Quazim0t0/Byrne-100M-Ultra-MC:F16 # Run inference directly in the terminal: llama cli -hf Quazim0t0/Byrne-100M-Ultra-MC:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Quazim0t0/Byrne-100M-Ultra-MC:F16 # Run inference directly in the terminal: llama cli -hf Quazim0t0/Byrne-100M-Ultra-MC:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Quazim0t0/Byrne-100M-Ultra-MC:F16 # Run inference directly in the terminal: ./llama-cli -hf Quazim0t0/Byrne-100M-Ultra-MC:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Quazim0t0/Byrne-100M-Ultra-MC:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Quazim0t0/Byrne-100M-Ultra-MC:F16
Use Docker
docker model run hf.co/Quazim0t0/Byrne-100M-Ultra-MC:F16
- LM Studio
- Jan
- vLLM
How to use Quazim0t0/Byrne-100M-Ultra-MC with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Quazim0t0/Byrne-100M-Ultra-MC" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-100M-Ultra-MC", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Quazim0t0/Byrne-100M-Ultra-MC:F16
- Ollama
How to use Quazim0t0/Byrne-100M-Ultra-MC with Ollama:
ollama run hf.co/Quazim0t0/Byrne-100M-Ultra-MC:F16
- Unsloth Desktop
- Docker Model Runner
How to use Quazim0t0/Byrne-100M-Ultra-MC with Docker Model Runner:
docker model run hf.co/Quazim0t0/Byrne-100M-Ultra-MC:F16
- Lemonade
How to use Quazim0t0/Byrne-100M-Ultra-MC with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Quazim0t0/Byrne-100M-Ultra-MC:F16
Run and chat with the model
lemonade run user.Byrne-100M-Ultra-MC-F16
List all available models
lemonade list
- Atomic Chat
Download audit_architecture_paths.py from Quazim0t0/Byrne-100M-Ultra-MC: direct link, hf CLI and curl.
- Browser
- Download file 2.17 kB
-
https://huggingface.co/Quazim0t0/Byrne-100M-Ultra-MC/resolve/main/audit_architecture_paths.py
- Command line
-
hf download hf://Quazim0t0/Byrne-100M-Ultra-MC/audit_architecture_paths.py
-
curl -L -o audit_architecture_paths.py https://huggingface.co/Quazim0t0/Byrne-100M-Ultra-MC/resolve/main/audit_architecture_paths.py
2.17 kB
| import json | |
| from pathlib import Path | |
| import torch | |
| from config import SpikeWhaleConfig | |
| from model_v2 import SpikeWhaleLM, reset_memory_cache | |
| torch.set_num_threads(2) | |
| torch.manual_seed(123) | |
| c = torch.load('checkpoints/dpo_3200.pt', map_location='cpu', weights_only=False) | |
| model = SpikeWhaleLM(SpikeWhaleConfig(**c['config'])).eval() | |
| model.load_state_dict(c['model_state'], strict=True) | |
| x = torch.tensor([[2, 41, 62, 81, 102, 121, 142, 161]]) | |
| changed = x.clone(); changed[:, 4:] += 31 | |
| def run(ids, mask=None): | |
| reset_memory_cache(model) | |
| return model(ids, attention_mask=mask).logits | |
| results = {} | |
| with torch.no_grad(): | |
| a, b = run(x), run(changed) | |
| results['no_mask_future_prefix_delta'] = (a[:, :4]-b[:, :4]).abs().max().item() | |
| mask = torch.ones_like(x) | |
| a, b = run(x, mask), run(changed, mask) | |
| results['binary_mask_future_prefix_delta'] = (a[:, :4]-b[:, :4]).abs().max().item() | |
| try: | |
| additive = torch.zeros(1,1,8,8).masked_fill(torch.triu(torch.ones(8,8,dtype=torch.bool),1), float('-inf')) | |
| additive_logits = run(x, additive) | |
| results['additive_binary_logits_delta'] = (additive_logits-a).abs().max().item() | |
| results['additive_mask_error'] = None | |
| except Exception as exc: | |
| results['additive_mask_error'] = str(exc) | |
| reset_memory_cache(model) | |
| results['all_masked_loss'] = str(model(x, labels=torch.full_like(x,-100)).loss.item()) | |
| baseline = run(x) | |
| model.train() | |
| training = run(x) | |
| results['train_eval_logits_delta'] = (baseline-training).abs().max().item() | |
| from model_v2 import _masked_ce | |
| empty_logits = torch.randn(5, 11, requires_grad=True) | |
| empty_loss = _masked_ce(empty_logits, torch.full((5,), -100)) | |
| empty_loss.backward() | |
| results['empty_ce_backward_finite'] = bool(torch.isfinite(empty_logits.grad).all()) | |
| assert results['no_mask_future_prefix_delta'] < 1e-4 | |
| assert results['binary_mask_future_prefix_delta'] < 1e-4 | |
| assert results['additive_mask_error'] is None | |
| assert results['additive_binary_logits_delta'] < 1e-4 | |
| assert results['all_masked_loss'] == '0.0' | |
| Path('architecture_path_audit.json').write_text(json.dumps(results, indent=2)) | |
| print(json.dumps(results, indent=2)) | |