Luau Coder 1.0 - 0.5B (Base) 🦭

Model Overview

  • Type: Text Causal Language Model
  • Training Stage: Pre-training
  • Language Model
    • Number of Parameters: 0.5B
    • Hidden Dimension: 1024
    • Token Embedding: 32768 (Padded)
    • Number of Layers: 32
    • Hidden Layout: 8 × (3 × (Kimi DeltaNet → FFN) → 1 × (Multi-head Latent Attention → FFN))
    • Kimi DeltaNet:
      • Number of Linear Attention Heads: 16
      • Head Dimension: 64
    • Multi-head Latent Attention:
      • Number of Attention Heads: 16
      • Head Dimension: 64 or 128
      • No Positional Embedding Dimension: 64
      • Use Output Gate: Yes
    • Attention Residuals:
      • Granularity: Sub-Layer
      • Block Size: 4
    • Feed Forward Network:
      • Intermediate Dimension: 2816
      • Activation: SiTU-GLU
    • LM Output: 32768 (Tied to token embedding)
  • Context Length: 65,536 natively.

License

This repository and Luau-Coder model weights are released under the Apache-2.0 License.

Downloads last month
369
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for khtsly/Luau-Coder-1.0-0.5B-Base

Unable to build the model tree, the base model loops to the model itself. Learn more.

Datasets used to train khtsly/Luau-Coder-1.0-0.5B-Base