BioRefusalAudit SAE β€” Gemma 2 2B-IT L12, 10K Steps Pairwise NT-Xent

Contrastive TopK Sparse Autoencoder trained on Gemma 2 2B-IT residual stream activations (layer 12, width 16K) as part of the BioRefusalAudit project (AIxBio Hackathon 2026, Track 3: Biosecurity Tools).

Training summary

Run Mode Steps L_recon final L_contrastive final
5000-step baseline centroid mean 5000 0.28 0.888 (+7%, worse)
this model pairwise NT-Xent 5000 0.0043 7e-6 (βˆ’99.999%)

Warm-started from the 5000-step mean-mode checkpoint. Pairwise mode uses NT-Xent cross-tier pairs (benign Γ— hazard-adjacent, benign Γ— dual-use) rather than centroid means β€” this is why it converges where mean mode plateaued.

Key finding: L_contrastive reaches near-zero by step ~4000 on the 75-prompt training corpus. This reflects corpus overfitting (75 prompts Γ— 3 tiers), not genuine tier separability β€” bio vocabulary is shared across tiers regardless of hazard level. Corpus size, not compute or training steps, is the binding constraint for contrastive SAE fine-tuning on biosecurity prompts.

Architecture

  • Type: TopK SAE (k=32)
  • d_model: 2304 (Gemma 2 2B residual stream)
  • d_sae: 16384 (width 16k)
  • Hook layer: 12 (residual stream post-MLP)
  • Base SAE family: Gemma Scope 1 (google/gemma-scope-2b-pt-res, JumpReLU)
  • Fine-tune architecture: TopK (replaces JumpReLU threshold with hard top-k)

Training config

--contrastive-mode pairwise
--lam-contrast 0.5
--lam-sparse 0.04
--steps 5000
--batch-size 4
--lr 3e-4
--init-from runs/sae-training-gemma2-5000steps/sae_weights.pt
--checkpoint-every 1000

Training corpus: 75 prompts from data/eval_set_public/eval_set_public_v1.jsonl (25 per tier: benign, dual-use biology, hazard-adjacent biology).

Files

  • sae_weights.pt β€” final TopK SAE weights (W_enc, W_dec, b_enc, b_dec)
  • training_log.jsonl β€” per-step loss breakdown (total, l_recon, l_sparsity, l_contrastive, L0)
  • training_log.txt β€” human-readable training log

Usage

import torch
from biorefusalaudit.models.sae_adapter import load_sae

sae = load_sae(
    source="custom",
    repo_or_path="Solshine/biorefusalaudit-sae-gemma2-2b-l12-10ksteps-pairwise",
    layer=12,
    architecture="topk",
    d_model=2304,
    d_sae=16384,
    k=32,
)

License

HL3-BDS-CL-ECO-EXTR-FFD-MEDIA-MIL-MY-SUP-SV-TAL-USTA-XUAR. See full license text.

Citation

@misc{deleeuw2026biorefusalaudit,
  title={BioRefusalAudit: Measuring Refusal Depth in Biosecurity-Sensitive LLM Responses},
  author={DeLeeuw, Caleb},
  year={2026},
  note={AIxBio Hackathon 2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Solshine/biorefusalaudit-sae-gemma2-2b-l12-10ksteps-pairwise

Finetuned
(1105)
this model

Collection including Solshine/biorefusalaudit-sae-gemma2-2b-l12-10ksteps-pairwise