BioRefusalAudit SAE β Gemma 2 2B-IT L12, 10K Steps Pairwise NT-Xent
Contrastive TopK Sparse Autoencoder trained on Gemma 2 2B-IT residual stream activations (layer 12, width 16K) as part of the BioRefusalAudit project (AIxBio Hackathon 2026, Track 3: Biosecurity Tools).
Training summary
| Run | Mode | Steps | L_recon final | L_contrastive final |
|---|---|---|---|---|
| 5000-step baseline | centroid mean | 5000 | 0.28 | 0.888 (+7%, worse) |
| this model | pairwise NT-Xent | 5000 | 0.0043 | 7e-6 (β99.999%) |
Warm-started from the 5000-step mean-mode checkpoint. Pairwise mode uses NT-Xent cross-tier pairs (benign Γ hazard-adjacent, benign Γ dual-use) rather than centroid means β this is why it converges where mean mode plateaued.
Key finding: L_contrastive reaches near-zero by step ~4000 on the 75-prompt training corpus. This reflects corpus overfitting (75 prompts Γ 3 tiers), not genuine tier separability β bio vocabulary is shared across tiers regardless of hazard level. Corpus size, not compute or training steps, is the binding constraint for contrastive SAE fine-tuning on biosecurity prompts.
Architecture
- Type: TopK SAE (k=32)
- d_model: 2304 (Gemma 2 2B residual stream)
- d_sae: 16384 (width 16k)
- Hook layer: 12 (residual stream post-MLP)
- Base SAE family: Gemma Scope 1 (
google/gemma-scope-2b-pt-res, JumpReLU) - Fine-tune architecture: TopK (replaces JumpReLU threshold with hard top-k)
Training config
--contrastive-mode pairwise
--lam-contrast 0.5
--lam-sparse 0.04
--steps 5000
--batch-size 4
--lr 3e-4
--init-from runs/sae-training-gemma2-5000steps/sae_weights.pt
--checkpoint-every 1000
Training corpus: 75 prompts from data/eval_set_public/eval_set_public_v1.jsonl
(25 per tier: benign, dual-use biology, hazard-adjacent biology).
Files
sae_weights.ptβ final TopK SAE weights (W_enc, W_dec, b_enc, b_dec)training_log.jsonlβ per-step loss breakdown (total, l_recon, l_sparsity, l_contrastive, L0)training_log.txtβ human-readable training log
Usage
import torch
from biorefusalaudit.models.sae_adapter import load_sae
sae = load_sae(
source="custom",
repo_or_path="Solshine/biorefusalaudit-sae-gemma2-2b-l12-10ksteps-pairwise",
layer=12,
architecture="topk",
d_model=2304,
d_sae=16384,
k=32,
)
License
HL3-BDS-CL-ECO-EXTR-FFD-MEDIA-MIL-MY-SUP-SV-TAL-USTA-XUAR. See full license text.
Citation
@misc{deleeuw2026biorefusalaudit,
title={BioRefusalAudit: Measuring Refusal Depth in Biosecurity-Sensitive LLM Responses},
author={DeLeeuw, Caleb},
year={2026},
note={AIxBio Hackathon 2026}
}