# CALIBRATION — what is actually in the file we calibrated on Quantizers ship an importance matrix and stop there. You are asked to trust a `.gguf` blob whose contents nobody states. This page states them, and the corpus itself is in this repository as `cal_archsloth.jsonl` — so you can read it, diff it, or replace it. --- ## Composition — measured, not intended | | samples | share | |---|---|---| | Korean | 249 | 44.9 % | | English | 196 | 35.3 % | | Source code | 110 | 19.8 % | | **total** | **555** | 7.64 MB | | by character | | |---|---| | Hangul | 25.2 % | | Latin letters | 42.7 % | | whitespace, punctuation, digits, symbols | the rest | | total characters | 4,995,000 | Every sample is exactly 9,000 characters. Median, minimum and maximum are all 9,000 — the corpus is built by fixed-width slicing, so no sample is quietly longer or shorter than another and no single document dominates. Sources: encyclopedic prose in Korean and English, and real C/C++/Python source files. No chat transcripts, no synthetic text, no benchmark material. --- ## The ordering is part of the recipe Samples are **interleaved at the sample level**: one Korean, one English, and a code sample every other pair. They are not shuffled and they are not concatenated by language. This is not cosmetic. Moving from a corpus with a similar overall character mix to strict alternation improved English by **11.4 %** with everything else held fixed. AutoRound's rounding search consumes samples in order; a run of same-language samples biases the search before it ever reaches the others. --- ## Why a calibration corpus moves anything at all An importance matrix reweights which weights matter, and llama.cpp's scale search then runs as it always does. The rounding search we use **replaces** that scale search: it decides, weight by weight, which direction to round. The text it reads therefore shapes the file directly rather than through a weighting term. That is why the language axis reopened for us after it had been closed on the imatrix side. They are different knobs on different parts of the pipeline, and only one of them moves. --- ## What is in the corpus vs what we measured Three of our ten evaluation axes appear in the corpus: **Korean, English, code.** The other seven — Spanish, Japanese, Chinese, Russian, Arabic, Hindi, Thai — are **not in the corpus at all.** They are there so that a build which buys one language with another has somewhere to show it. We keep them because we have seen exactly that failure in other people's weights: a build that held its English scores while losing a different language entirely, and looked healthy in aggregate the whole time. --- ## Things we changed and measured, that did not help All deltas below are against the corpus that shipped. Negative is better. | Change | Korean | English | |---|---|---| | 512 samples instead of 128 in the search | no movement, 4× the compute | | | 4,096-token samples instead of 2,048 | −2 % | **+23 %** | | Korean-only corpus | −3 % | **+14 %** | | English-only corpus, ours | **+65 %** | −10 % | | English-only corpus, `NeelNanda/pile-10k` — the tool's default | **+76 %** | +9 % | | adding code to the mix | +3 % | +5 % | The last row is the only trade we took: it moved the code axis **−52.9 %**, from a statistical tie into a clear lead, for three to five percent on two axes we already led by twenty to fifty. The two English-only rows are the ones worth sitting with. The tool's own default corpus is not only worse for Korean by 76 % — it is worse for **English** than a mixed corpus is, by 9 %. Reaching for the default is not a neutral choice. --- `cal_archsloth.jsonl` is in this repository. One JSON object per line, `{"text": ...}`. Re-run the recipe in the model card against it and you should land on the same file.