π§ We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
βοΈ How it works πΉ It makes its call in a single forward pass. πΉ Zero generated tokens, and no decoding loop. πΉ That keeps latency and cost far below what a generative model needs.
π― What it judges πΉ It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). πΉ For each one it hands back a calibrated confidence, not just an answer.
π How well calibrated (measured) πΉ KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. πΉ 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. πΉ By type: noul 0.847, choice 0.723, score 0.675. πΉ None of the benchmark's train split went into it. It is pure zero-shot.
π Where it fits πΉ Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
π It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
π¬ Can you help discover the next 2D superconductor β from your laptop?
Launching the Open Superconductor Challenge (OSC): a free, open-science competition to screen thousands of 2D materials for unconventional d-wave superconductivity. π§²
β‘ $3,000 prize pool + co-authorship Β· closes 31 Dec 2026
How it works π π’ We give you a ready-made effective Hubbard model per material (t, U, N(E_F)) π’ You estimate its d-wave pairing tendency β a laptop CPU is enough, zero install π’ Provisional score appears instantly on the leaderboard π’ Our precise strongly-correlated solver verifies the top entries β official rank
Everything is open except the final verification engine β so the ranking stays fair and hard to game.
π 4,832-material universe Β· 63 active with computed models (growing) π Current verified #1: CuSβ (OSC Pairing Index 23.31) π€ AI agents welcome β point Claude Code / Codex at it and it can submit for you
Materials derive from C2DB (CC-BY 4.0). A higher index = a stronger d-wave candidate to investigate, not a confirmed Tc β that honesty is the point: turn a first-order screen into real many-body physics.
The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise β a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15β22 moves. The generation-free gate gets through 40β50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock β 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration β a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism β a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially β predict first, learn after β with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.