vera-spark

vera-spark

vera-spark is a non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with calibrated probabilities in a single forward pass, over an 8K (8,192-token) context. It never generates text, so there is nothing to parse and nothing to hallucinate. It speaks the same wire format as Jev (POST /v1/systemone).

vera-spark is a generalist base: fine-tune it for your use case

Fine-tuning on your own decisions will yield better results.

The released checkpoint works zero-shot across many kinds of decisions, but it is built to be a starting point. Fine-tune it on decisions from your own domain and it will fit your labels, your wording and your calibration far better than the generalist base.

vera-spark is the compact sibling of vera-core. See vera-spark and vera-core.

Demo: JevPilot driving on vera-spark

The car is driven by JevPilot decisions served from vera-spark.

Use cases

vera-spark is a fast decision layer for anything that can be phrased as "given this state, pick from these options". Fine-tune it on your own labels for:

  • Routing and triage: send a ticket, email or request to the right team, queue or priority.
  • Intent and topic classification: map messages to your own intent set, with calibrated confidence for escalation.
  • Tool and function selection: choose which tool or API an agent should call next.
  • Agent control: pick the next action for a browser, driving or workflow agent in a few milliseconds.
  • Policy and compliance checks: decide whether a case follows a rule or breaks it.
  • Scoring and severity: rate on an ordinal scale (urgency, sentiment, risk) with a full probability distribution.
  • Fact and claim verification: check a statement against a document.
  • Judging and ranking: compare candidate responses and pick the better one.
  • Confidence-gated workflows: act automatically when confident, hand off to a slower model or a human when not.

Benchmarks

JevBench (accuracy %, higher is better; ECE lower is better)

model params easy original hard overall ECE (overall)
vera-core ~2B 100 93.1 58.6 77.9 0.098
vera-spark (this repo) ~630M 100 83.3 46.0 68.8 0.141
Jev (reference) – 100 99 74.1 – –
Laya (public leaderboard, not our run) 421M 94 73 34 – –
Julia-1 (our run) 144M 75.0 47.2 35.1 47.2 –
reflex-s1 (our run, fast profile) 23M/83M 81.2 45.8 39.6 50.2 –

ECE for vera-spark: 0.141 overall. Brier: 0.024 / 0.263 / 0.692. Median latency per question: 29-34 ms (MI300X, batch of one).

Hard-tier breakdown (correct / total):

family score family score
routing_hard 5/5 temporal_numeric 8/15
adversarial 5/6 tradeoff 4/6
trap 6/8 judge_hard 6/17
ambiguous 4/7 long_policy 6/19
multi_hop 5/18 probability 2/10

Internal selection set

set vera-spark
selection composite 75.3
held-out decisions (3.6k items, decontaminated) 75.1
judging (preference / RewardBench-style) 63.3
Open-Jev out-of-distribution 83.6
rule-following dev 86.9
code-verified rewritten cases 69.2
long documents (1.5-3k words) 72.7
typed decisions (zero-shot) 52.8

vera-spark and vera-core

The two models share the same System One interface, wire format and decision format, so you can swap one for the other.

  • vera-spark (this repo): the compact model (~630M parameters). Pick it when you want the smallest footprint.
  • vera-core: the larger model, built on Qwen3.5-2B (~2B parameters). Pick it when accuracy matters most, especially on hard, multi-step decisions.

Both are generalist bases meant to be fine-tuned for your use case.

Quickstart

pip install vera-s1

Python 3.10 or newer. The weights are downloaded once, on first use, and cached. It runs on a CUDA GPU or on CPU; the device is picked automatically.

import vera

agent = vera.load("scar-ai/vera-spark")

state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}

result = agent.predict(state, questions)
print(result["answers"]["department"]["choice"])       # billing
print(result["answers"]["department"]["probabilities"])
print(result["answers"]["churn_risk"]["noul"])         # probability the answer is yes

All questions in a call are answered together. The state can be plain text or any JSON-serialisable object. Question types: choice (2 to 255 options), score (2 to 10 ordered levels, returns an expected score and the full distribution) and noul (yes/no probability).

Weights: bf16 (default) and fp32

main holds the weights in bf16 (1.3 GB). The full-precision fp32 weights (2.6 GB) are on the fp32 branch:

agent = vera.load("scar-ai/vera-spark", revision="fp32")                                        # vera package
model = AutoModel.from_pretrained("scar-ai/vera-spark", revision="fp32", trust_remote_code=True)  # transformers

For vera-serve, download the branch first and point --model at the folder:

hf download scar-ai/vera-spark --revision fp32 --include "model.pt" --include "*.json" --local-dir vera-spark-fp32
vera-serve --model ./vera-spark-fp32 --port 8200

Self-hosting: Jev-compatible HTTP server

vera-serve exposes the model on the same POST /v1/systemone request and response shape as Jev, so existing clients work by changing their base URL:

vera-serve --model scar-ai/vera-spark --port 8200
curl -s localhost:8200/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"document": "I was charged twice. Please fix this ASAP."},
  "questions": {"billing": {"type": "noul", "instructions": "Is this ticket about billing?"}}
}'

It binds 127.0.0.1 by default (use --host 0.0.0.0 to expose it) and has no authentication, so put it behind your own gateway if you expose it. Concurrent requests are batched.

Use with transformers

The checkpoint also loads through transformers with AutoModel (custom code, so trust_remote_code=True). Tested with transformers 5.18.0 (vera-core needs a transformers release that includes Qwen3.5); outputs match the vera package exactly (max probability difference 0.0 on a three-question test).

from transformers import AutoModel

model = AutoModel.from_pretrained("scar-ai/vera-spark", trust_remote_code=True)   # add .to("cuda") if you have a GPU
state = "Hi, we were billed twice for March. Please refund the duplicate today."
questions = {"department": {"type": "choice", "instructions": "Which department should handle this?",
                            "criteria": {"billing": "invoices, payments, refunds", "other": "everything else"}}}
print(model.predict(state, questions)["answers"]["department"])    # same shape as POST /v1/systemone

For tensor-level use, model.encode_batch(state, questions) builds the inputs and model(**batch).logits returns the calibrated option scores (softmax(logits, -1) gives probabilities). The model is non-autoregressive and has no generate().

Inference engine and decision API status

vera answers decisions through the Jev-compatible POST /v1/systemone API, not through text generation. That decides which engines can serve it.

engine / API status notes
vera-serve (POST /v1/systemone, /api/v1/decide) Supported The reference decision API server; batches concurrent requests.
vera.load(...).predict(...) (Python) Supported Same result shape as the HTTP API.
transformers (AutoModel, trust_remote_code=True) Supported See above. predict() mirrors the decision API.
vLLM Not supported vera-spark's backbone has upcycled mixture-of-experts layers and a head that reads several backbone layers, which vLLM cannot run. Not tested.
SGLang Not supported Same reason as vLLM. Not tested.
Ollama (and llama.cpp / GGUF) Not supported There is no GGUF architecture for the decision head, and Ollama serves text generation only.

To serve vera-spark, use vera-serve or transformers. If you need vLLM or SGLang, use vera-core: it runs on both through an adapter (tested), with the same decision API.

Files

file description
model.pt weights (state dict, bf16), used by the vera package
config.json, model.safetensors* transformers-format config and weights (bf16) (AutoModel, trust_remote_code=True)
configuration_vera.py, modeling_vera.py, vera_model.py, vera_schema.py custom code for the transformers loader
arch.json architecture flags for the loader
tokenizer.json, tokenizer_config.json tokenizer
package/ source of the vera-s1 Python package
dist/vera_s1-0.1.2-py3-none-any.whl the installable package, also on PyPI
demo.mp4 JevPilot driving on vera-spark
Downloads last month
67
Safetensors
Model size
0.6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results