Small Jev-like decision models you can train and run yourself.
Kev is a family of small decision models built on Qwen3.5 and Qwen3.8 and based on the architecture described in Jev's Architecture Unmasked. You can use the pretrained weights or train your own. The API matches TypeSafe's System One, so you can point their Python SDK at your local server.
- Yes/no (
noul), multiple-choice (choice) and rating (score) questions in one request. The questions share the text but can't read each other. - Calibrated probabilities by default: each checkpoint ships with a temperature fitted on held-out data.
- Drop-in for Jev: the TypeSafe Python SDK works against a Kev server unchanged.
- Four sizes, from a 0.8B that runs on a laptop to a 27B for a single data-centre GPU.
- Fine-tune on your own labelled examples. A coding-agent skill runs the whole loop on Modal, from finding your questions to serving the result.
- Deploy your own HTTPS endpoint with one command. It scales to zero when idle.
- Try it in the browser first: huggingface.co/spaces/jaredpalmer/kev.
Start with Kev-4B. Move to Kev-9B if you have a bigger GPU, or to Kev-27B if you have an 80 GB GPU and want the most accurate Kev. Use Kev-0.8B when size matters more than accuracy.
| Model | Base | Accuracy: New Sources | Accuracy: Trained Sources | Brier: New Sources | Runs on | Model Card |
|---|---|---|---|---|---|---|
| Kev-0.8B | Qwen3.5-0.8B-Base | 0.648 / 0.697 | 0.827 / 0.838 | 0.481 / 0.416 | Any Apple Silicon Mac, L4 | Details |
| Kev-4B | Qwen3.5-4B-Base | 0.817 / 0.838 | 0.873 / 0.865 | 0.269 / 0.242 | 32 GB Mac, L40S, H100 | Details |
| Kev-9B | Qwen3.5-9B-Base | 0.820 / 0.852 | 0.874 / 0.873 | 0.289 / 0.217 | 32 GB Mac, L40S, H100 | Details |
| Kev-27B | Qwen3.8-27B (post-trained) | 0.851 / 0.889 | 0.865 / 0.866 | 0.225 / 0.156 | B200, H200, H100 80 GB; 96–128 GB Mac (expected) | Details |
| Jev | Hosted | 0.857 / – | 0.845 / – | 0.211 / – | TypeSafe's API | – |
Each cell is development / test. "New sources" means datasets and policy rules Kev never saw during training. It is the closest thing here to your own questions. "Trained sources" means held-out examples from the datasets Kev was trained on. We pick checkpoints using the development sets and read each test set only once per released model. Jev has only been run on the development sets. Brier scores the whole probability distribution, not just the top answer; lower is better.
On new sources Kev-27B is within a point of Jev (0.851 vs 0.857), and Kev-4B and Kev-9B are within four points. We don't know what Jev was trained on, so this isn't a controlled comparison of the two architectures. What to Expect says where Kev is as good as Jev and where it isn't.
Kev-0.8B, 4B and 9B start from Qwen base models and share one training recipe: a small adapter on a frozen base. Kev-27B starts from Qwen's post-trained release, and we don't know what that was trained on; every one of its weights is fine-tuned, so it ships as 51 GB of full weights rather than an adapter. Each model card has the full recipe, all results, and the earlier versions kept as Hub tags. The 0.8B, 4B and 9B weights are also in the GitHub release, with SHA-256 checksums; Kev-27B is too large for it and is on the Hub only.
The Hugging Face Space runs Kev-4B and Kev-0.8B, with nothing to install.
You'll need Python 3.12 or 3.13 and uv. The repo's .python-version makes uv sync use 3.13; torch has no wheels for 3.14 yet.
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009This starts Kev-4B on your machine: CUDA or ROCm if you have a GPU, MLX on Apple Silicon. The first run downloads the adapter and the base model. --run also accepts a local checkpoint directory or a Hub revision like jaredpalmer/kev-4b@qwen3.
In another terminal, send it a ticket:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
"model": "kev-latest",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"}},
"escalate": {"type": "noul", "instructions": "Does this need urgent human attention?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]}
}}'Example response from Kev-4B, running in bf16 on an Apple M5:
{
"model": "kev-latest",
"answers": {
"department": { "type": "choice", "choice": "returns", "confidence": 0.21,
"probabilities": { "returns": 0.47, "shipping": 0.28, "billing": 0.25 } },
"escalate": { "type": "noul", "noul": 0.93 },
"frustration": { "type": "score", "score": 1.44, "confidence": 0.34,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.00, "1": 0.56, "2": 0.44 } }
},
"usage": { "input_tokens": 101, "output_tokens": 161 },
"latency_ms": 495
}The ticket mentions a return, a late delivery and a billing problem, and the department probabilities say so. That's why Kev returns probabilities instead of a single label: your code can route the confident cases and send the rest to a person.
If you already call Jev, point your client at Kev and keep the rest of your code. The TypeSafe SDK is included in uv sync --extra serve:
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient(
api_key="local",
base_url="http://127.0.0.1:8009",
model="kev-latest",
)
response = client.system_one(
state="I was charged twice. Please fix this ASAP.",
questions={
"billing": Noul(instructions="Is this ticket about billing?"),
"tone": Choice(
instructions="What is the customer's tone?",
criteria={"calm": None, "frustrated": None, "angry": None},
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["can wait", "this week", "today"],
),
},
)
print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)The released models were trained on public datasets and generated policy examples. If your questions look different, like your own routing categories, your own escalation rules or another language, a short fine-tune usually helps more than any prompt change. It also fits the temperature to your data, so the confidence you set thresholds on is measured on your own labels.
What to expect: on an example support workload (three questions, 1,050 generated records, 15 minutes on an H100), fine-tuning took Kev-4B from 67.7% to 73.6% accuracy, and from automating 34% of decisions at a 5% error budget to 48% (details). On real data, one epoch on 5,219 labelled consumer-finance complaints took Kev-4B from 0.804 to 0.904 accuracy on complaints it had never seen. Gains like these are in distribution: they tell you how well Kev learns your task, not how it does on everything else. Size your dataset first. With 400 records, the gain on the example workload was inside the noise.
npx skills add jaredpalmer/kev@kev-finetuneThen ask your agent to "fine-tune Kev on my support tickets". The kev-finetune skill interviews you, finds the questions your code already asks Jev or TypeSafe, converts the labels you have or generates enough with any LLM to measure a gain, fine-tunes from a released checkpoint on Modal, fits the temperature on a held-out slice, scores the result against the untouched model, deploys an endpoint, and tears everything down at the end. You don't need a local GPU or a clone of this repo. A Kev-4B training run costs about $1 on an H100.
The skill's README is the same recipe for people: six short standard-library scripts and one Modal app. To train from this repo instead, put your examples in a JSONL file, one request per line. It's the same shape as an API request, plus a label on every question:
{"state": {"subject": "Charged twice", "body": "I see two charges for order #4411. Please refund one."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Payments and refunds", "shipping": "Delivery problems", "access": "Login and account access"}, "label": "billing"},
"angry": {"type": "noul", "instructions": "Is the customer angry?", "label": false},
"priority": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["low", "normal", "high"], "label": 1}}}For choice the label is the option name, for noul it's true or false, and for score it's the level's position starting at 0. Keep 10–20% of the file aside for evaluation.
Then start from a released checkpoint with --init_from:
uv run python -m kev.train --data train.jsonl --base Qwen/Qwen3.5-4B-Base --init_from jaredpalmer/kev-4b \
--epochs 2 --lr 2e-5 --batch 1 --accum 8 --dtype bf16 --checkpointing 1 --device cuda --out runs/mine
uv run python -m kev.benchmark --run runs/mine --data heldout.jsonl --out runs/mine-eval
uv run --extra serve python -m kev.serve --run runs/mine --port 8009--init_from loads the adapter and pointer head from the released model before training, so you keep what Kev already knows and add your domain on top. Starting from the base model instead throws that away: in one user's test on 836 support-tool decisions, a fine-tune from the base scored 0.33 on Kev's own evaluation set, against 0.84 for the released model; the same data with --init_from kept 0.83 there and reached 0.88 on the new domain. Use a smaller learning rate than the from-scratch recipe (2e-5 is a good start), and pick --base to match the checkpoint you start from; the trainer checks that the base, revision, LoRA rank and head size agree before it loads anything.
--batch 1 --accum 8 in bf16 fits the 0.8B model on a 4 GB GPU. The benchmark reports accuracy, Brier score and calibration per question type, so you can see which of your questions the fine-tune helped. The checkpoint you started from is recorded in runs/mine/training_config.json. On a Mac, run one training job at a time; two jobs on the same Apple GPU are much slower.
To get an HTTPS endpoint instead of a local server, you don't need this repo, just a Modal account:
pip install modal && modal setup
curl -LO https://raw.githubusercontent.com/jaredpalmer/kev/main/skills/kev-deploy/scripts/kev_serve.py
KEV_API_KEY=$(openssl rand -hex 24) modal deploy kev_serve.pyThat serves Kev-4B on an L40S at https://<your-workspace>--kev-api.modal.run, with the same API as above behind Authorization: Bearer <key>. It scales to zero when idle, so an unused endpoint costs nothing. The first request after idle waits about 35 seconds for a container to start. KEV_MODEL=jaredpalmer/kev-9b serves another model on the GPU that suits it; Kev-27B goes to a B200, falling back to an H200 or H100. If you use a coding agent, npx skills add jaredpalmer/kev@kev-deploy does the same and wires the URL into your code. skills/kev-deploy has the GPU and cost table.
A model you fine-tuned with the kev-finetune skill deploys the same way from its own Modal app (KEV_SERVE_SECRET=kev-serve-key KEV_SERVE_RUN=<run> modal deploy scripts/kev_modal.py; see its deploy guide). To host Kev on your own machines instead, run kev.serve from Run It Locally on a GPU box with --host 0.0.0.0 and put it behind your own proxy; Serving Performance says which GPU to pick.
Accuracy. Kev-27B is within three points of Jev, or ahead of it, on 9 of the 11 new-source categories in the chart below. Kev-4B and Kev-9B are about as close on classification-shaped sources like routing, entailment and science questions. Knowledge questions depend mostly on the base model: on MMLU Kev-9B scores 0.73 and Kev-27B matches Jev at 0.90, but on the harder MMLU-Pro Kev-27B scores 0.675 against Jev's 0.840. The smaller models also trail on day-precision date arithmetic.
Confidence. Each checkpoint ships with a fitted temperature, so its probabilities are calibrated by default. As served, Kev-9B puts at least 0.9 probability on a wrong answer for 2.4% of new-source questions, against Jev's 3.7%. Jev still ranks its answers better: at a 5% error budget, Kev-4B, 9B and 27B can automate 0.52–0.69 of new-source decisions, Kev-0.8B 0.14 and Jev 0.70. Check a threshold on your own data before you rely on it.
Speed. Kev-4B answers six questions about a new short text in 18.1 ms of model time on an H100 and 41.5 ms on an L40S, and a container serves around 101 requests per second on an H100. On an Apple M5, Kev-4B takes 721 ms for five questions, or 136 ms when the text repeats and comes from the cache. Serving Performance has every GPU and batch size.
Length. Kev-0.8B, 4B and 9B trained mostly on states of up to 384 tokens, with longer ones in their document and skill fine-tunes (up to 7,552 tokens for Kev-9B), Kev-27B on states of up to 32,768. The server accepts states of up to 65,536 tokens, and 8,192 more for each question. Longer inputs work, but the smaller models lose accuracy on long documents. On real contracts of up to 64k tokens (CUAD) Kev-27B scores 0.874, and its confidence there is less reliable than on short text; its model card has the numbers by length.
With the server running, open another terminal. You'll need Node 20.9+:
cd playground
npm install
npm run dev -- -p 3001Open localhost:3001, load a preset, and edit the text and questions. Press ⌘↵ to run it. "Packed vs separate" compares asking all questions at once with asking them one at a time. "Permute" runs a Choice question with six option orders. There are also presets for testing question isolation and fake delimiter tokens.
There's a chess demo, too. The board is the input, legal moves are Choice options, and a Score question rates the position. You can play against Kev or let it play itself. Games are saved in localStorage.
state is the text to evaluate. Each question has instructions and, where needed, a set of answers to choose from.
| Type | Criteria | Answer |
|---|---|---|
noul |
Optional descriptions for true and false |
noul: probability of yes |
choice |
1–255 option names, each with a description or null |
choice: most likely option; probabilities and confidence |
score |
1–255 descriptions, ordered from lowest to highest | score: mean level index, starting at 0; legend, probabilities, and confidence |
For Choice with K > 1 options, confidence is (p_max − 1/K) / (1 − 1/K). A single option has confidence 1. Score confidence is max(0, 1 − E|level − mode| / D): mode is the most likely level and D is the mean distance of a uniform distribution over the levels from its middle (2/3 for three levels), so all probability on one level gives 1 and a uniform or wider spread gives 0. Both formulas are the ones in TypeSafe's reference adapter (system-one-adapter 0.2.1). Neither field is a measured accuracy rate.
Objects and arrays are converted to labeled text. Delimiter-like strings in user input are escaped before tokenization. Invalid requests return 422, and so does a state longer than 65,536 tokens: the server never drops part of a document silently, and the error gives the state's token count and the limit. usage.output_tokens counts tokens in the serialized answers, not generated tokens.
| Method | Path | Purpose |
|---|---|---|
GET |
/v1/models |
Model cards (name, description, release_date) plus the loaded checkpoint's details |
POST |
/v1/systemone/permute |
Run one Choice question with different option orders (n_perm 1 to 64, default 6) |
POST |
/v1/systemone/separate |
Run each question in its own forward pass |
A request may carry any number of questions. The server runs them a token budget at a time (one maximal row of 16,384 tokens per forward pass, counting the cached document once per question in that pass), so memory does not grow with the question count and the answers do not depend on the split. Every response carries an x-typesafe-request-id header. The server binds to 127.0.0.1 (--host 0.0.0.0 to accept other machines) and is open by default; set KEV_API_KEY to require Authorization: Bearer <key> on /v1/*, as the TypeSafe clients always send it.
| Variable | Effect |
|---|---|
KEV_TEMPERATURE=1.0 |
Return raw probabilities instead of the calibrated ones |
KEV_DATE_FACTS=1 |
Append the number of days between any two dates in the state (see Benchmarks) |
KEV_TRUNCATE_STATES=1 |
Read the first 65,536 tokens of a longer state instead of refusing it; every response then has truncated and usage.state_tokens / state_tokens_used |
KEV_DTYPE=fp32 |
Serve the exact fp32 path the evaluations use (bf16 is the default on GPUs) |
KEV_API_KEY |
Require a bearer key |
Each checkpoint is a rank-16 LoRA adapter and a small pointer head on a Qwen base model. On an attention-only base (Qwen3), the state and questions go into one token sequence:
<state> …state…
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <decide>
<q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <decide>
The attention mask lets a token read the state and its own question, but not other questions or future tokens. Each question's position IDs restart just after the state. This lets the model process the state once and answer each question independently.
Qwen3.5 and Qwen3.8 mix attention layers with Gated DeltaNet layers, which are recurrent and ignore attention masks. For those models, which is every current Kev, each question runs as its own row: the state followed by that question, with the same positions as above. The rows are independent, so isolation is exact, and the server and DecisionModel.probs() compute the state once and reuse its cache for every row. forward(), which kev.benchmark scores and every published number comes from, keeps the plain rows and runs the state once per question; the two agree to fp32 rounding. On attention-only models the rows and the mask above give identical probabilities (tests/test_model.py).
Kev-27B uses the same design on Qwen/Qwen3.8-27B, with two differences. Its base is Qwen's post-trained release rather than a -Base checkpoint, and we don't know what it was post-trained on. And every backbone weight is trained, not just an adapter, and kept in bf16, so the checkpoint is the whole model: 51 GB of bf16 weights plus the pointer head. It serves in bf16 only (about 66 GB resident with the serving buffers), which is why it needs an 80 GB card. On Apple Silicon the MLX backend loads those weights as they are, with no merge (see Serving Performance); we expect that to fit a 96–128 GB Mac but haven't measured it. Its served probabilities stay within 0.022 of the evaluation path on an H200 (runs/serving-27b-r23).
The pointer head scores each option's </opt> hidden state against the question's <decide> hidden state. A softmax turns those scores into probabilities. Because <decide> comes last, it can attend to the full option list.
Training uses cross-entropy on the correct answer. The adapter and head are trained together; the rest of the base weights stay fixed (Kev-27B trains them all). Training examples and API requests use the same text format. No Jev outputs were used for training.
Asking questions together or separately produces probabilities within 4e-6 in the fp32 tests. This does not mean option order is irrelevant: options within a question can still affect one another. See the model code and parity tests.
The released models share one base training set, decision-v7: 10,000 examples from ten public datasets, 896 generated policy examples, and 1,680 examples from 60 generated rule structures. Kev-0.8B, 4B and 9B train on it for two epochs with LoRA rank 16 and cross-entropy. The learning rate is 1e-4 for 0.8B and 5e-5 for 4B and 9B. On these hybrid bases the adapter covers the attention, MLP and DeltaNet projections; kev.train picks the right targets from the model config.
Kev-0.8B, 4B and 9B then get short follow-up fine-tunes from their released checkpoints, through the same --init_from path you'd use for your own data: generated cases that state day counts or have the deciding evidence removed (all three), then real documents and generated skill data (all three; Kev-9B since v2, 2026-09-30). Kev-27B is trained differently. Every weight of the base is fine-tuned for one epoch on eight H200s (--full_ft 1, learning rate 2e-6) on a 145,840-record corpus: Kev's own data, the document, skill and developer-tooling suites, public datasets, licensed task families and generated long-document, tool-routing, agent-log and guardrail records, with states of up to 32,768 tokens. The result is then averaged with the earlier adapter-trained Kev-27B, 0.85 to 0.15. The model cards list every stage with its data and cost.
# sanity run, ~1 minute
uv run python -m kev.train --n_per_source 40 --accum 4 --out runs/smoke
# the first stage of Kev-0.8B (~20 min on one H100; the Mac path works but is slow for Qwen3.5 bases)
uv run python -m kev.train --suite evals/v7/decision-v7 --base Qwen/Qwen3.5-0.8B-Base --base_revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68 \
--epochs 2 --lr 1e-4 --batch 8 --dtype bf16 --p_none_pair 0.25 --device cuda --out runs/kev-0.8b
# the first stage of Kev-4B (one H100 via Modal, ~1 h; see below). Swap in Qwen/Qwen3-4B-Base for the previous generation.
uv run python -m kev.train --suite evals/v7/decision-v7 --base Qwen/Qwen3.5-4B-Base --base_revision 1001bb4d826a52d1f399e183466143f4da7b741b \
--epochs 2 --lr 5e-5 --batch 4 --accum 2 --dtype bf16 --checkpointing 1 --p_none_pair 0.25 --device cuda --out runs/kev-4bUse uv run python -m kev.train --help for all training options. The released models don't use the optional --perm_kl or --ord_w losses. PLAN.md records what was tried, what helped, and what didn't.
Each trial gets its own H100. The study keeps running if you disconnect, and you can download the results when it finishes:
uv run modal token new # once; opens the browser
KEV_GPU=T4 uv run modal run modal_app.py::smoke # end-to-end check, ~1 minute of GPU
uv run modal deploy modal_app.py # once; studies run on the deployed app and survive disconnects
uv run modal run modal_app.py::study \
--suite evals/v7/decision-v7 --plan experiments/v7-final.json \
--name my-study --transfer evals/v4/transfer-v4 --budget 30 --timeout 7200
uv run modal run modal_app.py::pull --name my-study # results -> runs/my-study, rankedStudy plans list training settings. Each trial saves the settings, code hashes, dataset hashes, and results. Choose models using the development results, not the locked test. After choosing a final candidate, you can read its test results once:
uv run modal run modal_app.py::locked_test --trial my-study/00-trial-0 --name my-candidate # one read, everThe evaluation data under evals/ is frozen: dataset versions and file checksums are recorded in each manifest. Large files are downloaded from the Hub mirror and checked against those hashes. Every model in the tables above is scored on the same items. The numbers in this README and the model cards are checked in CI against the committed reports they come from (docs/claims.json, uv run python scripts/verify_claims.py).
| Suite | What it measures |
|---|---|
decision-v7 |
Held-out examples from the ten training datasets, the generated policies and the rule structures ("trained sources") |
transfer-v4 |
764 records from datasets and policy and rule types Kev never trained on: QNLI, SciQ, PAWS, MMLU, Emotion, TweetEval, held-out policies and rules ("new sources") |
transfer-v9 |
transfer-v4 plus 10-way MMLU-Pro, records buried in unrelated text, and "unknowable" records whose deciding evidence was removed |
uv run python -m kev.benchmark --run jaredpalmer/kev-4b --suite evals/v4/transfer-v4 --out runs/my-eval # new sources
uv run python -m kev.benchmark --run jaredpalmer/kev-4b --suite evals/v9/transfer-v9 --out runs/my-eval-v9 # + MMLU-Pro, buried states, unknowable items
uv run python -m kev.benchmark --run jaredpalmer/kev-4b --suite evals/v7/decision-v7 --out runs/my-eval-id # trained sources
uv run python -m kev.benchmark --remote http://127.0.0.1:8009 --suite evals/v4/transfer-v4 --out runs/my-remote # any System One endpoint, Jev includedThese commands use development data. Test data requires --allow-test. The benchmark reports accuracy, Brier score, calibration error, the share of decisions you could automate at a 5% error budget, option-order changes, and question isolation. On the unknowable records it reports how often the model still answers with at least 0.9 confidence (Kev-9B 0%, Jev 9%). Published accuracy numbers use fp32 evaluation, not the bf16 serving path. kev.jev runs the same questions against Jev through Vercel AI Gateway, and kev.compare compares two saved runs with paired bootstrap confidence intervals.
Calibration. Each checkpoint stores a temperature, and the pointer head applies it when the model is loaded. Kev-4B (2.41) and Kev-0.8B (2.35) fitted theirs on their in-distribution development sets; Kev-27B (1.32) and Kev-9B (2.19) fitted theirs on held-out datasets they never trained on. It never changes which answer wins. On new sources it takes Kev-9B's calibration error from 0.103 to 0.041 and its confident errors (wrong answers with probability ≥ 0.9) from 8.2% to 2.4%, below Jev's 3.7%. The accuracy numbers above are the same either way; the Brier numbers are for the raw probabilities. scripts/calibrate_checkpoint.py also reports an out-of-fold estimate, so the in-sample fit can be checked against records it didn't see.
Dates. Kev can't subtract dates reliably, but it can use a day count it's given. KEV_DATE_FACTS=1 appends one sentence per pair of dates in the state ("June 26, 2026 is 8 days before July 4, 2026"). On the deadline policy questions this takes Kev-9B from 0.80 to 0.90 (Jev 0.93). None of the tables use it.
Other people's test sets. evals/external/ holds test sets from other projects, converted to this format, with their published live Jev results. Some were scored on earlier versions of the Kev weights, which the Kev column names. Three were removed because they cannot serve as a gate, and the model cards keep the numbers their releases were decided on: scienthoon's synthetic support tickets on 2026-09-27 (templated text; one of its three questions depends on a rule the text does not state), and on 2026-09-30 WANLI (wanli-v1, wanli-v2: a quarter of the pairs are ones WANLI's two annotators labelled differently, with the gold set to one of them) and TypeSafe's public evals (typesafe-v1: the gold is the averaged answer of two closed frontier models, and on 89 questions it cannot tell checkpoints apart).
| Suite | What it is | Jev | Kev |
|---|---|---|---|
| SemIf | 144 authored decisions | 0.965 | 0.917 (Kev-9B at v7-base) |
SemIf's labels hold up, but it is close to saturated: every Kev-27B checkpoint answers 130 of the 144 correctly, so it is a sanity check, not a way to rank models.
Pick the GPU by the model:
| Model | GPU ($/h) | 6 questions, short text | 5 questions, 2,200-token text | Requests/s, 64 clients |
|---|---|---|---|---|
| Kev-0.8B | L4 (0.80) | 22.7 / 16.1 ms | 108.6 / 32.3 ms | 62.8 |
| Kev-4B | L40S (1.95) | 41.5 / 27.7 ms | 145.2 / 43.0 ms | 51.4 |
| Kev-4B | H100 (3.95) | 18.1 / 12.9 ms | 89.4 / 22.5 ms | 100.8 |
| Kev-9B | L40S (1.95) | 66.4 / 42.7 ms | 235.6 / 57.5 ms | 32.7 |
| Kev-9B | H100 (3.95) | 24.0 / 16.6 ms | 88.5 / 26.4 ms | 79.5 |
| Kev-27B | B200 (6.25) | 46.5 / 32.2 ms | 178.0 / 52.1 ms | 44.2 |
| Kev-27B | H200 (4.54) | 67.2 / 50.0 ms | 274.8 / 73.8 ms | 28.6 |
| Kev-27B | H100 (3.95) | 75.0 / 52.0 ms | 277.5 / 79.3 ms | 28.9 |
Times are model time per request (the latency_ms the API returns), median of 20, for a new text / the same text again. The server caches the text, so asking more questions about a document you've already sent only pays for the questions. Requests per second are for 64 concurrent clients sending six questions about a new short text each; the server batches them. Kev-27B's B200 and H100 rows were measured on its previous version, the same architecture served in bf16 (runs/fused-27b-*); the H200 row is the current checkpoint (runs/serving-27b-r23). Network time is extra: about 65 ms per round trip through a Modal web endpoint in the same region.
An L4 is enough for Kev-0.8B but too slow for Kev-4B. The A100 is slower than the L40S here and costs more. Kev-9B needs about 17 GB of GPU memory and Kev-27B 51 GB of weights (about 66 GB with the batching buffers); under load Kev-27B is compute-bound, and a B200, H200 or H100 costs about the same per request. On CUDA, install flash-linear-attention for the Qwen3.5 models (kev_serve.py and the Modal images already do).
On Apple Silicon, uv sync --extra serve installs MLX and the server uses it automatically. Five questions about a ~270-token text on an M5 (32 GB):
| Model | New text | Same text again |
|---|---|---|
| Kev-0.8B | 149 ms | 28 ms |
| Kev-4B | 721 ms | 136 ms |
Adapter checkpoints are folded into the base as they load, which briefly holds a second copy of the weights. Full-weight checkpoints like Kev-27B load as saved, with nothing merged, so loading needs only the weights. We checked this on Kev-4B written out as full bf16 weights: loading peaked at 8.4 GB for 8.4 GB of weights, against 15.9 GB for the adapter path. Its answers matched the adapter path exactly once both hold the same bf16 values, and stayed within 0.015 of the fp32 path on 60 questions (runs/mlx-full-4b). Kev-27B's weights are 51 GB. By the same measurements it needs about 51 GB plus working memory, so a 64 GB Mac is borderline and a 96–128 GB Mac should fit. We haven't run it on a Mac that large yet. Kev-27B's first version, an adapter, did run this way on a 128 GB M5 Max, matching the published accuracy (thanks to Sean Connelly, #175).
The server runs in bf16 on GPUs and Macs. Its probabilities differ from the fp32 path the published evaluations use by at most about 0.03 on a GPU and 0.05 on a Mac, and the top answer changes on about one question in 300. Set KEV_DTYPE=fp32 for the exact path. /v1/models reports the backend and precision in use. uv run modal run modal_app.py::serving --run jaredpalmer/kev-4b --gpu L40S --name <name> measures a row of the table on your own account (the rows above: runs/serve-*, runs/grouping-4b-h100, runs/fused-27b-*, runs/serving-27b-r23).
- Calibration is a single temperature. It can't reorder confidences, so the share of new-source decisions you can automate at a 5% error budget (0.52–0.69 for Kev-4B, 9B and 27B) is still below Jev's 0.70. Test a probability threshold on your own data before you rely on it.
- Knowledge questions are set by the base model. MMLU is 0.73 for Kev-9B against Jev's 0.90, and MMLU-Pro 0.59 against 0.84.
- Fine-tuning can make the base model worse at individual tasks. Date arithmetic was the clearest case (issue #8); training on stated day counts plus
KEV_DATE_FACTS=1recovers it. - Changing option order can change an answer. Question isolation doesn't prevent this.
- Kev-0.8B, 4B and 9B trained mostly on at most 384 state tokens and 1,024 tokens for the state plus one question (their document and skill fine-tunes on longer states; up to 7,552 tokens for Kev-9B), Kev-27B on states of up to 32,768 tokens. Serving allows a 65,536-token state; longer context wasn't covered by training.
- On a Mac, answers take hundreds of milliseconds, not tens. Kev-27B needs an 80 GB GPU. On a Mac it needs about 51 GB plus working memory; we expect a 96–128 GB Mac to fit it but haven't measured one.
- Kev-27B starts from a post-trained model whose training data we don't know.
uv run --extra serve python -m pytest tests/test_unit.py tests/test_research.py tests/test_generators.py tests/test_conventions.py \
tests/test_documents_tools.py tests/test_hard_v1.py tests/test_devtools_v1.py tests/test_breadth_v1.py tests/test_rounds.py tests/test_skill_scripts.py -q # no weights, no server; what CI runs
KEV_BASE_URL=http://127.0.0.1:8009 uv run --extra serve python -m pytest tests/test_api.py -q # against a running server
cd playground && npm run lint && npx next typegen && npx tsc --noEmit -p .The API tests run TypeSafe's example requests and the official SDK against your local server. PLAN.md is the research plan: what we have learned, the rules every experiment follows, and one line per round. The full log (every experiment, the criteria set before it ran, and how it came out) is at the git tag research-archive-2026-09-24.
Previous generation (Qwen3) and the prototype
The first Kev family used Qwen3 bases with the same data and settings. Those weights stay published and run on plain PyTorch on a Mac, but they are no longer developed.
| Model | Base | Accuracy: Trained Sources | Accuracy: New Sources | Brier: New Sources | Model Card |
|---|---|---|---|---|---|
Kev-0.6B (Qwen3) — jaredpalmer/kev-0.6b |
Qwen3-0.6B-Base | 0.801 / 0.808 | 0.620 / 0.642 | 0.536 / 0.483 | Details |
Kev-4B (Qwen3) — jaredpalmer/kev-4b@qwen3 |
Qwen3-4B-Base | 0.854 / 0.856 | 0.790 / 0.806 | 0.328 / 0.294 | Details |
Kev-8B (Qwen3) — jaredpalmer/kev-8b |
Qwen3-8B-Base | 0.863 / 0.870 | 0.796 / 0.780 | 0.337 / 0.327 | Details |
The original Kev-0.5B used Qwen2.5-0.5B and is kept for reference; see its model card.
Troubleshooting
- If MPS runs out of memory during training, check that you're running only one job. Don't enable
output_hidden_statesor add tokens with peft'strainable_token_indices; both have caused memory problems here. - If the playground loads but buttons don't work, use
localhost:3001. Next.js checks development hostnames. Other hosts need an entry inallowedDevOriginsinplayground/next.config.ts. - If dataset loading reports
Dataset scripts are no longer supported, uselegacy-datasets/banking77. This repo already uses it.
- Jared Palmer (@jaredpalmer)
Built with Devin. Thanks to Archer Hume for the architecture write-up, TypeSafe for the API design, Qwen for the base models, 3x3xX3N0N for showing where the date-arithmetic failure really is, and Radexito for --init_from.
Related work: Hydragen, DeFT, FIRST.
Apache-2.0. The Qwen3, Qwen3.5 and Qwen3.8 base models are also Apache-2.0. Training datasets have their own licenses; see the model cards.


{ "state": "…", // string | object | array — the content to evaluate "model": "kev-latest", "questions": { "<id>": { // you choose the id; the model never sees it "type": "noul" | "choice" | "score", "instructions": "…", // string | object | array, optional "criteria": … // noul: {true?, false?} choice: {option: description|null} score: [level, …] } } }