# Introducing Kev

> An open source family of decision models you can train and run on your own.

- Author: Jared Palmer
- Published: 2026-10-01
- Canonical: https://jaredpalmer.com/blog/introducing-kev

Today I'm releasing [Kev 1.0](https://github.com/jaredpalmer/kev/releases/tag/kev-1.0), a family of four open source decision models, from 0.8B to 27B parameters. This release includes a new [**Kev-27B**](https://huggingface.co/jaredpalmer/kev-27b) and an updated [**Kev-9B**](https://huggingface.co/jaredpalmer/kev-9b).

TypeSafe's [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) takes text, questions, and possible answers and returns a probability for each answer. I wanted that interface with weights I could run and fine-tune myself. Kev uses the same API, so applications built with TypeSafe's SDK can switch by changing the endpoint and model name.

All four models have [Apache-2.0 weights on Hugging Face](https://huggingface.co/collections/jaredpalmer/kev-6aad9d0ea49f2589665e07cd). The [code is on GitHub](https://github.com/jaredpalmer/kev), including a server that runs on GPUs and on Apple Silicon through [MLX](https://github.com/ml-explore/mlx). You can [try Kev-4B in your browser](https://huggingface.co/spaces/jaredpalmer/kev) or use the [agent skills](#get-started) to deploy and fine-tune your own.

## The Kev family

Kev-27B leads overall in my evaluations. On consumer complaints, Kev-9B and Kev-4B are within a percentage point of it; on developer-tool decisions, Kev-9B matches it. The table separates tasks excluded from Kev's fine-tuning from held-out examples of task families it trained on.

*Chart: Overview of accuracy across six evaluation panels for Kev-27B, Kev-9B, Kev-4B and Kev-0.8B. Kev-27B leads overall; Kev-9B and Kev-4B are within one percentage point of it on consumer complaints, and Kev-9B matches it on developer-tool decisions. All results are from the author's evaluation harness, not an independent leaderboard.*

I recommend Kev-27B if you want the highest overall scores and have an 80 GB H100 for shorter inputs, or an H200 or B200 for [long documents](#limitations). Otherwise, start with Kev-4B and measure it on your own examples.

| Model | Best for | Runs on | Req/s | GPU $ / 1M | Model time |
|---|---|---|---|---|---|
| **[Kev-27B](https://huggingface.co/jaredpalmer/kev-27b)** | Best overall results in my evaluations | H200 (or H100, B200) | 29 | $44.09 | 67 ms |
| **[Kev-9B](https://huggingface.co/jaredpalmer/kev-9b)** | Workflow accuracy close to Kev-27B on a smaller GPU | H100 (or L40S) | 80 | $13.80 | 24 ms |
| [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) | A starting point for local use and fine-tuning | L40S, or a 32 GB Mac | 51 | $10.54 | 42 ms |
| [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) | High-throughput, narrowly defined decisions | L4, or any Apple Silicon Mac | 63 | **$3.54** | 23 ms |

*Measured on Modal using the first GPU listed in each row. Throughput uses 64 concurrent clients asking six questions about a new short text. GPU costs per million requests are estimates from Modal's October 2026 list prices and measured throughput, assuming a busy GPU. They exclude separately billed CPU, memory, and network; lightly used endpoints cost more per request. Model time is the median server time, measured separately from throughput, excluding network latency and cold starts. The 9B measurements use its previous version, with the same architecture and serving path. The [serving benchmark](https://github.com/jaredpalmer/kev#serving-performance) has the hourly prices and full setup.*

For local use, Kev-0.8B and Kev-4B run on Apple Silicon through MLX. On a 32 GB M5, Kev-4B answers five questions about a short text in 721 ms, or 136 ms when the text is cached; Kev-0.8B takes 149 ms for new text. Kev-9B is expected to fit a 32 GB Mac, and Kev-27B a 96–128 GB Mac. I haven't measured those two on Macs yet.

## How Kev works

Kev adds a pointer head, a layer that scores the supplied answer options, to a language model. It supports yes/no questions, multiple choice, and ordered scales. The document is processed once, and each question uses that cached computation without seeing the other questions.

*Chart: Kev-27B answers three questions about a duplicate-charge ticket using a shared document prefix and separate question branches. It generates no answer tokens. Its actual probabilities are billing 99%, not escalating 74%, and calm 78%. At an illustrative 80% threshold, only billing is accepted. The chat response shown for contrast is illustrative, not a measured model comparison.*

Caching saves more work on longer documents. On an H200, Kev-27B takes about 9.4 seconds of model time to read a new 64k-token document and about 0.7 seconds for subsequent questions using its cached state.

## Training the new models

### What's new in Kev-27B

The smaller Kevs use Qwen3.5 backbones with LoRA adapters: small sets of additional weights trained while the backbone stays fixed. For the new Kev-27B, I used Qwen3.8-27B, Qwen's post-trained release, and fine-tuned every weight of its text backbone.

Training the whole backbone on the previous Kev-27B's data did no better than the adapter-trained version. The new release also uses a broader corpus: about 146,000 examples and 337,000 questions, with documents up to 32k tokens. It combines Kev's existing data with licensed public tasks, synthetic decisions written by open-weight models, and documents and agent logs generated by code with exact answers.

I added examples for tone, grounding, prompt injection, and personal-data detection, and screened the training records against the frozen evaluation sets for exact and near matches. No Jev outputs were used. The [model card](https://huggingface.co/jaredpalmer/kev-27b) documents the data sources and training recipe, including the final weight blend with the previous checkpoint. The full training corpus is not public.

Processing each training question separately would repeat the document computation. I share that computation across the document's questions, which made training two to three times faster without changing the objective. One full run took about 16 hours on eight H200s, roughly $650 in GPU time for that run, not the whole research effort.

Compared with the previous Kev-27B, short-input accuracy is roughly unchanged. The new model is less accurate and more overconfident on long legal contracts. For that workload, compare it with `jaredpalmer/kev-27b@v1-lora` on your own documents.

### What's new in Kev-9B

Kev-9B had fallen behind Kev-4B on policy and rule reasoning, developer tools, and consumer complaints because it hadn't received the same additional training. One more fine-tuning pass, about three hours on one GPU, closes that gap.

| | Kev-9B (previous) | **Kev-9B (new)** |
|---|---|---|
| Policy and rule reasoning | 58% | **83%** |
| Developer-tool decisions | 64% | **79%** |
| Sorting consumer complaints | 83% | **90%** |
| Generalization index (0–100) | 40 | 41 |

Those gains are on held-out examples of the task families it trained on. On a separate held-out short-text decision test, accuracy remains 85%, while its probability estimates improve. The previous version remains available at `jaredpalmer/kev-9b@v1`.

## Evaluation

### Generalization

To test tasks excluded from Kev's fine-tuning, I assembled 14 public datasets spanning intent classification, retrieval, language understanding, tool selection, knowledge, and subjective judgments. They include choosing among 150 customer intents and deciding whether a contract supports a claim. These results use a separate test split from the development examples that guided training.

The family table reports ordinary accuracy. Here I use a chance-corrected index: random guessing is 0, perfect performance is 100, and the five task areas count equally. Kev-27B scores 52.3 versus Kev-9B's 41.0. These are index scores, not percentages of correct answers.

*Chart: Chance-corrected generalization index on 14 public datasets excluded from Kev's fine-tuning: Kev-27B 52.3, Kev-9B 41.0, Kev-4B 38.0, Kev-0.8B 23.3. Differences from Kev-27B with 95% paired bootstrap intervals: Kev-9B −11.3 (−14.2 to −7.6), Kev-4B −14.3 (−17.4 to −10.5), Kev-0.8B −29.0 (−32.3 to −25.2). Base-model exposure is unknown, and model selection was not fully blind.*

The gap is largest in language understanding and smallest in tool selection, where Kev-4B is already competitive within the family. The models differ in backbone and training recipe as well as size, so this comparison doesn't isolate the effect of size or establish an advantage over a prompted chat model.

### Calibration

Calibration measures whether a model's probabilities match how often its answers are correct. On the pooled 14-dataset test, Kev-27B's answers in the roughly 85%-confidence bin are correct about 85% of the time.

*Chart: Reliability diagram comparing stated confidence with observed accuracy on the 14-dataset test. Kev-27B's roughly 85%-confidence bin has about 85% accuracy. Pooled calibration error is 0.019; measuring within each dataset and then averaging gives 0.073. Calibration varies by workload.*

Pooling can hide overconfidence on one dataset and underconfidence on another, so I report both pooled and per-dataset calibration error.

To test how many questions could clear a confidence threshold, I chose a threshold for each model targeting 5% error on development data, then applied it unchanged to test. Kev-27B accepted 54% of questions with 5.5% observed error; Kev-9B accepted 36% with 6.7% error. Kev-4B accepted 31% with 3.5% error, and Kev-0.8B 15% with 5.4% error.

Those error rates won't necessarily hold on your data. Choose a threshold using representative examples, then measure it on a separate sample. Below the threshold, your code can call a larger model, retain a default, or decline the action.

To adjust the probabilities after training, I fit a temperature, a single scaling factor applied to the option scores before they're converted to probabilities. It changes confidence without changing which answer ranks first. For the new 27B and 9B, that factor is fitted on held-out datasets excluded from fine-tuning and ships with the checkpoint.

### Evaluation notes

The results come from my own harness, not an independent leaderboard. The family table identifies test and development panels. Development results guided the training recipes, and earlier candidates' test results were known before I selected the released 27B, so selection was not fully blind. Qwen's training data is not fully known either; excluding a task from Kev's fine-tuning doesn't rule out exposure in the base model. Uncertainty intervals describe variation across evaluation records, not repeated training runs.

Policy and rule labels are computed by code; developer-tool labels come from human annotations, heuristics, or construction. The complaints are real CFPB narratives with model-assisted label checks. In a 50-example spot check, I agreed with 47 labels. The [model cards](https://github.com/jaredpalmer/kev/tree/main/docs/model-cards) link the measurements, label audits, and selection history.

## Fine-tuning on your data

During Kev-4B's development, one pass over 5,219 labelled consumer complaints improved accuracy from 80% to 90% on held-out complaints. On a separate example support workload, about 1,000 generated examples improved accuracy from 68% to 74%. Those are workload-specific results; a few hundred examples can produce a gain too small to distinguish from noise.

I packaged that workflow into an official [`kev-finetune`](https://github.com/jaredpalmer/kev/tree/main/skills/kev-finetune) skill. Your coding agent helps define the questions, prepares labelled data or generates examples, fine-tunes a released Kev, fits its temperature on a separate calibration split, and compares the result with the starting checkpoint before deployment.

The companion [`kev-deploy`](https://github.com/jaredpalmer/kev/tree/main/skills/kev-deploy) skill makes it easy to deploy Kev to [Modal](https://modal.com) behind an authenticated HTTPS endpoint. Both [skills](https://skills.sh) work with Claude Code, Codex, Devin, and other coding agents. You don't need a local GPU or a clone of the Kev repository.

A small Kev-4B fine-tuning run costs about $1 in GPU time, excluding data generation and evaluation. Serving scales to zero when idle; the trade-off is a cold start, about 35 seconds in the measured setup.

## Limitations

Kev only scores the options you supply. It doesn't generate explanations or retrieve missing facts. You need to supply the relevant policy or evidence and handle the possibility that none of the options is correct. Answer order can also shift probabilities slightly.

Kev-27B was trained on documents up to 32k tokens and evaluated up to 64k. It runs on a single H100 for shorter inputs, but the measured 64k-token case used more than 80 GB of memory and needs an H200 or B200. The smaller Kevs were trained on documents up to about 7,500 tokens.

The code, weights, and evaluation reports are available, but the full 27B training corpus and some evaluation data are private. The model cards document what others can reproduce.

## Get started

The weights are on [Hugging Face](https://huggingface.co/collections/jaredpalmer/kev-6aad9d0ea49f2589665e07cd), and the code, model cards, and serving instructions are on [GitHub](https://github.com/jaredpalmer/kev). You can [try Kev-4B in your browser](https://huggingface.co/spaces/jaredpalmer/kev) before setting anything up.

Install the skills, then ask your coding agent to “deploy Kev-4B” or “fine-tune Kev on my support tickets”:

```bash
npx skills add jaredpalmer/kev@kev-deploy
npx skills add jaredpalmer/kev@kev-finetune
```

Thanks to TypeSafe for building Jev and an API worth being compatible with, to Archer Hume for [the write-up of Jev's architecture](https://archerhume.com/posts/jevs-architecture-unmasked) that started this, to Qwen for the base models, to [denis-pplx](https://github.com/denis-pplx/autojev) for AutoJev, and to the teams behind the open-weight models that wrote Kev's synthetic data.
