@jaredpalmer
BlogSpeakingProjects

IntroducingKev

An open source family of decision models you can train and run on your own.

GitHubHugging Face
October 1, 2026Open weights · Apache-2.0
  1. (1)The Kev family
  2. (2)How Kev works
  3. (3)Training the new models
  4. (4)Evaluation
  5. (5)Fine-tuning on your data
  6. (6)Limitations
  7. (7)Get started

Today I'm releasing Kev 1.0, a family of four open source decision models, from 0.8B to 27B parameters. This release includes a new Kev-27B and an updated Kev-9B.

TypeSafe's Jev takes text, questions, and possible answers and returns a probability for each answer. I wanted that interface with weights I could run and fine-tune myself. Kev uses the same API, so applications built with TypeSafe's SDK can switch by changing the endpoint and model name.

All four models have Apache-2.0 weights on Hugging Face. The code is on GitHub, including a server that runs on GPUs and on Apple Silicon through MLX. You can try Kev-4B in your browser or use the agent skills to deploy and fine-tune your own.

The Kev family

Kev-27B leads overall in my evaluations. On consumer complaints, Kev-9B and Kev-4B are within a percentage point of it; on developer-tool decisions, Kev-9B matches it. The table separates tasks excluded from Kev's fine-tuning from held-out examples of task families it trained on.

Overview of accuracy across six evaluation panels for Kev-27B, Kev-9B, Kev-4B and Kev-0.8B. Kev-27B leads overall; Kev-9B and Kev-4B are within one percentage point of it on consumer complaints, and Kev-9B matches it on developer-tool decisions. All results are from the author's evaluation harness, not an independent leaderboard.

Kev family results

Accuracy across six evaluation panels, measured in Kev's own harness. The best score in each row is in bold.

Kev-27BKev-9BKev-4BKev-0.8B
Tasks excluded from Kev's fine-tuning
Generalization14 public datasets, 3,089 test questions75.7%69.8%69.0%58.6%
Short-text decisionssources excluded from fine-tuning, 656 development questions85.1%82.0%81.7%64.8%
Knowledge and evidenceexams, buried facts and missing evidence, 1,046 development questions82.0%78.0%77.4%58.6%
Held-out examples of trained task families
Policy and rule reasoningpolicies, logic, numbers and missing evidence, 1,088 test questions91.8%83.4%80.3%66.5%
Sorting consumer complaintsproduct and issue, 936 questions90.8%90.0%90.3%85.1%
Developer-tool decisionscode review, commits and tools, 1,071 questions79.0%79.1%75.6%63.7%

Test splits: generalization and the three workflow panels. Development splits: short-text decisions, knowledge and evidence. Some labels are noisy or ambiguous; these are whole-panel results, not the audited subsets in the model cards. Qwen's training exposure is unknown, and checkpoint selection was not fully blind.

I recommend Kev-27B if you want the highest overall scores and have an 80 GB H100 for shorter inputs, or an H200 or B200 for long documents. Otherwise, start with Kev-4B and measure it on your own examples.

Model Best for Runs on Req/s GPU $ / 1M Model time
Kev-27B Best forBest overall results in my evaluations Runs onH200 (or H100, B200) Req/s29 GPU $ / 1M$44.09 Model time67 ms
Kev-9B Best forWorkflow accuracy close to Kev-27B on a smaller GPU Runs onH100 (or L40S) Req/s80 GPU $ / 1M$13.80 Model time24 ms
Kev-4B Best forA starting point for local use and fine-tuning Runs onL40S, or a 32 GB Mac Req/s51 GPU $ / 1M$10.54 Model time42 ms
Kev-0.8B Best forHigh-throughput, narrowly defined decisions Runs onL4, or any Apple Silicon Mac Req/s63 GPU $ / 1M$3.54 Model time23 ms
Measured on Modal using the first GPU listed in each row. Throughput uses 64 concurrent clients asking six questions about a new short text. GPU costs per million requests are estimates from Modal's October 2026 list prices and measured throughput, assuming a busy GPU. They exclude separately billed CPU, memory, and network; lightly used endpoints cost more per request. Model time is the median server time, measured separately from throughput, excluding network latency and cold starts. The 9B measurements use its previous version, with the same architecture and serving path. The serving benchmark has the hourly prices and full setup.

For local use, Kev-0.8B and Kev-4B run on Apple Silicon through MLX. On a 32 GB M5, Kev-4B answers five questions about a short text in 721 ms, or 136 ms when the text is cached; Kev-0.8B takes 149 ms for new text. Kev-9B is expected to fit a 32 GB Mac, and Kev-27B a 96–128 GB Mac. I haven't measured those two on Macs yet.

How Kev works

Kev adds a pointer head, a layer that scores the supplied answer options, to a language model. It supports yes/no questions, multiple choice, and ordered scales. The document is processed once, and each question uses that cached computation without seeing the other questions.

Kev-27B answers three questions about a duplicate-charge ticket using a shared document prefix and separate question branches. It generates no answer tokens. Its actual probabilities are billing 99%, not escalating 74%, and calm 78%. At an illustrative 80% threshold, only billing is accepted. The chat response shown for contrast is illustrative, not a measured model comparison.

Kev-27B’s actual probabilities
  1. A generated answer.

    A chat model generates a sequence of tokens, even when its output is constrained to JSON.

  2. Kev scores options directly.

    A pointer head returns probabilities without generating answer tokens.

  3. Kev reads the text once.

    The document prefix is computed once and cached.

  4. Every question branches off that one read.

    Ask three questions or thirty: the ticket is still read once.

  5. A probability for every answer.

    Each <decide> token scores the options. A softmax turns those scores into probabilities.

  6. Your code sets the threshold.

    Here, only billing clears the illustrative 80% threshold. The other decisions take the fallback path.

state01 · support ticket

I was charged twice for order #4821. Please refund the duplicate charge today.

I was charged twice for order #4821. Please refund the duplicate charge today.

cached stateread once
illustrative chat responsetokens generated 0

Sure! Looking at this ticket, the customer says they were charged twice for the same order, so I'd send it to the billing team. It sounds fairly urgent because they want it fixed today, so escalating could make sense. Their tone seems mildly frustrated. ```json {"department": "billing", "escalate": true, "frustration": "frustrated"} ```

department99% sure → act on billing
returns
shipping
billing
<decide>◆
escalate74% sure → fallback
yes
no
<decide>◆
frustration78% sure → fallback
calm
frustrated
very angry
<decide>◆

Caching saves more work on longer documents. On an H200, Kev-27B takes about 9.4 seconds of model time to read a new 64k-token document and about 0.7 seconds for subsequent questions using its cached state.

Training the new models

What's new in Kev-27B

The smaller Kevs use Qwen3.5 backbones with LoRA adapters: small sets of additional weights trained while the backbone stays fixed. For the new Kev-27B, I used Qwen3.8-27B, Qwen's post-trained release, and fine-tuned every weight of its text backbone.

Training the whole backbone on the previous Kev-27B's data did no better than the adapter-trained version. The new release also uses a broader corpus: about 146,000 examples and 337,000 questions, with documents up to 32k tokens. It combines Kev's existing data with licensed public tasks, synthetic decisions written by open-weight models, and documents and agent logs generated by code with exact answers.

I added examples for tone, grounding, prompt injection, and personal-data detection, and screened the training records against the frozen evaluation sets for exact and near matches. No Jev outputs were used. The model card documents the data sources and training recipe, including the final weight blend with the previous checkpoint. The full training corpus is not public.

Processing each training question separately would repeat the document computation. I share that computation across the document's questions, which made training two to three times faster without changing the objective. One full run took about 16 hours on eight H200s, roughly $650 in GPU time for that run, not the whole research effort.

Compared with the previous Kev-27B, short-input accuracy is roughly unchanged. The new model is less accurate and more overconfident on long legal contracts. For that workload, compare it with jaredpalmer/kev-27b@v1-lora on your own documents.

What's new in Kev-9B

Kev-9B had fallen behind Kev-4B on policy and rule reasoning, developer tools, and consumer complaints because it hadn't received the same additional training. One more fine-tuning pass, about three hours on one GPU, closes that gap.

Kev-9B (previous) Kev-9B (new)
Policy and rule reasoning Kev-9B (previous)58% Kev-9B (new)83%
Developer-tool decisions Kev-9B (previous)64% Kev-9B (new)79%
Sorting consumer complaints Kev-9B (previous)83% Kev-9B (new)90%
Generalization index (0–100) Kev-9B (previous)40 Kev-9B (new)41

Those gains are on held-out examples of the task families it trained on. On a separate held-out short-text decision test, accuracy remains 85%, while its probability estimates improve. The previous version remains available at jaredpalmer/kev-9b@v1.

Evaluation

Generalization

To test tasks excluded from Kev's fine-tuning, I assembled 14 public datasets spanning intent classification, retrieval, language understanding, tool selection, knowledge, and subjective judgments. They include choosing among 150 customer intents and deciding whether a contract supports a claim. These results use a separate test split from the development examples that guided training.

The family table reports ordinary accuracy. Here I use a chance-corrected index: random guessing is 0, perfect performance is 100, and the five task areas count equally. Kev-27B scores 52.3 versus Kev-9B's 41.0. These are index scores, not percentages of correct answers.

Chance-corrected generalization index on 14 public datasets excluded from Kev's fine-tuning: Kev-27B 52.3, Kev-9B 41.0, Kev-4B 38.0, Kev-0.8B 23.3. Differences from Kev-27B with 95% paired bootstrap intervals: Kev-9B −11.3 (−14.2 to −7.6), Kev-4B −14.3 (−17.4 to −10.5), Kev-0.8B −29.0 (−32.3 to −25.2). Base-model exposure is unknown, and model selection was not fully blind.

Generalization beyond Kev's fine-tuning tasks

Chance-corrected index: 0 is random guessing and 100 is perfect, not percent accuracy.

Knowledge, language, retrieval, tool use and taste count equally. Every model answered the same 3,089 test questions from 14 public datasets excluded from Kev's fine-tuning. Base-model exposure is unknown; earlier candidates' test results informed selection. Difference from Kev-27B, with 95% paired bootstrap interval: Kev-9B −11.3 [−14.2, −7.6], Kev-4B −14.3 [−17.4, −10.5], Kev-0.8B −29.0 [−32.3, −25.2].

The gap is largest in language understanding and smallest in tool selection, where Kev-4B is already competitive within the family. The models differ in backbone and training recipe as well as size, so this comparison doesn't isolate the effect of size or establish an advantage over a prompted chat model.

Calibration

Calibration measures whether a model's probabilities match how often its answers are correct. On the pooled 14-dataset test, Kev-27B's answers in the roughly 85%-confidence bin are correct about 85% of the time.

Reliability diagram comparing stated confidence with observed accuracy on the 14-dataset test. Kev-27B's roughly 85%-confidence bin has about 85% accuracy. Pooled calibration error is 0.019; measuring within each dataset and then averaging gives 0.073. Calibration varies by workload.

Kev's confidence against observed accuracy

Reliability on the pooled 14-dataset test. The dotted diagonal is perfect calibration, not a guarantee on a new workload.

  • Kev-27B
  • Kev-9B
  • Kev-4B
  • Kev-0.8B

Calibration error, all 14 datasets pooled: the average gap from the diagonal, lower is better

Kev-27B
0.019
Kev-9B
0.034
Kev-4B
0.029
Kev-0.8B
0.042

Calibration error measured within each dataset, then averaged

Kev-27B
0.073
Kev-9B
0.101
Kev-4B
0.098
Kev-0.8B
0.098

Each point is a group of answers given with similar confidence (ten equal-width groups); groups under 20 answers are left out. Each model's confidence uses its shipped temperature.

Pooling can hide overconfidence on one dataset and underconfidence on another, so I report both pooled and per-dataset calibration error.

To test how many questions could clear a confidence threshold, I chose a threshold for each model targeting 5% error on development data, then applied it unchanged to test. Kev-27B accepted 54% of questions with 5.5% observed error; Kev-9B accepted 36% with 6.7% error. Kev-4B accepted 31% with 3.5% error, and Kev-0.8B 15% with 5.4% error.

Those error rates won't necessarily hold on your data. Choose a threshold using representative examples, then measure it on a separate sample. Below the threshold, your code can call a larger model, retain a default, or decline the action.

To adjust the probabilities after training, I fit a temperature, a single scaling factor applied to the option scores before they're converted to probabilities. It changes confidence without changing which answer ranks first. For the new 27B and 9B, that factor is fitted on held-out datasets excluded from fine-tuning and ships with the checkpoint.

Evaluation notes

The results come from my own harness, not an independent leaderboard. The family table identifies test and development panels. Development results guided the training recipes, and earlier candidates' test results were known before I selected the released 27B, so selection was not fully blind. Qwen's training data is not fully known either; excluding a task from Kev's fine-tuning doesn't rule out exposure in the base model. Uncertainty intervals describe variation across evaluation records, not repeated training runs.

Policy and rule labels are computed by code; developer-tool labels come from human annotations, heuristics, or construction. The complaints are real CFPB narratives with model-assisted label checks. In a 50-example spot check, I agreed with 47 labels. The model cards link the measurements, label audits, and selection history.

Fine-tuning on your data

During Kev-4B's development, one pass over 5,219 labelled consumer complaints improved accuracy from 80% to 90% on held-out complaints. On a separate example support workload, about 1,000 generated examples improved accuracy from 68% to 74%. Those are workload-specific results; a few hundred examples can produce a gain too small to distinguish from noise.

I packaged that workflow into an official kev-finetune skill. Your coding agent helps define the questions, prepares labelled data or generates examples, fine-tunes a released Kev, fits its temperature on a separate calibration split, and compares the result with the starting checkpoint before deployment.

The companion kev-deploy skill makes it easy to deploy Kev to Modal behind an authenticated HTTPS endpoint. Both skills work with Claude Code, Codex, Devin, and other coding agents. You don't need a local GPU or a clone of the Kev repository.

A small Kev-4B fine-tuning run costs about $1 in GPU time, excluding data generation and evaluation. Serving scales to zero when idle; the trade-off is a cold start, about 35 seconds in the measured setup.

Limitations

Kev only scores the options you supply. It doesn't generate explanations or retrieve missing facts. You need to supply the relevant policy or evidence and handle the possibility that none of the options is correct. Answer order can also shift probabilities slightly.

Kev-27B was trained on documents up to 32k tokens and evaluated up to 64k. It runs on a single H100 for shorter inputs, but the measured 64k-token case used more than 80 GB of memory and needs an H200 or B200. The smaller Kevs were trained on documents up to about 7,500 tokens.

The code, weights, and evaluation reports are available, but the full 27B training corpus and some evaluation data are private. The model cards document what others can reproduce.

Get started

The weights are on Hugging Face, and the code, model cards, and serving instructions are on GitHub. You can try Kev-4B in your browser before setting anything up.

Install the skills, then ask your coding agent to “deploy Kev-4B” or “fine-tune Kev on my support tickets”:

npx skills add jaredpalmer/kev@kev-deploy
npx skills add jaredpalmer/kev@kev-finetune

Thanks to TypeSafe for building Jev and an API worth being compatible with, to Archer Hume for the write-up of Jev's architecture that started this, to Qwen for the base models, to denis-pplx for AutoJev, and to the teams behind the open-weight models that wrote Kev's synthetic data.

Jared Palmer headshot

Hey! I'm Jared Palmer. I'm the VP of Engineering at Cognition, the applied AI lab behind Devin, focused on building the future of software engineering with AI.

Next
Your CI is now the bottleneck
Jared Palmer © 2026