Jev AI complete guide: install the skill and download the kit

When I saw the two numbers sitting side by side, I thought the console was glitching. A model that makes a decision in 70 to 500 milliseconds, charges 0.042 dollars per million input tokens, and ships its outputs for free. And above all, a model that cannot hallucinate, because it does not generate words. It chooses.
It is called Jev. It is the first model of a new class called System One, and it changes how you treat the mountain of records that just needs an intelligent interface. Emails, tickets, leads, requests, anomalies: anything that sits on the fence between "process it" and "ignore it". This guide gives you the mechanics, the real numbers, and a downloadable starter kit that gets you running in ten minutes.
The problem Jev solves: AI that acts instead of talks
Language models have reached a level we could not imagine five years ago. They write, reason, translate, review. But the moment you plug them into a production pipeline, nobody can say exactly where they stop.
An email arrives at three in the morning. Should it trigger a reply, get forwarded to the billing team, or land in the junk folder? That is a decision. It repeats thousands of times a day, it is trivial to state, and yet it costs you either a human who drops their current task or an LLM whose ticket climbs past a dollar once you pair it with tool calls.
The team behind Jev comes from this gap. Diogo Almeida, co-founder and CEO, spent years at OpenAI before founding TypeSafe. The diagnosis is simple: the "chat with a genius" part is solved. The "operational decisions are taken on their own" part is not.
The name System One is no accident. It points straight at Daniel Kahneman: System one is fast, intuitive, almost instantaneous thought. System two is slow, deep reflection. LLMs are excellent System-two machines: slow, verbose, expensive. Jev is a System-one machine: fast, cheap, built to return a verdict, not a monologue.
What Jev actually is
Jev is a model of the System One class: it takes an unstructured state input (an email, a table, a page, a list of events) and produces a typed, probabilistic output. In practice, three output types, which I detail one notch below.
Three architecture decisions deserve your attention.
Training aimed at decisions. TypeSafe does not talk about prompt engineering. It talks about RLCD: Reinforcement Learning for Calibrated Decisions. The reinforcement learning focuses on the quality of the final decision and on the calibration of confidence. In other words, when Jev says it is spam with 97% confidence, it is genuinely right 97 times out of 100. Confidence is not a vague feeling. It is calibrated on training data.
A parallel sampler. A classic LLM generates text sequentially, one token after another. That is what makes it slow and expensive on small tasks. Jev uses a joint sampler that produces decisions in parallel. The official figures give 70 to 500 milliseconds per decision, versus 3 to 329 seconds for large chat models. Across ten thousand typical tasks, TypeSafe reports on average 193.6 times faster and 444.6 times cheaper than the reference models used in its benchmarks.
A structural guarantee against hallucination. This is the strongest point, and the one that needs the most nuance. Because the output is constrained by a type and a finite vocabulary, the model cannot invent free text. It cannot emit a word outside the contract, because the contract blocks generation. The result is valid by construction, 100% of the time. That is type-safe: the output matches the schema, period. The nuance, which I will get to, is that a valid output can still be wrong.
The pricing that changes everything
The model is pay-as-you-go, and this is what makes it viable in production.
- Input costs 0.042 dollars per million tokens, for the most expensive model in the lineup. Some models in the family are even cheaper.
- Outputs are free. Yes, free. Because an output is a choice, not text generation.
- Real latency sits between 70 and 500 milliseconds per decision.
To make it concrete: processing a thousand spam emails costs about one cent. Handing the same classification to a classic frontier model costs several dollars. When you multiply that by a company's inbound volume, the gap stops being a detail. It is the difference between a pipeline that runs and a pipeline that stays unfunded for accounting reasons.
The three Jev output types
Jev does not write. It classifies, and it states its certainty.
Noul, the yes/no output. A binary question, a probability between zero and one. Is this spam? Is this client about to churn? Does this message demand a reply? The probability is not a marketing score. It is calibrated. Between 0.51 and 0.99, it carries real information your code can act on.
Choice, picking among options. You define the list of options, plus an "other" fallback if the list is not exhaustive. Who handles this ticket: billing, support, engineering, leadership? What category is this partnership request? Jev returns the chosen option and its probability.
Score, a position on a written scale. The scale must be spelled out, not abstract. What is the urgency level, from 1 to 5, with each level defined? What is the churn risk, from 0 to 100? Why a written scale? Because each level has a precise definition, and the model knows how to position itself between two bounds.
Look at what this changes versus a prompt that says "answer in JSON". With Jev, the schema is not a request. It is the execution contract. Your code knows exactly what structure it will receive, and how much to trust each value.
The five use cases that justify the model
Here is the list of uses that keep coming up everywhere, with the orders of magnitude observed against a frontier LLM in reasoning mode.
Spam and scam filtering. A thousand emails processed in seconds for about one cent, versus several dollars with a big model. This is the canonical use: binary decision, huge volume, exploding costs elsewhere.
Triage, churn, support, inbound intake. The classic "who handles this message and with what priority". a few cents for a volume that would make a chat-model API call expensive, and the output drops straight into your queues.
AI-slop detection. Flagging generic AI-generated content at scale, for a ridiculous cost. Tests mention orders of magnitude around 4 cents versus 15 dollars for the same task given to a frontier model.
Design system selection. The example that struck me most: choosing the best component or the best library among hundreds of options, from a description of the need. Not keyword search, real semantic judgment. Three hundred systems evaluated for about 2 dollars, where the same job with a classic LLM ran into the hundreds of dollars.
Model routing. Deciding, for each request, which model should handle it. A perfect meta-decision for a System One: a choice, fast, cheap, recurring.
The common thread: a repeated decision, a finite set of possibilities, short latency. As soon as your problem looks like that, Jev is a candidate.
Another class of results: apps you used to build with vector databases
There is a narrative that says you need RAG, a vector database, and possibly hundreds of thousands of dollars of infrastructure for an application to answer with relevance. For classification problems, that story is expensive. Take the examples demonstrated in tools built with Jev.
An emoji picker and a "Netflix finder". You describe the need in plain language: "give me a movie that feels like this", "the emoji that matches this message", and the tool indexes results on the fly, with no vector database, no RAG pipeline. Instant indexing is the point: the search is computed live, not precomputed.
An ad library scorer. Reading four hundred ads, scoring the angles, the hooks, the format, and surfacing what has real force. All for a few cents, where a frontier model cost tens of dollars per batch. For a consultant or an agency that sifts through ad libraries every week, that is a permanent research tool that costs almost nothing.
A UI component searcher. The most impressive one. It scans 255 components from the 21st.dev catalog and returns the right one for a semantic need: "a monthly/annual toggle", "a pricing card with a promo badge". Roughly 750 calls per search, one to two seconds, and a judgment that operates on meaning, not keywords. That is exactly what a sequential LLM would do for far more money and far more time.
The strategic lesson: for any decision that has a finite set of possible answers, heavy search infrastructure is unnecessary. A System One model is enough, with latency in seconds and cost in cents.
The anti-hype section: what Jev does not do
Before the kit, the limits. There are some, and knowing them avoids disappointment.
Jev is not multimodal. It handles text and structured data. An image requires a separate vision step, with another model, before you feed the transcription to Jev. Design your pipeline accordingly.
Jev does not write content. It is not a writing tool, not a chatbot, not a generator. Do not try to make it draft a follow-up email: that is not its job. Its job is to choose.
Confidence is not measured accuracy. Calibration says that when displayed confidence is high, the chance of error is low. But low confidence does not tell you which option is right: it tells you to verify. That is a golden rule I restate below.
There are tasks where Jev fails, and that is normal. On micro counting tasks, performance drops: counting the letter N in a text lands around 59% confidence, counting words around 53%. On a recent date or an order of magnitude that the context cannot pin down, the model can be confidently wrong. Why? Because those tasks require the deep reasoning the model, calibrated to decide fast, never performed.
The instruction is therefore precise: if confidence drops below 70 to 75%, do not let the flow act. Send the record to a reasoning LLM for a second check, or to a human. That is the decision staircase: Jev at the bottom, reasoning above, a human at the top.
The Jev kit: installable in ten minutes
Enough theory. Here is the ultra-actionable kit I prepared, with ready-to-use forms and the script that runs the batch. It is downloadable, free, and it costs nothing to test thanks to the dry-run.
Download the Jev kit (zip, 16 files)
What it contains:
jev-handoff/SKILL.md: the complete skill to install in your agent,jev-handoff/references/integration.md: the real endpoints and parameters,jev-handoff/references/forms.json: five ready-to-use decision forms,jev-handoff/agents/openai.yaml: an agent configuration template,jev-handoff/examples/*: synthetic email batches per decision, to test without triggering anything,jev-handoff/scripts/jev_run.py: the runner that sends the batch to the model.
Installation, step by step
Two prerequisites. An OpenRouter API key (environment variable OPENROUTER_API_KEY), because the runner goes through the OpenRouter decisions endpoint, and Python 3.10 or newer.
Copy the skill into your agent. Under Claude Code, unzip and place the jev-handoff folder in ~/.claude/skills/. From a project, you invoke it with /jev-handoff. Under Codex, place it in ~/.codex/skills/ and use $jev-handoff. The skill carries the full procedure: choosing the right form, listing batches, reading outputs.
Run the first pass in dry-run. The script refuses to overwrite an existing results folder, which protects you from double runs. To test without spending a single cent:
python3 scripts/jev_run.py --input examples/spam-emails.jsonl \
--form examples/spam-form.json \
--output results/spam-01 \
--dry-run
The dry-run shows what would be sent, the number of records and the retained form, without calling the model. Once validated, you remove --dry-run and the batch goes out: four workers in parallel, a sixty-second timeout per call, and two output files, results.jsonl with each decision and its confidence, and summary.json with the aggregate.
The five decisions in the kit
The forms cover the flows that show up in nine companies out of ten.
Spam and scam filtering (Noul). A probability that the message is spam or a scam. It runs the whole inbound flow and keeps only the strong signals to delete or flag.
Who handles this message (Choice). You define the owners: billing, support, engineering, leadership, or no action. Each option described by its responsibilities, like a real internal routing rule. Typical example: partnerships and interviews go to one owner, bugs to another, billing to a third.
Reply urgency (Score). A written five-level scale, each level defined. Critical messages surface in a few hundred milliseconds, the rest waits for the next open.
Buyer signal (Noul). Does this lead or message show a buying signal? This is the form that feeds your prospecting campaigns without blowing up qualification cost.
Sponsorship type (Choice + Score). For partnership flows: what type of sponsorship, what budget range, with written tiers, for example one, five or fifteen thousand dollars. The kind of decision whose human handling cost far exceeds the value of the message.
All criteria are readable in forms.json, and you adapt them in two minutes to your own rules, since the form is a plain JSON file.
The real kit benchmark
I measured the behavior of the kit exactly as shipped, on September 19 2026, on the same batch of twenty emails, repeated five times.
| Metric | Jev 1.13 | GPT-6 Astra |
|---|---|---|
| Time per 20-email batch | 2.60 to 2.99 s | 10.04 to 17.81 s |
| Cost per batch | 0.00223 usd | 0.70583 usd |
| Correct labels | 100/100 | 100/100 |
With equal accuracy, Jev is here three to six times faster and about three hundred times cheaper. Mind the test limits: twenty synthetic emails, a single scenario, a 1.13 model version, roughly 32,000 tokens of state and questions. Your real volume may change the numbers. But the order of magnitude matches every independent test I have read.
The review rule is built into the kit: below a confidence threshold (for example 0.75 for the owner decision), the record leaves the batch and goes to a human. Nobody lets a machine decide everything on its own.
What it looks like under real conditions
Real-condition tests are the most convincing, because they remove the demo doubt.
Bulk email processing. A batch of twenty-four emails was classified in 55 milliseconds for 0.002 dollars. Easier to count bigger: about 1,700 emails per batch run around 0.19 dollars, where a classic chat model took about seven seconds per email, meaning roughly two hours for what Jev does in seconds.
Analyzing a 27-minute video. The flow: transcription and OCR of every frame, then into Jev to extract facts and decisions. Twenty-seven minutes of video in about a minute and a half, including automatic blurring of confidential data that appears on screen. A textbook hybrid pipeline where Jev handles the decision while another model handles vision.
Inbound lead prequalification. Answers land 200 to 300 milliseconds per prospect. A web form, a few fields, and qualification happens in real time instead of piling up in a tab.
SMS scam detection. A batch processed in 541 milliseconds for 0.002 dollars. When it is your insurer, your bank or your carrier writing to you, the cost of error is not the problem. The speed and reliability of the classification are.
Game and exploration agents. Internal demos show Jev playing continuously (a game session at roughly seven dollars an hour, where the equivalent with a big reasoning model was unaffordable), solving wiki racing runs by chaining choices, and holding 255 options in memory at once. The Jevons paradox here works in your favor: when a service becomes 400 times cheaper, its consumption explodes, and that is exactly the goal.
Accessing the model. You can use it through the TypeSafe console (the playground is sometimes saturated at peak hours), through the public OpenRouter decisions endpoint, or through TypeSafe's dedicated API. The free Vercel tier is plenty for testing, which pushes the experimentation cost to almost zero.
The comparison table that tells you when to switch
To decide whether Jev belongs in your stack, ask yourself four questions.
Switch to Jev when the task is a repeated decision, the set of possible results is finite and known, volume is high, latency matters, and long reasoning is not required. Filtering, triage, scoring, routing, prequalification: everything at the top of the table.
Keep a classic LLM when you need to generate text, reason in depth, write code, compose an email, analyze an image, or produce a long report. Jev does not replace your language model. It sits next to it, on the decision layer.
The hybrid case is the real case. You route with a System One, you draft the reply with an LLM, and you only escalate to a human the cases under the confidence threshold. It is the architecture I am rolling out with clients right now, and it costs the least while keeping maximum control.
One last reflex to keep: start by encoding your rule, not by bolting a model onto a model. Describe the decision as you would explain it to an assistant: the cases, the responsibilities, the thresholds. The kit's form is exactly that rules file, and it is what drives quality, far more than the choice of model.
The Jev mistake to avoid
The first failure with these models is almost never the model's fault. It is the schema's fault. If your form allows two overlapping options, if your score levels blur together, or if your scale is not written out, Jev answers fast and confidently in the wrong direction. Type-safe guarantees the form, not the content.
Three concrete rules:
- Describe each option by its responsibilities, not by a label. A form that says "what is it about" is weaker than a form that says "who handles it and why".
- Force a fallback option. "Other" is mandatory as soon as your list is not exhaustive, otherwise the model picks the least bad option instead of telling you it does not know.
- Write the scale out fully. "From 1 to 5" with no definition is a vague number. "5: immediate reply required, client incident" is a real decision.
FAQ
How much does Jev really cost?
Input is billed at 0.042 dollars per million tokens for the most expensive models in the lineup, and outputs are free. Classifying a thousand emails costs a few cents, and a 20-email batch measured in the kit costs 0.00223 dollars. The kit's dry-run costs nothing, and Vercel offers a free tier that is enough to test.
Does Jev replace a ChatGPT or a Claude?
No. Jev does not generate text: it returns typed decisions (yes/no, choice, score) with calibrated confidence. A classic LLM writes, reasons and generates, but it is slow and expensive for repeated decisions. The typical setup is a hybrid pipeline: Jev decides, the LLM writes, a human validates edge cases.
Why cannot Jev hallucinate?
Because its output is constrained by the type contract: the vocabulary and the schema are fixed, free text generation is blocked. The result is valid by construction. However, a valid output can still be wrong on substance, hence the confidence thresholds and the review of cases under 70 to 75%.
Does Jev handle images?
No. Jev is designed for text and structured data. For an image, you need a vision step (OCR, description, transcription) run by another model, then feed the result to Jev. That hybrid pipeline is what made it possible to analyze 27 minutes of video in about a minute and a half.
Where does the model run and how do I integrate it?
The OpenRouter decisions endpoint accepts calls with the typesafe/jev model, and TypeSafe exposes a dedicated API as well as a test console. The kit shipped with this article contains the full runner, the five decision forms and the install procedure for Claude Code and Codex: download the zip, copy the folder, run the dry-run, then the batch.
A final word
Jev is not the replacement for your language models. It is a new layer in the stack, the one for fast, repeated decisions, and it makes obsolete a good chunk of the infrastructure people built to imitate them.
The most profitable advantage is right in front of you: take a flow you currently handle by hand or with a too-expensive sequential LLM, write the rule as a JSON form, run the dry-run, measure. In an hour, you will know whether Jev fits your case. The downloadable kit above gives you exactly what you need for that test.
And if you want an outside look at automating your prospecting, your support, or your lead qualification, my AI audit is free and lasts forty-five minutes. We take one of your real flows, run it through this pipeline, and you leave with a concrete implementation plan.
New AI term along the way?
Clear, jargon-free definitions for business leaders.
Jev Kit (zip, 16 files)
Skill, decision forms, runner and ready-to-use examples for Claude Code and Codex.
Looking for direct guidance? Book your Express AI Audit (45 min) →
Weekly AI Strategy Log for Leaders
Every Saturday morning, receive our curated field notes, actual case studies, and production benchmarks. Concise, actionable, zero spam.
Fractional AI Director locations


