Jev 1.13
第一个 System One 决策模型,三原语的定义者。文本输入,返回类型化概率,不生成文本。
S1MB Task Avg
59.6
第 4 名 · 覆盖 137/137(100%)
Decision Index Full
60.1
第 3 名 · public 58.0
校准 ECE
0.074
越低越好 · Brier 0.356
模型信息
- 厂商
- TypeSafe AI
- 类别
- 托管 API
- 参数规模
- 未公开
- 权重
- 闭源 / 未公开
- 许可证
- 闭源 需自行核对
- 输入模态
- 文本
- 决策原语
- ChoiceNoulScore
- 上下文
- 32,000 token
- 输入价格
- $0.042/M tok
- 延迟
- 70–500 ms(厂商口径)
- 可微调
- 不支持
- 微调方法
- 官方明确:不对客户数据做微调或 LoRA,所有账号共用同一权重
- 训练方式
- hosted API
- OpenRouter 模型 ID
typesafe/jev-1.13- OpenRouter 计费
- 输入 $0.042/M tok · 输出 免费
- 上架状态
- Generally available
数据来源与链接
官网官方文档发布公告OpenRoutersystemonemodels.orgpypi.orgnpmjs.comGitHubyoutube.comX 帖子官方博客clef-evals.workers-ai-mle.worke…厂商页console.typesafe.aitechcrunch.comhpcwire.comDecision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、openrouter、s1mb、som。
分数与链接均指向原始出处。
Decision Index 领域得分
知识与推理
71.6
语言理解
82.5
检索与分类
66.8
工具与自动化
94.6
艺术与人类品味
57.0
S1MB · 137 个基准
Decision Index · 53 个基准
| 基准 | 类型 | 该模型 | 全场最佳 | 全场均值 | 对比 | 来源 |
|---|---|---|---|---|---|---|
| MiniWoB++ | 混合(Choice | 100.0 | 100.0 | 100.0 | 来源 ↗ | |
| ARC-Easy | 混合(Choice | 99.3 | 99.5 | 87.3 | 来源 ↗ | |
| ARC-Challenge | 混合(Choice | 97.8 | 98.2 | 79.4 | 来源 ↗ | |
| Email spam | 混合(Choice | 97.4 | 97.4 | 97.4 | 来源 ↗ | |
| BFCL | 混合(Choice | 95.8 | 98.8 | 80.5 | 来源 ↗ | |
| Phishing gradient | 混合(Choice | 94.6 | 94.6 | 94.6 | 来源 ↗ | |
| HellaSwag | 混合(Choice | 94.5 | 98.5 | 73.0 | 来源 ↗ | |
| OpenBookQA | 混合(Choice | 94.0 | 94.0 | 94.0 | 来源 ↗ | |
| BBH | 混合(Choice | 92.9 | 92.9 | 58.3 | 来源 ↗ | |
| WinoGrande | 混合(Choice | 92.0 | 97.5 | 69.8 | 来源 ↗ | |
| MMLU | 混合(Choice | 91.7 | 91.9 | 64.4 | 来源 ↗ | |
| BPoMP | 混合(Choice | 90.6 | 97.0 | 73.8 | 来源 ↗ | |
| Support tickets | 混合(Choice | 90.5 | 90.5 | 90.5 | 来源 ↗ | |
| CLINC150 | 混合(Choice | 89.3 | 97.4 | 67.8 | 来源 ↗ | |
| API-Bank | 混合(Choice | 88.2 | 93.1 | 56.3 | 来源 ↗ | |
| CommonsenseQA | 混合(Choice | 87.5 | 87.5 | 87.5 | 来源 ↗ | |
| FinEntity | 混合(Choice | 87.0 | 97.1 | 73.8 | 来源 ↗ | |
| NLI4CT | 混合(Choice | 84.1 | 86.2 | 69.3 | 来源 ↗ | |
| MMLU-Pro | 混合(Choice | 82.7 | 82.7 | 42.3 | 来源 ↗ | |
| When2Call | 混合(Choice | 81.0 | 91.7 | 57.8 | 来源 ↗ | |
| RouterBench | 混合(Choice | 79.9 | 80.1 | 73.2 | 来源 ↗ | |
| BANKING77 | 混合(Choice | 79.7 | 94.1 | 68.7 | 来源 ↗ | |
| GPQA Diamond | 混合(Choice | 78.3 | 78.3 | 37.7 | 来源 ↗ | |
| RAGTruth | 混合(Choice | 76.5 | 86.0 | 52.1 | 来源 ↗ | |
| ANLI | 混合(Choice | 74.8 | 98.6 | 54.0 | 来源 ↗ | |
| CRUXEval | 混合(Choice | 73.0 | 87.9 | 50.7 | 来源 ↗ | |
| HoVer | 混合(Choice | 72.8 | 89.4 | 63.8 | 来源 ↗ | |
| CLadder | 混合(Choice | 72.6 | 97.7 | 62.5 | 来源 ↗ | |
| GSM8K | 混合(Choice | 71.8 | 83.5 | 36.0 | 来源 ↗ | |
| ContractNLI | 混合(Choice | 71.7 | 86.5 | 59.9 | 来源 ↗ | |
| New Yorker | 混合(Choice | 70.1 | 82.0 | 52.9 | 来源 ↗ | |
| MuSR | 混合(Choice | 66.1 | 86.2 | 56.0 | 来源 ↗ | |
| ToolRet | 混合(Choice | 65.3 | 69.1 | 55.1 | 来源 ↗ | |
| cfcolor | 混合(Choice | 64.7 | 70.2 | 57.9 | 来源 ↗ | |
| VAST | 混合(Choice | 64.6 | 82.0 | 50.1 | 来源 ↗ | |
| PhishNChips | 混合(Choice | 62.5 | 99.9 | 62.1 | 来源 ↗ | |
| Humicroedit | 混合(Choice | 61.9 | 75.1 | 57.2 | 来源 ↗ | |
| Amazon ESCI | 混合(Choice | 55.2 | 61.6 | 40.9 | 来源 ↗ | |
| Home appliances | 混合(Choice | 52.3 | 98.9 | 23.2 | 来源 ↗ | |
| iSarcasmEval | 混合(Choice | 50.5 | 71.0 | 38.4 | 来源 ↗ | |
| BRIGHT | 混合(Choice | 47.5 | 50.9 | 36.1 | 来源 ↗ | |
| Habermas | 混合(Choice | 45.9 | 71.8 | 42.2 | 来源 ↗ | |
| SGD | 混合(Choice | 43.0 | 73.4 | 47.4 | 来源 ↗ | |
| ACOS | 混合(Choice | 29.5 | 57.3 | 14.6 | 来源 ↗ | |
| SATA-Bench | 混合(Choice | 26.4 | 37.2 | 20.3 | 来源 ↗ | |
| HLE | 混合(Choice | 20.4 | 20.4 | 12.3 | 来源 ↗ | |
| POP909 | 混合(Choice | 18.1 | 74.6 | 12.6 | 来源 ↗ | |
| ChessBench | 混合(Choice | 17.2 | 40.2 | 14.2 | 来源 ↗ | |
| RTFM | 混合(Choice | 17.0 | 17.0 | 17.0 | 来源 ↗ | |
| ScienceWorld | 混合(Choice | 6.7 | 6.7 | 6.7 | 来源 ↗ | |
| Boxoban | 混合(Choice | 0.0 | 0.0 | 0.0 | 来源 ↗ | |
| Hanabi | 混合(Choice | 0.0 | 0.0 | 0.0 | 来源 ↗ | |
| Codenames | 混合(Choice | 0.0 | 0.0 | 0.0 | 来源 ↗ |
来自 systemonemodels.org 的详细介绍
What Jev is
Jev is TypeSafe AI’s System One model, a classification model that reads a block of text and answers typed questions about it. You supply the state, which is the text you want judged, and the questions, which you define in your own code. It answers each one and stops. It does not write a reply, produce code, or explain itself.
Questions come in three shapes, Choice, Score and Noul, each walked through in Choice, Score and Noul. Several questions in one request are evaluated against the same state in parallel.
TypeSafe says it trained the model with Reinforcement Learning for Calibrated Decisions, so the probabilities are meant to match real outcome rates rather than match what a human rater prefers to read. No paper, dataset or training recipe has been published, which RLCD explained goes through. TechCrunch reported on 18 September 2026 that “Almeida says Jev is trained exclusively on synthetic data using a technique he calls ‘reinforcement learning from calibrated decisions.’” That is the only public statement about the training data, and TechCrunch renders the acronym as “from calibrated decisions” where TypeSafe’s own docs say “for calibrated decisions.”
Jev guides and comparisons
New to Jev: What is Jev AI? is the plain-English explainer.
Getting started: how to get Jev access and using Jev with Hermes.
Cost and speed: Jev speed and pricing and Jev benchmarks.
Internals: how Jev is built and RLCD explained.
What people build: Jev game demos.
Other options: Jev alternatives, plus Jev vs Clef, Jev vs pplx-decider, Jev vs Claude, Jev vs GPT, Jev vs BERT and Jev vs GLiNER.
Jev 1.13: the current version and its aliases
jev-1.13.0 is the only published version. Two aliases point at it, jev-latest and jev-preview, and the docs say why: “jev-preview currently points to the same model as jev-latest. There is no preview build available right now.”
There is no model changelog. TypeSafe’s llms.txt lists SDK changelogs only, and no version above 1.13.0 appeared in the docs, the blog, the press or search when this page was checked on 20 September 2026. Pinning jev-1.13.0 rather than an alias keeps the published weakness list matched to the model you are actually calling.
What Jev 1.13 does badly
TypeSafe publishes a jaggedness page per version. The one for 1.13 was last reviewed on 17 September 2026 and lists nine failure modes.
Literal reading: the model “answers the question you wrote, not the one you meant.”
Math and numbers: counting, numeric representations, and arithmetic attempted through a Score are all weak. The page ships a runnable Python sample that counts by asking one Noul per item instead.
Date and time comparison: Jev “reads dates as text, not as ordered quantities.”
Indirection: “instructions carrying double negatives or complex indirection are answered less reliably.”
Large state full of irrelevant detail: accuracy falls as the state grows with content unrelated to the decision. TypeSafe’s own name for the cause is context rot.
Adversarial content: “state is data, and jev-1.13 does not treat it as hostile by default.” Screening user-supplied text is your job, which is what the prompt injection screen recipe is for.
Contradictory instructions and criteria: answers degrade when the instructions and the criteria pull in different directions.
Common-sense structural invariants: the page works two numbers through, an idea that scores 0.22 as a Noul but comes back at 0.99 when asked as a Choice, and a Noul plus its negation summing to 1.19 instead of 1. TypeSafe’s line is that “there are many reasons that P(noul) and 1 - P(not noul) may not be directly comparable.”
Generation: “jev-1.13 is not trained to generate text.”
The page closes by telling you to avoid System Two tasks, which it describes as work with “more layers of indirection.”
How to call Jev
One POST to https://api.typesafe.ai/v1/systemone with an Authorization: Bearer header. Both SDKs read the key from TYPESAFE_API_KEY when you construct the default client.
The quickstart’s curl example, unchanged:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"urgency": {
"type": "noul",
"instructions": "Does this message express urgency?"
}
}
}
EOF
The Python SDK, typesafe-sdk, installs with uv add typesafe-sdk or pip install typesafe-sdk and needs Python 3.10 or newer.
from typesafe_sdk import Choice, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state={"document": "I was charged twice. Please fix this ASAP."},
questions={
"category": Choice(
instructions="What is this ticket about?",
criteria={"billing": None, "technical": None, "other": None},
),
},
)
print(response.choices["category"].choice)
Check that last line against the SDK you installed. TypeSafe’s quickstart page reads answers back as response.answers["department"].choice while the Python SDK’s own page uses response.choices["category"].choice. Both are TypeSafe’s text, and the two disagree.
Python is on 0.7.0, released 18 September 2026. That release swapped serialisation from msgspec to pydantic, which the release notes flag as a breaking change, and gave system_one() a response_model argument that takes a Pydantic model.
The JavaScript SDK, @typesafe-ai/sdk, installs with npm install @typesafe-ai/sdk and needs Node.js 20 or newer. It was still on 0.6.0 on 20 September 2026, with no 0.7.0 release published.
import { choice, TypeSafeClient } from "@typesafe-ai/sdk";
const client = new TypeSafeClient();
const response = await client.systemOne({
state: { document: "I was charged twice. Please fix this ASAP." },
questions: {
category: choice("What is this ticket about?", {
billing: n
Jev on OpenRouter
OpenRouter has listed Jev since 18 September 2026 as typesafe/jev-1.13, plus ~typesafe/jev-latest, an alias that points at the newest version. TypeSafe is the only provider. The price matches TypeSafe’s: $0.042 per million input tokens, with free output.
It is not a chat model there. OpenRouter serves it through its Decisions API, POST https://openrouter.ai/api/alpha/decisions, and warns that chat completions SDKs will not work with it. The request takes the same state and questions as TypeSafe’s API, and an OpenRouter key replaces the TypeSafe key. OpenRouter’s own SDK calls it through openrouter.alpha.decisions.create(). The alpha path means OpenRouter still treats the endpoint as unfinished. The OpenRouter integration page has more.
Languages and data handling
English first. The docs say “English is the primary training language and where accuracy is currently best. Other languages, including CJK scripts, are handled but not equally well.” They add that you should test on your own content before putting a non-English workload through the model. Input is text only, as a string, a JSON object, or an array of text values.
On data, TypeSafe states that “Jev is not trained on customer requests or responses” and that the model is “not fine-tuned or LoRA-adapted with customer data,” because “the same weights serve every account.” Zero data retention is offered to enterprise customers under a data processing agreement. None of that has been independently audited.
Access today: open sign-up, console and API key
There is no waitlist. TypeSafe removed it on 20 September 2026, five days after launch, posting on X that “Jev is now available to everyone. No waitlist.” Sign in at console.typesafe.ai with Google or an email code, create a key, and call the API. Getting access and making your first call covers the steps.
The weights are not released, so there is no self-hosted option and no local build.
What it is not for
Anything that ends in prose. Drafting, summarising and code generation sit outside what the model was trained to do, and TypeSafe says so plainly.
Arithmetic, counting and date comparison belong in your code rather than in a question. So does anything that needs two facts chained together, because indirection is one of the nine documented weak spots.
Unscreened user input is the other thing you handle yourself. The model reads the state as data and has no defence against instructions hidden in it, so a separate screen sits in front. The vendor speed and cost claims are unpacked in Jev speed and pricing, and what is known about the internals is in Jev architecture.
Specifications
Question types Choice Score Noul
Max Choice options255
Score levelsUp to 10
Questions per callNot documented
Total context64,000 tokens
State budget32,000 tokens
Rate limit100,000 tokens/sec and 40 requests/sec as of 2026-09-30, adjusting dynamically
EndpointPOST https://api.typesafe.ai/v1/systemone
SDKsPython: typesafe-sdkTypeScript: @typesafe-ai/sdk
64k tokens across the whole request. The state plus the longest single question must fit in 32k of that. OpenRouter lists Jev with a 32,000-token context, which is that second limit, not the whole request. A Score needs at least 2 levels. The maximum number of questions in one call is not documented. On rate limits the docs warn that the published numbers can change without notice while TypeSafe serves launch demand, and that higher limits come with custom and enterprise plans.
Versions
jev-1.13.0, 15 Sep 2026, The only published version. The aliases jev-latest and jev-preview both resolve to it, and the docs say there is no preview build available right now. TypeSafe ships a per-version jaggedness page listing what this version does badly. Release notes
Use cases
What people use Jev for, one page per pattern.
Workflow controlStarter