Jev:它是什么以及如何使用

来自 TypeSafe 的首个 System One 模型实用指南 · 供 @aizvestia · 2026年9月

🤖 给编码代理的模块(复制粘贴)
Install the TypeSafe skill. If you're in Claude Code, run `claude plugin marketplace add typesafe-ai/skills`, then `claude plugin install typesafe@typesafe-ai`. If you're in another agent, run `npx skills add typesafe-ai/skills --skill typesafe-ai` and select your agent. Use one installation method. You can read the skill directly at https://github.com/typesafe-ai/skills/blob/main/skills/typesafe-ai/SKILL.md (raw: https://raw.githubusercontent.com/typesafe-ai/skills/main/skills/typesafe-ai/SKILL.md). Then use the TypeSafe skill when working on this project.

它是什么

Jev 是 TypeSafe 推出的首个公开 System One 模型。它不是聊天模型,也不是「又一个带 JSON 模式的 LLM」:Jev 完全不生成文本。你发送一个 state(上下文:文本、JSON 或数组)和一组类型化问题,得到的是机器可直接使用的决策:选中的选项、量表上的分数或「是」的概率 - 外加每个答案的概率分布和 confidence。创始人用一句话概括:「smart if-statements」- 给代码用的智能 if-else。 最直观的例子是邮件:Jev 浏览每封邮件并判定「垃圾邮件,94% 确定」「需要回复」「可以等」。它不写任何回信,只做分拣:1700 封邮件只要 18 美分。

System One 这个名字来自卡尼曼的《思考,快与慢》:模型为快速直觉判断而设计,而不是慢速推理。有趣的细节:TypeSafe 联合创始人 Diogo Almeida 是 RLHF 的共同发明人 - 正是训练 ChatGPT 的那个方法。现在他押注自动化需要另一种训练方法。

它是什么

与「以前的东西」有何不同

语言模型的后训练经历了三种方法:

   
RLHF
ChatGPT 和所有聊天机器人
模型输出人们喜欢的回答谄媚、自信的幻觉、mode dropping - 答案分布变窄
RLVR
reasoning 推理模型
擅长数学、国际象棋、长链推理又慢又贵
RLCD
Reinforcement Learning for Calibrated Decisions - TypeSafe 的方法
带校准概率的决策不会写文本,System 2 任务弱

校准的意义:当模型给出 0.8,长期来看这类答案约有 80% 是对的。概率变成了可以用来搭建逻辑的数字。

与被「强迫」输出 JSON 的 LLM 的实际区别:

三个原语

三个原语
    
Choice选哪个?choice + probabilities + confidence工单路由:billing / technical / sales
Score量表上哪一级?score(可以是小数)+ probabilities + confidence客户沮丧度:0 = 平静,2 = 暴怒
Noul这是真的吗?noul:「是」的概率,0 到 1「客户在要求退款吗?」 - 0.95

能省下几个小时的经验:

5 分钟上手

0. 免排队:Jev 已上线 Vercel AI Gateway,模型名 typesafe-ai/jev($0.042/百万输入 token)- 在 Vercel 上可直接通过 ai 包调用,不用等 TypeSafe 的 waitlist。OpenRouter 上也有:~typesafe/jev-latest(兼容 OpenAI API)。

1. Playground。打开 console.typesafe.ai/playground,粘贴任意文本作为 state,添加问题,点 Run。(进入时有个小测验「能和 Jev 聊天吗?」 - 正确答案是 No。)

2. 直接调 API。在控制台拿 key,只有一个端点:

curl -X POST https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Trying to connect my Stripe for 3 days, keeps failing. Help ASAP.",
    "model": "jev-latest",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle this",
        "criteria": {
          "billing": "Payment or subscription issues",
          "technical": "Bugs or integration problems"
        }
      },
      "is_urgent": { "type": "noul", "instructions": "The message conveys urgency" }
    }
  }'

响应里每个问题都有选中的值、所有选项的概率和 confidence。

3. Python SDK(3.10+),客户端自动读取 TYPESAFE_API_KEY 并默认调用 jev-latest:

pip install typesafe-sdk
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()
response = client.system_one(
    state=ticket,
    questions={
        "department": Choice(
            instructions="Which team should handle this",
            criteria={
                "billing": "Payment or subscription issues",
                "technical": "Bugs or integration problems",
            },
        ),
        "is_urgent": Noul(instructions="The message conveys urgency"),
    },
)
print(response.answers["department"].choice)   # "billing"
print(response.answers["is_urgent"].noul)      # 0.999

JS 有 @typesafe-ai/sdk;给代理(Claude Code、Codex 等)有现成 skill:npx skills add typesafe-ai/skills --skill typesafe-ai。这个 skill 专门教代理把问题打包批量发 - 代理最喜欢一问一请求。

四个核心模式

  1. Speculative fan-out(投机扇出)。共享同一 state 的所有问题放进一个请求,包括偶尔才用得上的「投机」问题。多一个问题成本忽略不计,响应时间几乎不变。TypeSafe cookbook:13 个问题一次请求比 13 次单独调用便宜约 12 倍、快约 10 倍,答案不变。
  2. Confidence-gated routing(置信度路由)。答案告诉你「是什么」,confidence 告诉你「要不要行动」。三个区间:高 - 自动执行;中 - 确认或标记人工复核;低 - 上交人类或 reasoning 模型。阈值取决于风险:显示错屏幕 0.6 就行,执行支付最好 0.9+。
  3. Composite scoring(复合打分)。把复杂判断拆成原子 Score,在代码里加权合成。工单优先级 = severity * 0.5 + frustration * 0.3 + actionability * 0.2。优先级变了?改系数,不用重写 prompt。
  4. Intent routing(意图路由)。对进来的请求分类并路由:确定性代码、便宜模型、贵的 reasoning 模型,或者人。
四个核心模式

价格与限制(jev-1.13.0)

弱点(来自官方文档的诚实清单)

jev-1.13 有一份公开的「jaggedness」(毛边)页面 - 诚实地列出模型做不到的事:

适用场景

RAG rerankLLM guardrailscitation checkjailbreak detectticket routingmoderationcomplianceML featurescatalog matchingresume scoringlead scoring

一个活生生的例子:基于 Jev 的比特币 BUY/SELL/HOLD 终端 - jev-terminal.vercel.app。API 返回诚实的概率,决策在不到一秒内到达。

另一个例子 - 即时生成式 UI:实验 json-render + jev - Jev 从你的组件、操作和设计系统中选择,界面毫秒级渲染。

链接

在这里直接试 Jev

编辑这条客户消息,勾选问题,点 Run - 由真正的 jev-latest 回答。

2 · 问题与答案空间
Which team should handle this support message
billingPayment or subscription issues
technicalBugs or integration problems
salesPricing or account questions
How frustrated the customer appears
0Calm, just stating facts
1Frustrated but civil
2Very angry, strong language
The message conveys urgency or time-sensitivity
yesprobability that the statement is true

请求经服务器代理:state 最多 2000 字符,约每分钟 30 次。密钥存在 Vercel 环境变量,不会进浏览器。