decision.host
首页 / 模型 / Respan Span-01

Respan Span-01

Respan Noul curatedopenroutersom

面向 agent 轨迹的行为分类器:给定一段对话 span 与用自然语言描述的行为,返回每个行为 present / absent / not_observable 的概率。语义与标准三原语不同。

该模型目前没有出现在 S1MB 或 Decision Index 的公开评测结果中。

模型信息

厂商
Respan
类别
托管 API
参数规模
未公开
权重
闭源 / 未公开
许可证
闭源 需自行核对
输入模态
文本
决策原语
Noul
输入价格
$0.02/M tok
可微调
不支持
OpenRouter 模型 ID
respan/span-01
OpenRouter 计费
输入 $0.02/M tok · 输出 免费
上架状态
Generally available

数据来源与链接

本站聚合自: curated、openrouter、som。 分数与链接均指向原始出处。

来自 systemonemodels.org 的详细介绍

What Span-01 is

Span-01 is Respan’s classification model for agent traces. Like a System One model, it returns probabilities in one forward pass and writes no text. You send a conversation span, meaning the earlier messages plus the turn to judge, and a list of behaviors written in plain language, such as “the user expresses frustration.” Respan says the behaviors are not a fixed list. It trained the model for general classification reasoning with RLAIF, reinforcement learning from AI feedback, then specialised it for behavior detection.

What it returns

For each behavior you get three probabilities: p_present, p_absent and p_not_observable. They sum to about 1. Not observable means the trace cannot answer the question, for example when it ends before the customer replies. Respan says to read that as unknown, not absent. Every behavior in a request is scored in one forward pass, and output tokens are free. Respan’s own API takes POST /api/v1/scores with span.input, span.output and a behaviors list of ids and definitions. Its docs do not mention the /v1/systemone schema. The three-way output does not match Choice, Score or Noul exactly, and Noul is the closest. A third-party benchmark repo calls it through OpenRouter’s System One endpoint.

What it is good at

Respan aims it at evaluation, guardrails and monitoring of LLM and agent output. Its example scores escalation, user frustration and prompt injection in tool output, then uses thresholds in code to hand off, block or alert. That fits LLM guardrails and confidence-gated actions. On Respan’s own production behavior benchmark, Span-01 scores 0.806 overall F1, against 0.716 for Jev, 0.719 for Sonnet 5 and 0.885 for GPT-6 Sol. Its best domains are privacy and secrets at 1.000 and agent and tool reliability at 0.845. These are Respan’s numbers on its own dataset. The one outside test found is not about behaviors. On a 48-question sentiment and topic pool, the zero-shot-ie-bench author measured 85.4% for Span-01 and 79.2% for Lite, against 93.8% for Jev.

Span-01 Lite

Lite is the free, lighter tier of Span-01. On Respan’s overall behavior chart it scores 0.761 F1 against 0.843 for Span-01 and 0.715 for Jev. It is the default model on Respan’s API.

What it is not for

It does not write text or explain a score, and it reads text only. It is not a general classifier over a fixed option list. Respan publishes no context length or latency, so test long traces yourself.

Access

Respan’s launch post, dated 24 September 2026, says Span-01 is public. Its docs pages still describe an early access waitlist, with 403 errors until your organization is enabled. OpenRouter has listed it since 26 September, from Respan as the only provider.

Specifications

Question types Noul Max Choice optionsNot documented Score levelsNot documented Questions per callNot documented Total contextNot documented State budgetNot documented Rate limitSpan-01 Lite has a daily cap that resets at 00:00 UTC. Respan does not publish the size of the cap or any rate limit for Span-01. EndpointPOST https://api.respan.ai/api/v1/scores SDKs Respan does not publish a context length, and OpenRouter lists 0. Respan says there is no per-request cap on the number of behaviors, and that Span-01 scores text only, with each message's content as a string.

Versions

span-01-pro, 24 Sep 2026, Span-01 on Respan's API. OpenRouter lists it as respan/span-01, snapshot span-01-20260925, at $0.02 per million input tokens. Release notes span-01-free, 24 Sep 2026, Span-01 Lite, the free lighter tier and the default model on Respan's API. OpenRouter lists it as respan/span-01-lite and respan/span-01-lite:free. On Respan's overall behavior benchmark it scores 0.761 F1 against 0.843 for Span-01. Release notes

Use cases

What people use Span-01 for, one page per pattern. Safety and qualityIntermediate

LLM guardrails with System One models

Put one Jev request in front of an LLM and one behind it. Yes/no questions return the probability that each hazard holds, a Score rates how much harm complying would do, and your thresholds turn those numbers into pass, review, block, or a crisis path. Noul Score Workflow controlIntermediate

Confidence-gated actions with System One models

Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars. Choice Score VendorRespan TypeCommercial, hosted API StatusGenerally available Announced24 Sep 2026 AccessOpened 24 Sep 2026 Input price$0.02 per 1M tokens Output priceFree LatencyNot published ContextNot published Rate limitSpan-01 Lite has a daily cap that resets at 00:00 UTC. Respan does not publish the size of the cap or any rate limit for Span-01. API docsrespan.ai

Sources

01Span-01 launch post (Respan) respan.ai 02Span-01 concept page (Respan docs) respan.ai 03Span-01 quickstart (Respan docs) respan.ai 04Score behaviors with Span-01, API reference (Respan docs) respan.ai 05Span-01 on OpenRouter openrouter.ai 06Span-01 Lite on OpenRouter openrouter.ai 07Span-01 endpoints and pricing (OpenRouter API) openrouter.ai 08zero-shot-ie-bench README (GitHub) github.com