gliner2-large-v1
340M 编码器,既做类型化分类、也抽 span 与关系,并支持跨决策规则约束。这类「传统信息抽取模型转向类型化决策」的代表。
S1MB Task Avg
7.4
第 86 名 · 覆盖 137/137(100%)
模型信息
- 厂商
- Fastino Labs
- 类别
- 开源权重
- 参数规模
- 340M(活跃 355M)
- 基座模型
- fastino/gliner2-large-v1
- 权重
- 可下载(开源)
- 许可证
- Apache-2.0 可商用
- 输入模态
- 文本
- 决策原语
- ChoiceNoulScore
- 延迟
- 38-167 ms
- 可微调
- 未确认
- 训练方式
- full fine-tune
- 上架状态
- Generally available
数据来源与链接
systemonemodels.orgHugging Face 权重GitHubfastino.aimultimodalart-jev-decision-inde…X 帖子官方文档厂商页Decision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、s1mb、som。
分数与链接均指向原始出处。
S1MB · 137 个基准
同族条目(4)
| 型号 | 参数 | Decision Index | S1MB | Vision | 榜单 |
|---|---|---|---|---|---|
| gliner2-large-v1 当前 | 340M(活跃 355M) | — | 7.4 | — | s1mb |
| GLiNER2.5-Decide | 486M | 9.2 | — | — | di |
| gliner2.5-base-v1 | 194M(活跃 95M) | — | 6.4 | — | s1mb |
| gliner2.5-multi-v1 | 287M(活跃 95M) | — | 3.3 | — | s1mb |
来自 systemonemodels.org 的详细介绍
What GLiNER2.5-Decide is
GLiNER2.5-Decide is Fastino Labs’ System One model: an encoder-based decision model rather than a generative one. It evaluates a set of user-defined typed questions and rules against a piece of text in one pass, then jointly decodes the answers, returning probability distributions and confidence scores. Beyond typed Choice, Score and Noul questions, Fastino describes it returning character-level spans, relations, and structured records, and applying implications, exclusions, cardinality limits, and ordinal bounds across related decisions in one call.
The model card is direct about scope: “This release is not a general-purpose model. It does not reason, explain, or answer open questions.”
How it is built
The Hugging Face model card lists GLiNER2.5-Decide as a 340M-parameter model fine-tuned from Fastino’s own gliner2-large-v1, a DeBERTa-v3-large encoder. The Hugging Face API’s safetensors metadata for the published weights reports 486,444,053 parameters, which does not match the 340M figure. Neither source explains the gap.
It loads through the gliner2 Python package (AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")) and accepts label sets at runtime, so a new task needs no retraining or prompt template.
What it’s good at, per Fastino’s own benchmark
Fastino evaluated the model on “Fast Decisions,” an internal suite of 17 datasets covering intent routing, triage, sentiment and content understanding. This is Fastino’s own benchmark, not an independently run one. The blog post and company X post give GLiNER2.5-Decide 60.1% average accuracy, ahead of JevK5 (57.5%), SemIf (56.4%) and Laya (46.6%), leading on 9 of 17 datasets. The Hugging Face card’s own results table gives slightly different figures for the same run: 60.2% and JevK5 at 57.6%. Fastino highlights support intent classification at 75.3%, 18.6 points ahead of the next model, and banking intent at 64.3%, 8.6 points ahead.
The GitHub repository behind it, fastino-ai/GLiNER2, had 2,169 stars and was last pushed on 24 September 2026, the day of release.
What an independent benchmark shows
The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, is a broader test than Fastino’s. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, GLiNER2.5-Decide scores 11.21 against Jev’s 57.91. That is the best result among the models under 500M parameters it tested, but far behind the larger open models. JevK5, which Fastino’s suite ranks below GLiNER2.5-Decide, scores 38.81 here. Its expected calibration error, the average gap between stated confidence and actual accuracy, is 0.088 against Jev’s 0.074.
Access and running it
Weights are Apache 2.0 on Hugging Face with no waitlist. Fastino says the model is “efficient enough to run locally on consumer-grade CPUs or in air-gapped environments,” and co-founder George Maloney gives CPU latency around 167ms and GPU latency of 38 to 47ms for short requests. Fastino also runs hosted inference and fine-tuning at http://agent.fastino.ai, for calling from inside a coding agent; no pricing for that service appeared in the sources checked.
What it’s not for
It does not generate text, reason through multi-step problems, or answer open-ended questions; the model card says so outright. No context-window limit, maximum Choice count, or per-call question cap is documented anywhere checked, so those limits are unknown rather than unlimited.
Specifications
Question types Choice Score Noul
Max Choice optionsNot documented
Score levelsNot documented
Questions per callNot documented
Total contextNot documented
State budgetNot documented
Rate limitNone for self-hosted use. Fastino's hosted inference and fine-tuning API at http://agent.fastino.ai did not have published rate limits in the sources checked.
Endpointhttp://agent.fastino.ai
SDKsPython: gliner2
Neither the blog post, the Hugging Face model card, nor the GitHub repo README documented a fixed context window, a maximum number of Choice options, or a cap on questions per call as of this check.
Versions
fastino/GLiNER2.5-Decide, 24 Sep 2026, 340M-parameter encoder fine-tuned from fastino/gliner2-large-v1, a DeBERTa-v3-large base per the Hugging Face model card. The Hugging Face API's safetensors metadata reports 486,444,053 total parameters for the published weights, which does not match the 340M figure stated in the blog post and model card; this discrepancy is unresolved as of this check. Apache 2.0 license. Release notes
Use cases
What people use GLiNER2.5-Decide for, one page per pattern.
Workflow controlStarter
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Choice Score Noul
Workflow controlIntermediate
Intent and model routing with System One models
One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue.
Choice Score
Real-time and agentsIntermediate
Agent routing and skill selection with System One models
An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess.
Choice Noul
Data and operationsStarter