convaiinnovations/laya-multilingual
端侧首选。400M 级 ModernBERT 编码器,Apache-2.0,有 MLX / CoreML / C++ / ONNX 多套本地运行时,M3 Max 上短决策只需 5–14 ms。代价是 S1MB 上分数偏低(13–15 分)。
S1MB Task Avg
9.3
第 80 名 · 覆盖 137/137(100%)
模型信息
- 厂商
- Convai Innovations
- 类别
- 端侧小模型
- 参数规模
- 421M(活跃 370M)
- 基座模型
- ModernBERT-large(英文)/ mmBERT-base(多语言)
- 权重
- 可下载(开源)
- 许可证
- Apache-2.0 可商用
- 输入模态
- 文本
- 决策原语
- ChoiceNoulScore
- 延迟
- 32.8–39.5 ms(官方);MLX 短决策 7–14 ms(M3 Max)
- 可微调
- 可以
- 微调方法
- CLI: laya-train --data tickets.csv --out ./ft;RLCD 循环 + 温度校准,输出 train_report.json(含 acc / ECE / Brier)
- 微调硬件
- Kaggle 2×T4 或 MPS/CPU 笔记本均可
- 训练方式
- full fine-tune
- 上架状态
- Generally available
数据来源与链接
Hugging Face 权重HF · 多语言代码仓库MLX 运行时systemonemodels.orgpypi.org官方博客clef-evals.workers-ai-mle.worke…GitHubconvaiinnovations.commultimodalart-jev-decision-inde…dev.to官方文档厂商页Decision Index 排行榜S1MB 排行榜
本站聚合自:
curated、di、s1mb、som。
分数与链接均指向原始出处。
S1MB · 137 个基准
同族条目(2)
| 型号 | 参数 | Decision Index | S1MB | Vision | 榜单 |
|---|---|---|---|---|---|
| convaiinnovations/laya-multilingual 当前 | 421M(活跃 370M) | — | 9.3 | — | s1mb |
| convaiinnovations/laya-typed-decisions | 421M(活跃 370M) | 4.4 | 13.4 | — | s1mb, di |
来自 systemonemodels.org 的详细介绍
What Laya is
Laya is a System One model from Convai Innovations, an Indian company led by CEO Nandakishor M, who publishes the code. It reads a state, which can be text, an email, a ticket or a JSON document, and answers typed questions about it: Choice, Score and Noul, the same three shapes Jev uses. Every question in a call is answered in one forward pass, and nothing is generated, so there is no output to parse.
Unlike Jev, Laya is not a hosted API. It ships as Apache 2.0 weights on Hugging Face and a Python package, pip install laya, and runs on your own GPU or CPU.
How it is built
Laya is an encoder with a decision head on top. The English checkpoint fully fine-tunes ModernBERT-large (395M parameters) and adds a head trained from scratch: two transformer layers, a scorer that reads one [MASK] token per option, and an act-or-escalate output. That comes to 421M parameters. The multilingual checkpoint uses mmBERT-base and comes to 322M. Options are defined per request, so a new question schema needs no retraining.
Training uses what Convai also calls RLCD. The reward is a strictly proper scoring rule, so reporting honest probabilities earns the most reward, and updates use REINFORCE with a group-mean baseline. TypeSafe has not published its recipe. Convai has, along with a fine-tuning notebook that takes 4 to 5 hours on Kaggle’s free pair of T4 GPUs.
A Router picks between the checkpoints by detecting the script and language before the forward pass. The English checkpoint fails on other scripts without losing confidence: on Khmer it scored 0.000 at 0.952 confidence.
What it is good at
Its strengths are speed and ownership. Convai measures 32.8ms for one question on a T4, and you can keep, inspect and retrain the weights. After fitting a temperature per question type on held-out data, its calibration error falls from 0.466 to 0.081. The package ships ready-made question sets for model routing, prompt guardrails, moderation and ticket triage.
The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, tests Laya zero-shot across a broad suite. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, Laya scores 6.04 against Jev’s 57.91, close to random guessing. That fits Convai’s own advice to specialise it on your task rather than use it zero-shot.
What it is not for
The base checkpoints are close to chance on the typed-decisions benchmark without fine-tuning: 0.362 against a 0.318 random baseline. Convai’s own README calls Laya “a fast base to specialise, not a zero-shot decision engine.” Score questions are its weakest type, and large label sets need extra configuration. Probabilities are over-confident until you fit temperatures on your own data.
Convai’s comparisons with Jev set its own numbers against Jev figures published by third parties, not measured in the same run, so treat the head-to-head as indicative.
Specifications
Question types Choice Score Noul
Max Choice optionsNot documented
Score levelsNot documented
Questions per callNot documented
Total context512 tokens
State budget320 tokens
Rate limitNone. It runs on your own hardware.
SDKsPython: laya
The English checkpoint defaults to 512 tokens, 192 of them reserved for the question and its options, leaving about 320 for the state. laya-multilingual and laya-typed-decisions default to 1,024 tokens with 256 for options. The mmBERT encoder under the multilingual checkpoint accepts up to 8,192 if you raise max_len. There is no hard cap on Choice options, but they share the option budget: 77 labels get about 3 to 4 tokens each, and Banking77 accuracy falls to 0.425. Raise head_max_len or shortlist with predict_shortlist for large label sets.
Versions
convaiinnovations/laya, 18 Sep 2026, English checkpoint at the repo root. ModernBERT-large encoder, 421M parameters, 512-token context. Release notes
convaiinnovations/laya-typed-decisions, 18 Sep 2026, The same 421M English model fine-tuned on the typed-decisions benchmark's training split, with a 1,024-token context. Release notes
convaiinnovations/laya-multilingual, 19 Sep 2026, mmBERT-base encoder, 322M parameters, 1,024-token context, for text outside English. Release notes
Use cases
What people use Laya for, one page per pattern.
Workflow controlStarter
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Choice Score Noul
Workflow controlIntermediate
Intent and model routing with System One models
One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue.
Choice Score
Safety and qualityIntermediate
LLM guardrails with System One models
Put one Jev request in front of an LLM and one behind it. Yes/no questions return the probability that each hazard holds, a Score rates how much harm complying would do, and your thresholds turn those numbers into pass, review, block, or a crisis path.
Noul Score
Workflow controlIntermediate
Confidence-gated actions with System One models
Jev returns a confidence value from 0 to 1 alongside every Choice and Score answer. Your code treats it as a separate axis: act automatically when it's high, confirm or flag when it's middling, hand the decision to a person when it's low. Riskier actions get higher bars.
Choice Score
Examples built with Laya
The most-starred and most-viewed entries in the directory. Browse all examples.