The family
Laya is an open-weight decision model published by Convai Innovations under the Apache 2.0 licence. It answers typed questions about a text (a choice between options, a score on an ordered scale, a yes or no) with a probability per option, in a single pass of an encoder, and it generates no text at all. That makes it a different kind of model from the rest of the catalogue: there is no prompt to write, no reply to parse, and no output tokens to pay for.
One id fronts it here: laya. Behind that id sit two
published checkpoints, one for English and Latin-script text and one
for the rest of its more than 100 languages, and the model picks
between them per request by the text's script and language.
What we serve
| model | size | input | quantization | weights served |
|---|---|---|---|---|
| laya | 421M + 322M (two encoders) | text | unquantized | convaiinnovations/laya |
- Size and shape
- Two encoder checkpoints from the one published repository: a 421M-parameter encoder for Latin-script text with a 512-token window, and a 322M-parameter multilingual encoder with an 8,192-token window. Small enough that a whole request is answered in tens of milliseconds, and cheap enough that the batch rate is the lowest on the price list.
- Input
- Text only: a string, or a JSON object the model reads as text, plus the typed questions asked of it. No images, audio or video.
- Output
-
One answer per question. A
choiceanswer names the most probable option and carries a probability per option; ascoreanswer is the expected level on your scale with a probability per level; anoul(yes or no) answer is the probability that the answer is yes. Every answer carries a confidence. See decision jobs for the exact request and answer shapes. - What it is served from
- The published weights, unquantized, at a pinned revision. Both checkpoints are resident at once, so routing between them costs nothing per request.
Benchmark scores
Published figures, each from the source named beside it: the publisher's own runs, with both checkpoints answering the same questions. "Routed" is the configuration we serve, the model picking the checkpoint per request. Accuracy is a fraction, so 0.783 is 78.3% of questions answered with the expected option.
| benchmark | score (routed) | source |
|---|---|---|
| MASSIVE intent classification, English (20 options) | 0.783 | Laya benchmarks |
| MASSIVE intent classification, 13 other languages | 0.451 | Laya benchmarks |
| XNLI, English | 0.860 | Laya benchmarks |
| XNLI, 14 other languages | 0.731 | Laya benchmarks |
| Languages usable (above 3x random) out of 51 measured | 45 | Laya benchmarks |
| SST-5 sentiment on a five-level scale (score question) | 0.372 | Laya README |
For scale: these are zero-shot figures from a model with no example
of your data, on public benchmarks with many options per question.
The publisher names ordered scales (score questions) as
its weakest question type and reports that the English checkpoint
stays confident while wrong on non-Latin scripts, which is what the
routed configuration exists to avoid. Treat the probabilities as
ranking and doubt signals, and check a threshold on a labelled
sample of your own rows before relying on it.
When to pick it
- Routing, triage and tagging at volume
- Which team a ticket goes to, whether a review is positive, how urgent a message is, whether a document belongs to a category: anything where the answer is one of a short list you already know. A million rows through it cost cents, and the answer needs no parsing.
- Every answer carries its probability
- Every answer comes with its distribution, so a pipeline can act on the confident rows and send the uncertain ones to a person or to a larger model. That is the pattern it is built for.
- Query time as well as batch
- The same model answers on the realtime endpoint in tens of milliseconds, so a rule tuned on a batch run applies unchanged to live traffic.
- Not for
- Anything that needs a written answer, an extraction of free text, or reasoning over a long document: it reads a bounded window of the text and answers only with the options you gave it. Those jobs go to a generative model, such as Qwen or Gemma.
Pass laya exactly as written when you
submit a job, with rows shaped as
decision jobs describes. Its rates
are on the pricing page,
and because every job is estimated before it runs, you see what a
run costs before you approve it.