Laya on anex.sh

Convai Innovations' open-weight decision model, the one model here that answers without writing.

The family

Laya is an open-weight decision model published by Convai Innovations under the Apache 2.0 licence. It answers typed questions about a text (a choice between options, a score on an ordered scale, a yes or no) with a probability per option, in a single pass of an encoder, and it generates no text at all. That makes it a different kind of model from the rest of the catalogue: there is no prompt to write, no reply to parse, and no output tokens to pay for.

One id fronts it here: laya. Behind that id sit two published checkpoints, one for English and Latin-script text and one for the rest of its more than 100 languages, and the model picks between them per request by the text's script and language.

What we serve

modelsizeinputquantizationweights served
laya421M + 322M (two encoders)textunquantizedconvaiinnovations/laya
Size and shape
Two encoder checkpoints from the one published repository: a 421M-parameter encoder for Latin-script text with a 512-token window, and a 322M-parameter multilingual encoder with an 8,192-token window. Small enough that a whole request is answered in tens of milliseconds, and cheap enough that the batch rate is the lowest on the price list.
Input
Text only: a string, or a JSON object the model reads as text, plus the typed questions asked of it. No images, audio or video.
Output
One answer per question. A choice answer names the most probable option and carries a probability per option; a score answer is the expected level on your scale with a probability per level; a noul (yes or no) answer is the probability that the answer is yes. Every answer carries a confidence. See decision jobs for the exact request and answer shapes.
What it is served from
The published weights, unquantized, at a pinned revision. Both checkpoints are resident at once, so routing between them costs nothing per request.

Benchmark scores

Published figures, each from the source named beside it: the publisher's own runs, with both checkpoints answering the same questions. "Routed" is the configuration we serve, the model picking the checkpoint per request. Accuracy is a fraction, so 0.783 is 78.3% of questions answered with the expected option.

benchmarkscore (routed)source
MASSIVE intent classification, English (20 options)0.783Laya benchmarks
MASSIVE intent classification, 13 other languages0.451Laya benchmarks
XNLI, English0.860Laya benchmarks
XNLI, 14 other languages0.731Laya benchmarks
Languages usable (above 3x random) out of 51 measured45Laya benchmarks
SST-5 sentiment on a five-level scale (score question)0.372Laya README

For scale: these are zero-shot figures from a model with no example of your data, on public benchmarks with many options per question. The publisher names ordered scales (score questions) as its weakest question type and reports that the English checkpoint stays confident while wrong on non-Latin scripts, which is what the routed configuration exists to avoid. Treat the probabilities as ranking and doubt signals, and check a threshold on a labelled sample of your own rows before relying on it.

When to pick it

Routing, triage and tagging at volume
Which team a ticket goes to, whether a review is positive, how urgent a message is, whether a document belongs to a category: anything where the answer is one of a short list you already know. A million rows through it cost cents, and the answer needs no parsing.
Every answer carries its probability
Every answer comes with its distribution, so a pipeline can act on the confident rows and send the uncertain ones to a person or to a larger model. That is the pattern it is built for.
Query time as well as batch
The same model answers on the realtime endpoint in tens of milliseconds, so a rule tuned on a batch run applies unchanged to live traffic.
Not for
Anything that needs a written answer, an extraction of free text, or reasoning over a long document: it reads a bounded window of the text and answers only with the options you gave it. Those jobs go to a generative model, such as Qwen or Gemma.

Pass laya exactly as written when you submit a job, with rows shaped as decision jobs describes. Its rates are on the pricing page, and because every job is estimated before it runs, you see what a run costs before you approve it.