news_article.exe
📰

A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration

2026年10月7日1 次浏览来源:MarkTechPost 阅读原文

In this tutorial, we work with Laya, the open-source decision engine from Convai Innovations that became one of the most-starred machine-learning repositories of September 2026. Laya is a non-autoregressive System 1 model: instead of generating text, a 421-million-parameter encoder reads a piece of text and a set of typed questions, a choice between labels, a score on a scale, or a yes/no, and returns a probability for every option in a single forward pass with zero output tokens. Its pitch is speed and calibrated probabilities, the open answer to TypeSafes Jev. Rather than repeat the READMEs examples, we put those promises to work on real labelled data with known answers, the banking domain of the CLINC150 intent dataset, and measure what a production router actually gets: zero-shot...

In this tutorial, we work with Laya, the open-source decision engine from Convai Innovations that became one of the most-starred machine-learning repositories of September 2026. Laya is a non-autoregressive System 1 model: instead of generating text, a 421-million-parameter encoder reads a piece of text and a set of typed questions, a choice between labels, a score on a scale, or a yes/no, and returns a probability for every option in a single forward pass with zero output tokens. Its pitch is speed and calibrated probabilities, the open answer to TypeSafes Jev. Rather than repeat the READMEs examples, we put those promises to work on real labelled data with known answers, the banking domain of the CLINC150 intent dataset, and measure what a production router actually gets: zero-shot accuracy against a trained classifier, how much the wording and order of the options matter, how honest the shipped probabilities are, what fitting a temperature on validation data fixes and what it quietly breaks, an abstention gate fitted to an error budget, out-of-scope traffic, a yes/no question that temperature cannot repair, and typed outputs from a pydantic schema. Copy CodeCopiedUse a different Browser We install the released package, laya 0.3.27, and load the English checkpoint. Two choices here keep the run reproducible. By default laya.load follows the Hugging Face main branch, so we pin the revision the librarys own authors reviewed, which it exposes as laya.PINNED_REVISIONS. And on CUDA Laya autocasts to half precision, so we switch that off to keep every device in fp32 and let a GPU run reproduce the CPU numbers shown here. Printing the checkpoints shipped temperatures turns up the first finding before any prediction: the entry for choice questions with eleven or more options is 0.10, outside the valid range, so the loader clamps it to 0.5 and warns. A temperature below one sharpens probabilities, so every answer to a question with that many options will look twice as certain as the raw model is. Copy CodeCopiedUse a different Browser One call to predict answers three typed questions about a support ticket in a single forward pass: the department as a choice, the urgency as a score from 0 to 2, and the churn risk as a yes/no. The result carries a probability for every option and two confidence fields that are easy to confuse. answer_confidence is the probability of the reported answer, and it is the quantity that calibration, the abstention gate and every metric later in this tutorial use. confidence is one minus the normalized entropy, whose scale depends on how many options a question has. The usage block shows zero output tokens, because Laya scores the options it is given and never generates text. Copy CodeCopiedUse a different Browser Before building on Laya, we measure the cost of a forward pass on a single message. Each question becomes its own row, paired with the message, so sixteen yes/no questions take about eight times as long as one. All the options of a choice question share one row and its head budget, so a forty-option choice costs barely twice a three-option one and a quarter of what sixteen yes/no questions do on our CPU. That gives a design rule that shapes everything after it: ask one choice question with many options rather than many yes/no questions. Copy CodeCopiedUse a different Browser For real labeled data, we use CLINC150, a public intent-classification benchmark of 150 intents across ten domains plus a set of out-of-scope queries, read straight from the Hugging Face Hub as a parquet file. We take its banking domain, fifteen intents with 100 training, 20 validation, and 30 test queries each, and ask Laya to route the 450 test queries zero-shot, giving it each intents name and a one-line description of the kind a developer would write. It reaches 0.804 accuracy with no labeled examples. For scale, a TF-IDF and logistic-regression classifier reaches 0.651 with three labeled queries per intent, 0.848 with ten, and 0.904 with thirty. Copy CodeCopiedUse a different Browser Next we change only the wording of the options. Giving Laya the fifteen bare intent names, without our descriptions, lifts accuracy from 0.804 to 0.878 and halves the time, because the options take less than half the tokens. The descriptions blurred intents the names keep apart: account_blocked was routed to freeze_account ten times, and interest rate questions to balance. Reversing the order of the bare names changes 4.2 percent of individual answers even though overall accuracy barely moves, a sign of a position prior, so the option order you deploy should be the order you tested. Only labeled data could tell us either thing; from here on we route on the bare names. Copy CodeCopiedUse a different Browser We then ask how honest the probabilities are. A fifteen-option question falls into the checkpoints choice:11+ temperature bucket, the one clamped to 0.5. The reliability table on the 450 test queries shows the result: 92 percent of answers claim a confidence of at least 0.9, but 91.1 percent of those are right, and the mean confidence of 0.974 sits well above the accuracy of 0.878, an expected calibration error of 0.102. Layas training objective uses proper scoring rules, which is what the model card means by calibrated. Still, calibration is a property of a question on a distribution, and you have to measure it on your own labels. Copy CodeCopiedUse a different Browser Layas calibration module turns labeled examples into records of raw logits and targets, and fits a temperature to them. We build 300 records from the validation split and call agent.fit_temperatures, which fits a choice temperature of 1.258 and installs it, then prints the full temperature table next to what shipped. The fit replaced the entire map: every per-option-count entry is gone, because a bucket needs 2,000 records to keep its own temperature, and the score and yes/no temperatures were reset to 1.0 because there were no records of those types. One fit on a choice question silently changed how every yes/no question in the agent is calibrated. So we restore the shipped values and install only the bucket we measured. On the test set, calibration error falls from 0.102 to 0.059 with accuracy unchanged, since temperature never changes which option wins, and save_calibration writes the result to a JSON file that laya.load can read back. Copy CodeCopiedUse a different Browser laya.fit_abstention_thresholds takes the same validation records and, for each option-count bucket, returns the loosest confidence cut that keeps the validation error within a target. For a 5 percent target, it picks 0.602, which keeps 95.7 percent of the validation queries at 4.5 percent error, and Laya applies the same cut itself when it is passed to predict_batch as min_confidence, marking 35 of the 450 test answers as abstained. On the test set, though, that gate keeps 92.2 percent of the queries at 9.2 percent error, nearly twice the budget, and the 2 percent target realizes 5.3 percent. The data explains it: Laya is right on 92.7 percent of the validation queries but only 87.8 percent of the test queries, so an error budget fitted on one sample holds only for traffic that looks like it. The ranking itself is sound, with the most confident half of the test answers 97.8 percent correct, but an error target needs margin and periodic re-fitting on real traffic. The thresholds stay per option-count bucket because one number doesnt transfer between a two-way and a fifteen-way question. Copy CodeCopiedUse a different Browser Production traffic includes requests the router was never built for, so we add 150 out-of-scope CLINC queries and 150 queries from other CLINC domains. The calibrated confidence separates them sharply: banking queries average 0.912, the others about 0.25, and the 5 percent gate stops 89.3 percent of other-domain queries and 93.3 percent of out-of-scope ones while abstaining on 7.8 percent of banking queri

> 分享: