news_article.exe
📰

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

2026年9月24日1 次浏览来源:MarkTechPost 阅读原文

In this tutorial, we work with Jev, TypeSafe AIs first System One model, which does not generate text at all: we send it a piece of program state and a set of typed questions, and it returns choices, scores, and yes/no probabilities that our code can branch on directly. We install the official Python SDK, make a first call that uses all three question primitives at once, and look at how the shape of the state changes what the model can know. We then recompute the published confidence statistic from the returned probabilities, measure what batching ten questions into one call buys over ten separate calls, and build the patterns the API is designed for: confidence-gated routing, composite scoring with the weights kept in code, typed function calling, and counting done the way the model can...

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

In this tutorial, we work with Jev, TypeSafe AIs first System One model, which does not generate text at all: we send it a piece of program state and a set of typed questions, and it returns choices, scores, and yes/no probabilities that our code can branch on directly. We install the official Python SDK, make a first call that uses all three question primitives at once, and look at how the shape of the state changes what the model can know. We then recompute the published confidence statistic from the returned probabilities, measure what batching ten questions into one call buys over ten separate calls, and build the patterns the API is designed for: confidence-gated routing, composite scoring with the weights kept in code, typed function calling, and counting done the way the model can actually do it. We close with the production shape: Pydantic response models, an async client fanned out with asyncio, retry policies, typed errors, and a running ledger that prices the whole notebook.

Copy CodeCopiedUse a different Browser

We install typesafe-sdk, pinned to the version this notebook was written against, and load the API key from the environment, from Colabs Secrets tab, or from a hidden prompt, so it never appears in the notebook. TypeSafeClient reads TYPESAFE_API_KEY on its own and defaults to the jev-latest alias; listing the models shows which names and pinned versions the key can use. The small ask helper wraps system_one so that every call in the rest of the notebook is timed and its token usage lands in a ledger we total at the end.

A System One request has two parts: state, which is any text, JSON object or array describing the situation, and a dictionary of named questions. Choice selects one label from the criteria we define and returns a probability for every label; Score places the state on an ordered rubric and returns the probability-weighted level, so it can land between two levels; Noul returns a single probability that a statement is true. The question names are ours and never reach the model, which is why the instructions carry the full meaning and can point at nested fields with backticked paths. All four questions are evaluated in one request, in parallel and in isolation from one another, and the response reports the pinned model version that answered and the tokens it billed.

State is the only thing the model knows, so we ask one question, whether the customer is eligible for a refund under the companys written policy, over three shapes of state. A bare string contains the complaint and nothing else; an array adds the conversation; the JSON object adds the order with its two captured charges and the refund policy itself. The question never changes, so whatever difference appears in the returned probability is attributable to the state, and the token column shows what the extra context costs. Named fields are the documented recommendation whenever the context has several parts, because the instructions can then refer to them by name.

TypeSafe documents confidence as a statistic computed from the distribution that the answer already contains: the number of options times the peak probability, minus one, divided by the number of options minus one. We recompute it from a Choices probabilities and compare it with the confidence field, and we recompute the Score as the sum of each level times its probability. Running a blunt message and a deliberately vague one through the same two questions shows how the distribution, and therefore the confidence, responds to ambiguity. A Noul carries no confidence field at all, since its value already is the probability of yes, and a value near 0.5 means undecided rather than moderate.

Because questions in a request cannot see each other, we can ask everything we might need up front, including questions that only matter on one branch, and read only the relevant answers afterwards. We put ten questions about an incident postmortem, two Choices, two Scores and six Nouls, into one call, then ask each of them again in a call of its own, and compare wall time, input tokens and the answers. The state is sent once instead of ten times, which is where both the latency and the token savings come from, and the agreement column checks the isolation claim directly: a question should receive the same answer whether or not it travels with others.

Typed answers only matter if the code around them encodes how much certainty an action requires. We classify each message into an intent and route on two things: the intent itself and whether its confidence clears a bar that rises with the stakes, from 0.5 for reading a balance to 0.9 for closing an account. Anything classified as other, or below 0.5, goes to a person; an intent that is recognised but under its bar is confirmed with the user first. The thresholds are ordinary Python values, so risk tolerance is reviewed, versioned and tested like any other code rather than buried in a prompt.

Composite scoring keeps the models job narrow and the policy explicit. For each candidate we ask four Score questions, each describing concrete situations rather than degrees, normalise every score by its top level, and store the resulting table. The ranking is then plain arithmetic: one weight vector for a senior individual contributor, another for a team lead. Because the judgments are stored separately from the weights, changing what we value instantly re-ranks the candidates and requires no inference. You can trace every position in the ranking back to the dimension that produced it.

Function calling becomes a set of closed-set questions: one Choice selects the tool, including an explicit none option for commands we do not support, and one Choice per argument is asked speculatively in the same request. The code reads only the arguments that belong to the selected tool, reports the weakest judgment as the confidence of the whole call, and then executes an ordinary Python function with validated, enumerated values. The second half applies a documented workaround: Jev does not count reliably inside a single question, so we ask one Noul per item in one request and do the sum in code.

Four details turn the examples into a decision service. Subclassing SystemOneResponse and declaring the answers we expect gives attribute access validated by Pydantic, so the typed decision stays typed all the way into the application rather than becoming a dictionary lookup. Because every request is independent, a queue of tickets is a queue of independent decisions: AsyncTypeSafeClient with asyncio.gather sends them concurrently, and the run_async helper makes the same code work in a script and inside a notebook, where an event loop is already running. RetryPolicy bounds the retries, the backoff and the total time budget per call. Errors are typed as well: an empty question set is rejected before any request is made, and an unknown model name comes back from the API as a TypeSafeAPIError subclass carrying the HTTP status.

The summary prints the one-line result each section returned, then totals the ledger that every call has been feeding: the number of requests, the input and output tokens, and the cost at the published input price, with output tokens free.

In conclusion, we used Jev the way it is meant to be used: as a source of small, typed judgments that code composes, not as a text generator to be prompted and parsed. Every answer arrived as a label, a level, or a probability with its distribution attached, letting us set thresholds, weights, and routing rules in Python where they can be tested. Batching questions over a shared state reduces both latency and tokens because the state travels once per item; Nouls replace a count the model cannot be trusted to make, and closed-set Choices turn natural-language commands into validated function calls. The production p

> 分享: