1 Language models as application components
A file containing a language model’s learned parameters cannot complete a sentence by itself. Producing a continuation requires software to load those numbers and supply input text. The software runs the model’s computation to return likely pieces of text one at a time. Each piece is called a token. A product needs more than this bare completion. Application software is the code that forms requests from user input and company data, invokes the model, checks whether its answer is usable, and authorizes any next action. Each step has its own requirements for correctness, cost, and failure handling.
The application sends requests as structured messages and receives continuations as text. The model itself receives and produces tokens. During inference, data access, validation, and action happen outside the model, in application code or in services that the application calls and controls.
At each generation step, the model produces a probability distribution over the next token, and training sets the parameters that produce that distribution. Following a request as it moves from text to tokens, through generation, and back to application code shows where instructions, training, and software checks each affect the result.
Three operating modes recur throughout the book. During inference, a trained model receives a request and generates output while its learned parameters, the numbers stored in the model, stay fixed. During training, software compares predictions with example targets or preferences and updates those parameters. During evaluation, software sends held-out cases, compares the outputs with reference answers or stated criteria, and records measurements without updating the model. In each mode, the task description should state the inputs, expected outputs, and criteria used to judge the result.
The map follows the seven sections in reading order, from what a general model can do and how text becomes tokens, through generation, alignment, and the request cycle, to generation settings and evidence for choosing a model.
1.1 General language models can perform different text tasks
Many earlier language systems were trained for one job: one classifier might label reviews as positive or negative, while another system extracted names from documents. Adapting them to a different task usually required task-specific training data. A general language model can instead attempt different tasks using the same trained model. A prompt is the input text that describes the task and can include instructions, examples, or both.
A transformer is a neural-network design in which attention lets each token representation combine information from other positions.1 In continuation generation, attention at a position uses that position and earlier positions, not later text. Repeating attention and transformation layers produces the final token scores described in Section 1.3. In Language Models Are Few-Shot Learners2, Brown and colleagues tested GPT-3 (Generative Pre-trained Transformer 3), a large language model introduced in 2020, on tasks such as translation and question answering. They changed the instructions and examples supplied as text without retraining the model for each task. Their results showed both useful task transfer and important failures, so this flexibility still has to be tested on an application’s own inputs.
A foundation model is trained on broad data before it is used for a particular product. A large language model (LLM) is a foundation model that reads and writes text tokens. The same LLM may classify, extract, summarize, translate, or generate code. A bare model call does not fetch current facts, retain earlier calls, decide whether a user has permission, or verify an important answer. Surrounding services can add retrieval, stored state, tools, and checks, while application code remains responsible for deciding how those capabilities are used. The NIST Generative AI Profile3 gives guidance on governance, measurement, data, content, and security across the system lifecycle. It does not prescribe the particular software arrangement used in this book.
Some of the checks that application code applies are safety guardrails. A safety guardrail is a rule or validation step that limits unsafe inputs, disclosures, tool calls, or changes to stored data. For example, application code can reject a request for records the caller cannot access before sending any text to the model. It can also reject a tool call whose arguments exceed the caller’s permissions. Section 8.7 develops the input, output, tool, and state checks that enforce these limits.
At the model level, task behavior can change through the prompt, through training, or through the choice of saved model version:
In-context learning: Using instructions or examples in the prompt sent for a task to influence the model’s response. The model’s stored weights stay the same.
Parameter: A number learned during training and stored in the model. Parameters are also called weights. During ordinary inference these numbers stay fixed. Training updates them.
Checkpoint: A saved model version containing a particular set of parameters and its associated configuration. Loading another checkpoint changes the model used for inference even when the application sends the same request.
Product consequence
One general model can replace several early prototypes. The application still must define the task and validate the output against measured user-facing criteria.
One impressive demonstration does not establish reliability. Repeated tests on the intended inputs do.
Because an LLM accepts a token sequence, instructions and examples can describe a new task without loading a different checkpoint. This lowers prototype cost but does not guarantee success: results can still change with prompt wording, checkpoint version, sampling parameters, and input distributions.
A capability claim applies only to the conditions under which it was tested. A model that extracts names from short catalogue entries may fail when handling multilingual descriptions, long documents, or ambiguous entities. A repeatable comparison records the tested cases, prompt, model checkpoint, decoding settings, output schema, latency, and cost together. Later results can then be compared under the same conditions to identify what changed. One of those conditions is the token sequence the model receives.
1.2 Tokens as model input
A request begins as text, while the model processes token identifiers. Before inference, a tokenizer splits the request into pieces and gives each piece an identifying number. These pieces are tokens, which need not match whole words or individual characters. Tokenization and the resulting count affect the cost of a request, the amount of text that fits, and common failures with code, spelling, and non-English text.
Modern LLMs commonly use subword tokenization. A subword token can cover a whole common string, while a rare word, punctuation, code, or non-English text may split into several subword tokens. The same visible text can produce different token counts across different models.
One way to build subword pieces is byte-pair encoding (BPE). It starts with small units and repeatedly merges the pair that occurs most often in the training text. Sennrich, Haddow, and Birch4 used this subword method in neural machine translation and reported improvements on their tested translation tasks. For a toy vocabulary, frequent adjacent pieces play and ing might become the single token playing. The real merge list is learned from a large corpus, so this toy merge explains the idea rather than the exact vocabulary of any provider.
Other subword methods make different choices about which pieces become tokens. The Hugging Face Tokenizers documentation5 describes WordPiece, which takes the longest matching vocabulary pieces from left to right, and Unigram, which scores possible splits and selects a likely one. The model must use the tokenizer and vocabulary it was trained with. A different tokenizer would give the model a different token sequence.
The token number is only an identifier. An embedding table turns it into a learned vector: a row of numbers the transformer can process. The transformer also receives position information, so it sees numeric representations of token pieces in an order, not a clean list of dictionary words. Exact character counts, spelling, and unusual separators can behave unexpectedly when relevant characters sit inside a larger token.
Vocabulary: The fixed list of token identifiers a tokenizer and model share. Its entries may represent complete words, word fragments, punctuation, bytes, or other subword units.
Embedding: A learned vector representation associated with a token or another discrete object. Two embeddings can be numerically close, but that closeness does not guarantee identical meaning or factual equivalence.
Estimate the length of the request the model receives
To estimate cost or avoid a context-window limit, first build the complete request. Include the chat-role markers, system message, history, retrieved text, tool descriptions, and room for the answer. Then count that request with the selected model’s tokenizer and chat template when they are available. When a hosted provider does not publish them, use its token-counting endpoint, if available, or the usage it returns for a test request.
A word count is only a rough guess. It can be badly wrong for code, URLs, Hebrew, Chinese, unusual whitespace, or a different provider’s tokenizer. When the selected API defines capacity and billing in tokens, the budget must use the same tokenizer and accounting rules. Other services may also charge for requests, time, cached input, images, audio, or tool use.
Providers often charge one rate for input tokens and another for generated tokens. Count them separately before adding the two charges.
\[ C=r_{\mathrm{in}}N_{\mathrm{in}}+r_{\mathrm{out}}N_{\mathrm{out}} \tag{1.1}\]
\(N_{\mathrm{in}}\) and \(N_{\mathrm{out}}\) are the input and output token counts. \(r_{\mathrm{in}}\) and \(r_{\mathrm{out}}\) are their respective prices per token, so \(C\) is the charge for one call in the same currency. Divide a quoted price per million tokens by one million before using it as a rate in this formula. Multiply each count by its rate, then add the charges. Discounts, batch pricing, tool calls, evaluation calls, and infrastructure affect the cost of the full workflow. Appendix B collects the main formulas used in the book.
Example: Token cost for one million requests
Assume each turn makes one model request using 1,800 input tokens and 300 output tokens. Input costs $0.20 per million tokens and output costs $0.80 per million tokens.
1. Input cost per turn is 1,800 / 1,000,000 x $0.20 = $0.00036. 2. Output cost per turn is 300 / 1,000,000 x $0.80 = $0.00024. 3. Total cost per turn is $0.00060, so one million turns cost about $600 before retries, retrieval, or tool calls. Result: Output makes up about 14% of the tokens but contributes 40% of the token bill.
Interpretation: Small per-turn numbers become material at production scale. The complete workflow must be measured because agent retries and judge calls multiply the model-call count.
1.3 Generating an answer one token at a time
After tokenization, the model receives a sequence of token identifiers. The tokens already present form the prefix. At each generation step, the model computes one score, called a logit, for every token in its vocabulary. A logit can be positive or negative and has no fixed range.
To turn these scores into next-token probabilities, softmax first raises the number e (approximately 2.718) to each score, producing a positive weight. Exponentiation preserves the order of the scores while turning their differences into positive ratios. Softmax then divides each weight by the sum of all the weights. The resulting probabilities add up to 1 across all tokens in the vocabulary.6 They describe the model’s prediction of which token comes next. Suppose the prefix is The cat is on the and a toy vocabulary contains three candidates, mat, couch, and floor. Their scores of 2, 1, and 0 become probabilities of about 0.665, 0.245, and 0.090.
The generation software then chooses a token: it can take the one with the highest probability or sample a token using those probabilities. This selection procedure is called a decoding strategy. The chosen token is added to the prefix, and the model computes scores for the next position. Generation stops at an end token, a specified stop sequence, an output limit set by the application or service, or the end of the context window.
The probability assigned to a complete token sequence is the product of the conditional probabilities assigned at each position.
\[ P(x_{1:T} \mid c) = \prod_{t=1}^{T} P(x_t \mid c, x_{<t}) \tag{1.2}\]
Here, \(c\) is the input context, \(T\) is the number of generated tokens, \(x_t\) is the token at position \(t\), and \(x_{<t}\) means the generated tokens before that position. The product runs over positions 1 through \(T\). Each term conditions on the same input context and all earlier generated tokens. Their product is the probability of that particular continuation under the model. It predicts text, not whether the resulting answer is true or useful.
Example: Probability of a three-token continuation
Suppose the model assigns probabilities 0.50, 0.40, and 0.25 to the selected tokens at three successive positions.
1. Multiply the first two conditional probabilities: 0.50 × 0.40 = 0.20. 2. Include the third conditional probability: 0.20 × 0.25 = 0.05. Result: The complete three-token continuation has probability 0.05 under this model and prefix.
Interpretation: A fluent sequence can still have small total probability because every additional token contributes another factor.
The figure below applies softmax to the toy scores. The bars are possible next tokens, not facts stored in a database.
Once a token is selected, it becomes part of the next prefix. This feedback makes early choices affect every later probability.
The same model can support either an extractor whose output the application checks against a schema or an action loop whose tool calls can affect external systems. Application code determines what evidence the model receives, which outputs it validates, and whether an external action is permitted. Model behavior and the operating environment remain additional sources of risk.
This repeated prediction loop can produce fluent text, but a useful assistant must also respond to instructions and keep system, user, and assistant roles distinct. Further training makes those response patterns more likely. Request formatting then identifies the role the model should continue for the current request.
1.4 Model alignment
During pretraining, software updates a model’s parameters to improve its continuation of text that resembles the training data. An AI assistant has a more specific task. It should respond to a user’s request in the assistant role while following higher-priority instructions supplied by the application. Pretraining alone does not establish those expectations.
Model alignment is the process of making a model’s behavior better match intended instructions, preferences, or safety criteria in specified situations. This section focuses on alignment training, which adjusts the model’s parameters through further training. During later inference, the aligned model still uses the next-token loop from Section 1.3. Alignment changes the probabilities learned by the model. It does not add a separate answering engine or make every answer correct.
A base model mainly predicts tokens from patterns learned in broad text. If its prompt ends with Translate hello, it might continue with another translation exercise instead of translating the word. The continuation can fit the text pattern without serving the user’s intention.
Supervised Fine-Tuning (SFT) supplies prompts together with suitable responses. At each response position, training compares the model’s predicted token probabilities with the target token and updates the parameters to make the target response more likely. Across many examples, the model learns patterns such as answering a request directly, following an output format, and continuing in the assistant role.
Preference training addresses cases where several responses are fluent but differ in usefulness, safety, or style. Evaluators compare responses to the same prompt using stated criteria. In a common reinforcement learning from human feedback (RLHF) setup, demonstrations first train an initial model. The comparisons then train a reward model, a learned scorer that assigns a score to a candidate response in its prompt context. Reinforcement learning updates the initial model to make responses with higher reward-model scores more likely.7 The paper Direct Preference Optimization: Your Language Model Is Secretly a Reward Model8 introduces Direct Preference Optimization (DPO). DPO trains directly on preferred and rejected response pairs, without fitting that separate reward model or running reinforcement-learning optimization. In both cases, the comparison criteria and training data determine which behaviors become more likely.
An instruct model responds to instructions after further training on that behavior. A chat model responds to role-marked conversational turns after training on that structure. These labels describe useful differences in expected behavior. They do not require every provider to use the same training stages.
The alignment methods described above act during training. A chat template acts later, when software serves a request. The application sends separate system, user, and assistant messages as structured data. The serving software applies the template, which converts those messages into the control tokens and text expected by that model. With a hosted API, the provider’s service applies it. Code that calls a tokenizer library directly applies it itself. The model receives one formatted token sequence, not a collection of message objects. A template can end the request with a marker that places the next token after an assistant-role pattern encountered during training. The exact markers depend on the model.9
The chat-template figure below follows the structured messages into the token sequence used for generation.
A wrong template can merge or mislabel roles even when the application’s message objects look correct. The model then receives a different token sequence from the one the application intended.
From continuation to assistant reply
Application messages: system: Follow the return policy. Then user: Can I return this order?
Illustrative formatted prefix: <system> Follow the return policy. <end> <user> Can I return this order? <end> <assistant>
Generation: The first token selected after <assistant> is the next token in the sequence. For a model that does not first generate hidden reasoning, it is also the first token of the reply the application shows.
The assistant marker only selects a position for the existing next-token process. The template preserves role boundaries in the input, but it does not enforce permissions, identify false claims, or prevent the model from following instructions embedded in untrusted text.
Alignment can improve measured instruction following, safety, and usability. It can also cause false refusals or reduce performance on other tasks. Evaluation on the intended cases shows whether the new behavior meets the stated criteria.
Alignment changes which answers are likely. Whether an answer follows from the available information is a reasoning question. Reasoning means using available information through one or more intermediate steps to reach a conclusion. In an LLM system, reasoning can refer to three different kinds of work:
- Internal model computation. The model’s operations produce numeric activations and token scores. These values affect the output, but they are not a readable explanation or proof.
- Generated intermediate steps. The model can generate tokens before its final answer. A provider may display, hide, or summarize this sequence and may count a hidden portion as reasoning tokens. These generated tokens are different from the model’s internal activations, and they can contain mistakes.10
- Application-managed stages. Software can add records returned by retrieval or tools. It can then include user corrections, validate a result, and make another model call. These stages can supply evidence and checks that were absent from the first request.
Chain-of-thought prompting asks a model to generate intermediate steps as text before its final answer. In the cited research, this method improved accuracy on some arithmetic, commonsense, and symbolic tasks.11 Experiments have also shown cases in which a plausible step-by-step explanation omitted a factor that influenced the answer or justified an incorrect result.12 A displayed explanation needs the same evidence and result checks as a direct answer.
Suppose a return policy permits ordinary returns for 30 days and a verified order record dates the purchase 18 days ago. From those premises, the model can generate the conclusion, “The order is within the ordinary return window.” The conclusion was not written verbatim in either input, but it does not add external evidence. A user message saying that the item was final sale, a retrieved policy exception, or a tool result from the order system would supply new information. The model may still combine the premises incorrectly or overlook an exception. An important conclusion needs a check against the source, calculation, test, or authorized system result. Chapters 8 through 10 develop application-managed stages, tool observations, and replanning after new information arrives.
The application supplies conversation state
A bare stateless model call does not retain previous calls. On the next turn, the application must resend selected history or retrieve state stored by a surrounding service.
For a chat request, the application can now assemble the selected messages. It still needs a defined path to the model service. That interface is an API.
1.5 Request, response, and validation cycle
An Application Programming Interface (API) lets application code send a request to a model service and read the response. For each request, the application supplies a model identifier and the request data. It also configures token selection and rules for the returned format. These settings control which model runs, what information it receives, how it selects tokens, and how the application checks the result.
The model field selects the provider’s model identifier. Depending on the provider, that identifier can refer to a fixed checkpoint. It may instead refer to a provider-managed alias or a deployment. This choice affects capability, tokenizer, context capacity, and provider-specific behavior. Price and latency also depend on the endpoint, region, hardware, batching, service tier, current load, and commercial terms. The messages or input field supplies policy, user data, history, and evidence. Decoding fields control token selection. Output-schema settings let the provider or the application check that the returned data has the required fields and types.
Each field changes a different aspect of execution. Changing the model alters capability and cost. Changing messages updates evidence. Temperature shifts selection probabilities, while a small output limit can cut off an otherwise valid answer.
| Field or layer | Who sets it, and who applies or checks it | Changes | What it cannot guarantee |
|---|---|---|---|
| Model checkpoint | Application selects the identifier. The provider or self-hosted runtime resolves and serves the checkpoint | Capability, tokenizer, latency, cost, context capacity | Task success on the product distribution |
| Messages and context | Application | Instructions, evidence, state, examples, available tools | That the model will use every supplied fact correctly |
| Decoding controls | Application sets them. The inference API or runtime applies them | How probabilities become selected tokens | Factuality, safety, or deterministic semantics |
| Schema and validator | API plus application | Output syntax and accepted fields | That field values are supported by evidence |
Streaming changes when the application receives the output, not how the model chooses its next token. In OpenAI’s Responses API, the application receives a sequence of events labeled by kind while generation continues, instead of waiting for a complete response. A text-delta event adds another piece of output text. The API does not define this as a one-event-per-model-token interface.13 A partial sentence or incomplete JSON object remains provisional. The application should wait for the required completion event and validation before accepting a structured result or allowing an external action.
Mid-response steering supplies additional user input while the model is generating a response. Streaming alone does not provide this control. With an ordinary request interface, the application can stop or ignore the active response. It then appends the new message to conversation state and begins another call. A provider can also implement direct steering. As a dated example, OpenAI14 documents a WebSocket operation for its GPT-6 model family that queues the added input and automatically creates a continuation, unless the response needs a tool result or approval from the application. It does not rewrite output already delivered, undo earlier actions, or cancel tools that have started.
Product interfaces sometimes call this steering while the model is thinking. The update changes the input for later computation. It does not edit a hidden reasoning trace.
A streamed answer changes direction
1. A customer asks an AI assistant, “Can I return this order?”
2. The model begins generating a general answer. The application displays the partial text: Orders within 30 days can usually...
3. Before the response finishes, the customer adds, “The item was marked final sale.”
4. The application records that message. With an ordinary request interface, it stops or ignores the first response and starts another call with the updated conversation. With supported direct steering, it submits the update and continues reading provider events. The provider may finish the current output item before the continuation begins.
5. The customer’s statement that the item was marked final sale is new input. A later statement such as This order is not eligible for return would be a conclusion derived from the supplied facts. The application still needs the applicable policy and a verified order record before accepting that conclusion.
Conclusion: Steering changes subsequent work. A partial answer still needs the full checks before it becomes a decision.
Streaming and steering govern one interaction. A separate system choice is where the model runs and which organization operates it. A direct proprietary API reduces serving work, but it sends data to a vendor and creates rate-limit, pricing, and version dependencies. A managed platform adds catalogue and deployment controls. Self-hosting open weights can increase control over data location, retention, and customization. Privacy still depends on configuration, access control, logging, patching, and daily operations. The organization must also plan capacity, manage upgrades, enforce security, and maintain reliability.
Comparing access paths with the same task cases
A fair comparison uses the same evaluation cases and output requirements for a direct provider API, a managed platform, and self-hosted model weights. It measures answer quality, latency, data location, service work, and total cost.
Open weights can improve privacy and customization, but they do not remove licensing obligations, security responsibilities, hardware expenses, or potential model-quality gaps. A low token price can still cost more once capacity planning and reliability work are included.
An OpenAI-compatible client can send requests to Nebius Token Factory, the service formerly named Nebius AI Studio.15 Its result record can include the provider configuration, model identifier, token counts, and latency. The endpoint and model below are dated examples. Production use depends on current provider documentation.
This short Python client uses the openai package’s Chat Completions interface. It waits for one complete response and does not implement streaming, interruption, or steering. It assumes model_id, system_instructions, and user_request are supplied by the application, and that the environment contains a valid credential. The result is a dictionary containing the requested and served model identifiers, returned text, finish reason, decoding settings, token counts, and elapsed milliseconds. The client uses max_retries=0, so it does not hide retry waits in this measurement. The caller must handle authentication, rate-limit, timeout, missing-usage, and missing-text errors before accepting a record. A finish reason of length means the output limit cut the answer off.
This Nebius example uses max_tokens. An OpenAI-compatible endpoint does not ensure that every model accepts the same generation fields. Production code must check support for temperature and the output-limit field.
Code example: A Nebius Token Factory call records token usage and elapsed time.
from openai import OpenAI
import os
import time
client = OpenAI(
base_url="https://api.tokenfactory.nebius.com/v1/",
api_key=os.environ["NEBIUS_API_KEY"],
max_retries=0,
)
started = time.perf_counter()
response = client.chat.completions.create(
model=model_id,
messages=[
{"role": "system", "content": system_instructions},
{"role": "user", "content": user_request},
],
temperature=0.2,
max_tokens=300,
)
text = response.choices[0].message.content
if not isinstance(text, str) or not text.strip():
raise ValueError("The model returned no answer text")
record = {
"provider": "Nebius Token Factory",
"base_url": str(client.base_url),
"model": model_id,
"served_model": response.model,
"finish_reason": response.choices[0].finish_reason,
"temperature": 0.2,
"max_tokens": 300,
"text": text,
"input_tokens": response.usage.prompt_tokens,
"output_tokens": response.usage.completion_tokens,
"latency_ms": 1000 * (time.perf_counter() - started),
}A repeatable evaluation run retains model_id, endpoint configuration, prompt versions, and returned usage. Build and release comparisons also depend on current official documentation for prices, model availability, usage-field meanings, retention behavior, and rate limits.
The example sets temperature=0.2. This changes how the system chooses from the model’s possible next tokens. The next section explains how that choice affects variation between answers.
1.6 Token generation parameters
Section 1.3 followed logits through softmax to a selected token. Generation parameters adjust this process. Temperature changes how concentrated the probabilities are, while top-k and top-p remove some candidates before sampling. The final selection rule then takes the most probable token or samples from the remaining distribution. These settings help explain why identical requests can produce different answers. Most chat APIs return the chosen text rather than the full score vector.
In the sequence explained here, temperature divides logits before softmax, then truncation selects candidates. Runtime ordering can differ. With unequal scores, lowering a positive temperature increases the total probability of the highest-scoring token or tied tokens and lowers that of the lowest-scoring token or tied tokens. A larger temperature spreads probability more widely. At temperature zero, runtimes normally use greedy or near-greedy selection rather than dividing by zero.
\[ p_i = \frac{\exp(z_i / \tau)}{\sum_{j=1}^{V} \exp(z_j / \tau)} \tag{1.3}\]
\(z_i\) is the logit for vocabulary item \(i\), \(\tau\) is the positive temperature, \(V\) is vocabulary size, and \(p_i\) is the resulting probability. The index \(j\) runs over all vocabulary entries. Dividing the logits by temperature changes the ratios between the positive weights, and dividing by their sum again makes the probabilities add up to one. This changes which continuations are likely, without adding evidence or verifying claims.
Example: Temperature changes a two-token choice
Use logits 2 and 1 for two candidate tokens.
1. At temperature 1, exponentials are approximately 7.39 and 2.72, so the first token receives probability 7.39 / 10.11 = 0.73. 2. At temperature 0.5, the scaled logits are 4 and 2. Exponentials are approximately 54.60 and 7.39, so the first token receives probability 0.88. Result: Lower temperature increases concentration on the already higher-scored token.
Interpretation: Extraction and routing often benefit from low variation, while ideation may benefit from sampling. Task-specific evaluation results provide a better basis for the setting than a general preference for creativity or determinism.
Top-k keeps the k highest-probability tokens before sampling. Top-p sorts tokens by probability and keeps the smallest leading set whose total reaches the chosen threshold. The remaining probabilities are then rescaled to sum to one. Consider probabilities 0.60, 0.25, 0.10, and 0.05. Setting top-k with k = 2 retains two candidates. Setting top-p with threshold 0.90 retains three candidates, since 0.60 plus 0.25 totals only 0.85. Top-k fixes a candidate count, while top-p adjusts the count to the distribution.
The figure below compares top-k and top-p and places them in one example decoder sequence.
Temperature changes how concentrated the probabilities are, and truncation removes candidates. The final selection rule chooses what the user receives. Evaluating these controls needs representative task cases, not one plausible sample.
Even when the interface appears deterministic, the model, retrieval step, or tools can vary between runs. Repeated trials remain necessary whenever the complete system can vary.
Some APIs also return part of the probability distribution. When a request enables the option, the response lists a log-probability for each generated token, the natural logarithm of that token’s probability. It can also list the most likely alternatives at each position. In the OpenAI Chat Completions interface, logprobs: true returns the values and top_logprobs adds up to 20 alternatives per position.16 Nebius Token Factory accepts the same two fields.17
A log-probability of \(-0.31\) corresponds to a probability of \(e^{-0.31} \approx 0.73\). Log-probabilities add where probabilities multiply. For the three-token continuation in Section 1.3, the sum is \(\ln 0.50 + \ln 0.40 + \ln 0.25 \approx -3.00\), and \(e^{-3.00} \approx 0.05\).
Support differs by provider and model. A server may report probabilities before or after temperature and truncation, and vLLM reports the unadjusted values by default.18 The values describe the model’s prediction, not whether the output is correct. Section 5.1 checks them against labeled cases before using them as a confidence score.
Decoding settings apply to one request. The selected model, how the application accesses it, and its limits are larger choices that affect the whole application. Evidence from the intended task must determine those choices.
1.7 Evidence for model selection
Model type, how the application accesses the model, context capacity, and decoding settings affect different parts of the system. A leaderboard score cannot reveal whether one model fits its task, budget, data rules, and failure costs. The Holistic Evaluation of Language Models19 study, known as HELM, compared models across several scenarios and measurements and found trade-offs that a single aggregate score would hide. Even that broad evaluation does not represent one product’s users and operating constraints. Default and fallback choices depend on comparisons of checkpoints on representative tasks under those constraints.
A model-selection study defines the target behaviors, including extraction accuracy, citation fidelity, code execution success, latency, cost, privacy, and safety. A small representative evaluation can compare several checkpoints before extensive prompt tuning for one model. A stronger model may need less prompting, while a smaller model may be cheaper once the prompt and output are constrained.
- Capability fit: Does the model solve the intended task, including important slices (case subsets with their own scores, such as one language or request type)?
- Operational fit: Do latency, throughput, context capacity, rate limits, and reliability meet the service requirement?
- Data and legal fit: Are privacy, retention, data residency, model updates, and licensing acceptable?
- Economic fit: What is total workflow cost after retries, tool calls, retrieval, judge-model calls, and human review?
- Reversibility: Can the application switch checkpoints without rewriting state, prompts, schemas, and evaluations?
A selection study keeps the task set and output specification fixed while it compares checkpoints under realistic request settings. Error slices show whether a larger model reduces the need for detailed prompt guidance. Tests with retrieval or constrained decoding show whether a smaller model is adequate. The decision then weighs any quality gain against added latency and cost. The right answer depends on the application because an unsupported answer, a slow response, and a false tool action have different costs in different products.
A model call predicts tokens from a supplied prefix. Chapter 2 turns that prefix into a repeatable task specification: instructions, examples, an output format, and checks the caller can apply.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30. https://arxiv.org/abs/1706.03762. The paper introduces the Transformer as an architecture based on attention rather than recurrence or convolution. This chapter uses a simplified explanation.↩︎
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., . . . Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems 33 (pp. 1877–1901). https://arxiv.org/abs/2005.14165. Brown and colleagues evaluated GPT-3 on translation, question answering, and other tasks using instructions and examples without task-specific gradient updates. Their results do not establish reliable transfer to every application.↩︎
National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.600-1. The profile supports system-level and lifecycle risk management, but it does not prescribe the software architecture in this book.↩︎
Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 1715–1725). Association for Computational Linguistics. https://arxiv.org/abs/1508.07909.↩︎
Hugging Face. (n.d.). Components (Tokenizers documentation). Retrieved September 21, 2026, from https://huggingface.co/docs/tokenizers/components. The documentation lists BPE, WordPiece, and Unigram as distinct tokenizer models and describes how each selects subword pieces. Its examples describe this library’s implementations, not every provider tokenizer. ↩︎
PyTorch. (n.d.). Softmax (PyTorch 2.14 documentation). Retrieved September 16, 2026, from https://docs.pytorch.org/docs/2.14/generated/torch.nn.Softmax.html. The documentation gives the normalization formula and states that the values along the selected dimension lie between zero and one and sum to one.↩︎
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35 (pp. 27730–27744). https://arxiv.org/abs/2203.02155. The paper describes supervised fine-tuning from demonstrations followed by reward-model training from ranked outputs and reinforcement learning.↩︎
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2305.18290. DPO derives a classification-style objective that avoids explicit reward-model fitting and reinforcement-learning optimization. Its reported results do not make it preferable for every alignment task.↩︎
Hugging Face. (n.d.). Chat templates (Transformers v5.17.0 documentation). Retrieved September 19, 2026, from https://huggingface.co/docs/transformers/v5.17.0/en/chat_templating. The documentation explains that role-marked messages become one token sequence and that a generation prompt can mark the start of an assistant reply. Exact markers and template behavior depend on the model.↩︎
OpenAI. (n.d.). Reasoning models. Retrieved September 19, 2026, from https://developers.openai.com/api/docs/guides/reasoning. The documentation states that raw reasoning tokens are not exposed through the API and that supported models can return an optional reasoning summary. This describes current OpenAI API behavior rather than every model or provider.↩︎
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35. https://arxiv.org/abs/2201.11903. The paper reports gains from intermediate reasoning demonstrations on tested arithmetic, commonsense, and symbolic tasks. It does not establish that generated reasoning is always correct or useful.↩︎
Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2305.04388. In the tested settings, plausible explanations sometimes omitted biasing influences and rationalized incorrect answers. The result does not show that every generated explanation is unfaithful.↩︎
OpenAI. (n.d.). Streaming API responses. Retrieved September 19, 2026, from https://developers.openai.com/api/docs/guides/streaming-responses. The documentation describes typed server-sent events, including text deltas and response lifecycle events, while generation continues. The interface does not define each delta as one model token.↩︎
OpenAI. (n.d.). Mid-turn steering. Retrieved September 19, 2026, from https://developers.openai.com/api/docs/guides/steering. As documented on the access date, direct steering is available for the GPT-6 model family over a Responses API WebSocket connection. The update creates a continuation unless the response needs a tool result or application approval. It does not rewrite delivered output, undo actions, or cancel tools already started. Earlier OpenAI model families and other providers can behave differently.↩︎
Nebius. (n.d.). Quickstart (Nebius Token Factory documentation). Retrieved September 23, 2026, from https://docs.tokenfactory.nebius.com/quickstart. The documentation gives the OpenAI-compatible base URL
https://api.tokenfactory.nebius.com/v1/and theNEBIUS_API_KEYenvironment variable. Endpoints and model identifiers can change.↩︎OpenAI. (n.d.). Create chat completion (API reference). Retrieved September 24, 2026, from https://developers.openai.com/api/reference/resources/chat/subresources/completions/methods/create. The reference defines
logprobsandtop_logprobs(0 to 20) and states that parameter support can differ by model, particularly for newer reasoning models.↩︎Nebius. (n.d.). Create chat completion (Nebius Token Factory documentation). Retrieved September 24, 2026, from https://docs.tokenfactory.nebius.com/api-reference/inference/create-chat-completion. The endpoint accepts
logprobsandtop_logprobsfor output tokens. Support for an individual model can change.↩︎vLLM. (n.d.). vLLM V1 (vLLM documentation). Retrieved September 24, 2026, from https://docs.vllm.ai/en/latest/usage/v1_guide.html. By default, returned log-probabilities come from the raw model output, before temperature, penalties, or top-k and top-p. A server option selects processed values instead.↩︎
Liang, P., et al. (2023). Holistic evaluation of language models. Transactions on Machine Learning Research. https://arxiv.org/abs/2211.09110. HELM supports multi-scenario, multi-measurement evaluation, not the chapter’s exact product-selection procedure. ↩︎