Appendix B — Formula and measurement reference

This table collects formulas and their uses for quick reference. Core marks a formula used directly in the book’s main engineering decisions. Supporting marks a useful diagnostic or implementation measure that serves a narrower part of the workflow. The links lead to the first explanation, including its definitions, assumptions and limits.

Table B.1: The formula registry preserves the connection between notation and an engineering decision.
Quantity Role First explained in Notation Engineering meaning
Autoregressive sequence probability Core @sec:M1-C01-S03 \(P(x_{1:T}\mid c)=\prod_{t=1}^{T}P(x_t\mid c,x_{<t})\) The product multiplies one conditional probability per generated token. The model repeats this generation step within a request.
Temperature-scaled softmax Core @sec:M1-C01-S06 \(p_i=\frac{\exp(z_i/\tau)}{\sum_{j=1}^{V}\exp(z_j/\tau)}\) For positive temperature, dividing logits changes the sharpness of the normalized distribution. Zero-temperature behavior is a separate API convention.
Context-budget constraint Core @sec:M1-C03-S01 \(B_{sys}+B_{hist}+B_{evidence}+B_{tools}+B_{user}+B_{out}\le C\) For a shared context limit, input components plus planned output must fit within that limit. Endpoints with separate input and output limits require both checks.
Precision Core @sec:M1-C05-S01 \(\mathrm{Precision}=\frac{TP}{TP+FP}\) The numerator counts correct positive predictions; the denominator counts every predicted positive.
Recall Core @sec:M1-C05-S01 \(\mathrm{Recall}=\frac{TP}{TP+FN}\) The numerator counts detected positives; the denominator counts every actual positive.
F1 score Core @sec:M1-C05-S01 \(F_1=\frac{2PR}{P+R}\) The harmonic mean penalizes a system when either precision or recall is small.
Independent-attempt pass at k Supporting @sec:M1-C05-S02 \(P(\mathrm{success\ by\ }k)=1-(1-p)^k\) Under independent attempts with the same success probability p, subtract the probability of all failures from one. This differs from the finite-sample pass@k estimator.
Bootstrap percentile interval Supporting @sec:M1-C04-S05 \(CI_{1-\alpha}=[q_{\alpha/2}(m_{1:B}),q_{1-\alpha/2}(m_{1:B})]\) Resampling evaluation cases with replacement estimates sampling uncertainty. Coverage depends on the sampling assumptions and chosen interval method.
Cosine similarity Core @sec:M1-C06-S05 \(\cos(q,d)=\frac{q\cdot d}{\sqrt{\sum_{j=1}^{n}q_j^2}\sqrt{\sum_{j=1}^{n}d_j^2}}\) Division by both nonzero vector lengths removes magnitude from the dot product. Cosine is undefined for a zero vector.
Retrieval hit at k Core @sec:M1-C07-S02 \(\mathrm{hit}@k=\mathbb{1}(R_k\cap E\ne\emptyset)\) The indicator is one when at least one labeled evidence unit appears among the first k results. Both sets use distinct IDs at the same source-unit level.
Mean reciprocal rank Supporting @sec:M1-C07-S02 \(MRR=\frac{1}{N}\sum_{i=1}^{N}\begin{cases}\frac{1}{r_i},&\text{if a relevant result is retrieved}\\0,&\text{otherwise}\end{cases}\) Earlier first hits receive higher scores. A question with no hit within the recorded assessment depth contributes zero. N is positive.
Repeated-trial agent success rate Core @sec:M1-C09-S07 \(\hat{p}=\frac{\sum_{i=1}^{N}s_i}{N}\) Repeated binary outcomes estimate reliability instead of treating one trajectory as representative.
Dependency-graph critical path Supporting @sec:M1-C10-S03 \(T_{DAG}\ge\max_{\pi\in\mathrm{paths}(G)}\sum_{v\in\pi}t_v\) Parallel execution cannot finish sooner than the longest dependency path.
Request token cost Core @sec:M1-C01-S02 \(C=r_{\mathrm{in}}N_{\mathrm{in}}+r_{\mathrm{out}}N_{\mathrm{out}}\) Rates and token counts use matching units. Provider-specific caching, reasoning and tool charges may require additional billing terms.
Observed human-judge agreement Supporting @sec:M1-C05-S04 \(A_o=\frac{1}{N}\sum_{i=1}^{N}\mathbb{1}(h_i=j_i)\) Agreement averages case-level label matches while preserving the need for a confusion matrix and class-aware interpretation.
Retrieval precision at k Core @sec:M1-C07-S02 \(\mathrm{Precision}@k=\frac{\lvert R_k\cap E\rvert}{k}\) For positive k, the intersection counts distinct relevant units. This fixed-k convention treats unfilled result positions as misses.
Retrieval recall at k Core @sec:M1-C07-S02 \(\mathrm{Recall}@k=\frac{\lvert R_k\cap E\rvert}{\lvert E\rvert}\) Divide the intersection by a nonempty labeled evidence set at the same unit level. Empty labels leave recall unscored unless another convention is declared.
All-trial reliability Supporting @sec:M1-C09-S07 \(R_m=p^m\) Independent trials with the same success probability p give the probability that all m required trials succeed.
Agent trial cost Core @sec:M1-C09-S07 \(C_{trial}=\sum_{t=1}^{T}(r_{\mathrm{in}}L_{\mathrm{in},t}+r_{\mathrm{out}}L_{\mathrm{out},t}+c_{tool,t})\) The equation accumulates model input, model output, and tool charges across the complete trajectory.
Dependency-graph ready set Supporting @sec:M1-C10-S03 \(\mathrm{Ready}(t)=\{v:v\notin\mathrm{Done}(t),\mathrm{Pred}(v)\subseteq\mathrm{Done}(t)\}\) A node is eligible when every predecessor is complete. An asynchronous dispatcher also excludes running nodes to prevent duplicate submission.
Reciprocal rank fusion Supporting @sec:M1-C06-S06 \(S(d)=\sum_{j:d\in L_j}\frac{1}{k+\mathrm{rank}_j(d)}\) Lists L_j use ranks starting at one. A missing candidate contributes nothing. The positive smoothing constant k is distinct from top-k; fused scores rank candidates and are not probabilities.
Memory retrieval score Supporting @sec:M1-C11-S05 \(S(m,q)=w_r\mathrm{relevance}(m,q)+w_f\mathrm{freshness}(m)+w_s\mathrm{reliability}(m)\) The weighted sum ranks eligible records after deterministic identity, permission, and deletion filters have passed.

Formula interpretation rule

A formula’s variable definitions, assumptions, worked calculation, and limitations determine whether it applies to a particular measurement.