Appendix A — Formula reference

A formula gives a useful estimate only under its assumptions. The table collects relationships for cost, capacity, memory, communication, latency, evaluation, storage, retrieval, and pipeline duration. Each row links to the explanation and worked example that establish its conditions.

Table A.1: The formulas are paired with a short explanation of the relationship each one expresses.
Topic Relationship Meaning
Total cost over one decision horizon \(C_{\mathrm{total}} = C_{\mathrm{fixed}} + C_{\mathrm{compute}} + C_{\mathrm{storage}} + C_{\mathrm{network}} + C_{\mathrm{operations}} + C_{\mathrm{risk}} - V_{\mathrm{residual}}\) A common horizon and workload make the alternatives comparable. Section 1.5 defines the cost terms and comparison limit.
Replica capacity with utilization and redundancy headroom \(R_{\mathrm{needed}} = \left\lceil \frac{Q_{\mathrm{peak}}}{Q_{\mathrm{replica}} \cdot u} \right\rceil + R_{\mathrm{redundancy}}\) Peak demand becomes a whole-number replica estimate with spare capacity. Section 1.5 gives the worked sizing assumptions.
Approximate per-device training memory \(M_{\mathrm{train}} ≈ P\left(b_w + b_g + b_{\mathrm{master}} + b_{\mathrm{opt}}\right) + A\) Persistent model state and batch-dependent activations explain the memory requirement. Section 2.1 defines the state and activation terms.
Global gradient averaging \(g_{\mathrm{global}} = \frac{\sum_{r=0}^{N-1} g_r}{N}\) Local batches contribute to one shared data-parallel update. Section 2.3 gives the gradient-averaging mechanism.
Approximate ring all-reduce sent bytes per rank \(B_{\mathrm{ring}} ≈ \frac{2 \cdot \left(N-1\right)}{N} \cdot G\) Sent volume for the illustrated ring algorithm; received volume is the same, not included in this value. Section 2.3 gives the cited derivation and limits.
Pipeline bubble fraction \(\beta_{\mathrm{bubble}} \approx \frac{p-1}{m+p-1}\) Idle share of a pipeline schedule with p stages and m micro-batches, assuming equal stage times. Section 2.4 gives the worked example.
Model FLOPs Utilization \(\mathrm{MFU} = \frac{F_{\mathrm{observed}}}{F_{\mathrm{peak}}}\) Useful model forward/backward arithmetic, excluding recomputation, divided by stated precision/dense-or-sparse peak. Section 2.8 gives the cited definition and limits.
Storage bandwidth from operation rate and transfer size \(B = \mathrm{IOPS} \cdot S_{\mathrm{block}}\) Operation rate and transfer size together determine byte throughput. Section 3.5 gives the workload and measurement context.
Request latency by serving stage \(T_{\mathrm{request}} ≈ \mathrm{TTFT} + \left(N_{\mathrm{out}}-1\right) \cdot \mathrm{ITL}\) Queueing, prompt processing, and token generation contribute separately to latency. Section 4.1 defines TTFT, ITL, and the request-time relationship.
Serving cost per million output tokens \(C_{\mathrm{1M}} = \frac{P_{\mathrm{hour}} \cdot 10^{6}}{3600 \cdot R_{\mathrm{tok}}}\) Replica price and sustained output-token rate at the latency objective set the compute cost. Section 4.1 converts the same rate to requests per second.
KV-cache memory \(M_{\mathrm{KV}} = 2 \cdot B \cdot L \cdot T \cdot d_{\mathrm{KV}} \cdot b\) Retained attention state determines context and concurrency capacity. Section 4.5 defines B, L, T, d_KV, and b.
Cosine similarity \(\cos\left(q,x\right) = \frac{q^{T}x}{\left\lVert q\right\rVert \cdot \left\lVert x\right\rVert}\) For nonzero vectors, direction is compared under an explicit score convention. Section 5.2 gives the geometric calculation and convention.
Reciprocal rank fusion \(\operatorname{RRF}\left(d\right) = \sum_{r \in \left\{\mathrm{dense},\mathrm{sparse}\right\}} \frac{1}{k+\operatorname{rank}_r\left(d\right)}\) Reciprocal rank fusion combines ranked lists without assuming their raw scores share a scale. Section 5.4 gives the cited formula and limits.
Retrieval recall at k \(\mathrm{Recall}@k = \frac{\lvert R_k \cap G \rvert}{\lvert G \rvert}\) For a nonempty \(G\), the value shows how much known relevant evidence appears in the retrieved set. Section 5.6 defines the evaluation population and relevance labels.
Mean reciprocal rank \(\mathrm{MRR} = \frac{1}{\lvert Q \rvert} \sum_{i=1}^{\lvert Q \rvert} \frac{1}{\operatorname{rank}_i}\) Average of one over the position of the first relevant result; a question without one contributes zero. Section 5.6.
Normalized discounted cumulative gain \(\mathrm{nDCG}@k = \frac{\mathrm{DCG}@k}{\mathrm{IDCG}@k}, \quad \mathrm{DCG}@k = \sum_{i=1}^{k} \frac{2^{\mathrm{rel}_i}-1}{\log_2\left(i+1\right)}\) Graded relevance discounted by position and divided by the best possible top-\(k\) ordering from the query’s judged candidates. When \(\mathrm{IDCG}@k > 0\), the ratio is from 0 to 1. Section 5.6.
Classification accuracy \(\mathrm{Accuracy} = \frac{\mathrm{TP}+\mathrm{TN}}{\mathrm{TP}+\mathrm{TN}+\mathrm{FP}+\mathrm{FN}}\) For a nonempty evaluation set, the value summarizes agreement with labeled reference cases. Section 6.5 defines the labeled-case calculation.
F1 score \(\mathrm{F1} = \frac{2\mathrm{TP}}{2\mathrm{TP}+\mathrm{FP}+\mathrm{FN}}\) The count form balances classification precision and recall and exposes its zero-denominator case. Section 6.5 gives the detector example and the undefined-slice policy.
Repeated-run evaluator consistency \(C_i = \frac{\max_y \sum_{r=1}^{R} 1\left[\hat{y}_{i,r}=y\right]}{R}\) Modal vote share measures repeat stability, not independence or calibrated correctness confidence. Section 6.6 gives the cited evaluator context.
Cohen’s kappa for chance-adjusted agreement \(κ = \frac{p_o-p_e}{1-p_e}\) Observed agreement is adjusted for agreement expected from the two label distributions. Section 6.6 gives the cited definition and validity limit.
DAG critical-path lower bound \(T_{\mathrm{DAG}} \ge \max_{p \in P} \sum_{i \in p} t_i\) Fixed task durations and finish-before-start dependencies give a lower bound; resource waits can increase actual duration. Section 7.2 gives the cited method and assumptions.

Comparisons depend on population, aggregation window, and version context. Mean request latency and p99 latency answer different questions, the relevant-precision GPU peak differs from a nominal peak, retrieval scores change across embedding spaces, and throughput changes with workload shape.

A.1 Measurement details

The equations provide compact relationships. A metric remains ambiguous when its population, aggregation, scale, exclusions, or version context is missing. The following details make a dashboard or benchmark interpretable.

  • Decision: The measurement specifies the action or hypothesis it informs.
  • Population: The workload, tenant, model, release, hardware, corpus, or evaluation slice is explicit.
  • Scale and direction: Time or byte scales are explicit, together with whether larger or smaller is better.
  • Aggregation: Latency uses distributions and tail percentiles, rates retain counts and denominators.
  • Versions: Bundle, dataset, evaluator, embedding, index, code, and configuration identities stay attached.
  • Uncertainty: Sample count, variance, intervals, repeated runs, or sensitivity describe uncertainty where applicable.
  • Raw evidence: Task, trace, or evaluator output remains available under policy to explain aggregate changes.
  • Actionability: Each alert or release check has a responsible component and a plausible response.