19 Comparing AI deployment options
A deployment comparison holds task, data access, quality target, and action limits constant, then measures results and full operating cost for each option.
An organization choosing an AI model for a security task has three broad options. It can call a supplier’s model, run its own model on a managed platform, or run it on its own hardware. The labels hide the questions that matter. Who can inspect the model, who patches the runtime, and who sees the prompts and logs? What happens when the supplier fails, and what does each accepted result really cost? A choice made on purchase price or on a vendor’s model list can leave those duties unassigned.
Suppose the organization runs the same defensive task and workload under three arrangements. An external model API is operated by a supplier. A managed platform runs an organization-controlled model on rented infrastructure. An internal deployment runs the same model on the organization’s own infrastructure. Task inputs, expected results, allowed actions, data access, quality target, and oversight requirements stay constant where possible. Holding the task constant makes the options easier to compare. If the hosted service uses a different model, a quality difference can reflect the model as well as the hosting arrangement. The result then compares complete deployment options rather than isolating the effect of hosting alone.
19.1 Deployment rights and operating duties
A deployment label leaves inspection rights and operating duties unclear. A hosted model service provides model access through a supplier-operated interface. A self-hosted model deployment runs model software that loads specified weights on infrastructure operated or controlled by the adopting organization. The managed-platform option here is one form of self-hosting: the organization controls the model deployment while the supplier operates the underlying hosts and platform. The selected service contract determines who patches each software layer. For this comparison, the internal option specifically means infrastructure operated by the organization. Access to the weights, license rights, runtime control, dependencies, and update duties remain separate facts in all three arrangements.
CrowdStrike1 announced Charlotte AI AgentWorks on March 25, 2026, describing it as an ecosystem for building security agents. The announcement says customers can select among several models, including NVIDIA’s Nemotron models. A list of selectable models names what can run. It does not say who operates each model, who can inspect it, or who patches it. Choosing among those options requires facts that such an announcement leaves out.
Modification rights are permissions granted by a license or contract to alter and reuse supplied material. The Open Source Initiative definition2 goes further than access to model weights. It requires freedom to use, study, modify, and share an AI system, together with access to the preferred form for modification. This definition does not classify every downloadable model or settle how a license applies in a license dispute.
Security also depends on the operator who patches the runtime and controls the network, including who can observe prompts, retrieved context, model outputs, logs, and credentials. Identity, isolation, logging, incident response, and access to evidence follow the same operator boundary. Chapter 7 (Supply chain compromise), Chapter 8 (Isolation on shared infrastructure), and Chapter 9 (Runtime intrusion and resource abuse) describe these controls. Guidelines for Secure AI System Development3 treats model choice as a comparison across security, privacy, performance, and transparency, without ranking one hosting arrangement as generally best. These rights and duties define the feasible choices for the intervention comparison that follows, before quality or cost decides the outcome.
Example
Rights trace. The example organization compares the three arrangements for the fixed alert triage task introduced in Chapter 18 (Evaluating AI for defense). The input is redacted alert text, the target is a cited priority recommendation, and none of the three options may execute a response. Under the assumed external service contract, the organization may use the interface but receives no model weights. The supplier retains runtime patching and network control. The managed option permits organization control of the model while the supplier retains host patching and physical control. The internal option permits local runtime inspection and modification under its selected license, while the organization assumes patching, network, and recovery duties. The observed result is a deployment record with four fields: operating control, access to weights, modification rights, and update responsibility. The conclusion is that the words external, managed, and internal do not answer those four questions without the record.
Architecture reviewers check those four fields, the allowed data, the permitted model and platform operators, and the recovery duty before approving any option. They postpone the choice when a required right or operator is missing. The review reports complete deployment records over all candidates and the annual labor required to maintain each accepted option. Contract ambiguity and undocumented subcontracting can defeat the check. Recovery uses the tested alternate service or manual path while the organization resolves the missing right. A complete record does not prove that the supplier or internal operator will meet the stated duty.
19.2 Model and application adaptation choices
The deployment decision becomes confused when the organization treats every missing capability as a reason to train a model. The useful question is what is missing, because each possible change brings different security work. Prompting supplies task instructions with a request. Retrieval adds selected external material to the current context, which adds two questions: which sources are allowed, and whether retrieved data can be mistaken for instructions.
Tool access allows the application to call another function or service, which adds identity and action authorization. Fine-tuning changes model parameters through further training on task data. In supervised fine-tuning, each training example pairs an input with a target response, and a training objective guides parameter updates. Later inference uses the resulting fixed parameters to generate outputs. Evaluation compares those outputs with reference answers or failure rules without updating parameters. Fine-tuning therefore adds training-data controls, versioned parameters, and regression testing.
NIST’s generative AI profile4 says that intended-purpose analysis includes internal or external use, narrow or broad scope, fine-tuning, and retrieval-based data sources. The list identifies factors to examine but does not establish the right intervention for a task. No study cited in this chapter compares all four choices under the same task, model, data, and evaluation conditions.
These changes act at different points and can be combined. Instructions can specify how to use retrieved evidence, a read-only tool can supply a current record, and a fine-tuned model can still receive both. Lewis et al.5 combined retrieval with a trained generator and demonstrated updating answers by replacing its document index. Their Wikipedia question-answering experiment supports that mechanism, not a security or quality ranking for alert triage.
Example
Choosing a change for read-only triage. The example organization needs a cited priority recommendation from redacted alert text. Suppose development cases show that the model returns the required format but cites an obsolete escalation rule. Supplying the current approved rule in the same request produces the expected recommendation on those cases. This suggests a missing-information problem. The rules change weekly, and the full runbook exceeds the context allowance. Because these cases helped select the change, they cannot also serve as untouched evidence of its success.
The comparison keeps the same base model, redacted inputs, quality target, and prohibition on response actions. It asks how each change supplies current information or changes behavior, what data it exposes, and whether it adds authority.
| Change | Information and behavior | Data exposure | Authority in this pilot |
|---|---|---|---|
| Prompting | Instructions and examples can guide classification or format. Current rules must be supplied and updated with the request. | Included alerts, rules, and examples reach the selected model operator and any configured request logs. | No new service access. The application still prohibits response actions. |
| Retrieval | An updated index can supply current runbook passages and source versions. Retrieval quality and index freshness need tests. | Selected passages enter model context. Source permissions must limit selection before disclosure. | Read access to approved sources only. Retrieved instructions cannot grant authority. |
| Read-only tool access | A checked lookup can obtain current alert status or an exact record when the model requests it. | The service receives query fields, and its returned fields enter model context. Both need data limits. | A scoped identity permits only the approved reads. Response actions remain prohibited. |
| Fine-tuning | Training examples can change recurring classification or formatting behavior. Changing rules still need fresh context or a later update. | Task examples enter training and may be retained in records or learned by the model, adding privacy and regression checks. | Parameter changes grant no service permission. The same action prohibition remains. |
For this gap, the organization first tests retrieval of approved runbook passages together with an instruction to cite the version used. The application checks source access before placing passages in the model’s context. A current excerpt supplied directly in the prompt remains a simpler candidate if the rules fit and can be kept current reliably. A read-only lookup becomes useful if the recommendation also depends on live alert status. Fine-tuning becomes a candidate when correct context and adequate instructions still leave a recurring behavior problem, such as inconsistent priority classification. It does not provide an ongoing feed of changing rules.
Conclusion. Retrieval plus prompting is a design judgment to test for this information gap. Untouched cases must check current-rule use, correct citations, authorized retrieval, hostile retrieved instructions, and continued absence of response actions. No measured winner follows from the development trace. The selected change also determines the versions, evidence, and rollback material the adaptation workflow must keep.
Change reviewers record the observed gap and check whether the chosen change alters data access, parameters, or action authority. A pilot goes ahead only when the change has a matching test and rollback point. Results report accepted task cases and security failures over all pilot cases. More capable changes add integration, review, and recovery work. The selection can fail when the original gap is measured poorly or several changes are bundled without separate tests. Recovery disables the added capability or restores the previous model and application state. The pilot supports only the change it tested.
19.3 Adaptation and security regression
An adapted model or prompt can improve the target task while weakening a security property. That loss is a security regression, which occurs when a security property becomes worse after a change. Fine-tuning can weaken model safeguards even when task quality improves. Qi et al.6 observed this effect after adversarial and ordinary fine-tuning in the models and data they tested. The organization therefore needs a controlled path from baseline to candidate version.
NIST7 recommends renewed evaluation after a third-party model is fine-tuned and after adaptation to a new domain where earlier assumptions may fail. Guidelines for Secure AI System Development8 also states that changes to data, models, or prompts can alter system behavior. It recommends versioning major updates and retaining a route to a known-good state. A known-good state is the recorded version accepted for restoration. It extends the model baseline of Chapter 3 (Training poisoning and model tampering) to the complete system state. These recommendations do not specify one validation set or prove that re-evaluation prevents failure.
The process compares the candidate with the accepted baseline on a held-out test set. The clean holdout in Chapter 3 checks ordinary task performance after poisoning, while the evaluation holdout in Chapter 6 checks behavior on unseen cases. Here the held-out set compares versions after a model or application change. It contains examples excluded from training, prompt selection, and other tuning, and reserved for later evaluation. Cases repeatedly consulted while developing a fix become development evidence. A new estimate then needs untouched cases, as explained in Chapter 14. The comparison records every security regression. Its results cover each adversarial case, a test in which an actor intentionally tries to cause a security failure. The comparable evaluation in the next section reuses these cases. The release decision also references a rollback condition. A support assistant might, for example, pass ordinary answer tests after fine-tuning but reveal restricted case details under adversarial prompts. The fine-tuned version remains a candidate until the confidentiality regression is corrected and the version passes every mandatory threshold on the held-out task set and fixed security regression set. Understanding the failure alone does not satisfy these release conditions.
The deployment record links the base model, adaptation method, data version, and both task and safety regression results. When training occurs, it also identifies the training job and output weights or adapter, a separate set of learned parameters used with the base model. Prompt or retrieval changes instead record the changed instructions or index version and do not invent a training job. Isolation protects one organization’s adaptation data and credentials from another job. A comparison of parameter or adapter changes can help locate an unexpected change. Behavior tests remain necessary, because a small numeric difference does not state its effect.
The release pipeline enforces the thresholds. It checks the candidate version against the held-out task set and a fixed security regression set, which includes adversarial cases matched to the deployment’s allowed data and action limits. It accepts the version only when every mandatory threshold passes. The report gives failures over all cases for each set and records review time and training cost. Data leakage into the held-out set or an untested attack can create a false pass, so the recorded sets apply only within their data and attacker assumptions. Recovery restores the recorded known-good model, prompt, index, and runtime versions only where those versions remain available and the responsible operator can deploy them. If a supplier no longer offers the accepted version or the recorded state cannot be restored, the service uses the approved fallback only if it can enforce the required data access and action limits; otherwise, affected work pauses. The workflow produces versioned alternatives for a comparable evaluation under shared permissions and quality targets.
19.4 Comparable deployment evaluations
A deployment can look better simply because it received easier work, broader permissions, or a lower bar. A valid comparison gives each deployment the same task difficulty, permissions, and context allowance and measures it against the same quality target. A comparable evaluation tests alternatives with the same task set and decision criteria.
The AI Risk Management Framework9 calls for validity and reliability under expected deployment conditions and for security and resilience to be evaluated and documented. It does not provide a complete head-to-head protocol. A worked comparison under constant workload, permissions, context size, quality target, and scoring rules remains an open need. Chapter 14 (Evidence for release decisions) supplies the general evaluation method reused here.
For the example organization’s incident-triage task, all three candidates would receive the same redacted alerts, the same action prohibition, the same context allowance, and the same time window. Reviewers would measure accepted classifications against the shared quality target. They would also count unsafe disclosures, unauthorized action attempts, latency, and unavailable results, including results on the same adversarial cases. Each kind of failure then receives a failure cost, the operational consequence assigned to an incorrect, unavailable, or unauthorized result. The resulting rates remain specific to that test set, permissions, and operating setup. They supply the quality and failure evidence needed for the cost model.
The evaluation harness enforces the shared conditions. It checks that every deployment receives the same case identifiers, context allowance, permissions, time window, acceptance rules, and adversarial cases, and rejects the comparison when any of them differ. Each rate identifies the cases eligible for that outcome. For example, quality uses all assigned task cases, while attack success uses the assigned adversarial cases. Latency is reported for completed cases with failed or timed-out requests counted separately. Reviewer effort includes rejected results and corrections as well as accepted work. Cost per accepted result uses this total effort, so an option cannot look efficient by hiding the work spent on failures. A shared quality target is itself a test input: without it, a deployment could appear better by accepting lower quality. Comparable inputs may still omit a feature that matters in production or fail to represent future work. A failed comparison returns the pilot to the accepted baseline.
19.5 Cost per accepted result
A choice made on purchase price alone can be reversed once the system is running. A service with a low price per request can still cost more per useful result once review, failures, and staff time are counted. Unit cost divides the total cost recorded for a period by a defined unit of useful work.
FinOps Foundation10 describes unit economics as a relationship between technology cost and business outcomes. Example units include cost per token, request, customer, transaction, or case resolved. Its more mature view includes fully loaded costs. The FinOps for AI guidance11 notes costs from data centers and enterprise agreements. Software services, model suppliers, and several cloud providers can add further costs. These categories organize the calculation. Actual prices and economic value require deployment evidence.
Fully loaded cost adds assigned infrastructure, labor, and operating cost to direct service cost. For this deployment, the calculation adds service charges or annual hardware depreciation, storage, network, and energy. It allocates the paid capacity to the workload, including any capacity reserved for it but left idle. Idle capacity already included in service charges or depreciation is not added a second time. It also adds engineering labor, patching, monitoring, human review, and incident handling, including the staff and supplier cost of maintaining a managed platform. Capacity use is the share of available compute capacity doing useful work, so unused owned hardware still incurs depreciation and operating costs. The denominator counts only results that meet the shared quality target and acceptance criteria. If no results are accepted, cost per accepted result is undefined, and the record reports the total spend and zero accepted results. Actual prices, staffing effort, capacity use, and incident cost remain organization-specific inputs that require current evidence. This full cost view exposes the availability, fallback, and exit assumptions examined next.
Finance and service records supply the period cost, while the evaluation record supplies accepted work for the same period and the same quality target. A unit-cost comparison is rejected when the period, acceptance rule, or quality target differs. Collection and allocation add accounting effort. Shared infrastructure, deferred maintenance, and rare incident costs can remain hidden. When records expose a missing cost or uncertain allocation, the analyst recalculates a range and the service owner reopens the deployment decision if the cost exceeds the approved limit. A lower unit cost does not establish lower risk or better continuity.
Example
Comparable cost calculation. Suppose the external and internal options each process 600,000 alerts during one year of the same read-only triage workload. Each accepted result must meet the same review criteria, and both options meet the exercise’s assumed minimum completion requirement. The managed option has no cost or accepted-result record in this exercise, so its cost ranking is unknown. All costs are invented currency units, not current supplier prices. The internal hardware amount is annual depreciation allocated to this workload, not a purchase price mixed with annual charges. Labor for all processed alerts, including corrections and rejected results, is allocated to the listed categories once.
| Annual item | Hosted service | Internal self-hosted deployment |
|---|---|---|
| Annual service charge or allocated hardware depreciation | 180,000 | 150,000 |
| Storage, network, and energy | 24,000 | 72,000 |
| Engineering and patching | 55,000 | 145,000 |
| Monitoring and recovery tests | 35,000 | 60,000 |
| Human review | 120,000 | 120,000 |
| Total annual cost | 414,000 | 547,000 |
| Accepted alert results | 570,000 | 576,000 |
The hosted unit cost is 414,000 divided by 570,000, or about 0.73 currency units per accepted result. The internal self-hosted unit cost is 547,000 divided by 576,000, or about 0.95. Acceptance is 570,000/600,000 = 95 percent for the hosted option and 576,000/600,000 = 96 percent for the internal option. The internal option accepts 6,000 more results but costs 133,000 more. The listed costs favor the hosted option if both completion levels are acceptable. The table does not price the consequences of the remaining unaccepted alerts or supply a managed-platform comparison. Supplier charges, staffing, capacity use, quality, and exit work would need their own managed-option record. None of these factors implies that its total falls between the illustrated totals. Data-transfer limits, inspection needs, or a failed continuity test can still change the choice.
19.6 Deployment continuity and exit
An option with good test results and acceptable cost can still fail an operating requirement. A task may rely on an outbound dependency, an external service required during operation.
NIST12 recommends documenting dependence on third-party AI data and systems. It also recommends identifying a fallback, an alternate service or manual path used when the primary path fails, and testing rollover or manual procedures. Guidelines for Secure AI System Development13 also recommends recording where models, data, prompts, software, logs, and assessments are stored and retaining versions that permit restoration. The selected organization still needs legal analysis for location rules and system tests for continuity. A managed platform changes but does not remove these dependencies, because the host layer remains supplier-operated.
The final deployment record joins task quality under the shared quality target, rights, action limits, full cost per accepted result, required service level, fallback test results, and exit cost. Exit cost is the work and cost required to leave a supplier or deployment design. An internal model may remove an online supplier dependency while adding scarce operations staff and hardware recovery duties. A managed option reduces local hardware work but retains a supplier host dependency and its exit cost. A hosted service may reduce local maintenance while adding transfer and supplier-exit constraints. This bounded record becomes the input to organization-wide adoption in Chapter 20 (Shadow AI and unmanaged use).
The service manager tests supplier loss and internal runtime loss against the same workload. The application routes work to the approved manual or deterministic fallback when the health check fails only if that path can still enforce the required data access and action limits. If it cannot, affected work pauses. The test reports completed fallback cases over all cases arriving during the disruption, time to switch, backlog, and total staff hours. Standby capacity and rehearsals raise ongoing cost. A shared identity, network, or data dependency can disable both paths. Once the approved service is back, recovery reconciles duplicate work, and normal routing resumes after credential verification. One exercise does not prove continuity for every outage duration.
Note
Chapter checkpoint. Compare the two deployments in the cost table for a pilot that permits redacted alerts and prohibits response actions. It requires a successful manual-fallback test under the same quality target and permissions. The preceding text describes the test procedure but gives no result. What decision is supported now, what evidence is missing, and what change would reopen a later approval?
Answer. The hosted option is the lower-cost candidate under the exercise assumptions, but neither option is ready for approval on the evidence given. The required fallback result is missing. A recorded pass must show that the manual path handled the agreed outage workload within the required time, quality, data-access, and action limits, with staffing and any remaining backlog reported. If that test and the other release criteria pass, the hosted option can be chosen for this redacted, read-only pilot. A failed supplier-exit test, changed transfer rules, a need for local inspection, or a cost change would reopen that approval. The managed option remains unranked because its comparable cost and accepted-work records are absent. No general claim that hosted service is safer or cheaper follows.
CrowdStrike, “CrowdStrike Launches the Charlotte AI AgentWorks Ecosystem for Building Secure Agents,” press release (March 25, 2026), source. A vendor press release describes intended features and does not state the contract terms, operators, or update duties for each model.↩︎
Open Source Initiative, The Open Source AI Definition, version 1.0 (2024), “What is Open Source AI” and “Preferred form to make modifications to machine-learning systems,” source. The definition is a community standard. A legal conclusion about a selected license requires separate analysis.↩︎
UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (November 27, 2023), section 1, “Consider security benefits and trade-offs when selecting your AI model,” printed pp. 10–11, source.↩︎
NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (2024), MAP 1.1, action MP-1.1-001, printed p. 22, source.↩︎
Patrick Lewis et al. (2020), “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems 33, 9459-9474, sections 2 and 4.5, especially “Index hot-swapping,” PDF pp. 7-8, published paper. The study combines learned retrieval and generation and tests replacing a Wikipedia index. The triage selection and access restrictions are this guide’s worked design, not an experimental result from that paper.↩︎
Xiangyu Qi et al., “Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!,” International Conference on Learning Representations (2024), abstract and sections 3–5, source.↩︎
NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (2024), MAP 4.1, actions MP-4.1-007 and MP-4.1-008, printed p. 26, source.↩︎
UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (November 27, 2023), section 4, “Follow a secure by design approach to updates,” printed p. 16; section 2, “Identify, track and protect your assets,” pp. 12–13, source.↩︎
Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), Table 3, MEASURE 2.5 and MEASURE 2.7, printed pp. 29–30, source.↩︎
FinOps Foundation, Unit Economics, living framework (n.d.), “Definition,” “Define Unit Metrics which support Organizational Goals,” and maturity level “Run,” source. Actual costs depend on workload, supplier, and date.↩︎
FinOps Foundation, FinOps for AI, living framework (n.d.), “FinOps Considerations for AI,” source.↩︎
NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (2024), GOVERN 6.2, actions GV-6.2-001 and GV-6.2-006, printed pp. 21–22, source.↩︎
UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (November 27, 2023), section 2, “Identify, track and protect your assets,” printed pp. 12–13, source.↩︎