Self-Operated GenAI Reflex Area
2026-09-25
Table of Contents
- Defining the Need and Requirements for Self-Operated GenAI
- Choosing the Operating and Control Model
- Designing the GenAI Platform Architecture
- Planning Compute, Capacity, and Facilities
- Selecting and Managing GenAI Models
- Managing Data, Knowledge, and Retrieval
- Securing the Self-Operated GenAI Capability
- Meeting Classification, Privacy, and Compliance Requirements
- Planning the Sourcing and Procurement Approach
- Managing Suppliers, Licenses, and External Dependencies
- Implementing and Integrating the Capability
- Testing, Accrediting, and Accepting the Capability
- Operating, Monitoring, and Supporting the GenAI Service
- Defining Operational Responsibilities Across the GenAI Service
- Establishing Normal Day-to-Day Operation of the GenAI Service
- Monitoring Service Health Across Infrastructure, Platform, Models, and Integrations
- Monitoring Performance, Capacity, and Emerging Demand
- Monitoring Model, Retrieval, and Knowledge Quality
- Monitoring Operating Cost and Resource Efficiency
- Supporting Users and Resolving GenAI Service Issues
- Detecting and Triaging Production Anomalies
- Responding to Infrastructure and Platform Failures
- Responding to Model Failures
- Responding to Unsafe or Materially Harmful Outputs
- Responding to Suspected Security Compromise
- Responding to Data Exposure or Unauthorized Information Flow
- Responding to Integration or Tool Failures
- Coordinating Infrastructure, Security, Model, Data, and Application Teams During Complex Incidents
- Escalating to External Support Without Surrendering Operational Control
- Managing Updates, Maintenance, and Configuration Changes
- Maintaining Resilience and Continuity
- Operating GenAI in Restricted and Disconnected Environments
- Planning Capability Evolution and Technology Refresh
- Replacing and Decommissioning the Capability
Defining the Need and Requirements for Self-Operated GenAI
Determining Whether Self-Operation Is Actually Required
- What problem would self-operation solve that an externally operated GenAI service would not?
- Which information, mission, sovereignty, continuity, or control requirements actually require self-operation?
- What external access or dependency would be unacceptable for the intended uses?
- Could a managed private or externally operated solution satisfy the same requirements with less organizational burden?
- What additional responsibilities would the organization assume by self-operating the capability?
- Which claimed reasons for self-operation are genuine requirements, and which are preferences or assumptions?
- What evidence would justify choosing self-operation despite its lifecycle cost and complexity?
Clarifying the Mission Need the Capability Should Serve
- What mission or organizational problem should this capability improve?
- Which outcomes matter more than simply making GenAI available?
- Who experiences the current problem, and how does it affect their work?
- What existing process, tool, or service is currently used instead?
- Which limitations of the current approach are significant enough to justify a new capability?
- What would successful GenAI support change in observable terms?
- What need would remain even if current GenAI technology changed substantially?
Identifying GenAI Use Cases That Belong Inside the Capability
- Which recurring activities are strong candidates for the self-operated capability?
- What information does each use case require?
- What role should GenAI play in each use case?
- Which use cases share enough infrastructure, models, controls, and service requirements to belong on the same platform?
- Which use cases create materially different security, quality, latency, or availability demands?
- What volume and frequency make each use case operationally significant?
- Which proposed uses should be treated as experiments rather than production requirements?
Defining Which Uses Should Remain Outside the Capability
- Which GenAI uses do not require the controls provided by this capability?
- Which workloads would be cheaper or more effective on an external service?
- Which activities would create unacceptable legal, safety, security, or operational exposure if supported?
- Which proposed uses depend on capabilities the platform is not intended to provide?
- Where would supporting an additional use case distort the architecture or operating model?
- What criteria should determine when a request is redirected elsewhere?
- How should exceptions be handled without gradually expanding the capability beyond its intended scope?
Identifying the Users and Organizations the Capability Must Serve
- Who will use the capability directly?
- Which groups differ meaningfully in tasks, information access, language, location, or technical needs?
- How many people are likely to be active rather than merely eligible?
- Which organizational units need separate administration, policy, or resource controls?
- Are contractors, partners, or other organizations expected to use the capability?
- Which users require priority access during constrained operation?
- How could changes in the user population alter capacity, security, support, or governance requirements?
- What information types must users be able to submit or retrieve?
- What classification, sensitivity, privacy, or proprietary restrictions apply to that information?
- Which derived artifacts such as embeddings, logs, histories, or evaluation traces inherit sensitivity from the source material?
- Does the capability need to combine information with different access restrictions?
- Which information must never enter the capability?
- What retention, deletion, provenance, or handling requirements follow from the information involved?
- How would the required information boundary change the architecture or operating model?
Defining Availability and Continuity Expectations for Different Uses
- Which use cases require continuous availability and which can tolerate interruption?
- What operational consequences follow if the service is unavailable for minutes, hours, days, or longer?
- Which users or workloads should receive priority during degraded operation?
- What recovery time and data recovery expectations are actually justified?
- Which functions can be reduced or disabled while preserving useful service?
- What non-GenAI alternatives must remain available if the capability cannot be restored?
- How should availability requirements differ between normal operations, surge conditions, and major disruption?
Translating Expected Use Into Realistic Workload Requirements
- How many requests are likely to arrive during normal, peak, and surge periods?
- What concurrency matters more than the total number of registered users?
- How large are typical prompts, retrieved contexts, attachments, and generated outputs?
- Which workloads use long conversations, retrieval, tools, multimodal inputs, or other expensive features?
- How variable is demand across times, locations, user groups, and mission conditions?
- Which assumptions about workload are based on evidence and which still need measurement?
- What representative workload profile should later capacity and performance testing reproduce?
- What does an acceptable result look like for each important use case?
- Which errors are tolerable and which should count as critical failures?
- What response latency is acceptable for interactive and noninteractive work?
- Which tasks require strong factual grounding, source attribution, language quality, or structured output?
- How much variability between repeated responses can the workflow tolerate?
- What human review is required before outputs can influence consequential work?
- How will quality and performance requirements be measured rather than described only in general terms?
Reconciling Competing Mission, Security, Cost, and Service Requirements
- Which requirements currently conflict with one another?
- What additional cost or complexity follows from the strongest security and continuity requirements?
- Which service expectations would become unrealistic under the available infrastructure or budget?
- Where could a lower model tier, reduced context, narrower scope, or different operating mode resolve the conflict?
- Which requirements are mandatory and which are negotiable?
- What downside is accepted when one requirement is prioritized over another?
- Which unresolved tradeoffs must be decided before architecture or procurement begins?
Choosing the Operating and Control Model
Determining What the Organization Must Control Directly
- Which operational decisions must remain under organizational authority?
- Which systems, credentials, configurations, and artifacts require direct internal access?
- What must internal staff be able to diagnose, restore, change, or disable without supplier action?
- Which functions could be delegated without weakening the required level of control?
- What would happen if the organization lost access to a supplier for an extended period?
- Which control claims are meaningless unless the organization also has the skills and tooling to exercise them?
- What minimum set of authorities defines self-operation for this capability?
Distinguishing Self-Operation From Local Hosting
- Who actually operates the service even if the hardware is located on organizational premises?
- Who controls privileged administration, identity, keys, updates, recovery, and configuration?
- Does any external control plane remain necessary for normal operation?
- Can the organization continue operating if the vendor loses connectivity?
- Are licensing or activation mechanisms effectively giving an external party operational control?
- Which responsibilities remain external despite local infrastructure?
- Is the proposed arrangement genuinely self-operated or merely locally hosted?
Deciding Which Responsibilities Can Remain With External Suppliers
- Which responsibilities benefit materially from specialist external expertise?
- Which responsibilities would expose the organization to unacceptable dependence if outsourced?
- Can suppliers advise or support without holding routine privileged control?
- What information would external personnel need to perform their role?
- Which activities must remain possible when external support is unavailable?
- How will responsibility be divided during incidents, upgrades, and recovery?
- What internal capability is still required even when a supplier performs substantial work?
Defining Who Holds Privileged Administrative Authority
- Which roles require privileged access to infrastructure, platforms, models, data, and applications?
- Who can create, grant, revoke, and audit those privileges?
- Which privileged actions require dual control or additional approval?
- How are supplier and contractor privileges constrained?
- What emergency privileges are needed and how are they governed?
- How will administrative access work in disconnected or degraded environments?
- What evidence will show who exercised privileged authority and when?
Determining Whether External Control-Plane Dependencies Are Acceptable
- Which external services are required to administer or operate the capability?
- What happens if those services become unreachable?
- Does the external dependency control licensing, identity, keys, configuration, monitoring, recovery, or updates?
- How long could the capability continue safely without the external control plane?
- Can the dependency be hosted or replicated internally?
- Does the dependency undermine the intended security or sovereignty boundary?
- What compensating benefit would justify retaining it?
Defining How External Support Can Operate Without Taking Operational Control
- What support activities can suppliers perform without direct administrative access?
- What diagnostics can internal staff collect and share safely?
- Can suppliers provide procedures or remediation packages for internal execution?
- How should remote support be handled where remote access is allowed?
- What approvals and monitoring are required for temporary supplier access?
- How will external support work when the environment is disconnected?
- Can the organization resolve ordinary incidents without escalating to the supplier?
Determining Which Internal Competencies the Operating Model Requires
- Which technical and operational skills must exist internally from the first day of production?
- Who understands the complete architecture well enough to diagnose cross-layer failures?
- Can internal staff administer the platform, models, security controls, data services, and integrations?
- Who can evaluate model behavior after changes or incidents?
- Which skills currently exist and which depend entirely on suppliers?
- How many people are required to avoid critical knowledge being held by one individual?
- What knowledge transfer is necessary before the operating model can be considered viable?
Designing Operational Authority for Normal and Emergency Conditions
- Who can make routine operational decisions?
- Which changes require approval from another role or authority?
- Who can isolate systems, disable models, block integrations, or reduce service during an incident?
- What authority changes during a major security or continuity event?
- Who decides when degraded service is acceptable?
- How are conflicting operational, security, and mission priorities resolved?
- Are emergency authorities explicit enough to act quickly without creating uncontrolled privileges?
Assessing Whether the Proposed Operating Model Provides Meaningful Organizational Control
- Could internal teams start, stop, monitor, diagnose, restore, patch, and reconfigure the service themselves?
- Can the organization deploy and roll back models without supplier intervention?
- Can hardware be replaced and capacity added with internal authority?
- Can the organization maintain the service if supplier connectivity disappears?
- Are documentation, credentials, tools, and knowledge actually available to exercise these rights?
- Which important activities still require an external party to act?
- Does the remaining dependence fit the organization's intended meaning of self-operation?
- Which functions are required to deliver the intended GenAI services?
- What belongs in physical infrastructure, infrastructure software, model serving, GenAI services, applications, and enterprise integration?
- Which capabilities need dedicated components and which can use existing organizational services?
- What artifact repositories, evaluation services, security controls, and lifecycle tooling are required?
- Which proposed components exist only because a particular vendor architecture includes them?
- What can be omitted without reducing required capability or control?
- Where are responsibilities between components currently unclear?
- Which functions are common enough to provide once for multiple applications?
- Should model serving, retrieval, identity, observability, evaluation, or guardrails be shared?
- Where would shared services improve efficiency and operational consistency?
- Where would sharing create unacceptable coupling, contention, or security exposure?
- Which applications require independent scaling or change cycles?
- How should shared resources be governed when workloads have different priorities?
- What platform boundary would reduce duplication without creating a monolith?
Separating Applications From Model-Serving Infrastructure
- Which application functions should remain independent of the selected model runtime?
- Can an application change models without being substantially rewritten?
- Which model-specific assumptions have leaked into application logic?
- Should applications call a stable internal inference interface?
- What behavior differences would still need application-level handling even behind a common API?
- How can model routing or fallback occur without changing every application?
- What separation would make future model replacement materially easier?
Separating Control-Plane Functions From Inference Workloads
- Which functions administer the environment and which process user content?
- What security boundary should exist between administration and inference?
- Can compromise of an inference workload expose control-plane authority?
- Which management services need access to production content, if any?
- How should identity, secrets, deployment, monitoring, and configuration flow between the two planes?
- What must remain operable if the inference plane is degraded?
- Does the proposed architecture create unnecessary privileged pathways into production?
- Which interfaces are likely to remain stable even as models or runtimes change?
- What internal API should applications depend on?
- Which model-specific features should be exposed and which should be abstracted?
- How will version differences in tools, context handling, streaming, structured output, or embeddings be represented?
- What assumptions would make nominally compatible APIs behave differently?
- How will interfaces be versioned and retired?
- Can a replacement component be introduced without forcing simultaneous changes across the stack?
- Which functions should be centralized at platform level and which belong inside applications?
- Where should access control be enforced so it does not depend on model behavior?
- Which retrieval capabilities are reusable across applications?
- Who should control tool definitions and permissions?
- Where should prompts, routing, safety controls, and workflow logic be managed?
- What would become difficult to change if too much logic is embedded in a single layer?
- Which placement gives the clearest ownership and security boundary?
Designing Separate Environments for Development, Evaluation, and Production
- Which environments are needed to develop, integrate, evaluate, accredit, and operate the capability safely?
- What data may be used in each environment?
- How should artifacts move between environments?
- Which production controls must be reproduced during evaluation?
- What differences between test and production could invalidate results?
- Who can promote models, software, prompts, and configurations into production?
- Are additional environments required for classified, sensitive, or disconnected work?
Deciding Which Components Should Be Modular and Which Should Be Integrated
- Which components are likely to change at different rates?
- Where would modularity create useful substitution or independent scaling?
- Where would it merely increase integration and operational burden?
- Which tightly integrated components provide a meaningful reliability or support advantage?
- What lock-in follows from choosing an integrated stack?
- Does the organization have enough engineering capability to operate a more modular design?
- What level of modularity is justified by realistic replacement needs rather than theoretical flexibility?
Identifying Architectural Dependencies That Could Prevent Later Replacement
- Which component currently depends on proprietary interfaces, formats, or control planes?
- What data or configuration would be difficult to export?
- Are indexes, prompts, tools, policies, and evaluation assets portable?
- Which replacement would force changes in several other layers?
- Does any supplier hold information or access necessary for migration?
- What part of the architecture has never been tested with an alternative component?
- What changes now would materially reduce future replacement cost?
Reviewing Whether the Architecture Has Become More Complex Than the Mission Requires
- Which components and interfaces are essential to the required capability?
- What complexity was introduced for hypothetical future needs?
- Which layers duplicate functions already available elsewhere?
- Where is operational effort growing without a corresponding mission benefit?
- Has modularity become fragmentation?
- Could a simpler architecture meet the same security, resilience, and replacement requirements?
- Which complexity should be removed before it becomes an enduring operational dependency?
Planning Compute, Capacity, and Facilities
Estimating Compute Requirements From Real Workload Characteristics
- Which representative workloads should capacity planning be based on?
- How do prompt length, output length, concurrency, retrieval, tools, and modality affect compute demand?
- Which models and numerical precisions are expected to run?
- What latency and throughput must the service achieve?
- How much demand comes from supporting workloads such as embeddings or evaluation?
- Which workload assumptions still need measurement through benchmarking?
- What validated serving profile should replace simple estimates based on user count?
Determining Whether a Model Configuration Fits the Available Accelerator Memory
- How much memory do the model weights require at the intended precision?
- What additional memory is needed for KV cache, activations, buffers, and runtime overhead?
- How does expected context length affect memory use?
- How many concurrent sequences must fit?
- Does a mixture-of-experts model require substantially more resident memory than its active parameter count suggests?
- What parallelism or quantization would be required to fit the configuration?
- Does fitting the model leave enough headroom for reliable production operation?
Sizing Capacity for Concurrent Users and Long Contexts
- What concurrency should the platform support under realistic usage?
- How long are prompts and conversations likely to become in practice?
- Which requests create unusually large KV-cache requirements?
- How does increasing context affect throughput and latency?
- Would batching improve utilization without violating interactive latency targets?
- Should long-context workloads be isolated or governed separately?
- What limits should be set if unrestricted context use would undermine service for other users?
Determining the Capacity Needed for Expected Peak Demand
- When and why does demand peak?
- How much higher is peak concurrency than normal concurrency?
- Which workloads dominate during peak periods?
- What queueing delay is acceptable before the service is considered degraded?
- Can noncritical work be deferred during predictable peaks?
- What reserve should exist for uncertainty in the forecast?
- How should peak requirements be validated before committing to infrastructure?
Planning for Mission-Critical Surge Demand
- What events could create demand far above normal peaks?
- Which workloads and users must retain service during a surge?
- How much capacity should be reserved rather than shared normally?
- Could lower-cost models, shorter contexts, or reduced functionality preserve priority service?
- Which noncritical workloads can be queued or suspended?
- How long must the surge mode be sustainable?
- What exercise would demonstrate that the planned surge controls actually work?
Choosing How Models Should Be Replicated or Distributed Across Accelerators
- Can the model run efficiently on one accelerator or one server?
- When is replication preferable to splitting a model across devices?
- Does tensor, pipeline, or expert parallelism introduce unacceptable communication overhead?
- What interconnect performance does the selected topology require?
- How does the topology affect failure domains and maintenance?
- Can replicas be assigned to different workload priorities or security boundaries?
- Which topology performs best under the actual request distribution rather than synthetic maximum throughput?
Determining Whether Multiple Model Tiers Could Reduce Infrastructure Demand
- Do all workloads genuinely require the largest available model?
- Which tasks can meet their quality requirements with smaller models?
- Could routing different requests to different models reduce cost or improve latency?
- What errors could result from choosing the wrong model tier?
- How would routing decisions be governed and evaluated?
- Could a smaller model preserve useful capability during degraded operation?
- Does maintaining several model tiers create more complexity than the saved capacity justifies?
Assessing Whether Existing Power, Cooling, Racks, and Networking Can Support the Design
- What power density will the proposed accelerator systems require?
- Can existing cooling handle sustained production load?
- Are rack space, weight, cabling, and physical security adequate?
- What internal network bandwidth and latency are needed between accelerators, storage, and supporting services?
- Does the facility have enough electrical and cooling redundancy?
- What upgrades would be required before additional compute can be installed?
- Are facility constraints likely to become the real limit on future model or capacity choices?
Planning Capacity Expansion Under Uncertain Future Demand
- Which demand drivers are most likely to change?
- How much infrastructure can be added without redesigning the platform or facility?
- What lead times apply to accelerators, servers, power, cooling, and procurement?
- Which headroom is worth maintaining now and which capacity can be deferred?
- Could technology improvements reduce future hardware needs?
- What scenarios should be planned instead of relying on one growth forecast?
- What expansion path preserves options without buying speculative inventory too early?
Reassessing Capacity as Models, Workloads, and Serving Efficiency Change
- Which assumptions from the original capacity model are no longer true?
- Have request volumes, contexts, outputs, or concurrency changed?
- Has a new runtime or model changed memory and compute efficiency?
- Are observed bottlenecks where the original design expected them?
- Is current utilization evidence of genuine capacity shortage or inefficient scheduling?
- Should the response be additional hardware, workload controls, model changes, or software optimization?
- When should the validated serving profile be rebenchmarked?
Selecting and Managing GenAI Models
Comparing Candidate Models Against Real Mission Tasks
- Which real tasks should candidates perform during comparison?
- What quality dimensions matter for those tasks?
- Which candidate behaves best across normal cases, difficult cases, and important failure cases?
- How do results differ across relevant languages, domains, and input lengths?
- What infrastructure and latency does each candidate require?
- Which weaknesses can be mitigated elsewhere in the system and which are fundamental?
- What evidence is strong enough to narrow the candidates for deeper evaluation?
Determining Whether Public Benchmarks Are Relevant to the Intended Use
- What capability does each benchmark actually measure?
- How similar are the benchmark tasks to our real workloads?
- Does the benchmark evaluate the language, domain, context length, or tool use we need?
- Could benchmark contamination or optimization distort the apparent result?
- Are small score differences operationally meaningful?
- Which benchmarks are useful only for initial screening?
- What internal evaluation is still required before any model decision can be made?
Evaluating a Model Across Required Languages and Specialist Domains
- Which languages and specialist domains matter in actual use?
- Does quality remain stable across those languages and domains?
- Where does terminology or domain reasoning break down?
- Are safety behavior and refusals consistent across languages?
- Does retrieval compensate for missing domain knowledge without fixing deeper reasoning weaknesses?
- Which users can judge specialist quality reliably?
- What minimum performance is required before the model can serve a particular user group?
Deciding Whether Prompting, Retrieval, Adapters, or Fine-Tuning Are Needed
- What specific deficiency are we trying to correct?
- Does the model already have the capability but need clearer instructions?
- Is the problem missing or changing knowledge that retrieval could supply?
- Does the desired behavior require systematic adaptation rather than prompt changes?
- Would a lightweight adapter be sufficient?
- What evidence would justify the additional complexity of fine-tuning or continued pretraining?
- What is the least invasive intervention that can solve the actual problem?
Qualifying a New Model for Production Use
- Which mission tasks must the model pass before approval?
- Which critical failure modes are disqualifying regardless of average performance?
- Has the exact model package and runtime configuration been evaluated?
- How does the model behave with our prompts, retrieval, safety controls, and tools?
- What security and provenance evidence exists for the model artifacts?
- What operational limits or approved use cases should accompany qualification?
- What evidence must be recorded so the qualification can later be reproduced or challenged?
Managing Model Weights, Tokenizers, Templates, and Runtime Dependencies Together
- Which artifacts are required to reproduce the approved model behavior?
- Is the tokenizer version controlled with the weights?
- Which chat template, system prompt, inference settings, and runtime are part of the approved configuration?
- What happens if one dependency changes while the weights remain the same?
- Are all required artifacts stored internally and integrity protected?
- Can the complete package be redeployed without retrieving anything from an external service?
- How will provenance and compatibility be recorded across the package?
Maintaining Approved Model Versions for Different Workloads
- Which model versions are approved for which use cases?
- Why are multiple versions being retained?
- How are applications prevented from silently using an unapproved model?
- Which version should be the default and which remain fallback options?
- How long must previous versions remain deployable?
- What evidence is needed before an approved version is withdrawn?
- How will version ownership and review dates be maintained?
Investigating a Model That Behaves Differently After a Change
- What behavior changed and when was it first observed?
- Which model, runtime, prompt, retrieval, adapter, tool, or configuration changed at the same time?
- Can the difference be reproduced on a known evaluation set?
- Is the change systematic or limited to particular tasks or inputs?
- Did infrastructure or serving settings alter output behavior?
- Does rollback restore the previous behavior?
- What evidence identifies the actual changed layer rather than assuming the model itself is responsible?
Deciding Whether a Newer Model Is Actually Better for the Mission
- What concrete improvement does the newer model claim to provide?
- Does that improvement appear on our own mission tasks?
- What new weaknesses or behavioral changes appear?
- How do infrastructure, latency, licensing, and support requirements change?
- Would adoption force changes to prompts, retrieval, applications, or accreditation?
- What migration and regression-testing effort would the change create?
- Is the operational benefit large enough to justify replacing the approved model?
Removing a Model From the Approved Operational Portfolio
- Why is the model no longer suitable for approved use?
- Which applications and users still depend on it?
- Is there an approved replacement for every required workload?
- What evidence supports withdrawal rather than continued restricted use?
- How should the model be removed from routing, deployment, and administrative interfaces?
- Must the artifacts remain retained for rollback, audit, or reproducibility?
- What downstream evaluations or documentation need updating after removal?
Managing Data, Knowledge, and Retrieval
- Which information would materially improve the intended GenAI tasks?
- Who owns each source and can authorize its use?
- How sensitive, current, complete, and reliable is the information?
- Which sources contain information users should not receive through GenAI?
- Are there restrictions on indexing, transformation, adaptation, or derived representations?
- Which sources are important enough to justify continuous ingestion and maintenance?
- What information should remain outside the GenAI environment even if technically accessible?
Determining Which Sources Are Authoritative Enough for Retrieval
- Which source should be treated as authoritative when several contain similar information?
- Who is responsible for keeping the source correct and current?
- Does the source distinguish approved information from drafts or informal material?
- How should conflicting sources be ranked or presented?
- Can retrieved content retain enough provenance for users to judge it?
- Which sources should be excluded because their status is unclear?
- What process should change retrieval priority when organizational authority changes?
Preparing Documents and Data for Reliable GenAI Use
- What structure does the source material currently have?
- Which formatting, duplication, scanning, or parsing problems would reduce retrieval quality?
- What metadata is needed to preserve source, date, owner, classification, and access?
- How should documents be divided without separating important context?
- Which content should be normalized and which must remain unchanged?
- How will preparation avoid creating unauthorized copies or losing provenance?
- What sample evaluation can show whether the prepared data is actually usable?
Preparing and Governing Data for Model Adaptation and Evaluation
- What data is needed for adaptation or evaluation and why?
- Who has authority to permit that use?
- Does the dataset contain sensitive, personal, copyrighted, classified, or otherwise restricted material?
- How will provenance, version, inclusion criteria, and transformations be recorded?
- Could training and evaluation data overlap in a way that invalidates results?
- How will datasets be corrected, withdrawn, retained, or superseded?
- What evidence allows a future model result to be traced back to the dataset version used?
Preserving Source Permissions Through the Retrieval Pipeline
- Which user permissions apply to every source?
- Are those permissions preserved when content is copied, chunked, indexed, or embedded?
- Can retrieval enforce access before content is exposed to the model?
- What happens when a user's access changes after indexing?
- Could shared indexes leak information across teams or compartments?
- How are group membership and identity synchronized with retrieval controls?
- What test would demonstrate that users cannot retrieve information they are not authorized to see?
Deciding When Structured Access Is Better Than Vector Retrieval
- Is the user asking for semantic context or precise current data?
- Does the source already expose reliable structured queries or APIs?
- Would vector retrieval introduce ambiguity where exact fields are required?
- Does the task depend on current values that change frequently?
- Would a tool call provide stronger authorization and auditability?
- Can structured and semantic retrieval be combined safely?
- Which access method gives the most reliable answer for the actual task?
Managing Conflicting, Outdated, or Superseded Knowledge
- How can the system detect that two sources disagree?
- Which source should prevail and why?
- Can users see when retrieved information is old or superseded?
- How quickly must a changed source propagate into retrieval?
- Should obsolete content be deleted, archived, or retained with lower priority?
- Who resolves conflicts when authority is unclear?
- How will the system avoid presenting a mixture of old and current rules as one coherent answer?
Detecting When Retrieval Quality Has Deteriorated
- What signs indicate that relevant information is no longer being retrieved reliably?
- Have source content, permissions, chunking, embeddings, or ranking changed?
- Are users seeing irrelevant, incomplete, or stale context more often?
- Which retrieval evaluation cases should remain stable over time?
- Can the problem be separated from a generation-quality problem?
- What operational metrics or user reports should trigger investigation?
- What known-good retrieval configuration can be used for comparison or rollback?
- What source information can be inferred from the embeddings or index?
- What classification or sensitivity should the derived artifacts inherit?
- Who may access, copy, export, or administer them?
- Are indexes shared across security or organizational boundaries?
- What happens to derived artifacts when the source is deleted or reclassified?
- Can the index be reconstructed from retained source data if necessary?
- Which security and retention controls should apply even though the artifacts are not human-readable documents?
- Where has the information been copied or transformed?
- Which chunks, indexes, embeddings, caches, prompts, histories, and backups may still contain it?
- How quickly must removal take effect?
- Can the item be removed without rebuilding the entire knowledge store?
- How will deletion be verified rather than assumed?
- What retained copies are legally or operationally required?
- How should future ingestion prevent a withdrawn source from reappearing?
Rebuilding or Migrating Retrieval Systems Without Losing Knowledge Control
- Which source data, metadata, permissions, and configurations must move?
- Can indexes be regenerated rather than transferred?
- How will old and new retrieval results be compared?
- What permission differences could appear after migration?
- Which provenance and audit records must remain intact?
- How will applications transition without losing access to required knowledge?
- What evidence is needed before the old retrieval system can be retired?
Securing the Self-Operated GenAI Capability
Building a Threat Model for the Complete GenAI Capability
- What assets would an attacker most want to access, manipulate, or disrupt?
- Which entry points exist across users, applications, APIs, models, data, suppliers, and administrators?
- What can an attacker influence through prompts, retrieved content, tools, or imported artifacts?
- Which model behavior could turn manipulated information into operational impact?
- What authority would the compromised workflow have beyond generating text?
- Which threats are ordinary infrastructure threats and which are specific to GenAI?
- What attack paths deserve priority because they combine plausible access with serious consequence?
- Which hosts, services, containers, APIs, and management interfaces are exposed?
- Are systems patched and configured against an approved secure baseline?
- How are network boundaries and service-to-service access restricted?
- Where are credentials, certificates, secrets, and encryption keys stored?
- Which unnecessary services or administrative paths can be removed?
- How are platform vulnerabilities detected and remediated?
- What evidence shows that the surrounding system is secure even before GenAI-specific controls are considered?
Protecting Model Artifacts Against Tampering and Unauthorized Replacement
- Where do model weights, tokenizers, adapters, and related artifacts come from?
- How is provenance verified before import?
- Are hashes or signatures checked at each controlled transfer?
- Who can approve, store, modify, or deploy model artifacts?
- Could a malicious or accidental replacement occur without detection?
- How are approved and quarantined artifacts separated?
- Can the deployed artifact always be traced back to the exact approved package?
Limiting the Consequences of Prompt Injection
- Where can untrusted instructions enter the model context?
- What authority could a successful injection cause the model to misuse?
- Which security decisions are currently delegated to natural-language instructions?
- Can tools and data access be constrained independently of the model's response?
- How should suspicious or adversarial content be handled without assuming it can always be recognized?
- What actions require explicit confirmation or deterministic policy checks?
- If prompt injection succeeds, what limits the resulting impact?
Preventing Retrieved Content From Becoming Trusted Instructions
- Which retrieved sources may contain untrusted or attacker-controlled text?
- Does the system distinguish source content from trusted system instructions?
- Can retrieved content alter tool use or security-sensitive behavior?
- What data and actions remain inaccessible regardless of retrieved instructions?
- Should particular sources be treated as higher risk?
- How can citations or provenance help users recognize the origin of generated claims?
- What tests would show whether malicious retrieved content can influence privileged behavior?
- Which tools can the model call?
- What permissions does each tool actually hold?
- Can a user or malicious prompt cause actions outside the user's own authority?
- Which operations require deterministic validation or human approval?
- Are read and write capabilities separated where possible?
- How are tool calls logged, rate limited, and revoked?
- What is the maximum plausible consequence if the model chooses the wrong tool or arguments?
- At which points can sensitive information leave its intended boundary?
- Can prompts, outputs, retrieval context, tool results, logs, or telemetry expose protected data?
- Does one user's conversation influence another user's context or memory?
- Are administrative and support interfaces able to view sensitive content?
- Which external integrations receive data from the GenAI platform?
- How are output controls and user permissions enforced?
- What test cases would expose cross-user, cross-domain, or unintended external leakage?
Protecting the Capability Against Poisoned Models, Data, and Dependencies
- Which externally sourced artifacts can alter system behavior?
- How is each artifact's provenance established?
- What inspection is possible before an artifact enters the trusted environment?
- Could malicious data influence retrieval, adaptation, or evaluation results?
- How are software dependencies and containers checked for known vulnerabilities or tampering?
- Which changes require renewed behavioral or security testing even when malware scans are clean?
- What rollback path exists if a trusted artifact later proves compromised?
Restricting Privileged Access to GenAI Infrastructure and Administration
- Which administrative functions require the highest privilege?
- Are privileges limited to the people and systems that genuinely need them?
- Can normal user accounts ever reach administrative interfaces?
- How are privileged sessions authenticated, logged, and reviewed?
- What separation of duties is required for sensitive changes?
- How are temporary supplier or emergency privileges revoked?
- Could compromise of one administrative account expose the entire capability?
Protecting GenAI Logs, Prompts, Outputs, and Other Sensitive Operational Artifacts
- What sensitive information may appear in operational logs?
- Which prompts, outputs, evaluation traces, and incident records need retention?
- Who can access those records?
- Can monitoring remain useful without collecting unnecessary content?
- How are logs protected against alteration and unauthorized export?
- What retention and deletion rules apply to different artifact types?
- How will incident investigation obtain necessary evidence without turning logging into a new information-exposure risk?
Meeting Classification, Privacy, and Compliance Requirements
Determining Which Classification Levels the Capability May Process
- Which classification levels must the capability support?
- Does each level require a separate environment or security domain?
- What infrastructure and accreditation requirements follow from the highest permitted level?
- Which model, data, logging, and support practices change with classification?
- Can the same capability safely support differently classified workloads?
- What information must never cross between domains?
- How will users know which environment is authorized for the material they are handling?
- Which users may access each information compartment or source?
- Can people with the same classification still have different access rights?
- How should retrieval enforce compartment-specific permissions?
- Could prompts or outputs combine information from compartments the user should not see together?
- How are changes in assignments and access reflected in the GenAI system?
- What audit evidence is needed for compartmented information use?
- Where is physical or logical separation preferable to complicated permission logic?
Designing Controlled Information Flows Across Security Domains
- What information genuinely needs to move between security domains?
- In which direction may information flow?
- What review, transformation, or release controls apply before transfer?
- Which GenAI artifacts such as prompts, outputs, indexes, models, or logs may cross the boundary?
- Can any automated interface preserve the required security properties?
- Who authorizes exceptions or changes to the flow?
- How will the organization detect information crossing outside the approved path?
Determining How Generated Outputs Should Be Classified or Handled
- What source information influenced the output?
- Can the output reveal classified or compartmented information even if it contains no direct quotation?
- Who is responsible for determining the output's handling level?
- Can users reliably distinguish generated content from formally reviewed material?
- What markings or metadata should follow the output?
- When must generated material be reviewed before leaving the environment?
- How should uncertain classification be handled without relying on the model to decide?
Assessing Whether Personal Data Can Be Used in the Intended GenAI Workflow
- What personal data enters the workflow and for what purpose?
- Is that use necessary for the intended outcome?
- What legal basis or organizational authority applies?
- Which parties or systems can access the data?
- Does the workflow create new profiles, inferences, or sensitive derived information?
- How long must prompts, outputs, logs, or retrieval artifacts containing personal data be retained?
- What alternative design could reduce personal data use without undermining the mission?
Designing Retention and Deletion Rules for Prompts, Outputs, and Logs
- Which records are operationally necessary to retain?
- Which records are required for security, audit, legal, or mission purposes?
- What sensitive information is likely to appear in each artifact type?
- How long should each category remain available?
- Can a deletion request be propagated through histories, indexes, caches, backups, and replicas?
- Which records must be immutable and which should be routinely purged?
- How will the organization prove that the configured retention policy is actually being enforced?
Determining Which AI-Specific Regulatory Requirements Apply
- Which jurisdiction and sector govern the intended use?
- What legal role does the organization hold in relation to the GenAI system?
- Does the use fall into a regulated or specially restricted category?
- Which obligations depend on the intended purpose rather than the underlying model alone?
- Are transparency, documentation, risk management, human oversight, or evaluation obligations triggered?
- Which requirements apply to the organization and which remain with suppliers?
- What legal uncertainty needs specialist interpretation before deployment proceeds?
Integrating GenAI Into Existing Cybersecurity and Compliance Frameworks
- Which existing controls already apply to this capability?
- Where do GenAI-specific risks require additional controls rather than a separate governance structure?
- Which existing asset, change, incident, access, supplier, and continuity processes can be reused?
- What evidence will existing auditors or security authorities expect?
- Which current framework assumptions do not fit probabilistic model behavior?
- How can responsibilities be integrated without creating duplicate approval processes?
- Where is a dedicated GenAI control necessary because existing processes leave a real gap?
Preparing the Evidence Required for Privacy, Audit, or Regulatory Review
- What decision or obligation must the evidence support?
- Which system description, data flow, model information, controls, tests, and records are required?
- Can the evidence be traced to the exact deployed configuration?
- Which claims depend on supplier documentation and which have been independently verified?
- Are residual risks and limitations stated clearly?
- What operational records need to be retained continuously rather than assembled later?
- Could an independent reviewer reproduce the basis for the organization's conclusion?
Resolving Conflicts Between Operational Use and Compliance Requirements
- Which operational need conflicts with which specific requirement?
- Is the conflict real or based on an overly broad interpretation?
- Could the workflow, data boundary, retention rule, or human review process be redesigned?
- What operational value would be lost by complying with the stricter interpretation?
- What risk would be accepted by using an exception?
- Who has authority to decide the tradeoff?
- What evidence should be recorded if the capability proceeds under a constrained or exceptional arrangement?
Reassessing Compliance When Models, Data, or Intended Uses Change
- What changed in the model, data, workflow, user group, or intended purpose?
- Does the change alter the organization's legal or regulatory role?
- Are previous privacy and compliance assessments still applicable?
- Does the change introduce new personal data, automated decisions, or sensitive outputs?
- Which documentation and controls need updating?
- Is renewed review required before the change reaches production?
- What change threshold should automatically trigger compliance reassessment?
Planning the Sourcing and Procurement Approach
Deciding What the Organization Should Build, Buy, or Integrate
- Which parts of the capability are commodity and which are mission-specific?
- Where does the organization need direct architectural or intellectual control?
- Which components can be purchased with acceptable dependency?
- What would internal development cost to sustain over the full lifecycle?
- Where does integration capability matter more than component ownership?
- Which responsibilities should remain internal even when technology is purchased?
- What combination of internal development, commercial products, and external services best fits the intended operating model?
Choosing Between an Integrated Solution and a Modular Procurement
- What integration burden would a modular approach create?
- What supplier dependence would an integrated solution create?
- Which components are likely to need independent replacement?
- Does one supplier have credible responsibility for end-to-end performance?
- Can the organization manage interfaces between multiple vendors?
- How would each approach affect accreditation, support, and future competition?
- Which risks are reduced and which are merely transferred under each option?
Deciding Whether a Prime Contractor or Multiple Specialist Suppliers Are More Appropriate
- Does the organization need one party accountable for system integration?
- Which specialist capabilities would be weakened by routing everything through a prime contractor?
- Can internal teams coordinate multiple suppliers directly?
- How would responsibility be assigned when failures cross supplier boundaries?
- What commercial margin and dependency does a prime arrangement create?
- Could a prime contractor restrict direct access to component vendors or technical knowledge?
- Which arrangement preserves sufficient organizational ownership of architecture and operations?
Structuring Procurement Lots Without Creating Integration Gaps
- Which components form natural procurement boundaries?
- Which interfaces must be specified across lots?
- Who is responsible for end-to-end integration?
- Could separate contracts leave important work outside every supplier's responsibility?
- Which dependencies need coordinated schedules and acceptance criteria?
- How will changes in one lot be managed when they affect another?
- Does the lot structure preserve competition without fragmenting accountability?
Using Market Engagement to Test Whether Requirements Are Realistic
- Which requirements are most uncertain from a market perspective?
- Can suppliers meet the required security, disconnected operation, support, and lifecycle conditions?
- Which requirements significantly narrow the supplier market?
- What alternative architectures or delivery models are available?
- Are suppliers interpreting key requirements differently?
- Which claims require evidence rather than marketing statements?
- What should be changed in the specification after market engagement, and what must remain non-negotiable?
Designing Competitive Evaluations Around Mission Workloads
- Which real workloads should suppliers demonstrate?
- What information and environment are needed for a fair comparison?
- Which quality, latency, security, operability, and resilience dimensions should be measured?
- How will probabilistic model behavior be evaluated consistently?
- What prevents a supplier from optimizing only for a scripted demonstration?
- Which claims should require hands-on verification by the organization?
- How will evaluation results influence the award rather than remain an informal impression?
Deciding Whether Prototypes or Pilot Deployments Are Needed Before Award
- Which uncertainties cannot be resolved from documentation or demonstrations?
- What should a prototype prove about technology, integration, operations, or supplier capability?
- How representative must the pilot environment be?
- What data and workloads can safely be used?
- What exit criteria distinguish useful learning from an endless pilot?
- How will pilot results be incorporated into the procurement decision?
- Does the pilot create unfair incumbent advantage or technical lock-in before competition is complete?
Comparing Acquisition Options Using Whole-Lifecycle Cost
- What acquisition, implementation, licensing, support, staffing, facility, and energy costs does each option create?
- What internal integration and operating effort is required?
- How do expected hardware and model refreshes affect cost?
- What does accreditation, testing, training, and documentation add?
- What exit or migration costs are likely later?
- Which costs are fixed, usage dependent, or uncertain?
- Does the apparently cheaper option remain cheaper over the period the capability is expected to operate?
Defining Requirements Across Rapid Technology Change
- Which requirements express enduring mission outcomes rather than current technical implementations?
- Where is technical specificity necessary for interoperability, security, or compatibility?
- Which named technologies could become obsolete before deployment?
- Can performance requirements be expressed through workload and service outcomes?
- How much substitution should suppliers be allowed during delivery?
- What changes should require organizational approval even if they meet nominal specifications?
- How can the procurement remain adaptable without making requirements unverifiable?
Designing the Procurement So Future Competition Remains Possible
- Which contractual or technical choices would make the incumbent difficult to replace?
- What interfaces, data, documentation, and configuration must remain under organizational control?
- Are future competitors able to understand and integrate with the delivered architecture?
- Which rights are needed to continue operating while a replacement is procured?
- Could proprietary formats or bundled licenses block later competition?
- What knowledge transfer must occur during the contract rather than only at exit?
- What evidence would show that a credible replacement procurement remains possible?
Managing Suppliers, Licenses, and External Dependencies
Mapping the Suppliers and Dependencies Behind the Delivered Capability
- Which suppliers contribute to each layer of the capability?
- Who owns the underlying technology when a reseller or integrator is the contractual supplier?
- Which external services are required for operation, support, updates, or licensing?
- Who holds privileged access or unique technical knowledge?
- Which dependencies sit behind the visible supplier?
- What happens operationally if each dependency disappears?
- Which dependencies are critical enough to require specific continuity plans?
Reviewing Licensing and Entitlement Terms Across the Capability
- How is each model, platform, runtime, driver, or other component licensed?
- Are licenses based on users, processors, accelerators, nodes, sites, environments, or another metric?
- Do development, evaluation, production, disaster-recovery, and disconnected environments require separate rights?
- Are modification, quantization, adaptation, merging, or derivative models allowed?
- What restrictions apply to outputs, internal use, contractors, geography, or redistribution?
- Which entitlements depend on continued subscriptions or external services?
- Could normal intended operation accidentally breach a license condition?
Checking Whether Licenses Permit Disconnected and Long-Term Operation
- Does the software require online activation or periodic entitlement checks?
- How long can the capability operate without contacting the supplier?
- Can licenses be transferred to replacement hardware?
- Are offline update packages and license files contractually available?
- What happens to operating rights after contract termination?
- Can approved versions continue to run after the vendor stops selling or supporting them?
- Which licensing term could unexpectedly make a technically self-contained system unusable?
Identifying Supplier Dependencies That Could Stop the Service
- Which supplier-controlled function is necessary for continued operation?
- How quickly would loss of that supplier affect the service?
- Is the dependency technical, commercial, legal, operational, or knowledge based?
- Can the organization temporarily continue on the last approved version?
- Is there an alternative supplier or internal workaround?
- What inventories, spares, rights, or documentation could extend safe operation?
- Which dependency represents the largest single point of external failure?
Determining Whether External Support Creates Hidden Operational Dependence
- Which incidents can internal staff resolve without supplier assistance?
- Does routine maintenance require supplier personnel?
- Are diagnostic tools or documentation withheld from the organization?
- Does the supplier hold unique credentials or configuration knowledge?
- Can the organization recover the service if support is unavailable for months?
- Is supplier involvement advisory or operationally mandatory?
- What internal capability would need to be built to turn support into optional assistance rather than dependence?
Assessing Whether Important Components Can Be Replaced by Another Supplier
- What would need to change technically if this component were replaced?
- Are compatible alternatives actually available?
- Which data, configuration, interfaces, or integrations would have to migrate?
- Do licenses permit the transition?
- What new testing, accreditation, or training would be required?
- Does the organization have enough knowledge to perform the replacement?
- Has replaceability ever been demonstrated rather than merely assumed?
Managing Supplier Obligations for Security, Support, and Lifecycle Change
- What vulnerability and security notification duties does the supplier have?
- How quickly must critical fixes or workarounds be provided?
- What support response and resolution commitments are required?
- How much advance notice is required for end-of-support or incompatible changes?
- What technical documentation and knowledge transfer must remain current?
- What audit or assurance evidence must the supplier provide?
- How will the organization enforce these obligations when the capability is already operationally dependent on the supplier?
Planning for End-of-Support and Product Discontinuation
- When does support end for each critical component?
- How much notice will the organization receive before that date?
- Can the current version continue operating securely afterward?
- Are extended support, source access, alternate suppliers, or migration options available?
- What dependencies must be replaced before support actually ends?
- How long will migration, testing, and reaccreditation take?
- What trigger should start replacement planning before the deadline becomes urgent?
Preparing for Supplier Failure or Loss of External Support
- Which supplier failures would materially affect the capability?
- What can the organization continue operating without that supplier?
- Which updates, licenses, spares, documentation, and technical knowledge are already held internally?
- Can another provider assume support?
- What contractual rights survive supplier insolvency or termination?
- How long can safe operation continue under the contingency?
- What conditions should trigger accelerated replacement rather than continued operation?
Managing Export-Control and Jurisdictional Dependencies
- Which hardware, software, models, or support services are subject to foreign legal restrictions?
- Could future export controls block updates, spare parts, or replacement systems?
- Which supplier decisions depend on another jurisdiction?
- Is support access affected by location, nationality, classification, or end use?
- What alternative supply paths exist?
- How quickly could geopolitical change become an operational problem?
- What level of dependency is acceptable given the mission and expected service life?
Maintaining Contractual Rights Needed for Continued Operation and Exit
- What rights are required to continue using the capability after contract expiry?
- Can the organization retain and deploy the last approved software and model versions?
- Does it have access to configuration, documentation, build artifacts, and operational data?
- Are data export and migration assistance guaranteed?
- Can another supplier legally maintain or replace the system?
- What supplier cooperation is required during transition?
- Are the exit rights useful in practice, or do technical and knowledge dependencies make them nominal?
Implementing and Integrating the Capability
Establishing the Technical Baseline Before Implementation Begins
- What exact architecture and component set has been approved for implementation?
- Which versions, interfaces, network zones, security controls, and environments are in scope?
- What assumptions remain unresolved?
- Who owns each implementation workstream?
- Which acceptance requirements must the implementation be designed to prove later?
- What dependencies and prerequisites must exist before installation starts?
- How will changes to the baseline be controlled during delivery?
- Are power, cooling, racks, cabling, storage, and network capacity ready?
- Which secure areas and physical controls are required?
- Are accelerator servers and supporting systems compatible with the facility?
- What infrastructure must be installed before platform software can be deployed?
- How will resilience and maintenance requirements affect physical layout?
- Which facility constraints discovered during implementation require design changes?
- What evidence shows that the environment is ready for the production workload rather than only initial installation?
- Which infrastructure and platform components must be installed first?
- What approved configuration baseline should each component follow?
- How will administrative access, secrets, certificates, and service identities be established?
- Which dependencies must remain internal for self-operation?
- What configuration should be automated rather than performed manually?
- How will the deployed state be documented and reproduced?
- Which supplier defaults must be changed before the platform is considered secure or operational?
Integrating Organizational Identity and Access Management
- Which organizational identities and groups should map to platform roles?
- How are authentication and authorization enforced across applications, models, data, and administration?
- Which service identities are required?
- How will joiner, mover, and leaver changes propagate?
- What happens if the central identity service becomes unavailable?
- Are privileged roles sufficiently separated from ordinary user access?
- How will access behavior be tested before production use?
Connecting the Capability to Internal Networks and Security Services
- Which network zones must communicate with the GenAI platform?
- What ports, protocols, proxies, DNS, time, certificate, and load-balancing services are required?
- Which security monitoring and logging systems must receive telemetry?
- What connections should remain blocked by default?
- Are management and user traffic separated appropriately?
- How will connectivity work in restricted or disconnected environments?
- What network evidence is required before security approval?
Integrating Document Repositories and Organizational Knowledge Sources
- Which repositories should the capability access?
- How will source permissions be preserved?
- How frequently must changes, deletions, and permission updates synchronize?
- What metadata is needed for reliable retrieval and provenance?
- How will failed ingestion or stale indexes be detected?
- Who owns operational responsibility for each integration?
- What happens when a source system changes its interface or access model?
- Which systems need to exchange data or actions with GenAI?
- What user authority should each integration inherit?
- Which tool operations are read-only and which can change organizational systems?
- How are inputs and outputs validated at the interface?
- What happens if the downstream system is slow, unavailable, or returns unexpected data?
- How can an individual integration be disabled without taking down the entire GenAI service?
- What audit trail is required for model-initiated tool use?
Onboarding Models Through a Controlled Implementation Process
- Where will model artifacts be acquired from?
- How are provenance, integrity, licensing, and compatibility checked?
- What quarantine or staging environment is used before deployment?
- Which serving runtime and configuration are required?
- What evaluation evidence must accompany the model into production?
- How are approved artifacts promoted into internal repositories?
- Can the complete onboarding process be repeated without hidden supplier steps?
Resolving Integration Problems Across Supplier and System Boundaries
- Which component is actually producing the observed failure?
- What evidence can separate interface problems from component defects?
- Which supplier owns the affected boundary?
- Are interface assumptions documented consistently on both sides?
- Can the organization reproduce the problem without giving suppliers uncontrolled access?
- What temporary workaround risks becoming permanent technical debt?
- How will the final resolution be incorporated into configuration and operational documentation?
Preparing Documentation and Internal Teams for Operational Handover
- What knowledge must the receiving operations team possess before handover?
- Are architecture, configuration, recovery, support, security, and maintenance procedures complete?
- Can internal staff perform routine privileged tasks themselves?
- Which supplier knowledge remains undocumented?
- Have operational contacts, escalation paths, and responsibilities been established?
- Can the team demonstrate recovery and change procedures rather than merely describe them?
- What unfinished implementation work would become an operational burden if handover occurred now?
Testing, Accrediting, and Accepting the Capability
Designing an End-to-End Test and Acceptance Strategy
- Which requirements must be verified before the capability can enter production?
- What separate evidence is needed for functionality, integration, security, performance, resilience, GenAI quality, usability, and operability?
- Which tests require production-representative conditions?
- What thresholds constitute acceptance?
- Who performs and independently reviews each type of test?
- How will the exact tested configuration be recorded?
- Which approvals remain separate from contractual acceptance?
Testing Technical Components Before Testing the Integrated Capability
- Does each component perform its intended function in isolation?
- Are configuration, security, logging, and management functions working as specified?
- Which vendor claims can be directly verified?
- Are component interfaces behaving as documented?
- What defects should block further integration testing?
- Is the tested component version identical to the one intended for integration?
- What evidence should be retained before moving to end-to-end testing?
Testing End-to-End Integration Across the Delivered Capability
- Do identity, networks, models, retrieval, tools, logging, and applications work together correctly?
- Does information move only across intended interfaces?
- Are permissions preserved across system boundaries?
- How does the service behave when one integrated component is unavailable?
- Can failures be traced across the complete request path?
- Do monitoring and audit records cover the full transaction?
- Which integration defects would prevent production operation even if every component works individually?
- Does the system meet latency and throughput requirements under representative demand?
- What happens as concurrency and context size increase?
- Where do queues and resource bottlenecks appear?
- Can the platform sustain peak and surge profiles for the required duration?
- How do retrieval, tools, and supporting workloads affect inference performance?
- Does performance remain acceptable during component failure or maintenance?
- What tested workload envelope can operations rely on after acceptance?
Testing GenAI Security Against Realistic Attack Scenarios
- Which threat scenarios should be exercised rather than assessed only on paper?
- Can prompt injection cause unauthorized retrieval or tool use?
- Can malicious retrieved content influence privileged behavior?
- Can one user obtain another user's information?
- Can manipulated artifacts or data enter the trusted deployment path?
- Do containment controls limit impact when the model behaves adversarially or unexpectedly?
- What security claims remain unproven after testing?
Evaluating the Integrated GenAI Configuration on Representative Mission Tasks
- Does the exact production configuration perform well on the tasks it is intended to support?
- How do prompts, retrieval, tools, adapters, safety controls, and runtime settings affect results?
- Which critical cases fail even when average performance is acceptable?
- How variable are results across repeated trials?
- Does quality hold across required languages, user groups, and information conditions?
- What human expert judgment is needed to assess outputs?
- What baseline should be retained for future regression testing?
Testing Resilience, Failure, and Recovery Scenarios
- Which component and site failures should be deliberately introduced?
- Does the system fail in the way the continuity design expects?
- Can priority workloads continue under reduced capacity?
- Can backups and known-good configurations actually restore the service?
- How long does recovery take in practice?
- Which dependencies prevent recovery when they are unavailable?
- What resilience claim must be revised if the exercise does not match the design assumption?
Testing Whether Real Users Can Work Effectively With the Capability
- Can intended users complete representative tasks successfully?
- Where do they misunderstand the interface or the role of GenAI?
- Does the service fit existing workflows?
- Are latency, output quality, and retrieval behavior acceptable in real use?
- Which failure modes are users able to recognize and handle?
- Do users understand information-handling and oversight requirements?
- What operational or design changes are needed before broader rollout?
Demonstrating That Internal Staff Can Operate the Capability Without Supplier Intervention
- Can internal staff start, stop, monitor, and administer the service?
- Can they diagnose ordinary infrastructure, platform, and model failures?
- Can they deploy and roll back an approved model or software release?
- Can they restore the capability from backup?
- Can they replace failed hardware and rejoin it to the service?
- Can they collect diagnostics and work with external support without granting uncontrolled access?
- Which task still reveals an unacceptable dependence on the supplier?
Preparing Evidence for Security Authorization and Accreditation
- What security claims and boundaries must the authorizing authority assess?
- Is the architecture and data flow documentation complete?
- Does the evidence cover conventional cybersecurity and GenAI-specific threats?
- Are vulnerability findings and residual risks recorded accurately?
- Has the deployed configuration been tested rather than a simplified reference environment?
- Which supplier claims require independent verification?
- What unresolved issue would prevent authorization despite contractual acceptance?
Completing Required Safety and Mission Assurance Reviews
- What mission or safety consequences could result from incorrect GenAI behavior?
- Which professional authorities need to review the intended use?
- Are human oversight and fallback arrangements appropriate to the consequence?
- Have representative harmful or misleading outputs been evaluated?
- Are system limitations communicated to users and mission owners?
- Which use cases need stronger controls than the platform baseline?
- What restrictions or conditions should accompany mission approval?
Deciding Whether Remaining Defects and Risks Are Acceptable Before Handover
- Which defects remain open?
- Which are contractual defects, operational risks, security findings, or future enhancements?
- What consequence could each unresolved issue create in production?
- Is there a tested workaround?
- Who has authority to accept the residual risk?
- What deadline and owner exist for post-handover remediation?
- Would accepting the capability now transfer unresolved implementation problems into operations without a credible solution?
Operating, Monitoring, and Supporting the GenAI Service
Defining Operational Responsibilities Across the GenAI Service
- Who owns the service as a whole?
- Who owns infrastructure, platform, models, data, retrieval, security, integrations, support, and evaluation?
- Which decisions can each role make independently?
- Where do responsibilities currently overlap or leave gaps?
- Who coordinates issues that cross several technical layers?
- Which responsibilities remain with suppliers?
- Are ownership and escalation paths clear enough for incidents and changes?
Establishing Normal Day-to-Day Operation of the GenAI Service
- What routine activities keep the service healthy?
- Which checks and operational tasks occur daily, weekly, or on another cycle?
- What does normal service behavior look like?
- Which operational thresholds require intervention?
- How are routine jobs, repositories, certificates, storage, and queues maintained?
- What records must operations keep?
- Which recurring task still depends unnecessarily on informal knowledge or supplier assistance?
- Which signals show whether the service is actually available end to end?
- Can operations distinguish infrastructure failure from platform, model, retrieval, or integration failure?
- Which dependencies are not currently visible in monitoring?
- What health checks should represent real user journeys?
- How are alerts correlated across layers?
- Which alerts indicate genuine service impact rather than harmless technical variation?
- Can the monitoring system itself operate during restricted or degraded conditions?
- How are request volume, concurrency, queueing, latency, and generation performance changing?
- Which workloads consume the most resources?
- Are peak patterns approaching the validated service envelope?
- Is demand growing in expected or unexpected user groups?
- Which limits are affecting users before infrastructure is technically exhausted?
- When should capacity planning be reopened?
- Could workload management solve the problem before additional hardware is required?
Monitoring Model, Retrieval, and Knowledge Quality
- What production signals can indicate behavioral degradation?
- Are users reporting different failures from those represented in formal evaluations?
- Is retrieval returning current and relevant sources?
- Have source changes introduced stale or conflicting knowledge?
- Which evaluation cases should run periodically against production versions?
- Can quality changes be traced to model, prompt, retrieval, tool, or data changes?
- What threshold should trigger deeper investigation or rollback?
Monitoring Operating Cost and Resource Efficiency
- What does the service cost to operate by workload or user group?
- Which hardware is underused or persistently saturated?
- How much capacity is consumed by oversized models or unnecessary context?
- Are software, support, and licensing costs aligned with actual use?
- Could routing, scheduling, quantization, or smaller models reduce cost without harming required quality?
- Which cost increases reflect useful adoption rather than inefficiency?
- At what point does the operating model need to change because the service is economically unsustainable?
Supporting Users and Resolving GenAI Service Issues
- What problem is the user actually experiencing?
- Is it a service defect, model limitation, retrieval problem, permission issue, or misunderstanding of intended use?
- Can support reproduce the issue with safe test data?
- What information should be collected without exposing unnecessary sensitive content?
- Which issues can first-line support resolve and which require specialist escalation?
- Are recurring support cases revealing a design, guidance, or training problem?
- How should resolved issues feed back into documentation and service improvement?
Detecting and Triaging Production Anomalies
- What changed from normal behavior?
- Which users, workloads, components, or environments are affected?
- Is the anomaly technical, behavioral, security-related, or data-related?
- What recent changes could be relevant?
- Does the issue require immediate containment before the cause is known?
- Which evidence should be preserved before restarting or changing systems?
- Who needs to take ownership as the incident becomes better understood?
- Which infrastructure or platform component failed?
- What services and workloads are affected?
- Can traffic fail over or priority service continue?
- Is restart, replacement, rollback, or recovery the safest response?
- What evidence should be captured before restoring service?
- Does the incident reveal a missing redundancy or unsupported dependency?
- What follow-up change is needed to reduce recurrence?
Responding to Model Failures
- What model behavior is failing?
- Is the problem limited to particular tasks, languages, contexts, or users?
- Has the model, runtime, prompt, retrieval, adapter, or serving configuration changed?
- Can an approved alternate model or previous version restore service?
- Should the affected workload be restricted while investigation continues?
- What evaluation cases can reproduce the failure?
- Does the incident require changing qualification criteria or regression tests?
Responding to Unsafe or Materially Harmful Outputs
- What output was produced and what real-world consequence could follow?
- Was the output acted on or shared beyond the original interaction?
- Which context, retrieval, tools, or instructions influenced it?
- Is the problem isolated or reproducible across similar tasks?
- Should the affected feature, model, or workflow be restricted immediately?
- Which safety, professional, legal, or mission authority needs to be involved?
- What evidence should be preserved so the event can improve future controls and evaluations?
Responding to Suspected Security Compromise
- What evidence suggests a compromise rather than an ordinary service failure?
- Which accounts, systems, models, data stores, or integrations may be affected?
- What should be isolated before restoration is attempted?
- Which evidence must be preserved for investigation?
- Could the attacker have modified model or software artifacts?
- Which credentials and trust relationships need rotation or revalidation?
- What must be proven clean before the service returns to operation?
- What information may have been exposed?
- Who was able to receive or access it?
- Did the exposure occur through retrieval, prompts, outputs, logs, tools, or configuration?
- How far did the information propagate?
- What containment or access changes are required immediately?
- Which security, privacy, classification, or legal reporting obligations apply?
- What control failed and how will recurrence be tested after remediation?
- Which integration or tool is failing?
- Is the failure affecting only information access or also causing incorrect actions?
- Can the integration be disabled without removing the rest of the service?
- Are retries or fallback behavior creating duplicate or unsafe actions?
- Has the external system changed its API, permissions, or data format?
- What user workflows require an alternative while the integration is unavailable?
- What interface or monitoring change would make similar failures easier to detect and contain?
Coordinating Infrastructure, Security, Model, Data, and Application Teams During Complex Incidents
- Which parts of the capability are involved?
- Who is coordinating the incident across specialist teams?
- What facts are known and what remains inference?
- Which team controls each containment or recovery action?
- Could one team's remediation destroy evidence or worsen another layer?
- What operational priority governs restoration?
- How should decisions and evidence be recorded so responsibility remains clear after the incident?
Escalating to External Support Without Surrendering Operational Control
- What specific expertise or information is required from the supplier?
- What diagnostics can be provided without exposing unnecessary sensitive data?
- Can the supplier analyze evidence without direct access?
- If temporary access is necessary, what scope and duration are justified?
- Who authorizes and monitors supplier activity?
- Can internal staff execute the proposed remediation themselves?
- What should be learned from the escalation so the same dependency is reduced next time?
Managing Updates, Maintenance, and Configuration Changes
Assessing the Impact of a Proposed Change Before Deployment
- What component or configuration is changing?
- Could the change alter GenAI behavior even if it appears technically minor?
- Which applications, models, data flows, controls, or dependencies could be affected?
- What testing and evaluation should the change trigger?
- Does the change affect accreditation, licensing, privacy, or operational documentation?
- What rollback path exists?
- Who needs to approve the change before production deployment?
Applying an Urgent Security Patch Without Bypassing Control
- What vulnerability or threat makes the patch urgent?
- Which components and environments are affected?
- What risk comes from delaying the patch?
- What risk comes from deploying it with reduced testing?
- Which minimum compatibility, security, and behavioral checks must still occur?
- Is a temporary mitigation available while testing continues?
- How will the accelerated decision and any residual risk be documented?
Introducing a New Model Version Into Production
- What changed between the approved model and the proposed version?
- Which mission evaluations need to be rerun?
- Does the new version require different runtime, prompts, context handling, or hardware?
- Are security, licensing, and provenance still acceptable?
- Which applications or user groups should receive the model first?
- How will behavior be monitored after deployment?
- Can the previous model be restored quickly if production behavior differs from evaluation?
- Which runtime or platform component is changing?
- Does the update affect model compatibility or behavior?
- What vulnerabilities or operational improvements does it address?
- Which drivers, libraries, APIs, or integrations could break?
- Can the update be staged on the same model and configuration used in production?
- How will state and configuration be preserved during rollback?
- What evidence is required before the new software becomes the production baseline?
Maintaining Hardware and Firmware Without Disrupting Service
- Which hardware or firmware requires maintenance?
- Can workload be moved away from the affected equipment?
- Does the maintenance change accelerator compatibility or performance?
- What firmware and driver combinations are approved?
- How will failed maintenance be recovered?
- Does the change affect resilience or capacity during the maintenance window?
- What post-maintenance checks are needed before returning the equipment to service?
- What behavior is the change intended to improve?
- Which users and workflows could be affected?
- Can a prompt or retrieval change alter security-sensitive behavior?
- What representative evaluations should be rerun?
- Could the change affect access permissions or tool authority?
- How will the previous configuration be versioned and restored?
- What production signals should be watched immediately after deployment?
Managing Compatibility Across Hardware, Drivers, Runtimes, and Models
- Which versions are currently validated together?
- What dependency prevents one component from being upgraded independently?
- Does a new model require a newer runtime or driver?
- Does new hardware force software changes elsewhere?
- Which compatibility combinations are officially supported and which have only been observed to work?
- How many legacy combinations must operations continue maintaining?
- What sequencing would reduce compatibility risk during a multi-layer upgrade?
Moving Approved Updates Through Development, Test, and Production Environments
- Where should the update first enter the internal lifecycle?
- What checks occur before promotion to the next environment?
- Is the same artifact promoted, or is it rebuilt at each stage?
- How is integrity preserved during transfer?
- Which evaluation and approval evidence follows the artifact?
- Who can authorize production promotion?
- Can the organization reconstruct exactly how the production version reached its current state?
Rolling Back a Change That Causes Technical or Behavioral Problems
- What evidence shows that the recent change is responsible?
- Which previous state is known to be good?
- Can the rollback restore all affected components consistently?
- Will data or configuration created after the change remain compatible?
- What operational or security risk comes from returning to the older version?
- How will rollback success be verified?
- What must be learned before the change is attempted again?
Detecting and Correcting Configuration Drift
- Which deployed settings differ from the approved baseline?
- How did the drift occur?
- Is the difference intentional, temporary, or unauthorized?
- Could the drift affect security, behavior, resilience, or supportability?
- Can configuration be reconciled automatically without disrupting service?
- Which manual changes should be converted into controlled configuration management?
- What monitoring should detect similar drift earlier?
Maintaining Resilience and Continuity
Identifying the Failures That Could Make the GenAI Service Unavailable
- Which components can independently stop the service?
- Which apparently redundant components share the same underlying failure domain?
- What external dependencies could make an internally hosted service unusable?
- What happens if identity, storage, networking, model repositories, or licensing fails?
- Which failures affect only some workloads?
- How long can each dependency remain unavailable before service is materially affected?
- Which failure scenarios deserve design changes rather than procedural recovery alone?
Designing Redundancy Around Real Service Failure Domains
- What physical and logical failure domains exist?
- Does redundancy actually cross those boundaries?
- Could one power, network, storage, or management failure take out all replicas?
- Which components need active redundancy and which can be restored from spares?
- What additional complexity does redundancy create?
- Can the organization maintain and test the redundant path regularly?
- Does the design improve service resilience or merely duplicate hardware inside the same failure domain?
Determining Which Workloads Must Survive Reduced Capacity
- Which workloads remain mission-critical during disruption?
- Which users should receive priority?
- What minimum service level is still useful?
- Which tasks can wait until full capacity returns?
- Are priorities based on mission consequence rather than organizational status?
- How will the platform enforce priority under contention?
- What decision changes when the disruption lasts much longer than expected?
Designing Useful Degraded Modes for GenAI Service
- What reduced service would still create meaningful value?
- Could smaller models maintain critical functions?
- Which context lengths, tools, modalities, or retrieval features can be restricted?
- Can noncritical workloads be queued or disabled?
- How will users know the service is operating in a degraded mode?
- What risks appear when lower-capability models replace the normal model?
- What conditions trigger entry into and exit from degraded operation?
Reserving Capacity for Priority Work During Major Disruption
- How much capacity should remain unavailable to normal demand so it exists during disruption?
- Which users or workloads may consume the reserve?
- Could dynamic preemption or scheduling replace static reservation?
- What happens if priority demand exceeds the reserve?
- How is the reserve tested without wasting capacity permanently?
- Does the reserved capacity share hidden failure domains with normal capacity?
- What balance between efficiency and assured availability is acceptable?
Planning Backup and Restoration for the Complete Capability
- Which parts of the capability require backup?
- Which components can be reconstructed instead of backed up?
- Are model artifacts, configuration, prompts, indexes, credentials, and application state covered?
- Which source data must remain the system of record?
- What recovery-point objective applies to each artifact?
- Are backups protected to the same security level as production?
- Can the organization restore the complete service rather than individual files?
Recovering the GenAI Service From a Known-Good State
- What state is considered known good?
- How is that state preserved and protected from the incident affecting production?
- Which components must be restored in what order?
- Can models, indexes, configuration, identities, and integrations be restored consistently?
- How long does full recovery take under realistic conditions?
- What validation is required before users regain access?
- How frequently must the recovery procedure be exercised to remain credible?
Preparing Spare Hardware and Replacement Components
- Which hardware failures cannot wait for normal procurement lead times?
- What spares should be held locally?
- Which components can be substituted with compatible alternatives?
- How long can stored spares remain technically and supportably useful?
- Are firmware, drivers, and configuration available to commission the replacement?
- How does inventory policy differ for central facilities and remote sites?
- What shortage would create the longest service interruption?
Planning for Loss of a Site or Major Infrastructure Component
- Which services disappear if the primary site is lost?
- Is another site capable of running the required workloads?
- What data, models, configuration, and credentials are available there?
- How will users reach the alternate service?
- Does the alternate site depend on the same power, network, identity, or supplier infrastructure?
- What service level can realistically be restored?
- Which business or mission activities need a non-GenAI fallback if site recovery takes longer?
Exercising Continuity and Disaster-Recovery Procedures Before They Are Needed
- Which failure scenario should the exercise simulate?
- Are teams required to operate without the resources assumed lost?
- Can recovery be completed using existing documentation?
- Which hidden dependencies become visible only during the exercise?
- Does actual recovery time match the stated objective?
- Are degraded-mode priorities workable under real demand?
- What plan, architecture, staffing, or documentation changes should follow from the exercise?
Operating GenAI in Restricted and Disconnected Environments
Determining How Disconnected the Environment Actually Needs to Be
- What threat or policy requirement drives the connectivity restriction?
- Which external connections are prohibited and which may be controlled?
- Is the environment restricted, intermittently connected, fully disconnected, or genuinely air-gapped?
- What operational benefits would be lost by stronger isolation?
- Which external dependencies would have to be recreated locally?
- How will information and updates enter and leave the environment?
- Is the chosen connectivity model stronger than necessary for the actual risk?
Designing GenAI Operation Without External Runtime Dependencies
- Which external services does the current design assume during normal operation?
- Can identity, DNS, time, certificates, secrets, monitoring, repositories, and licensing all function locally?
- Are model and software artifacts stored inside the environment?
- Does any component require vendor telemetry or activation?
- Can administrators diagnose and recover the service without internet access?
- Which documentation and tooling must be available locally?
- What dependency would fail first during a prolonged disconnection?
Establishing Local Repositories for Models, Software, and Updates
- Which artifacts must be available inside the restricted environment?
- How are approved versions organized and retained?
- Can the repositories support rollback to previous known-good states?
- Who may add or remove artifacts?
- How is integrity verified after import?
- Are dependencies complete enough to reinstall systems without external package sources?
- How will repository content remain supportable without becoming an uncontrolled archive?
Moving Models, Software, Firmware, and Updates Into Controlled Environments
- Where are external artifacts first acquired?
- How is provenance and integrity established before transfer?
- Which malware, vulnerability, compatibility, licensing, and GenAI-specific checks occur?
- Who approves release into the controlled environment?
- How is the artifact transferred without bypassing the security boundary?
- Is integrity revalidated after transfer?
- What evidence links the deployed artifact to the externally acquired and internally approved package?
Operating Identity, Monitoring, and Administration Without External Services
- Which identity services must exist locally?
- How are credentials, certificates, and privileged accounts maintained?
- Can monitoring and logging function without cloud backends?
- Where are administrative tools and documentation hosted?
- How are time synchronization and audit integrity maintained?
- What happens if the central enterprise identity system is unreachable?
- Can the environment remain securely manageable for the full expected period of disconnection?
Supporting the Capability When Vendors Cannot Connect Remotely
- What diagnostics can local staff collect?
- How can evidence be exported without exposing unnecessary protected information?
- Can suppliers analyze sanitized logs or reproduced failures externally?
- How are recommended fixes converted into controlled import packages?
- Which incidents require expertise that the organization does not currently possess?
- What local tools and documentation reduce the need for remote access?
- How long can the organization sustain operations if external support becomes unavailable entirely?
Deploying GenAI to Remote, Edge, or Tactical Locations
- What mission need requires local GenAI at the remote location?
- What power, cooling, space, weight, and connectivity constraints apply?
- Which models can provide useful capability within the local compute envelope?
- What data and knowledge must travel with the deployment?
- How much operational skill is available on site?
- How will hardware failure, patching, and model replacement be handled?
- What capability should remain usable when the site is completely isolated?
Synchronizing Distributed GenAI Environments After Connectivity Returns
- Which data, configuration, logs, models, and updates need synchronization?
- Which side should be authoritative when both changed independently?
- Could synchronization move information across an inappropriate security boundary?
- How are version conflicts detected?
- Should operational logs synchronize differently from user content or knowledge data?
- What happens if one environment has missed several update cycles?
- How is restored connectivity prevented from creating uncontrolled automatic change?
Maintaining Restricted Operations When External Ecosystems Become Unavailable
- Which external ecosystem has become unavailable?
- What local versions, documentation, repositories, and skills remain usable?
- How long can the current approved stack operate safely without new external inputs?
- Which vulnerabilities or compatibility issues will accumulate over time?
- What mission capability can be preserved if no replacement arrives?
- Which external dependency should be substituted first?
- At what point does continued operation become less safe than planned degradation or shutdown?
Planning Capability Evolution and Technology Refresh
Reassessing Whether the Existing Capability Still Meets Mission Needs
- Have the original mission needs changed?
- Are users relying on the capability for work that was not originally anticipated?
- Which quality, capacity, security, or continuity requirements are no longer being met?
- Have alternative technologies made an important limitation avoidable?
- Is dissatisfaction caused by the core capability or by fixable implementation problems?
- Which metrics and operational evidence support a change?
- Does the capability still justify its cost and complexity?
Identifying Which Parts of the Capability Are Becoming Operationally Obsolete
- Which component is becoming unsupported, insecure, insufficient, or uneconomic?
- Is the problem physical age, software support, capability, capacity, compatibility, or supplier dependence?
- How long can the component continue safely?
- Which other components depend on it?
- Would replacement solve a real operational problem or only modernize the stack?
- What lead time is required before obsolescence becomes disruptive?
- Which lifecycle risk should be addressed first?
Evaluating New Technology Without Assuming Newer Is Better
- What problem would the new technology solve?
- How does it perform on our actual mission tasks and workloads?
- What new infrastructure or dependencies does it introduce?
- Does it improve security, resilience, operability, or only headline capability?
- What migration, training, testing, and reaccreditation effort would adoption require?
- Which current strengths would be lost?
- What evidence would justify adoption instead of continuing with the current technology?
Deciding Whether Incremental Improvement Is Still Sufficient
- Which limitations can still be addressed through ordinary updates or targeted replacement?
- Which limitations originate in the underlying architecture?
- How much technical and operational debt is accumulating?
- Are repeated fixes preserving a design that no longer fits the mission?
- What improvement can be achieved without major redesign?
- What future change becomes harder if redesign is postponed?
- At what point is continued incremental investment less rational than a broader refresh?
Determining When Architecture Redesign Has Become Necessary
- Which architectural assumptions no longer hold?
- Are component boundaries preventing substitution, scaling, or secure operation?
- Has the platform become too coupled to update safely?
- Do new mission or classification requirements demand a different structure?
- Could a redesign simplify the capability rather than merely modernize it?
- What services must continue during architectural transition?
- What evidence shows that redesign is necessary rather than desirable?
Planning Hardware Refresh Around Capacity and Support Needs
- Which hardware is approaching support, capacity, reliability, or facility limits?
- What future workload should the replacement support?
- Could software or model changes extend the useful life of existing hardware?
- How do new accelerators affect power, cooling, networking, and rack requirements?
- What driver and runtime changes follow from the new hardware?
- Should refresh occur incrementally or by cluster?
- What headroom is justified without betting on speculative long-term demand?
- When does the current platform become difficult or unsafe to support?
- What applications, models, data, and operating procedures depend on it?
- Is a compatible successor available?
- Which interfaces or configurations must be recreated?
- How much parallel operation is required?
- What testing and accreditation must be completed before cutover?
- When must migration start to avoid turning end-of-support into an emergency?
Coordinating Refresh Cycles Across Models, Software, Hardware, and Facilities
- Which layers are reaching lifecycle triggers at similar times?
- Which changes depend on another change happening first?
- Could combining refreshes reduce duplicated integration and accreditation effort?
- Would combining too many changes make testing and troubleshooting unsafe?
- Which facility upgrades have the longest lead times?
- How can model and software choices avoid prematurely stranding hardware?
- What sequence provides the best balance between continuity and modernization?
Introducing New GenAI Capabilities Without Destabilizing Existing Services
- What new capability is being added and who needs it?
- Does it require new models, modalities, tools, data, or infrastructure?
- Which shared services could be affected?
- Could the new workload consume capacity needed by existing mission services?
- What security and evaluation requirements differ from the current baseline?
- Should the capability begin in an isolated environment or limited user group?
- What evidence is needed before it becomes part of the standard service?
Maintaining a Technology Roadmap Without Locking the Organization to Forecasts
- Which lifecycle events can be predicted with reasonable confidence?
- Which future technologies are too uncertain to plan as commitments?
- What architectural options should remain open?
- Which facility or contract decisions have long lead times and need early action?
- What triggers should cause the roadmap to change?
- How can budget planning preserve flexibility instead of preselecting future products?
- Does the roadmap describe decisions to revisit rather than pretend to know future technology?
Replacing and Decommissioning the Capability
Deciding When the Existing Capability Should Be Replaced
- What condition is making replacement necessary?
- Can the issue be solved by updating or replacing only one component?
- Does the existing capability still meet mission, security, support, and economic requirements?
- What happens if replacement is delayed?
- Is a viable successor available?
- What migration and reaccreditation effort would replacement require?
- What evidence justifies a full capability replacement rather than continued evolution?
Planning Migration to a Successor Capability
- What must move from the current capability to the successor?
- Which applications, models, data, prompts, indexes, integrations, and operational processes are affected?
- What cannot be migrated directly?
- Which dependencies must be recreated or redesigned?
- How will users transition without losing required service?
- What testing and approvals must the successor complete before cutover?
- What exit criteria determine when the old capability can begin decommissioning?
Running Old and New Capabilities in Parallel During Transition
- Which workloads should move first?
- How long is parallel operation genuinely necessary?
- How will data and configuration consistency be maintained?
- Are users likely to continue using the old service longer than intended?
- What additional security and operating burden comes from running both?
- How will results from the two systems be compared?
- What conditions trigger final cutover rather than indefinite coexistence?
- Which assets must be transferred and which should be rebuilt?
- Are licenses compatible with the new platform?
- Can model behavior be reproduced with the new runtime?
- Should retrieval indexes be migrated or regenerated from authoritative sources?
- How will permissions, provenance, prompts, tools, and policies be preserved?
- What differences must be accepted because the successor behaves differently?
- How will migrated assets be validated before production reliance?
Exiting a Supplier Relationship Without Losing Operational Continuity
- Which services or knowledge currently depend on the outgoing supplier?
- What contractual rights allow continued operation during transition?
- Which documentation, credentials, configuration, artifacts, and data must be transferred?
- Can another supplier or internal team assume support?
- What cooperation does the outgoing supplier still owe?
- What dependency could the supplier withdraw before migration is complete?
- When is the organization genuinely independent enough to end the relationship?
Decommissioning Retired Models and Software Safely
- Where are retired model and software artifacts still deployed or stored?
- Which copies must remain for audit, rollback, legal, or forensic reasons?
- Which execution paths could still invoke the retired component?
- Have repositories, deployment pipelines, and configuration references been updated?
- What security risk comes from leaving unsupported software installed?
- How should retained artifacts be isolated from accidental production use?
- What evidence shows that operational use has actually ended?
Revoking Credentials, Interfaces, and Dependencies From Retired Components
- Which user accounts, service identities, certificates, keys, and tokens belong to the retired capability?
- Which firewall rules, APIs, routes, and trust relationships remain open?
- Are external systems still sending data to retired interfaces?
- Which monitoring, backup, or automation jobs still reference the old environment?
- What supplier access should be revoked?
- Could removing a dependency unexpectedly affect the successor?
- How will the organization verify that residual access paths have been closed?
Determining What Records and Artifacts Must Be Retained After Retirement
- Which records are required for audit, legal, security, mission, or contractual purposes?
- Must model versions, prompts, evaluation evidence, logs, configuration, or datasets be retained?
- How long should each category remain?
- Which retained material contains sensitive information?
- Can evidence be retained without preserving a runnable vulnerable system?
- Who owns the archive after the operating team disbands?
- How will future reviewers understand which configuration produced historical outputs?
Sanitizing or Disposing of Hardware and Storage Securely
- What sensitive information may remain on each device?
- Which storage media can be sanitized and which must be destroyed?
- Do accelerator, server, management, or embedded devices contain persistent storage?
- What sanitization standard applies to the information involved?
- How will leased or supplier-owned equipment be handled?
- What evidence of sanitization or destruction must be retained?
- Could reusable hardware be safely repurposed instead of disposed of?
Verifying That Decommissioning Has Removed the Old Capability Without Losing Required Evidence
- Are any production workloads still reaching the retired capability?
- Have all integrations, routes, credentials, repositories, and scheduled processes been checked?
- Is sensitive data removed from systems that no longer need to retain it?
- Are required records preserved in an accessible and protected form?
- Has monitoring confirmed that the retired service is no longer active?
- Do inventories and architecture documentation reflect the new state?
- Can the organization demonstrate both that the capability is gone and that necessary historical evidence remains?