Archive of Christian Ullrich

Self-Operated GenAI Reflex Area

2026-09-25

Table of Contents

Defining the Need and Requirements for Self-Operated GenAI

Determining Whether Self-Operation Is Actually Required

Clarifying the Mission Need the Capability Should Serve

Identifying GenAI Use Cases That Belong Inside the Capability

Defining Which Uses Should Remain Outside the Capability

Identifying the Users and Organizations the Capability Must Serve

Determining What Information the Capability Must Be Able to Process

Defining Availability and Continuity Expectations for Different Uses

Translating Expected Use Into Realistic Workload Requirements

Defining Quality and Performance Requirements for Important Workloads

Reconciling Competing Mission, Security, Cost, and Service Requirements

Choosing the Operating and Control Model

Determining What the Organization Must Control Directly

Distinguishing Self-Operation From Local Hosting

Deciding Which Responsibilities Can Remain With External Suppliers

Defining Who Holds Privileged Administrative Authority

Determining Whether External Control-Plane Dependencies Are Acceptable

Defining How External Support Can Operate Without Taking Operational Control

Determining Which Internal Competencies the Operating Model Requires

Designing Operational Authority for Normal and Emergency Conditions

Assessing Whether the Proposed Operating Model Provides Meaningful Organizational Control

Designing the GenAI Platform Architecture

Defining the Components the GenAI Platform Actually Needs

Deciding Which Platform Capabilities Should Be Shared Across Use Cases

Separating Applications From Model-Serving Infrastructure

Separating Control-Plane Functions From Inference Workloads

Designing Stable Interfaces Between Models, Platforms, and Applications

Choosing Where Retrieval, Tools, Policy Enforcement, and Orchestration Belong

Designing Separate Environments for Development, Evaluation, and Production

Deciding Which Components Should Be Modular and Which Should Be Integrated

Identifying Architectural Dependencies That Could Prevent Later Replacement

Reviewing Whether the Architecture Has Become More Complex Than the Mission Requires

Planning Compute, Capacity, and Facilities

Estimating Compute Requirements From Real Workload Characteristics

Determining Whether a Model Configuration Fits the Available Accelerator Memory

Sizing Capacity for Concurrent Users and Long Contexts

Determining the Capacity Needed for Expected Peak Demand

Planning for Mission-Critical Surge Demand

Choosing How Models Should Be Replicated or Distributed Across Accelerators

Determining Whether Multiple Model Tiers Could Reduce Infrastructure Demand

Assessing Whether Existing Power, Cooling, Racks, and Networking Can Support the Design

Planning Capacity Expansion Under Uncertain Future Demand

Reassessing Capacity as Models, Workloads, and Serving Efficiency Change

Selecting and Managing GenAI Models

Comparing Candidate Models Against Real Mission Tasks

Determining Whether Public Benchmarks Are Relevant to the Intended Use

Evaluating a Model Across Required Languages and Specialist Domains

Deciding Whether Prompting, Retrieval, Adapters, or Fine-Tuning Are Needed

Qualifying a New Model for Production Use

Managing Model Weights, Tokenizers, Templates, and Runtime Dependencies Together

Maintaining Approved Model Versions for Different Workloads

Investigating a Model That Behaves Differently After a Change

Deciding Whether a Newer Model Is Actually Better for the Mission

Removing a Model From the Approved Operational Portfolio

Managing Data, Knowledge, and Retrieval

Identifying Which Organizational Information Should Be Available to GenAI

Determining Which Sources Are Authoritative Enough for Retrieval

Preparing Documents and Data for Reliable GenAI Use

Preparing and Governing Data for Model Adaptation and Evaluation

Preserving Source Permissions Through the Retrieval Pipeline

Deciding When Structured Access Is Better Than Vector Retrieval

Managing Conflicting, Outdated, or Superseded Knowledge

Detecting When Retrieval Quality Has Deteriorated

Managing Embeddings and Indexes as Sensitive Derived Information

Removing Information Reliably From GenAI Knowledge Systems

Rebuilding or Migrating Retrieval Systems Without Losing Knowledge Control

Securing the Self-Operated GenAI Capability

Building a Threat Model for the Complete GenAI Capability

Hardening the Infrastructure, Platform, and Interfaces Around GenAI

Protecting Model Artifacts Against Tampering and Unauthorized Replacement

Limiting the Consequences of Prompt Injection

Preventing Retrieved Content From Becoming Trusted Instructions

Controlling What GenAI Can Do Through Tools and External Systems

Preventing Sensitive Information From Leaking Through GenAI Workflows

Protecting the Capability Against Poisoned Models, Data, and Dependencies

Restricting Privileged Access to GenAI Infrastructure and Administration

Protecting GenAI Logs, Prompts, Outputs, and Other Sensitive Operational Artifacts

Meeting Classification, Privacy, and Compliance Requirements

Determining Which Classification Levels the Capability May Process

Handling Information With Different Need-to-Know Restrictions

Designing Controlled Information Flows Across Security Domains

Determining How Generated Outputs Should Be Classified or Handled

Assessing Whether Personal Data Can Be Used in the Intended GenAI Workflow

Designing Retention and Deletion Rules for Prompts, Outputs, and Logs

Determining Which AI-Specific Regulatory Requirements Apply

Integrating GenAI Into Existing Cybersecurity and Compliance Frameworks

Preparing the Evidence Required for Privacy, Audit, or Regulatory Review

Resolving Conflicts Between Operational Use and Compliance Requirements

Reassessing Compliance When Models, Data, or Intended Uses Change

Planning the Sourcing and Procurement Approach

Deciding What the Organization Should Build, Buy, or Integrate

Choosing Between an Integrated Solution and a Modular Procurement

Deciding Whether a Prime Contractor or Multiple Specialist Suppliers Are More Appropriate

Structuring Procurement Lots Without Creating Integration Gaps

Using Market Engagement to Test Whether Requirements Are Realistic

Designing Competitive Evaluations Around Mission Workloads

Deciding Whether Prototypes or Pilot Deployments Are Needed Before Award

Comparing Acquisition Options Using Whole-Lifecycle Cost

Defining Requirements Across Rapid Technology Change

Designing the Procurement So Future Competition Remains Possible

Managing Suppliers, Licenses, and External Dependencies

Mapping the Suppliers and Dependencies Behind the Delivered Capability

Reviewing Licensing and Entitlement Terms Across the Capability

Checking Whether Licenses Permit Disconnected and Long-Term Operation

Identifying Supplier Dependencies That Could Stop the Service

Determining Whether External Support Creates Hidden Operational Dependence

Assessing Whether Important Components Can Be Replaced by Another Supplier

Managing Supplier Obligations for Security, Support, and Lifecycle Change

Planning for End-of-Support and Product Discontinuation

Preparing for Supplier Failure or Loss of External Support

Managing Export-Control and Jurisdictional Dependencies

Maintaining Contractual Rights Needed for Continued Operation and Exit

Implementing and Integrating the Capability

Establishing the Technical Baseline Before Implementation Begins

Preparing Facilities and Infrastructure for the GenAI Platform

Installing and Configuring the Core Platform

Integrating Organizational Identity and Access Management

Connecting the Capability to Internal Networks and Security Services

Integrating Document Repositories and Organizational Knowledge Sources

Connecting GenAI to Internal APIs, Applications, and Tools

Onboarding Models Through a Controlled Implementation Process

Resolving Integration Problems Across Supplier and System Boundaries

Preparing Documentation and Internal Teams for Operational Handover

Testing, Accrediting, and Accepting the Capability

Designing an End-to-End Test and Acceptance Strategy

Testing Technical Components Before Testing the Integrated Capability

Testing End-to-End Integration Across the Delivered Capability

Testing Performance and Capacity Under Realistic Workloads

Testing GenAI Security Against Realistic Attack Scenarios

Evaluating the Integrated GenAI Configuration on Representative Mission Tasks

Testing Resilience, Failure, and Recovery Scenarios

Testing Whether Real Users Can Work Effectively With the Capability

Demonstrating That Internal Staff Can Operate the Capability Without Supplier Intervention

Preparing Evidence for Security Authorization and Accreditation

Completing Required Safety and Mission Assurance Reviews

Deciding Whether Remaining Defects and Risks Are Acceptable Before Handover

Operating, Monitoring, and Supporting the GenAI Service

Defining Operational Responsibilities Across the GenAI Service

Establishing Normal Day-to-Day Operation of the GenAI Service

Monitoring Service Health Across Infrastructure, Platform, Models, and Integrations

Monitoring Performance, Capacity, and Emerging Demand

Monitoring Model, Retrieval, and Knowledge Quality

Monitoring Operating Cost and Resource Efficiency

Supporting Users and Resolving GenAI Service Issues

Detecting and Triaging Production Anomalies

Responding to Infrastructure and Platform Failures

Responding to Model Failures

Responding to Unsafe or Materially Harmful Outputs

Responding to Suspected Security Compromise

Responding to Data Exposure or Unauthorized Information Flow

Responding to Integration or Tool Failures

Coordinating Infrastructure, Security, Model, Data, and Application Teams During Complex Incidents

Escalating to External Support Without Surrendering Operational Control

Managing Updates, Maintenance, and Configuration Changes

Assessing the Impact of a Proposed Change Before Deployment

Applying an Urgent Security Patch Without Bypassing Control

Introducing a New Model Version Into Production

Updating Model-Serving and Platform Software Safely

Maintaining Hardware and Firmware Without Disrupting Service

Changing Prompts, Retrieval Settings, or Tool Configuration

Managing Compatibility Across Hardware, Drivers, Runtimes, and Models

Moving Approved Updates Through Development, Test, and Production Environments

Rolling Back a Change That Causes Technical or Behavioral Problems

Detecting and Correcting Configuration Drift

Maintaining Resilience and Continuity

Identifying the Failures That Could Make the GenAI Service Unavailable

Designing Redundancy Around Real Service Failure Domains

Determining Which Workloads Must Survive Reduced Capacity

Designing Useful Degraded Modes for GenAI Service

Reserving Capacity for Priority Work During Major Disruption

Planning Backup and Restoration for the Complete Capability

Recovering the GenAI Service From a Known-Good State

Preparing Spare Hardware and Replacement Components

Planning for Loss of a Site or Major Infrastructure Component

Exercising Continuity and Disaster-Recovery Procedures Before They Are Needed

Operating GenAI in Restricted and Disconnected Environments

Determining How Disconnected the Environment Actually Needs to Be

Designing GenAI Operation Without External Runtime Dependencies

Establishing Local Repositories for Models, Software, and Updates

Moving Models, Software, Firmware, and Updates Into Controlled Environments

Operating Identity, Monitoring, and Administration Without External Services

Supporting the Capability When Vendors Cannot Connect Remotely

Deploying GenAI to Remote, Edge, or Tactical Locations

Synchronizing Distributed GenAI Environments After Connectivity Returns

Maintaining Restricted Operations When External Ecosystems Become Unavailable

Planning Capability Evolution and Technology Refresh

Reassessing Whether the Existing Capability Still Meets Mission Needs

Identifying Which Parts of the Capability Are Becoming Operationally Obsolete

Evaluating New Technology Without Assuming Newer Is Better

Deciding Whether Incremental Improvement Is Still Sufficient

Determining When Architecture Redesign Has Become Necessary

Planning Hardware Refresh Around Capacity and Support Needs

Planning Platform Migration Before Support Becomes Critical

Coordinating Refresh Cycles Across Models, Software, Hardware, and Facilities

Introducing New GenAI Capabilities Without Destabilizing Existing Services

Maintaining a Technology Roadmap Without Locking the Organization to Forecasts

Replacing and Decommissioning the Capability

Deciding When the Existing Capability Should Be Replaced

Planning Migration to a Successor Capability

Running Old and New Capabilities in Parallel During Transition

Moving Models, Data, Knowledge, and Configuration to a Replacement Platform

Exiting a Supplier Relationship Without Losing Operational Continuity

Decommissioning Retired Models and Software Safely

Revoking Credentials, Interfaces, and Dependencies From Retired Components

Determining What Records and Artifacts Must Be Retained After Retirement

Sanitizing or Disposing of Hardware and Storage Securely

Verifying That Decommissioning Has Removed the Old Capability Without Losing Required Evidence