








Infrastructure
License
This work is licensed under the Creative Commons Attribution-NoDerivatives 4.0 International License. To view a copy of this license, visit http://creativecommons.org/licenses/by-nd/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.
https://bsamm.org
Infrastructure
Infrastructure is the ground everything else stands on. If your organization runs any servers, operates a website, or keeps even a modest presence in the cloud — and most organizations now do — then you have an infrastructure to secure, and its condition quietly determines the security of everything built upon it. This is the domain of the platforms, networks, and accounts that are exposed to the internet day and night, that mix long-lived on-premises systems with cloud services spun up in minutes, and that fail most often not through exotic attacks but through ordinary drift and misconfiguration.
That is exactly why attention here pays off. A misconfigured setting — a storage bucket left open, an over-permissioned account — is behind a large and rising share of cloud incidents, and serious cloud breaches surged more than one and a half times in a single year as attackers followed workloads into the cloud. The reassuring part is that infrastructure rewards discipline more than spend: an organization that hardens its estate to a known baseline, watches it for drift, and keeps it patched removes the commonest causes of compromise and gives every other domain a solid floor to stand on.
- Applies to almost every organization. If you run servers, a website, or any cloud service, this domain is yours — very few modern organizations have no infrastructure at all.
- Misconfiguration, not wizardry, is the usual cause. Open storage, loose permissions, and unpatched systems drive most infrastructure breaches — which means most are preventable.
- Weak infrastructure undermines everything above it. Every other domain is only as sound as the platforms it runs on, so a solid floor multiplies the value of all your other work.
- The cloud raised the stakes. Serious cloud breaches jumped sharply as workloads moved, and a growing share of all attacks now target cloud infrastructure specifically.
- Discipline beats budget. Hardening to a baseline, monitoring for drift, and patching promptly remove the commonest causes of compromise without exotic tooling.
The pages that follow apply the Security Assurance Maturity Model to the systems an organization operates; the shared framework they build on is set out in the Introduction.
InfrastructureThis is the Infrastructure volume of the Security Assurance Maturity Model. It assumes you have read the Introduction, which explains the framework every volume shares: the four Business Functions, the twelve Security Practices beneath them, the three Maturity Levels, and how to assess an organization and build an assurance program. That material is not repeated here.
What follows is the infrastructure security-specific detail. Each of the twelve Security Practices is given at all three Maturity Levels — with its activities, results, success metrics, costs and personnel — followed by the assessment worksheets you use to score this domain. For anything about how the model itself works, refer back to the Introduction.
Governance
Construction
Verification
OperationsMaturity Levels

Strategy & Metrics
The Strategy & Metrics (SM) Practice is focused on establishing the framework within an organization for an infrastructure security program. This is the most fundamental step in defining security goals for the estate in a way that's both measurable and aligned with the organization's real business risk.
By starting with a lightweight profile of the estate, an organization grows into more advanced classification schemes for environments and the workloads they carry. With additional insight on relative risk measures, an organization can tune its per-environment security goals and develop granular roadmaps to make the program more efficient.
At the more advanced levels within this Practice, an organization draws upon many data sources, both internal and external, to collect metrics and qualitative feedback on the program. This allows fine tuning of cost outlay versus the realized benefit at the program level.
Policy & Compliance
The Policy & Compliance (PC) Practice is focused on understanding and meeting external legal and regulatory requirements that apply to the systems an organization operates, while also driving internal standards to ensure compliance in a way that's aligned with the business purpose of the organization.
Infrastructure carries an unusual complication: where systems run on a cloud provider, responsibility for controls is divided between the provider and the customer, and the line moves depending on the service consumed. An organization that has not established where that line falls for each service will have gaps it does not know about.
In a sophisticated form, provision of this Practice entails organization-wide understanding of both internal standards and external compliance drivers while also maintaining low-latency checkpoints with the teams that build and run systems, so that no part of the estate operates outside expectations without visibility.
Education & Guidance
The Education & Guidance (EG) Practice is focused on arming the people who design, build and operate infrastructure with knowledge and resources to keep it secure. With improved access to information, teams will be better able to proactively identify and mitigate the specific risks that apply to their organization.
One major theme for improvement across the Objectives is providing training, either through instructor-led sessions or self-paced modules. As an organization progresses, a broad base of training is built by starting with the engineers who build systems and extending to everyone who holds privileged access, culminating with role-based certification to ensure comprehension of the material.
In addition to training, this Practice also requires pulling security-relevant information into guidelines and runbooks that serve as reference material. This builds a foundation for establishing a baseline expectation for how systems are built and run, and later allows for incremental improvement once usage of the guidelines has been adopted.
Infrastructure| Strategy & Metrics | SM1 | SM2 | SM3 |
| OBJECTIVE | Establish unified strategic roadmap for infrastructure security within the organization | Measure relative value of environments and workloads and choose risk tolerance | Align infrastructure spend with relevant business indicators and estate value |
| ACTIVITIES |
|
|
|
| Policy & Compliance | PC1 | PC2 | PC3 |
| OBJECTIVE | Understand governance and compliance drivers relevant to the estate | Establish security and compliance baseline and understand per-environment risks | Require compliance and measure adherence across the whole estate |
| ACTIVITIES |
|
|
|
| Education & Guidance | EG1 | EG2 | EG3 |
| OBJECTIVE | Offer engineering staff awareness training on infrastructure security | Educate all privileged personnel and provide role-specific guidance | Mandate comprehensive competency and centralize guidance |
| ACTIVITIES |
|
|
|

Threat Assessment
The Threat Assessment (TA) Practice is centered on identification and understanding of the risks to an organization's infrastructure based on how it is built, connected and administered. From details about threats and likely attacks against each environment, the organization operates more effectively through better decisions about prioritization.
Infrastructure threat modeling has a particular character: the most consequential attacks rarely target a single system. They target the credentials, trust relationships and management planes that grant authority over many systems at once, which means the interesting question is usually not how a host is compromised but how far the attacker travels afterwards.
By starting with simple threat models and building toward weighted analysis, an organization improves over time. Ultimately it maintains this information tightly coupled to the compensating controls deployed and the residual risk carried by systems it does not fully control.
Security Requirements
The Security Requirements (SR) Practice is focused on proactively specifying the expected behavior of infrastructure before it is committed to. Through analysis at the platform and provider selection stage, requirements are initially gathered from the business purpose of the workload. As an organization advances, more advanced techniques surface requirements that would not otherwise have been obvious.
Requirements are unusually consequential here because platform decisions are durable and expensive to reverse. A provider without the isolation model a regulated workload needs, or an appliance whose vendor ships firmware updates twice a decade, constrains every other Practice for as long as it remains in service.
In a sophisticated form, this Practice entails pushing the organization's requirements into its supplier relationships and then auditing the estate to ensure all parties adhere to expectations.
Secure Architecture
The Secure Architecture (SA) Practice is focused on proactive steps for an organization to build secure infrastructure by default. By enhancing the design process with reusable reference architectures and centrally maintained automation, the overall risk from the estate can be dramatically reduced.
Beginning with simple recommendations about approved platforms and explicit consideration of design principles, an organization evolves toward consistently using hardened baselines derived from recognized benchmarks and toward shared services for identity, secrets and logging rather than per-system implementations.
As an organization evolves, sophisticated provision of this Practice entails building reference architectures and automation modules covering the generic patterns it deploys. These become the path of least resistance, which is the only reliable way to make secure construction the default.
Infrastructure| Threat Assessment | TA1 | TA2 | TA3 |
| OBJECTIVE | Identify and understand high-level threats to the organization's estate | Increase granularity of threat understanding and weight threats for comparison | Concretely tie compensating controls to each threat against the estate |
| ACTIVITIES |
|
|
|
| Security Requirements | SR1 | SR2 | SR3 |
| OBJECTIVE | Consider security explicitly during platform and provider selection | Increase granularity of requirements and derive from known risks | Mandate security requirements process for all platforms and suppliers |
| ACTIVITIES |
|
|
|
| Secure Architecture | SA1 | SA2 | SA3 |
| OBJECTIVE | Insert consideration of proactive security guidance into the design process | Direct the design process toward known-secure services and baselines | Formally control the build process and validate utilization |
| ACTIVITIES |
|
|
|

Design Review
The Design Review (DR) Practice is focused on assessment of an infrastructure design before it is built. Beginning with lightweight review of what a proposed environment is intended to do, an organization improves to a formal process capable of catching serious problems while they are still cheap to fix.
The economics here are stark. A segmentation decision or an identity model chosen at design time costs little to change on a whiteboard and a great deal to change once workloads depend on it. Infrastructure designs also tend to be replicated, so a mistaken assumption is inherited by everything built afterwards.
In an advanced form, review is driven by the threat models and trust model already built, and its results feed the routine audit programme so that a design cannot reach production without having been examined.
Implementation Review
The Implementation Review (IR) Practice is focused on inspection of infrastructure as it is actually configured, rather than as it was designed. Where Design Review examines intent, this Practice examines what was built and what has happened to it since.
The gap between the two is drift, and it is the central problem of running an estate. Systems are built correctly and then diverge: a rule is opened during an incident and never closed, a permission is broadened to unblock a release, an instance is launched by hand outside the automation. A system compliant on the day it was built tells you little about its state a year later.
Beginning with point checks of high-risk systems, an organization improves toward continuous automated evaluation of the whole estate, and ultimately toward gating change so that non-compliant configuration cannot be deployed in the first place.
Security Testing
The Security Testing (ST) Practice is focused on testing infrastructure in its running state, in order to discover weaknesses that inspection of configuration will not reveal. Configuration review establishes that settings are as intended; testing establishes whether those settings withstand a determined attempt to defeat them.
For infrastructure the Practice carries an additional obligation that other domains do not: testing that the organization can recover. A backup that has never been restored is a hypothesis, and a failover that has never been exercised is an assumption. Both are routinely discovered to be wrong at the worst possible moment.
In a sophisticated form, this Practice establishes a minimum standard that must be met before an environment reaches production, and generates test cases from the organization's own attack paths rather than from a generic catalogue.
Infrastructure| Design Review | DR1 | DR2 | DR3 |
| OBJECTIVE | Support ad hoc reviews of infrastructure designs to ensure baseline mitigations | Offer assessment services and increase review granularity | Require review of infrastructure designs and audit against expectations |
| ACTIVITIES |
|
|
|
| Implementation Review | IR1 | IR2 | IR3 |
| OBJECTIVE | Opportunistically find configuration problems in deployed infrastructure | Make configuration review more accurate and efficient through automation | Mandate comprehensive configuration review and gate change against a baseline |
| ACTIVITIES |
|
|
|
| Security Testing | ST1 | ST2 | ST3 |
| OBJECTIVE | Establish process to perform basic security tests based on requirements | Make infrastructure testing more complete and efficient | Mandate infrastructure testing and establish a production release standard |
| ACTIVITIES |
|
|
|

Issue Management
The Issue Management (IM) Practice is focused on establishing consistent processes for handling what goes wrong in the estate: compromised systems, exploited vulnerabilities, exposed services, and reports from staff or outsiders that something is not right.
Infrastructure incidents differ from most others in that containment and availability pull against each other. Isolating a compromised system may take a business service down, and the decision to do so cannot sensibly be made for the first time during the incident. Deciding in advance who holds that authority is one of the highest-value activities in this Practice.
Beginning with a known route for reporting and a named contact, an organization improves toward a consistent response process with defined containment actions, and ultimately to root-cause analysis that feeds the assurance program.
Environment Hardening
The Environment Hardening (EH) Practice is focused on the controls applied to systems and the environment around them once they are running. Where Secure Architecture establishes what infrastructure should look like when built, this Practice keeps it that way and tightens it over time.
Two activities dominate. Keeping software current addresses the majority of opportunistic compromise, since most successful attacks exploit conditions for which a fix already existed. Constraining privilege and connectivity addresses what happens next, and largely determines whether an incident is a single system or the whole estate.
The Practice also covers the organization's ability to survive an attack that succeeds anyway. Backups that an attacker can reach and delete are not backups, and this is the specific failure that turns a ransomware incident into an existential one.
Monitoring & Maintenance
The Monitoring & Maintenance (MM) Practice is focused on the information an operator needs to run an estate: what systems are doing, whether they are healthy, and what must happen to each as it moves through its life to decommissioning.
The Practice has two halves that support each other. Monitoring covers the telemetry infrastructure produces and the detections derived from it. Maintenance covers the procedures that keep the estate coherent over time — change control, operational documentation, and the disciplined retirement of systems.
Decommissioning deserves particular attention because it is where infrastructure programmes most often fail quietly. A system left running after its purpose ended is unpatched, unowned and still connected, and it appears in incident reports far out of proportion to its usefulness.
Infrastructure| Issue Management | IM1 | IM2 | IM3 |
| OBJECTIVE | Identify and handle infrastructure issues in an ad hoc manner | Elaborate the response process for consistency and speed | Improve the assurance program through analysis of infrastructure incidents |
| ACTIVITIES |
|
|
|
| Environment Hardening | EH1 | EH2 | EH3 |
| OBJECTIVE | Understand and maintain the baseline operating environment | Improve confidence in operation through hardening and access control | Validate resilience continuously and harden the surrounding environment |
| ACTIVITIES |
|
|
|
| Monitoring & Maintenance | MM1 | MM2 | MM3 |
| OBJECTIVE | Capture the operational information an operator needs | Improve expectations for continuous operation through detailed procedures | Mandate monitoring of estate state and validate the full system lifecycle |
| ACTIVITIES |
|
|
|









The Security
Practices

SM1 | SM2 | SM3 | |
| OBJECTIVE | Establish unified strategic roadmap for infrastructure security within the organization | Measure relative value of environments and workloads and choose risk tolerance | Align infrastructure spend with relevant business indicators and estate value |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Estimate overall infrastructure risk profile
Interview business owners, platform operators and stakeholders and create a list of worst-case scenarios across the organization's estate. Based on the way in which your organization builds and runs systems, the list can vary widely, but common issues include ransomware reaching production storage, compromise of a cloud control plane granting authority over every environment at once, loss of a facility, exposure of a database to the public internet, and an outage that outlasts the business's tolerance.
After broadly capturing worst-case scenario ideas, collate and select the most important based on collected information and knowledge about the core business. Any number can be selected, but aim for at least 3 and no more than 7 to make efficient use of time and keep the exercise focused.
Elaborate a description of each of the selected items and document details of contributing worst-case scenarios, potential contributing factors, and potential mitigating factors for the organization.
The final infrastructure risk profile should be reviewed with business owners and other stakeholders for understanding.
B. Build and maintain assurance program roadmap
Understanding the main business risks to the organization, evaluate the current performance of the organization against each of the twelve Practices. Assign a score for each Practice from 1, 2, or 3 based on the corresponding Objective if the organization passes all the cumulative success metrics. If no success metrics are being met, assign a score of 0 to the Practice.
Once a good understanding of current status is obtained, the next goal is to identify the Practices that will be improved in the next iteration. Select them based on the infrastructure risk profile, other business drivers, compliance requirements, budget tolerance, etc. Once Practices are selected, the goals of the iteration are to achieve the next Objective under each.
Iterations of improvement should be approximately 3-6 months, but a strategy session should take place at least every 3 months to review progress on activities, performance against success metrics and other business drivers that may require program changes.
RESULTS
- Concrete list of the most critical business-level risks caused by infrastructure
- Tailored roadmap that addresses the security needs of the estate with minimal overhead
- Organization-wide understanding of how the assurance program will grow over time
SUCCESS METRICS
- >80% of stakeholders briefed on infrastructure risk profile in past 6 months
- >80% of staff briefed on assurance program roadmap in past 3 months
- >1 assurance program strategy session in past 3 months
COSTS
- Buildout and maintenance of infrastructure risk profile
- Quarterly evaluation of assurance program
PERSONNEL
- Infrastructure Engineers (1 day/yr)
- Architects (4 days/yr)
- Managers (4 days/yr)
- Business Owners (4 days/yr)
- Platform Operators (1 day/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Policy & Compliance - 1
- Threat Assessment - 1
- Security Requirements - 2

ACTIVITIES
A. Classify environments and workloads based on business risk
Establish a simple classification system to represent risk-tiers for environments and the workloads they host. In its simplest form, this can be a High/Medium/Low categorization. More sophisticated classifications can be used, but there should be no more than seven categories and they should roughly represent a gradient from high to low impact against business risks.
Working from the organization's risk profile, create evaluation criteria that map each environment to one of the risk categories. The sensitivity of data processed, the tolerance for downtime, internet exposure, and the blast radius of the credentials that operate the environment are all common inputs.
Evaluate collected information about each environment and assign a risk category based upon overall evaluation criteria. Non-production environments deserve deliberate treatment rather than automatic classification as low risk: they frequently hold copies of production data and share identity with production.
An ongoing process for classification should be established to assign categories to new environments and keep the existing information updated at least biannually.
B. Establish and measure per-classification security goals
With a classification scheme for the estate in place, direct security goals and roadmap choices can be made more granular.
The roadmap should be modified to account for each risk category by specifying emphasis on particular Practices for each. For each iteration, this would typically take the form of prioritizing more higher-level Objectives on the highest risk environments and progressively less stringent Objectives for lower categories.
This process establishes the organization's risk tolerance since active decisions must be made as to what specific Objectives are expected of each category. By choosing to keep lower risk environments at lower levels of performance, resources are saved in exchange for acceptance of a weighted risk. However, it is not necessary to arbitrarily build a separate roadmap for each category since that can lead to inefficiency in management of the program itself.
RESULTS
- Customized assurance plans per environment tier based on core value to the business
- Organization-wide understanding of security-relevance of environments and workloads
- Better informed stakeholders with respect to understanding and accepting risks
ADD’L SUCCESS METRICS
- >90% of environments evaluated for risk classification in past 12 months
- >80% of staff briefed on relevant environment risk ratings in past 6 months
- >80% of staff briefed on relevant assurance program roadmap in past 3 months
ADD’L COSTS
- Buildout or license of environment classification scheme
- Program overhead from more granular roadmap planning
ADD’L PERSONNEL
- Architects (2 days/yr)
- Managers (2 days/yr)
- Business Owners (2 days/yr)
- Security Auditors (2 days/yr)
RELATED LEVELS
- Policy & Compliance - 2
- Threat Assessment - 2
- Design Review - 2
InfrastructureACTIVITIES
A. Conduct periodic industry-wide cost comparisons
Research and gather information about infrastructure security costs from intra-industry communication forums, business analyst and consulting firms, or other external sources.
First, use collected information to identify the average security effort being applied by similar types of organizations in your industry. This can be done top-down from estimates of total percentage of infrastructure budget, or bottom-up by identifying the controls and operational activities considered normal for your type of estate.
The next goal is to determine whether there are savings available on the tooling your organization currently licenses. Consolidation is often possible where capabilities have been bought piecemeal, but account for hidden costs such as re-instrumenting environments or running two agents in parallel during migration.
These exercises should be conducted at least annually prior to the subsequent strategy session, and comparison information presented to stakeholders in order to better align the program with the business.
B. Collect metrics for historic infrastructure spend
Collect information on the cost of past infrastructure incidents. Time and money spent rebuilding compromised systems, revenue lost during outages, regulatory fines, emergency tooling purchases, and the cost of engineering effort diverted to response are the usual components.
Using the environment risk categories and the respective roadmaps for each, a baseline security cost per environment can be initially estimated from the costs associated with the corresponding category.
Combine the per-environment cost information with the general cost model, then evaluate for outliers, i.e. sums disproportionate to the risk rating. These indicate either an error in classification or the necessity to tune the program to address root causes more effectively.
Tracking should be done quarterly at the strategy session, and the information reviewed by stakeholders at least annually.
RESULTS
- Information to make informed case-by-case decisions on infrastructure expenditures
- Estimates of past loss due to infrastructure incidents and outages
- Per-tier consideration of security expense versus loss potential
- Industry-wide due diligence with regard to infrastructure security
ADD’L SUCCESS METRICS
- >80% of environments reporting security costs in past 3 months
- >1 industry-wide cost comparison in past 1 year
- >1 historic infrastructure spend evaluation in past 1 year
ADD’L COSTS
- Buildout or license industry intelligence on infrastructure programs
- Program overhead from cost estimation, tracking, and evaluation
ADD’L PERSONNEL
- Architects (1 day/yr)
- Managers (1 day/yr)
- Business Owners (1 day/yr)
- Security Auditors (1 day/yr)
RELATED LEVELS
- Issue Management - 1

PC1 | PC2 | PC3 | |
| OBJECTIVE | Understand governance and compliance drivers relevant to the estate | Establish security and compliance baseline and understand per-environment risks | Require compliance and measure adherence across the whole estate |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Identify and monitor external compliance drivers
Gather the regulations, contractual obligations and industry standards that place requirements on the systems your organization operates. Payment, health care and public-sector standards are the usual sources, along with customer contracts that specify control requirements, data residency and audit rights.
For each driver, extract the obligations that land on infrastructure rather than on the wider organization — encryption at rest and in transit, access control and separation of duties, log retention, network isolation, vulnerability remediation windows, and recovery objectives are the usual set.
Assign an owner for each driver and establish a lightweight review to catch changes, since standards and their interpretations shift over time.
B. Establish the shared responsibility boundary
For every platform and provider the organization uses, document which controls are the provider's responsibility and which are yours. The division differs by service model, and it differs again between services from the same provider.
The failure mode this prevents is assuming a control is inherited when it is not. A provider securing its facilities and hypervisors says nothing about the configuration of the network, the identity policy, or the encryption settings your teams chose, all of which remain yours.
Record the boundary alongside the compliance obligations so that an auditor's question can be answered with evidence rather than an assumption, and review it when a new service is adopted.
RESULTS
- Concrete list of the compliance drivers that apply to each environment
- Documented division of responsibility for every platform in use
- Assurance that obligations are known before an audit rather than during one
SUCCESS METRICS
- >75% of relevant staff briefed on compliance obligations in past 12 months
- >80% of platforms with a documented responsibility boundary
- >1 review of external compliance drivers in past 12 months
COSTS
- Ongoing research into applicable regulation and standards
- Buildout and maintenance of the responsibility boundary record
PERSONNEL
- Managers (2 days/yr)
- Architects (3 days/yr)
- Business Owners (2 days/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Strategy & Metrics - 1
- Education & Guidance - 1

ACTIVITIES
A. Build policies and standards for the estate
Translate the compliance obligations into internal policies and standards stating what the organization requires of its infrastructure. Where the obligations record what outside parties demand, the policy records what your organization has decided, which is usually a superset.
Cover the durable decisions: which platforms and regions are approved, how environments are separated, what may be exposed to the internet, how privileged access is obtained, where secrets live, what must be encrypted, and the retention period for logs and backups.
Write the standards so they can be checked mechanically wherever possible. A standard expressed as an observable configuration can later become policy-as-code; one expressed as an aspiration cannot.
B. Establish compliance gates for infrastructure change
Define checkpoints where compliance is confirmed rather than assumed. For infrastructure the highest-value gate is at deployment, since a non-compliant system that reaches production tends to stay there.
Specify what each gate checks and who can grant an exception. Legacy systems, vendor appliances that cannot be hardened, and acquisitions not yet integrated will all need exceptions, so the process needs an owner, a compensating control and an expiry date rather than an indefinite waiver.
Track the exception population as a metric. A steadily growing exception list is usually a sign that the standard is wrong or the platform is unsupportable, and both are programme problems.
RESULTS
- Concrete set of internal standards for the estate
- Standards written so they can later be checked mechanically
- Documented exception process with owners and expiry dates
- Exception population visible as a trend
ADD’L SUCCESS METRICS
- >80% of new environments passing the compliance gate in past 3 months
- <20% of the estate operating under a documented exception
- >80% of standards expressed as observable configuration
ADD’L COSTS
- Buildout and maintenance of policy and standards documentation
- Program overhead from operating compliance gates
ADD’L PERSONNEL
- Infrastructure Engineers (3 days/yr)
- Architects (3 days/yr)
- Managers (2 days/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Secure Architecture - 1
- Implementation Review - 1
InfrastructureACTIVITIES
A. Conduct periodic compliance audits of the estate
Move from confirming compliance at deployment to confirming it continuously. Infrastructure drifts through emergency changes, manual intervention during incidents, and resources created outside the standard process.
Audit against the internal standards rather than the tooling's own defaults, and include the resources nobody claims. Orphaned systems — created for a project that ended, or inherited through acquisition — are consistently over-represented in incidents because no team considers them theirs.
Each environment should undergo audit at least biannually, with the highest risk tiers audited more often.
B. Collect and control compliance evidence
Establish a repository of evidence sufficient to satisfy an external auditor without a scramble: configuration state over time, access reviews, change records, vulnerability remediation timelines, backup and restore test results, and exception records.
Automate collection from the platforms themselves wherever possible. Evidence assembled by hand at audit time is expensive, tends not to survive scrutiny, and diverts the people who would otherwise be fixing the findings.
Report adherence to stakeholders at least quarterly, trending over time rather than presenting a single snapshot.
RESULTS
- Organization-wide visibility of infrastructure compliance
- Unclaimed and orphaned resources identified rather than invisible
- Evidence available on demand rather than assembled under pressure
- Stakeholders able to see adherence trends across environments
ADD’L SUCCESS METRICS
- >95% of environments audited for compliance in past 6 months
- >90% of required evidence collected automatically
- >1 compliance report delivered to stakeholders in past 3 months
ADD’L COSTS
- Buildout or license of compliance evidence repository
- Ongoing overhead from audit and evidence review
ADD’L PERSONNEL
- Infrastructure Engineers (2 days/yr)
- Managers (2 days/yr)
- Security Analysts (2 days/yr)
- Security Auditors (6 days/yr)
RELATED LEVELS
- Implementation Review - 3
- Monitoring & Maintenance - 3

EG1 | EG2 | EG3 | |
| OBJECTIVE | Offer engineering staff awareness training on infrastructure security | Educate all privileged personnel and provide role-specific guidance | Mandate comprehensive competency and centralize guidance |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Conduct technical infrastructure security awareness training
Offer the staff who build and run systems access to training covering the fundamentals of infrastructure security, reflecting the platforms your organization actually operates rather than a generic curriculum.
Cover the mechanisms that matter: the identity and authorization model of each platform, how network isolation is achieved and commonly defeated, where secrets are supposed to live, how credentials are escalated and rotated, and the common paths by which an estate is compromised — which are usually credential theft followed by lateral movement rather than a novel exploit.
Aim to reach the whole engineering and operations population within a year, and refresh at least every two years as platforms change.
B. Build and maintain technical guidelines
Assemble reference guidelines for the teams who work on infrastructure, covering the specific decisions and settings your organization expects rather than restating general advice available elsewhere.
Practical topics include how to request and use privileged access, which baseline applies to which environment tier, how to expose a service safely, where to store and retrieve secrets, and how to hand a system over to another team.
Keep the guidelines lightweight and current. A short document that is accurate beats a comprehensive one that describes a platform you migrated off last year.
RESULTS
- Increased staff awareness of the mechanisms protecting an estate
- Baseline expectation for how systems are built and operated
- Reference material covering the organization's own platforms
SUCCESS METRICS
- >50% of engineering and operations staff briefed on security topics in past 12 months
- >1 technical guideline document published or updated in past 12 months
- >50% of relevant staff able to locate the guidelines
COSTS
- Buildout or license of training materials
- Ongoing maintenance of technical guidelines
PERSONNEL
- Infrastructure Engineers (2 days/yr)
- Platform Operators (2 days/yr)
- Managers (1 day/yr)
- Security Auditors (2 days/yr)
RELATED LEVELS
- Policy & Compliance - 1
- Secure Architecture - 1

ACTIVITIES
A. Conduct role-specific infrastructure security training
Extend training beyond the core platform team to everyone who holds privileged access, and tailor the material to what each audience can act upon.
Application teams deploying their own infrastructure need depth on the platform's identity model and on what their configuration exposes. Network engineers need segmentation and device hardening. Database administrators need encryption, access control and backup integrity. On-call staff need the incident material, since under pressure people fall back on training rather than documentation.
The population holding privileged access is usually larger than the platform team expects, and enumerating it is often the most useful output of this activity.
B. Utilize guidance to establish infrastructure expectations
Turn the technical guidelines into expectations that appear at the moments they matter — in the templates teams start from, in code review, in the deployment pipeline, and in the runbook opened during an incident.
Guidance embedded in the path of work is followed; guidance filed on an intranet is not. A secure default in a starter template will change more outcomes than a policy document, because it changes what happens when nobody is thinking about security.
Establish a route for engineers to ask questions and feed problems back, and use what comes back to improve both the guidance and the baselines.
RESULTS
- Role-appropriate understanding across everyone holding privileged access
- Enumerated population of privileged users
- Guidance embedded in templates and pipelines rather than filed away
- Feedback loop from engineers into the programme
ADD’L SUCCESS METRICS
- >80% of privileged staff trained on infrastructure security in past 12 months
- >1 role-specific training track delivered in past 12 months
- >80% of starter templates carrying secure defaults
ADD’L COSTS
- Buildout of role-specific training tracks
- Program overhead from embedding guidance in the path of work
ADD’L PERSONNEL
- Infrastructure Engineers (3 days/yr)
- Platform Operators (3 days/yr)
- Architects (2 days/yr)
- Managers (2 days/yr)
RELATED LEVELS
- Threat Assessment - 2
- Issue Management - 2
InfrastructureACTIVITIES
A. Establish role-based examination and certification
Introduce assessment so that competency is demonstrated rather than assumed, particularly for the roles that hold broad authority over the estate.
Tie certification to access. An engineer with administrative rights over the cloud control plane holds more effective authority than almost anyone else in the organization, and it is reasonable to require demonstrated competency before granting it and to review that competency periodically.
Practical examination against your own platforms and baselines is more informative than a vendor certificate, though the two can complement each other.
B. Establish centralized guidance control
Bring guidance and runbooks under central control with clear ownership, review cycles and version history, so the organization can state with confidence what its current expectations are.
Distribute from a single authoritative source and retire superseded material actively. Stale runbooks are worse than absent ones, because they are followed during incidents when nobody has time to verify them.
Measure whether guidance is reaching people and being used, and feed that back into the roadmap.
RESULTS
- Demonstrated competency for privileged infrastructure roles
- Single authoritative source for guidance and runbooks
- Superseded material actively retired rather than left to circulate
ADD’L SUCCESS METRICS
- >90% of privileged infrastructure staff certified in past 12 months
- >90% of runbooks reviewed in past 12 months
- All guidance centrally controlled with a named owner
ADD’L COSTS
- Buildout or license of examination and certification program
- Ongoing overhead from central guidance management
ADD’L PERSONNEL
- Infrastructure Engineers (3 days/yr)
- Platform Operators (3 days/yr)
- Security Analysts (2 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Security Testing - 2
- Monitoring & Maintenance - 2

TA1 | TA2 | TA3 | |
| OBJECTIVE | Identify and understand high-level threats to the organization's estate | Increase granularity of threat understanding and weight threats for comparison | Concretely tie compensating controls to each threat against the estate |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Build and maintain environment-specific threat models
For each environment the organization runs, build a lightweight model of how it could be attacked. A workshop and a page of notes per environment is enough to start.
Work through the ways an estate is exposed. Services are published to untrusted networks. Systems trust one another, often more than intended. Administrators hold credentials that work everywhere. Management planes and automation pipelines can deploy anywhere. Backups are reachable from the systems they protect. Vendors have remote access.
Record for each threat what currently stands in the way. Organizations frequently discover at this point that a control they assumed was universal covers only production, or only one cloud account.
Review the models with the teams who operate each environment and refresh at least annually.
B. Develop attacker profile from estate exposure
Characterize who would attack your infrastructure and what they would want. The commodity ransomware operator scanning for an exposed service, the group specifically targeting your sector, and the insider with legitimate credentials call for very different defenses.
Ground the profile in how the estate is actually exposed and operated. An organization running systems on-premises with a small operations team carries different risk from one where dozens of application teams provision cloud resources directly, even where the workloads are identical.
Document the profiles alongside the threat models so the two are read together.
RESULTS
- Concrete list of the threats facing each environment
- Better understanding of which controls actually mitigate which threats
- Shared vocabulary for discussing infrastructure risk across teams
SUCCESS METRICS
- >80% of environments with a documented threat model in past 12 months
- >1 threat model review in past 12 months
- >80% of relevant staff briefed on the attacker profile
COSTS
- Buildout and maintenance of threat models
- Ongoing overhead from annual review
PERSONNEL
- Architects (3 days/yr)
- Infrastructure Engineers (2 days/yr)
- Security Analysts (2 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Strategy & Metrics - 1
- Security Requirements - 1

ACTIVITIES
A. Build and maintain attack path models per environment
Move from listing threats to tracing paths. An attack path describes how an intruder gets from an initial foothold to something that matters, and it exposes gaps that a list of threat categories hides.
Trace a realistic path end to end. A vulnerable service in a low-tier environment yields code execution; a credential cached on that host is valid in the management network; that credential can read the automation pipeline's secrets; the pipeline can deploy to production. Each hop is a place a control could have interrupted the chain, and organizations commonly find their controls concentrated at the first hop only.
Pay particular attention to paths that cross a boundary the organization believes is solid — between non-production and production, between tenants, or between the corporate network and the management plane.
B. Adopt a weighting system for measurement of threats
Introduce a consistent scheme for rating threats so that they can be compared rather than merely enumerated. Simple qualitative scales suffice provided they are applied uniformly.
Rate on likelihood and impact at minimum, and consider adding a factor for blast radius: a moderate threat that reaches every environment usually deserves attention before a severe one confined to a single system.
Use the ratings to order remediation and to inform which Practices the next roadmap iteration should advance. Publish the ratings so the reasoning behind prioritization is visible to stakeholders.
RESULTS
- Concrete attack paths showing how a foothold becomes an incident
- Comparable ratings enabling prioritization across threats
- Visibility of where existing controls interrupt an attack chain
- Evidence about boundaries the organization believed were solid
ADD’L SUCCESS METRICS
- >80% of environments with documented attack paths in past 12 months
- >90% of identified threats carrying a rating
- >1 prioritization exercise driven by threat ratings in past 6 months
ADD’L COSTS
- Buildout of attack path library
- Program overhead from threat rating and review
ADD’L PERSONNEL
- Architects (3 days/yr)
- Security Analysts (4 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Strategy & Metrics - 2
- Design Review - 2
- Security Testing - 2
InfrastructureACTIVITIES
A. Explicitly evaluate risk from providers and third parties
Extend threat assessment to the infrastructure your organization depends upon but does not operate: cloud providers, hosting and colocation partners, managed service providers with administrative access, and the software supply chain feeding your images and automation.
For each, establish what is actually known and what is merely assumed. Where assurance cannot be obtained directly, decide whether to reduce what the party can reach, require an intermediating control, or accept the residual risk explicitly and record who accepted it.
Give particular weight to parties holding standing administrative access, since their compromise is indistinguishable from your own until it is investigated.
B. Elaborate threat models with compensating controls
Complete the mapping from each identified threat to the specific controls that mitigate it, and record the residual risk remaining after those controls apply.
This mapping is what lets an organization answer the question executives and auditors actually ask — not 'what controls do you have' but 'what happens if this occurs'. It also exposes controls mitigating nothing in the current threat model, which are candidates for retirement.
Maintain the mapping as the estate changes, and review residual risk with business owners at least annually so acceptance is renewed deliberately rather than inherited silently.
RESULTS
- Complete mapping of threats to the controls that mitigate them
- Explicit, owned acceptance of residual risk
- Understanding of exposure through providers and third parties
- Identification of controls that no longer mitigate a live threat
ADD’L SUCCESS METRICS
- >90% of identified threats mapped to compensating controls
- >1 residual risk review with business owners in past 12 months
- >90% of third parties with administrative access assessed in past 12 months
ADD’L COSTS
- Ongoing maintenance of threat-to-control mapping
- Program overhead from residual risk review
ADD’L PERSONNEL
- Architects (2 days/yr)
- Security Analysts (3 days/yr)
- Business Owners (1 day/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Secure Architecture - 3
- Environment Hardening - 2

SR1 | SR2 | SR3 | |
| OBJECTIVE | Consider security explicitly during platform and provider selection | Increase granularity of requirements and derive from known risks | Mandate security requirements process for all platforms and suppliers |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Derive security requirements from workload purpose
For each class of workload, work from what it does and what it handles to the security properties its infrastructure must provide. Start from the data and the availability expectation rather than a generic checklist.
A short set per class is enough at this level. Typical entries include encryption at rest and in transit with control over keys, isolation from other tenants and environments, authentication that integrates with the organization's identity provider, audit logging exportable to your own retention, and recovery objectives the platform can actually meet.
Have the requirements reviewed by the teams who will run the workload. A requirement nobody can operate is not yet a requirement.
B. Evaluate security and compliance guidance for requirements
Draw on the compliance obligations, the responsibility boundary and the threat models already produced to catch requirements that purpose alone would not surface.
Compliance drivers frequently mandate specific properties — data residency, key custody, separation of duties — and the threat models will suggest others, particularly around administrative access and logging.
Consolidate into a single requirement set per workload class so that selection conversations have one document to work from.
RESULTS
- Concrete security requirements for each class of workload
- Requirements grounded in purpose, compliance and threat rather than vendor material
- Shared reference for platform and provider selection
SUCCESS METRICS
- >80% of workload classes with documented security requirements
- >80% of requirements reviewed against compliance obligations
- >1 requirements review in past 12 months
COSTS
- Buildout and maintenance of requirement sets
- Ongoing overhead from requirements review
PERSONNEL
- Architects (4 days/yr)
- Infrastructure Engineers (2 days/yr)
- Business Owners (2 days/yr)
- Security Auditors (2 days/yr)
RELATED LEVELS
- Threat Assessment - 1
- Policy & Compliance - 1

ACTIVITIES
A. Build an access and trust model for the estate
Build an explicit model of which principals may reach which environments, and which environments may reach one another. This is where infrastructure security stops being about individual systems and starts being about the trust relationships between them.
Express it as a matrix of principal or environment against target, with the conditions in each cell — direct, via a broker with session recording, only through an approved pipeline, or not at all. Include the automation identities, which are frequently more privileged than any human and rarely reviewed.
Reviewing the matrix reliably reveals trust nobody intended: a non-production environment able to reach production data, a build system with standing administrative rights, or a legacy network path left open after a migration.
B. Specify requirements based on known risks
Feed the rated threats and attack paths from Threat Assessment directly into the requirement sets, so that requirements answer identified risks rather than restating generic good practice.
Where an attack path shows a chain running from foothold to production, specify the requirement that interrupts it at the earliest practical point. This frequently produces requirements about credential lifetime, session brokering and network egress rather than about the host itself.
Record the risk each requirement answers, so requirements whose originating risk has been retired can be retired with confidence rather than accumulating.
RESULTS
- Explicit model of trust between principals and environments
- Automation identities enumerated and their privilege made visible
- Requirements traceable to the specific risks they answer
- Unintended trust relationships surfaced
ADD’L SUCCESS METRICS
- >80% of environments covered by the access and trust model
- >80% of requirements traceable to a rated threat
- >1 trust model review in past 6 months
ADD’L COSTS
- Buildout of the access and trust model
- Program overhead from traceability maintenance
ADD’L PERSONNEL
- Architects (4 days/yr)
- Security Analysts (3 days/yr)
- Infrastructure Engineers (2 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Threat Assessment - 2
- Secure Architecture - 2
InfrastructureACTIVITIES
A. Build security requirements into supplier agreements
Carry the organization's requirements into contracts with cloud providers, hosting partners, hardware vendors and managed service providers.
The commitments worth securing in writing are those impossible to retrofit: breach notification obligations and timescales, audit and evidence rights, data residency and deletion guarantees, defined firmware and platform support periods, and constraints on the provider's own administrative access to your environments.
For managed service providers, specify the standard required of the systems they administer and the evidence they must produce, since their access makes their posture effectively part of yours.
B. Expand audit program for infrastructure requirements
Extend routine audit to cover whether systems in service actually meet the requirements specified for their class, closing the loop between what was specified and what was built.
Audit the supplier commitments too. A stated support period is only useful if someone notices when it lapses, and appliances or platform versions approaching end of support need to enter the replacement cycle before they stop receiving fixes.
Report findings into the strategy session so requirement failures inform the roadmap rather than being handled solely as individual exceptions.
RESULTS
- Supplier commitments captured contractually rather than assumed
- Assurance that systems in service meet their specified requirements
- Early warning of platforms approaching end of support
- Requirement failures feeding back into programme planning
ADD’L SUCCESS METRICS
- >90% of infrastructure suppliers under agreements specifying security requirements
- >80% of workload classes audited against requirements in past 12 months
- >90% of the estate with a known platform support end date
ADD’L COSTS
- Legal and procurement overhead from supplier agreements
- Ongoing audit of requirements adherence
ADD’L PERSONNEL
- Architects (2 days/yr)
- Managers (2 days/yr)
- Business Owners (2 days/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Policy & Compliance - 3
- Monitoring & Maintenance - 3

SA1 | SA2 | SA3 | |
| OBJECTIVE | Insert consideration of proactive security guidance into the design process | Direct the design process toward known-secure services and baselines | Formally control the build process and validate utilization |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Maintain a list of recommended platforms and services
Publish the platforms, services and regions the organization supports, and keep it short. Every additional platform multiplies the baselines to write, the tooling to integrate and the expertise the operations team must hold.
Base inclusion on the security requirements already specified — isolation model, identity integration, exportable audit logging, encryption with controllable keys, and a support commitment outlasting the intended service life.
Record what is not supported as clearly as what is, and give the list an owner and review cadence so it tracks vendor and provider lifecycle announcements.
B. Identify and promote secure design principles
Establish the handful of principles every design should honor, and make them explicit so design decisions can be checked against something.
The durable ones are: grant least privilege and prefer short-lived credentials; separate environments so that compromise of one does not imply compromise of another; deny network traffic by default and permit deliberately; keep secrets out of images, code and configuration; make systems reproducible rather than hand-tuned; and ensure that anything protecting the estate cannot itself be reached from the systems it protects.
Circulate the principles to design teams and use them as review criteria.
RESULTS
- Published set of supported platforms, services and regions
- Explicit design principles to check decisions against
- Reduced platform sprawl and the support cost that follows it
SUCCESS METRICS
- >80% of new systems built on recommended platforms
- >80% of design staff aware of the design principles
- Recommended platform list reviewed in past 12 months
COSTS
- Buildout and maintenance of recommended platform list
- Ongoing review of provider lifecycle announcements
PERSONNEL
- Architects (4 days/yr)
- Infrastructure Engineers (3 days/yr)
- Managers (1 day/yr)
- Security Auditors (2 days/yr)
RELATED LEVELS
- Security Requirements - 1
- Education & Guidance - 1

ACTIVITIES
A. Establish and advertise shared security services
Stand up the shared services individual designs should consume rather than reimplement: centralized identity with short-lived credentials, a secrets manager, certificate issuance and rotation, brokered privileged access with session recording, and centralized log collection outside the environments it observes.
Advertise these with clear guidance on consumption. A shared service nobody knows how to use produces the same sprawl as no shared service at all, and teams will build their own.
Instrument the services so adoption can be measured, and treat low adoption as a signal that the service is hard to consume rather than that teams are uncooperative.
B. Establish hardened baselines from recognized benchmarks
Derive the organization's baselines for operating systems, network devices, databases and cloud services from published benchmarks rather than assembling them from scratch. Recognized sources give a defensible starting point and a shared vocabulary with auditors.
Tailor rather than adopt wholesale. A benchmark's strictest profile will break workloads in most organizations, so the work is deciding which settings to relax, recording why, and holding the line on the rest.
Version the resulting baselines and express them as automation wherever possible, so that a system can be said to have been built to a specific, identifiable standard rather than to a document someone read.
RESULTS
- Shared security services available to every design
- Baselines derived from recognized benchmarks with documented deviations
- Baselines expressed as automation rather than prose
- Ability to state which baseline version a system was built to
ADD’L SUCCESS METRICS
- >80% of new systems consuming the shared identity and secrets services
- >80% of platforms with a versioned baseline derived from a recognized benchmark
- >1 baseline review in past 6 months
ADD’L COSTS
- Buildout or license of shared security services
- Ongoing tailoring and maintenance of baselines
ADD’L PERSONNEL
- Architects (4 days/yr)
- Infrastructure Engineers (6 days/yr)
- Managers (2 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Security Requirements - 2
- Implementation Review - 2
InfrastructureACTIVITIES
A. Build reference architectures and automation modules
Produce complete reference architectures for the patterns the organization deploys — not documents describing an architecture, but working automation modules that yield a compliant environment without manual steps.
Reproducibility is the goal: an environment created from the module is identical to every other, and a change is made by altering the module and redeploying rather than by adjusting the running system. This removes the largest single source of drift, which is manual intervention.
Maintain the modules under change control with review, versioning and their own test suite, and make them genuinely easier to use than building by hand. Adoption follows convenience far more reliably than it follows policy.
B. Validate usage of reference architectures
Verify that systems in service were in fact built from the reference modules and continue to match them, rather than assuming that publishing a module ensures its use.
Resources created outside the standard process — during an incident, by a team in a hurry, or inherited through acquisition — are the ones that will not match, and they are disproportionately represented in findings.
Report the proportion of the estate built from reference modules as a programme metric, and route exceptions through the documented process rather than letting them accumulate silently.
RESULTS
- Reference modules producing compliant environments without manual steps
- Substantially reduced drift through reproducible construction
- Measured proportion of the estate built from reference architecture
- Resources created outside the standard process identified rather than invisible
ADD’L SUCCESS METRICS
- >90% of new environments built from a reference module
- >90% of the estate matching its reference architecture
- >1 reference module review and test in past 6 months
ADD’L COSTS
- Buildout and maintenance of reference architectures and modules
- Ongoing validation of module utilization
ADD’L PERSONNEL
- Architects (4 days/yr)
- Infrastructure Engineers (8 days/yr)
- Platform Operators (2 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Threat Assessment - 3
- Implementation Review - 3
- Environment Hardening - 2

DR1 | DR2 | DR3 | |
| OBJECTIVE | Support ad hoc reviews of infrastructure designs to ensure baseline mitigations | Offer assessment services and increase review granularity | Require review of infrastructure designs and audit against expectations |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Identify the attack surface of a design
For each proposed environment, document what it exposes and where its boundaries lie. This need not be elaborate; a diagram and a page of notes is enough to start.
Record what is reachable from the internet, what is reachable from other environments, which identities can administer it, what credentials and secrets it holds, where its data is stored and encrypted, and what it can reach outbound. Egress is routinely omitted and is how most compromises communicate.
The exercise is most valuable where a design diverges from the standard pattern, since that is exactly where an exception gets made quietly.
B. Check the design against known security risks
Review each proposed design against the organization's threat models and design principles, working through the identified threats and asking what in this design addresses each one.
Concentrate on properties that are hard to change later: how environments are separated, how administrative access is obtained, where secrets originate, whether logging leaves the environment it describes, and whether backups can be reached and destroyed from the systems they protect.
Record the review outcome with the design. Even an informal note stating who reviewed it and what was raised gives the next reviewer a starting point.
RESULTS
- Documented attack surface for each proposed environment
- Early identification of designs that diverge from the standard pattern
- Baseline expectation that designs are reviewed before build
SUCCESS METRICS
- >50% of new environment designs reviewed before build in past 6 months
- >80% of designs with a documented attack surface
- >1 design review conducted in past 3 months
COSTS
- Ongoing overhead from design review
- Buildout of review criteria and checklists
PERSONNEL
- Architects (4 days/yr)
- Infrastructure Engineers (3 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Threat Assessment - 1
- Secure Architecture - 1

ACTIVITIES
A. Deploy a formal design review process
Establish a defined review with named reviewers, entry criteria and a recorded outcome, so that review is a step in the process rather than a favor asked of a colleague.
Make the route obvious and the turnaround short. A review process with a long queue will be bypassed, and the designs that bypass it will be the urgent ones that most needed looking at.
Define what a review can conclude — approved, approved with conditions, or rejected — and who can overrule a rejection. An escalation path exercised openly is healthier than one worked around.
B. Analyze designs against the trust model
Extend review beyond the environment to its relationships with everything else, using the access and trust model built under Security Requirements.
For each design, establish what it will be able to reach and what will be able to reach it, including automation identities and management paths. A design that grants a build pipeline standing administrative rights should not pass review merely because the systems themselves are hardened.
This is also where credential lifetime and secret origin should be examined, since most serious infrastructure compromises become serious through credentials rather than through the initially compromised host.
RESULTS
- Formal review process with named reviewers and recorded outcomes
- Designs assessed against what they can reach, not only what they expose
- Automation identities and management paths considered at design time
- Consistent turnaround that keeps review from being bypassed
ADD’L SUCCESS METRICS
- >80% of new designs passing through formal review in past 6 months
- >80% of reviews completed within the defined turnaround
- >80% of designs assessed against the trust model
ADD’L COSTS
- Program overhead from operating a formal review process
- Reviewer time and training
ADD’L PERSONNEL
- Architects (5 days/yr)
- Infrastructure Engineers (3 days/yr)
- Security Analysts (3 days/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Threat Assessment - 2
- Security Requirements - 2
- Implementation Review - 2
InfrastructureACTIVITIES
A. Develop data-level review of infrastructure designs
Deepen review to cover the data an environment will hold and what happens to it under adverse conditions.
Trace specific data classes through the design. Where does regulated data come to rest, and in which jurisdiction? Who holds the keys, and can the provider read it? What is replicated, and where to? What does a backup contain, how long is it kept, and who can delete it? What remains after a system is decommissioned?
These questions are answerable at design time and expensive to answer after a regulator or an incident asks them.
B. Require reviews and audit design compliance
Make review mandatory for designs destined for the higher environment tiers, and verify through routine audit that the requirement is met rather than assuming it.
Audit both that reviews happened and that their conditions were implemented. A review concluding 'approved provided the management network is not routable from the workload subnet' is worth nothing if nobody checked.
Feed audit findings into the strategy session so systematic review failures are addressed as programme problems rather than individually.
RESULTS
- Data-level understanding of what an environment holds and retains
- Mandatory review for high-tier environments with verified conditions
- Audit evidence that review is happening and its conditions implemented
ADD’L SUCCESS METRICS
- >90% of high-tier designs formally reviewed before build
- >90% of review conditions verified as implemented
- >1 audit of the review programme in past 6 months
ADD’L COSTS
- Ongoing overhead from mandatory review and audit
- Buildout of data-level review methodology
ADD’L PERSONNEL
- Architects (4 days/yr)
- Security Analysts (3 days/yr)
- Business Owners (1 day/yr)
- Security Auditors (5 days/yr)
RELATED LEVELS
- Policy & Compliance - 3
- Secure Architecture - 3

IR1 | IR2 | IR3 | |
| OBJECTIVE | Opportunistically find configuration problems in deployed infrastructure | Make configuration review more accurate and efficient through automation | Mandate comprehensive configuration review and gate change against a baseline |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Create review checklists from known requirements
Build a checklist per platform from the security requirements and baselines already defined, expressed as things that can actually be observed on a running system.
Keep it to the settings carrying real weight rather than every line of a benchmark. Public exposure of services and storage, administrative access paths and their authentication, encryption at rest and in transit, logging enabled and leaving the environment, patch currency, and default or shared credentials will identify most of what needs attention.
Have the checklist reviewed by the teams who operate each platform, since they know which settings are commonly relaxed and why.
B. Perform point review of high-risk environments
Sample systems from the higher risk tiers and check them directly against the checklist, rather than relying solely on a console's own compliance view.
Direct inspection matters because tooling reports what agents and APIs tell it, and the systems of most concern are often those outside the tooling's scope entirely — the appliance nobody can install an agent on, the account created outside the standard process, the environment inherited from an acquisition.
Start with internet-facing systems, anything holding regulated data, and the management plane itself. Record findings and route them for remediation with an owner and a date rather than filing them as a report.
RESULTS
- Checklists expressed as observable system configuration
- Direct evidence of the configuration of high-risk environments
- Identification of systems outside the tooling's scope
SUCCESS METRICS
- >50% of high-risk environments point-reviewed in past 6 months
- >80% of platforms with a review checklist
- >80% of findings assigned an owner and remediation date
COSTS
- Buildout of platform review checklists
- Ongoing overhead from manual review
PERSONNEL
- Infrastructure Engineers (4 days/yr)
- Platform Operators (3 days/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Security Requirements - 1
- Secure Architecture - 1

ACTIVITIES
A. Utilize automated configuration and posture assessment
Move from sampling by hand to continuous automated evaluation of the estate against the baselines, using configuration management tooling, benchmark scanners and cloud posture assessment against the provider's control plane.
Tune the tooling to your tailored baselines rather than running default profiles. An assessment reporting thousands of deviations against a benchmark the organization never adopted trains everyone to ignore it.
Scope matters more than depth at this stage. Coverage gaps — an unmonitored account, a region nobody enabled scanning in, a network the collector cannot reach — are more dangerous than a missed setting, because they are invisible rather than merely open.
B. Detect and remediate configuration drift
Compare running configuration against the automation that is supposed to define it, and treat divergence as a finding in its own right rather than an inconvenience.
Drift detection answers a question that compliance scanning does not: not 'is this setting acceptable' but 'did someone change this outside the process'. A manual change that happens to be compliant still indicates a path around the controls, and it will not survive the next redeployment.
Establish what happens when drift is found. Reverting automatically suits reproducible environments; for systems that cannot be safely redeployed, an alert and an owner is the realistic answer. Either way the decision should be deliberate.
RESULTS
- Continuous automated evaluation across the estate
- Coverage gaps treated as findings rather than silence
- Divergence from automation detected as a signal of process bypass
- Defined response when drift is found
ADD’L SUCCESS METRICS
- >80% of the estate automatically assessed in past 1 month
- >80% of assessment findings remediated within the defined window
- <5% of the estate outside automated assessment coverage
ADD’L COSTS
- Buildout or license of posture assessment and drift detection tooling
- Ongoing tuning of assessment rules to tailored baselines
ADD’L PERSONNEL
- Infrastructure Engineers (6 days/yr)
- Platform Operators (4 days/yr)
- Security Analysts (3 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Secure Architecture - 2
- Design Review - 2
- Environment Hardening - 2
InfrastructureACTIVITIES
A. Customize assessment for organization-specific concerns
Extend automated assessment beyond benchmark settings to the configurations that matter specifically to your organization and would not appear in any published standard.
These are usually the interesting ones: the network path a previous incident showed to be dangerous, the legacy protocol that must not be re-enabled, the specific role assignment that would let a build pipeline reach production, or the provider service the organization has decided not to use.
Maintain these custom checks with the same discipline as the baselines, since a check written after an incident years ago may now be enforcing something obsolete.
B. Gate infrastructure change with policy as code
Move enforcement earlier by expressing the standards as machine-checkable policy evaluated before a change is applied, so that non-compliant infrastructure is prevented rather than detected.
Evaluating proposed changes in the pipeline turns a finding that would have taken weeks to remediate into a failed check the author fixes immediately, and it produces an auditable record of what was permitted and why.
Introduce the gate in warning mode first to discover what would have been blocked, and enforce once the noise is understood. Provide an exception route with a named approver, since a gate with no legitimate way past it will be circumvented rather than respected.
RESULTS
- Assessment covering organization-specific configuration concerns
- Non-compliant infrastructure prevented rather than detected after the fact
- Auditable record of what changes were permitted and why
- Exception route with named approvers rather than circumvention
ADD’L SUCCESS METRICS
- >90% of the estate assessed against custom and baseline checks monthly
- >90% of infrastructure changes passing through policy evaluation
- >1 audit requiring a configuration baseline in past 6 months
ADD’L COSTS
- Ongoing maintenance of organization-specific policy
- Program overhead from gate operation and exception handling
ADD’L PERSONNEL
- Infrastructure Engineers (8 days/yr)
- Architects (3 days/yr)
- Security Analysts (3 days/yr)
- Security Auditors (5 days/yr)
RELATED LEVELS
- Policy & Compliance - 3
- Secure Architecture - 3
- Monitoring & Maintenance - 3

ST1 | ST2 | ST3 | |
| OBJECTIVE | Establish process to perform basic security tests based on requirements | Make infrastructure testing more complete and efficient | Mandate infrastructure testing and establish a production release standard |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Derive test cases from known requirements
Turn the security requirements for each environment class into tests that can actually be run, so a requirement can be shown to hold rather than asserted.
Start with the controls the organization is most relying upon. If the design rests on network separation, the test is to attempt to reach a production service from the environment that should not be able to. If it rests on encryption, the test is to read the data directly from storage.
Document the expected result alongside each test so that a change in behavior after a platform update or a configuration change is noticed.
B. Conduct penetration testing on infrastructure
Have a representative environment tested by someone attempting to defeat its controls, working from a realistic starting position rather than with administrative access.
Useful scenarios at this level include an attacker reaching the environment from the internet, an attacker with a foothold on one host inside it, and an attacker holding a low-privilege credential. The question in each case is what they reach and how far they travel.
Test an environment built by the standard process rather than one prepared for the test, and record findings against the reference architecture so remediation lands in the pattern rather than on a single system.
RESULTS
- Test cases derived from the requirements the organization relies upon
- Evidence that key controls withstand attempted defeat
- Findings landing in the reference architecture rather than on individual systems
SUCCESS METRICS
- >50% of environment classes with derived test cases in past 12 months
- >1 penetration test of a standard environment in past 12 months
- >80% of test findings routed to the reference architecture
COSTS
- Buildout of infrastructure test cases
- External or internal penetration testing effort
PERSONNEL
- Infrastructure Engineers (3 days/yr)
- Security Analysts (4 days/yr)
- Security Auditors (4 days/yr)
RELATED LEVELS
- Security Requirements - 1
- Design Review - 1

ACTIVITIES
A. Utilize automated testing and detection validation
Automate the tests that should run repeatedly, and validate that the detection controls covering the estate actually fire.
Detection validation is the higher-value half. Running realistic techniques — credential access, lateral movement, privilege escalation, unusual egress — against a representative environment and confirming each produces the expected alert tells you whether your monitoring works on your estate, which is not the same question as whether the product works.
Run validation after significant platform changes as well as on a schedule, since a migration that quietly stops shipping a log source is a common and silent failure.
B. Test recovery and failover regularly
Exercise the organization's ability to recover, rather than trusting that backups and failover mechanisms will work when needed.
Restore from backup to a clean environment and measure how long it takes and whether the result is usable. Fail over the systems that are supposed to fail over. Compare the measured times against the recovery objectives the business was promised, and report the difference honestly — measured objectives are worth more than aspirational ones.
Include the scenario where the primary environment is not merely unavailable but compromised, since recovering into an environment an attacker still controls is a well-documented way to prolong an incident.
RESULTS
- Automated regression testing of infrastructure controls
- Evidence that detection fires on your own estate
- Measured rather than assumed recovery times
- Recovery tested against compromise, not only against failure
ADD’L SUCCESS METRICS
- >80% of environment classes covered by automated testing in past 6 months
- >1 detection validation exercise in past 3 months
- >80% of critical systems with a tested restore in past 6 months
ADD’L COSTS
- Buildout or license of automated testing and validation tooling
- Program overhead from recovery exercises
ADD’L PERSONNEL
- Infrastructure Engineers (4 days/yr)
- Platform Operators (4 days/yr)
- Security Analysts (5 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Threat Assessment - 2
- Environment Hardening - 2
- Issue Management - 2
InfrastructureACTIVITIES
A. Employ organization-specific test cases
Generate test cases from your own threat models and attack paths rather than a generic catalogue, so testing addresses the paths that matter for your estate.
Where an attack path describes a chain from foothold to production, build a test that attempts the chain and establishes where it breaks. This produces findings expressed in business terms, which is what makes them actionable outside the security team.
Include the response process in scope. Testing whether a compromised system is detected, contained and recovered within the target time measures the whole system rather than any single product.
B. Establish a production release standard
Define the security testing that must pass before an environment or significant change reaches production, and hold the line on it.
Require the standard to be met by the automation rather than by the instance, so that passing is inherited by everything built from the same pattern. This is the difference between testing that scales and testing that becomes a bottleneck.
Require audit evidence that the standard was met, and route exceptions through the documented process with a named approver.
RESULTS
- Test cases generated from the organization's own attack paths
- Whole-system testing covering detection, containment and recovery
- Defined standard that must be met before production release
- Standard met by the automation so that passing is inherited
ADD’L SUCCESS METRICS
- >90% of high-tier environment classes covered by organization-specific test cases
- >90% of production releases meeting the defined testing standard
- >1 end-to-end response and recovery test in past 6 months
ADD’L COSTS
- Buildout of organization-specific test case library
- Program overhead from release standard enforcement
ADD’L PERSONNEL
- Infrastructure Engineers (4 days/yr)
- Architects (2 days/yr)
- Security Analysts (6 days/yr)
- Security Auditors (5 days/yr)
RELATED LEVELS
- Threat Assessment - 3
- Issue Management - 3
- Monitoring & Maintenance - 2

IM1 | IM2 | IM3 | |
| OBJECTIVE | Identify and handle infrastructure issues in an ad hoc manner | Elaborate the response process for consistency and speed | Improve the assurance program through analysis of infrastructure incidents |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Identify point of contact for infrastructure issues
Establish and publish a single obvious route for reporting a suspected infrastructure problem, and make sure it works from outside the affected environment.
Cover the paths reports actually arrive by: an internal engineer noticing something odd, an automated alert, a customer report, a provider notification, and an external researcher who has found an exposed service. The last of these needs a published address that does not require an account to use.
Publish the route where people will find it under pressure, and make sure it reaches a human out of hours.
B. Create informal infrastructure response capability
Assemble a group who can act when an issue is reported, with the access needed to do something about it.
The capabilities that matter most are the ability to isolate a system from the network, revoke credentials and sessions, and reach the provider or vendor. Many organizations discover during their first real incident that the person taking the report holds none of these and must find someone who does.
Write down who holds these permissions and how they are reached outside working hours, and confirm that the mechanism works when the environment it depends on is the one that is broken.
RESULTS
- Published reporting route reachable from outside the affected environment
- Named individuals able to isolate systems and revoke credentials
- Out-of-hours response path that does not depend on the affected estate
SUCCESS METRICS
- >80% of staff aware of how to report a suspected infrastructure issue
- >80% of issues reaching a named responder in past 6 months
- Out-of-hours response path documented and tested in past 12 months
COSTS
- Buildout of reporting routes and contact material
- Ongoing availability of responders
PERSONNEL
- Platform Operators (3 days/yr)
- Security Analysts (3 days/yr)
- Managers (1 day/yr)
RELATED LEVELS
- Education & Guidance - 1
- Strategy & Metrics - 1

ACTIVITIES
A. Establish a consistent infrastructure response process
Define what happens for each class of incident so that response does not depend on who happens to be on call.
Write a short runbook per scenario — suspected host compromise, credential compromise, exposed service, ransomware, provider outage — each stating the immediate containment action, what evidence to preserve, who to notify, and the decision point for isolating or rebuilding.
The containment decision deserves explicit treatment because it trades availability for security. State who can authorize taking a production system offline, what evidence is captured first, and what the fallback is if the business declines. Deciding this in advance converts an argument into a procedure.
Define target times for each step, particularly from detection to containment, which is the interval that determines how far an intrusion spreads.
B. Adopt a vulnerability handling and disclosure process
Establish how infrastructure vulnerabilities are received, triaged and remediated — both those found internally and those reported from outside.
Define severity criteria and a remediation window for each, and hold to them. The common failure is a backlog where everything is important and nothing is scheduled, which is functionally the same as having no process.
Publish a route for external reporters and respond to them. Researchers who cannot find a contact, or who are ignored, tend to escalate in ways the organization controls less well.
RESULTS
- Runbooks covering the common classes of infrastructure incident
- Explicit authority and criteria for the containment decision
- Target times for detection to containment, measured
- Vulnerability handling with severity-based windows that are honored
ADD’L SUCCESS METRICS
- >80% of incidents handled through the defined process in past 6 months
- >80% of incidents meeting the target time from detection to containment
- >80% of vulnerabilities remediated within their severity window
ADD’L COSTS
- Buildout and maintenance of response runbooks
- Program overhead from process operation and triage
ADD’L PERSONNEL
- Platform Operators (5 days/yr)
- Security Analysts (6 days/yr)
- Managers (2 days/yr)
- Business Owners (1 day/yr)
RELATED LEVELS
- Education & Guidance - 2
- Security Testing - 2
- Monitoring & Maintenance - 2
InfrastructureACTIVITIES
A. Conduct root-cause analysis of infrastructure incidents
Investigate what allowed each significant incident to occur and what allowed it to spread, rather than closing on the immediate remediation.
The causes worth finding are structural: the system was unpatched because it was outside the automation, the credential worked everywhere because roles were never separated, the intrusion went unnoticed because that log source stopped shipping months ago, or the environment was reachable because a rule opened during a previous incident was never closed. Each is a programme finding rather than an incident finding.
Route causes to the Practice that owns them and track them to closure. An analysis programme producing recommendations nobody implements erodes willingness to participate.
B. Collect and report infrastructure incident metrics
Instrument the process so its performance can be measured and trended, and report results into the strategy session.
The measurements that drive behavior are time from compromise to detection, detection to containment, the proportion of incidents involving systems outside the standard automation, and the proportion involving conditions that were already known and unremediated. The last is usually the most uncomfortable and the most useful.
Trend over time rather than reporting snapshots, and use the data to argue for the roadmap rather than to allocate blame.
RESULTS
- Structural causes identified and routed to the owning Practice
- Measured time from compromise to detection and detection to containment
- Visibility of how many incidents involve unmanaged or known-vulnerable systems
- Incident data informing programme planning
ADD’L SUCCESS METRICS
- >90% of significant incidents receiving root-cause analysis
- >80% of identified causes closed within the agreed period
- >1 incident metrics report to stakeholders in past 3 months
ADD’L COSTS
- Program overhead from root-cause analysis
- Buildout of incident metrics collection and reporting
ADD’L PERSONNEL
- Security Analysts (7 days/yr)
- Managers (2 days/yr)
- Architects (2 days/yr)
- Security Auditors (3 days/yr)
RELATED LEVELS
- Strategy & Metrics - 3
- Security Testing - 3
- Environment Hardening - 3

EH1 | EH2 | EH3 | |
| OBJECTIVE | Understand and maintain the baseline operating environment | Improve confidence in operation through hardening and access control | Validate resilience continuously and harden the surrounding environment |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Maintain an inventory of systems and their software
Establish and maintain an accurate inventory of the systems the organization runs, along with their operating system versions, installed software and owners.
Inventory is foundational because every other control is scoped by it. An organization cannot patch, assess or decommission what it does not know exists, and the systems missing from inventory are disproportionately those created outside the standard process.
Reconcile sources rather than trusting one. The configuration management database, the cloud provider's resource list, the network's own view and the billing record will each know about systems the others do not, and the differences are the finding. Billing is often the most complete source, since resources that cost money are hard to hide.
Record an owner for every system. Ownerless systems do not get patched, and they are the ones nobody dares turn off.
B. Apply security updates to systems and their software
Establish a consistent process for delivering operating system, firmware and software updates across the estate, with a defined window by environment risk tier.
Set the window from the risk tier rather than uniformly, and measure against it. Latency is the meaningful number: how long, in practice, between a fix becoming available and the estate having it. Most organizations find their real latency is considerably longer than their policy states.
Handle the long tail deliberately. Systems that cannot be restarted, appliances with vendor-controlled update cycles, and anything supporting a business process nobody will schedule downtime for will account for most of the residual risk, and each needs an owner and a compensating control rather than a recurring reminder.
RESULTS
- Reconciled inventory of systems, their software and their owners
- Defined update windows by environment risk tier
- Measured patch latency rather than assumed compliance
- Named ownership of the systems that resist updating
SUCCESS METRICS
- >90% of systems present in a reconciled inventory with a named owner
- >80% of systems within the defined update window for their tier
- >1 inventory reconciliation in past 3 months
COSTS
- Buildout or license of inventory and patch management tooling
- Ongoing overhead from update operation and reconciliation
PERSONNEL
- Infrastructure Engineers (6 days/yr)
- Platform Operators (4 days/yr)
- Managers (1 day/yr)
- Security Auditors (2 days/yr)
RELATED LEVELS
- Secure Architecture - 1
- Implementation Review - 1

ACTIVITIES
A. Establish routine patch and configuration management
Move from delivering updates to managing the process as a service with measurement, exception handling and escalation.
Deploy in stages: a non-production population receives updates first, and wider deployment follows once stable. This is the main protection against an update that breaks a critical workload everywhere simultaneously.
Extend coverage beyond the operating system to the layers routinely omitted — firmware and management controllers, hypervisors, network device software, container base images, and the dependencies baked into golden images. An image rebuilt monthly is patched; one built two years ago and cloned since is not.
Report latency by tier to stakeholders and escalate populations that persistently miss their window.
B. Constrain privilege and network connectivity
Limit what can reach what, and what can be done once reached. This is what determines whether a compromise stays local.
Segment the estate so that environments are separated and the management plane is not reachable from general workloads. Deny network traffic by default and permit deliberately, including egress, which is how most intrusions communicate and exfiltrate.
Remove standing administrative access in favor of brokered, time-bound elevation with session recording, and hold secrets in a managed store rather than in images, configuration or code. Enumerate and constrain automation identities too — they are frequently the most privileged principals in the estate and the least reviewed.
RESULTS
- Staged deployment protecting against estate-wide update failure
- Patch coverage extended to firmware, images and network devices
- Environments and management plane separated by default-deny networking
- Standing administrative access replaced by brokered, time-bound elevation
ADD’L SUCCESS METRICS
- >90% of systems within the defined update window for their tier
- >80% of privileged access obtained through a broker with session recording
- >80% of environments enforcing default-deny egress
ADD’L COSTS
- Buildout of segmentation, privileged access and secrets management
- Program overhead from exception handling during privilege reduction
ADD’L PERSONNEL
- Infrastructure Engineers (10 days/yr)
- Platform Operators (6 days/yr)
- Architects (3 days/yr)
- Security Analysts (3 days/yr)
RELATED LEVELS
- Secure Architecture - 2
- Implementation Review - 2
- Threat Assessment - 3
InfrastructureACTIVITIES
A. Establish resilient backup and recovery
Ensure the organization can recover even when an attacker has had administrative access, which is the scenario ordinary backup arrangements are not designed for.
The properties that matter are isolation and immutability: recovery points that cannot be altered or deleted within their retention window, held outside the identity and network domain of the systems they protect, so that compromise of the estate does not imply compromise of its backups. Attackers target backup infrastructure first precisely because this is so often untrue.
Set recovery objectives per workload from business need, and verify them by measurement rather than assertion. Include the ability to rebuild the foundation itself — identity, networking, the automation — since restoring workloads into an environment that no longer exists is not recovery.
B. Expand audit program for infrastructure hardening
Bring the hardening controls under routine audit so their continued operation is verified rather than assumed.
Audit should confirm both that controls are present and that they are effective. A segmentation rule permitting any-to-any is present but protecting nothing, and a privileged access broker that most engineers bypass is worse than none because it creates false confidence.
Review privileged accounts and automation identities on a defined cycle, removing what is unused. Access accumulates silently, and the accounts of people who changed role years ago are a recurring finding.
Report into the strategy session so systematic weaknesses drive the roadmap.
RESULTS
- Recovery points isolated and immutable against administrative compromise
- Recovery objectives verified by measurement rather than asserted
- Ability to rebuild the foundation, not only the workloads
- Privileged and automation identities reviewed and pruned on a cycle
ADD’L SUCCESS METRICS
- >95% of systems within the defined update window for their tier
- >90% of critical workloads with immutable, isolated recovery points
- >1 privileged access review in past 6 months
ADD’L COSTS
- Buildout or license of immutable backup and recovery capability
- Ongoing audit of hardening effectiveness and access review
ADD’L PERSONNEL
- Infrastructure Engineers (8 days/yr)
- Platform Operators (5 days/yr)
- Security Analysts (4 days/yr)
- Security Auditors (5 days/yr)
RELATED LEVELS
- Issue Management - 3
- Monitoring & Maintenance - 3
- Secure Architecture - 3

MM1 | MM2 | MM3 | |
| OBJECTIVE | Capture the operational information an operator needs | Improve expectations for continuous operation through detailed procedures | Mandate monitoring of estate state and validate the full system lifecycle |
| ACTIVITIES |
|
|
|
| ASSESSMENT |
|
|
|
| RESULTS |
|
|
|
InfrastructureACTIVITIES
A. Capture critical infrastructure telemetry
Identify the information required to establish what is happening in the estate and ensure it is collected and retained centrally.
The core set is well established: authentication and authorization events, administrative and configuration changes, control plane and API activity, network flow records, and the systems' own security logs. Control plane activity deserves emphasis in cloud environments, since that is where the most consequential actions occur and where they are least visible from inside a host.
Collect to a destination outside the environment being observed, so that an attacker who compromises a system cannot erase the record of how they got there.
Review the collected set with the people who respond to incidents, since they will know which source they most often wish they had.
B. Document procedures for typical infrastructure alerts
For the alerts the estate routinely produces, document what they mean and what the recipient should do.
Cover the common ones: administrative access outside expected hours, a change made outside the automation, a new principal granted broad privilege, unusual egress, a log source falling silent, and backup failure. For each, record the likely benign explanation as well as the malicious one, since most alerts have both.
Keep the procedures where responders work, and review them when alert volumes change materially — a rule producing hundreds of alerts a day is training people to ignore it.
RESULTS
- Central record of the telemetry needed to reconstruct events
- Logs held outside the environment they describe
- Documented meaning and response for the common alerts
SUCCESS METRICS
- >80% of systems shipping the core telemetry sources
- >80% of routine alert types with a documented procedure
- >1 review of collected sources with responders in past 12 months
COSTS
- Buildout of telemetry collection and retention
- Ongoing maintenance of alert procedures
PERSONNEL
- Infrastructure Engineers (5 days/yr)
- Platform Operators (4 days/yr)
- Security Analysts (4 days/yr)
RELATED LEVELS
- Environment Hardening - 1
- Issue Management - 1

ACTIVITIES
A. Establish change management for the estate
Establish change management proportionate to impact, recognizing that infrastructure changes can affect every dependent service at once.
Require, for each change: a statement of what it does, the blast radius, a test or pilot result, a rollback plan, and a named owner. Rollback deserves the most scrutiny, since some infrastructure changes are genuinely hard to reverse — a network change that severs the management path removes the channel needed to withdraw it.
Define an emergency path that is faster but still recorded. Emergency changes are legitimate and unavoidable; undocumented ones are the leading source of drift and the reason nobody can explain why a rule exists.
B. Maintain formal operational documentation
Produce and keep current the operational documentation the teams running the estate depend upon, so that operation does not rest on individuals' memory.
Cover build and provisioning per platform, the escalation path for each alert type, recovery procedures including how to obtain credentials when the usual systems are unavailable, dependency maps showing what breaks when a system stops, and the decommissioning sequence.
Assign ownership and a review cadence. Operational documentation decays faster than most, because the systems it describes change continuously and the people who know are too busy to write it down.
RESULTS
- Change management proportionate to blast radius
- Tested rollback plans, including for changes that sever the management path
- Emergency changes recorded rather than invisible
- Current operational documentation with named owners
ADD’L SUCCESS METRICS
- >90% of infrastructure changes passing through change management
- >80% of changes with a tested rollback plan
- >90% of emergency changes documented within the agreed period
ADD’L COSTS
- Program overhead from change management
- Ongoing maintenance of operational documentation
ADD’L PERSONNEL
- Infrastructure Engineers (6 days/yr)
- Platform Operators (6 days/yr)
- Managers (2 days/yr)
- Architects (2 days/yr)
RELATED LEVELS
- Issue Management - 2
- Environment Hardening - 2
- Security Testing - 3
InfrastructureACTIVITIES
A. Expand audit program for operational information
Bring the monitoring itself under audit, verifying that the telemetry the organization believes it has is actually arriving and is actually being watched.
Audit for gaps rather than volume. The questions that matter are which systems are not reporting, which sources have fallen silent, which detections have never fired in a period when they plausibly should have, and which alerts were raised and never actioned. A collector that stopped receiving produces no alerts, which is easily mistaken for an absence of problems.
Verify retention meets both compliance obligations and realistic investigation needs. Intrusions are frequently discovered months after they begin, and retention shorter than that period converts an investigation into speculation.
B. Validate decommissioning and data destruction
Establish and verify the end-of-life process, so that systems leave the estate without leaving anything behind.
The sequence needs to be explicit and confirmed: revoke the system's credentials and certificates, remove its access and trust relationships, confirm data destruction to the standard the data class requires, release addresses and DNS entries, remove it from inventory and monitoring with a recorded disposition, and retain evidence of disposal for owned hardware.
Audit the outcome rather than the intent. Reconciling systems marked decommissioned against those still holding credentials, still resolving in DNS, or still incurring cost reliably finds systems retired in name only. Dangling DNS entries pointing at released addresses deserve particular attention, since they are trivially reclaimed by someone else.
RESULTS
- Audited assurance that telemetry is arriving and being actioned
- Retention matched to realistic investigation timescales
- Verified end-of-life sequence with recorded disposition
- Reconciliation catching systems decommissioned in name only
ADD’L SUCCESS METRICS
- >95% of systems reporting telemetry within the expected interval
- >90% of decommissioned systems with confirmed data destruction and revoked access
- >1 audit of monitoring coverage and retention in past 6 months
ADD’L COSTS
- Ongoing audit of monitoring coverage and retention
- Program overhead from lifecycle reconciliation
ADD’L PERSONNEL
- Infrastructure Engineers (5 days/yr)
- Platform Operators (5 days/yr)
- Security Analysts (5 days/yr)
- Security Auditors (5 days/yr)
RELATED LEVELS
- Policy & Compliance - 3
- Implementation Review - 3
- Environment Hardening - 3









Assessment
Worksheets

| Is there an infrastructure security assurance program already in place? | ||
| Do most of the business stakeholders understand your organization's infrastructure risk profile? | ||
| Is most of your operations staff aware of future plans for the assurance program? | SM1 | |
| Are most of your environments and workloads categorized by risk? | ||
| Are risk ratings used to tailor the required assurance activities? | ||
| Does most of the organization know about what's required based on risk ratings? | SM2 | |
| Is per-environment data for cost of assurance activities collected? | ||
| Does your organization regularly compare your infrastructure spend with other organizations? | SM3 |
| Do most stakeholders know their infrastructure compliance obligations? | ||
| Is it documented which controls are your responsibility and which the provider's? | PC1 | |
| Does the organization utilize a set of policies and standards to control infrastructure? | ||
| Are teams able to request an exception for systems that cannot meet baseline? | PC2 | |
| Are environments periodically audited to ensure a baseline of compliance with policies and standards? | ||
| Does the organization systematically use audits to collect and control compliance evidence? | PC3 |
| Have most engineering and operations staff been given security awareness training? | ||
| Does each team have access to infrastructure security best practices and guidance? | EG1 | |
| Are most roles given role-specific infrastructure security training and guidance? | ||
| Are most teams able to pull in security expertise when they need it? | EG2 | |
| Is infrastructure guidance centrally controlled and consistently distributed? | ||
| Are most privileged staff tested to ensure a baseline skill-set? | EG3 |
Infrastructure| Do most environments in your organization have documented likely threats? | ||
| Does your organization understand and document the types of attackers it faces? | TA1 | |
| Do teams regularly analyze how an attacker would move through an environment? | ||
| Do teams use a method of rating threats for relative comparison? | ||
| Are stakeholders aware of relevant threats and ratings? | TA2 | |
| Do teams specifically consider risk from providers and third parties? | ||
| Are all protection mechanisms and controls captured and mapped back to threats? | TA3 |
| Do most workload classes have specified security requirements? | ||
| Do teams pull requirements from best practices and compliance guidance? | SR1 | |
| Are stakeholders reviewing which principals may reach which environments? | ||
| Are requirements being specified based on feedback from other security activities? | SR2 | |
| Are stakeholders reviewing supplier agreements for infrastructure security requirements? | ||
| Are the security requirements specified for infrastructure being audited? | SR3 |
| Are teams provided with a list of recommended platforms and services? | ||
| Are most teams aware of secure design principles and applying them? | SA1 | |
| Do you advertise shared security services with guidance for design teams? | ||
| Are teams provided with prescriptive baselines based on their platform? | SA2 | |
| Are systems built from centrally controlled reference architectures? | ||
| Are environments being audited for usage of secure architecture components? | SA3 |

| Do teams document what an infrastructure design exposes? | ||
| Do teams check infrastructure designs against known security risks? | DR1 | |
| Do most teams specifically analyze designs for security mechanisms? | ||
| Are most stakeholders aware of how to obtain a formal design review? | ||
| Does the review process incorporate analysis of trust relationships? | DR2 | |
| Does the review process incorporate detailed data-level analysis? | ||
| Does routine audit require a baseline for design review results? | DR3 |
| Do teams have checklists for reviewing deployed infrastructure configuration? | ||
| Are high-risk environments reviewed against their expected configuration? | IR1 | |
| Is automation used to evaluate infrastructure configuration across the estate? | ||
| Do teams detect when running configuration diverges from its definition? | IR2 | |
| Are assessment rules customized for organization-specific concerns? | ||
| Does routine audit require a configuration baseline prior to deployment? | IR3 |
| Are environment classes tested against the requirements specified for them? | ||
| Do you perform penetration testing on standard infrastructure builds? | ST1 | |
| Are teams using automation to evaluate infrastructure security test cases? | ||
| Do you regularly test restoring from backup and measure how long it takes? | ||
| Are most stakeholders aware of test status prior to production release? | ST2 | |
| Are test cases comprehensively generated for organization-specific risks? | ||
| Do routine audits demand minimum standard results from infrastructure testing? | ST3 |
Infrastructure| Do most staff have a point of contact for infrastructure security issues? | ||
| Does your organization have people able to isolate systems and revoke credentials? | IM1 | |
| Does the organization utilize a consistent process for incident reporting and handling? | ||
| Is there a defined remediation window for infrastructure vulnerabilities by severity? | IM2 | |
| Are most incidents inspected for root causes to generate further recommendations? | ||
| Do teams consistently collect and report data and metrics related to incidents? | IM3 |
| Do you maintain an inventory of the systems you operate and who owns them? | ||
| Do you check for and apply security updates to systems and their software? | EH1 | |
| Is a consistent process used to apply upgrades and patches across the estate? | ||
| Is privileged access brokered and time-bound rather than standing? | EH2 | |
| Are backups isolated and immutable against an attacker with administrative access? | ||
| Does routine audit check the estate for baseline environment health? | EH3 |
| Do you collect the telemetry needed to reconstruct what happened on a system? | ||
| Are security-related alerts and error conditions documented for most system types? | MM1 | |
| Are most infrastructure changes managed through a process that's well understood? | ||
| Do teams maintain operational documentation for the platforms they run? | MM2 | |
| Are systems audited to check that expected telemetry is arriving and actioned? | ||
| Is decommissioning and data destruction routinely validated using a consistent process? | MM3 |

https://bsamm.org